Skip to content
@AI4LIFE-GROUP

AI4LIFE-GROUP

The AI4LIFE group at Harvard is led by Hima Lakkaraju. We study interpretability, fairness, privacy, and reliability of AI and ML models.

Popular repositories Loading

  1. OpenXAI OpenXAI Public

    OpenXAI : Towards a Transparent Evaluation of Model Explanations

    JavaScript 259 48

  2. SpLiCE SpLiCE Public

    Sparse Linear Concept Embeddings

    Python 134 16

  3. med-safety-bench med-safety-bench Public

    MedSafetyBench: Evaluating and Improving the Medical Safety of LLMs, NeurIPS 2024

    Python 50 5

  4. LLM_Explainer LLM_Explainer Public

    Code for paper: Are Large Language Models Post Hoc Explainers?

    Jupyter Notebook 34 5

  5. temporal-saes temporal-saes Public

    Codebase for Temporal SAEs paper

    Python 26 2

  6. ROAR ROAR Public

    Jupyter Notebook 5 2

Repositories

Showing 10 of 29 repositories
  • sae_robustness Public

    Official Codebase for Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders (EACL 2026)

    AI4LIFE-GROUP/sae_robustness's past year of commit activity
    Python 3 MIT 0 0 0 Updated Feb 14, 2026
  • med-safety-bench Public

    MedSafetyBench: Evaluating and Improving the Medical Safety of LLMs, NeurIPS 2024

    AI4LIFE-GROUP/med-safety-bench's past year of commit activity
    Python 50 MIT 5 1 0 Updated Dec 4, 2025
  • temporal-saes Public

    Codebase for Temporal SAEs paper

    AI4LIFE-GROUP/temporal-saes's past year of commit activity
    Python 26 Apache-2.0 2 1 0 Updated Nov 14, 2025
  • SpLiCE Public

    Sparse Linear Concept Embeddings

    AI4LIFE-GROUP/SpLiCE's past year of commit activity
    Python 134 Apache-2.0 16 3 0 Updated Mar 27, 2025
  • RLHF_Trust Public
    AI4LIFE-GROUP/RLHF_Trust's past year of commit activity
    Python 5 0 1 0 Updated Dec 21, 2024
  • interp_interv Public

    Code for "Towards Unifying Interpretability and Control: Evaluation via Intervention"

    AI4LIFE-GROUP/interp_interv's past year of commit activity
    Python 2 0 0 0 Updated Nov 8, 2024
  • rocerf_code Public

    Source code for ROCERF

    AI4LIFE-GROUP/rocerf_code's past year of commit activity
    Jupyter Notebook 0 MIT 0 0 0 Updated Sep 2, 2024
  • OpenXAI Public

    OpenXAI : Towards a Transparent Evaluation of Model Explanations

    AI4LIFE-GROUP/OpenXAI's past year of commit activity
    JavaScript 259 MIT 48 7 1 Updated Aug 17, 2024
  • LLM_Explainer Public

    Code for paper: Are Large Language Models Post Hoc Explainers?

    AI4LIFE-GROUP/LLM_Explainer's past year of commit activity
    Jupyter Notebook 34 MIT 5 1 0 Updated Jul 22, 2024
  • average-case-robustness Public

    Characterizing Data Point Vulnerability via Average-Case Robustness, UAI 2024

    AI4LIFE-GROUP/average-case-robustness's past year of commit activity
    Python 0 MIT 0 0 0 Updated May 7, 2024

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…