AI Model Interpretability
Showing 12 of 12 AI Model Interpretability tools on page 1 of 1.
AI Model Interpretability results
Attention-Viewer
A plug-and-play tool for visualizing attention-score heatmap in generative LLMs. Easy to customize for your own…
Captum
Model interpretability and understanding for PyTorch - meta-pytorch/captum
captum.ai
Captum · Model Interpretability for PyTorch Docs Tutorials API Reference GitHub Captum Model Interpretability for…
CEBRA (Nature 2023)
Learnable latent embeddings for joint behavioral and neural analysis - Official implementation of CEBRA -…
EasyEdit
ACL 2024] An Easy-to-use Knowledge Editing Framework for LLMs. - zjunlp/EasyEdit
Llm_Surprisal
Simple tool for generating tokens with open source transformers and/or calculate per-token surprisal. -…
Llm-Transparency-Tool
LLM Transparency Tool (LLM-TT), an open-source interactive toolkit for analyzing internal workings of…
nnsight
The nnsight package enables interpreting and manipulating the internals of deep learned models. - ndif-team/nnsight
SAELens
Training Sparse Autoencoders on Language Models. Contribute to decoderesearch/SAELens development by creating an…
SHAP
Game theoretic approach to explain the output of any machine learning model. Industry standard for model…
Shapash
Shapash: User-friendly Explainability and Interpretability to Develop Reliable and Transparent Machine Learning…
TransformerLens
A library for mechanistic interpretability of GPT-style language models - TransformerLensOrg/TransformerLens