Merged Interpretability Reading List

Deduplicated union of three source lists — A frontier-lab posts (OpenAI · DeepMind · Anthropic), B venue papers 2024–2026 (ACL/EMNLP/ICLR/ICML/NeurIPS), C team reading list. 2810 raw entries → 2678 unique after removing 132 duplicates.

2678
Unique works
991
2026
898
2025
600
2024
29
in List A (labs)
2293
in List B (venues)
402
in List C (team)
#TitleYearVenue / FieldAffiliationLists
1 xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
Intrinsically Interpretable Models Interpretability
2026 ICLR Inria; Massachusetts Institute of Technology; UC San Diego; University of Califo B
2 xHC: Expanded Hyper-Connections 2026 Shanghai Jiao Tong University / Xiaohongshu (RED) C
3 scCBGM: Single-Cell Editing via Concept Bottlenecks
Concept Models Counterfactual Concept Bottleneck Interpretability
2026 ICML Genentech; Genentech Inc.; Guide Labs; New York University; Yale University B
4 daVinci-Dev: Agent-native Mid-training for Software Engineering 2026 SII / Shanghai Jiao Tong University / GAIR C
5 Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language
Probing Analysis Probing
2026 ICML KAIST; Korea Advanced Institute of Science & Technology; POSTECH; Pohang Univers B
6 Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
Explanation Methods Explainability
2026 ICLR Purdue University B
7 Your Language Model Secretly Contains Personality Subnetworks
Mechanistic Interpretability Interpretability
2026 ICLR Northwestern University; Northwestern University, NVIDIA; University of Arizona; B
8 Why These Documents? Explainable Generative Retrieval with Hierarchical Category Paths. 2026 ACL Yonsei University; Korea University B
9 Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics. 2026 ACL Zhejiang University; Alibaba Group (China) BC
10 Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Other Mechanistic Interpretability Explainability
2026 ICLR Tsinghua University B
11 Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention 2026 Stanford/Harvard Kempner/Anthropic C
12 Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations. 2026 ACL University of Illinois Chicago; University of Illinois Urbana-Champaign B
13 Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models. 2026 ACL Tianjin University; German Research Centre for Artificial Intelligence; Saarland B
14 Why Attention Patterns Exist: A Unifying Temporal Perspective Analysis
Mechanistic Interpretability Explainability
2026 ICLR Huawei Technologies Ltd.; Tianjin University; University of Science and Technolo B
15 Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons
Mechanistic Interpretability
2026 ICML Harbin Institute of Technology; Nanjing University; Nanjing University of Scienc B
16 Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents. 2026 ACL Chinese Academy of Sciences; Sinoma Science & Technology Co., Ltd. (China); Stat B
17 Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
Attribution Methods Attribution
2026 ICLR Singapore Management University B
18 Where Concept Erasure Should Occur: Concept–Layer Alignment in Text-to-Video Diffusion Models
Mechanistic Interpretability Concept-based
2026 ICML Huazhong University of Science and Technology; University of Nevada Reno B
19 Where CoT Reasoning Commits: Entropy Traces Identify Interpretable Attention Heads. 2026 ACL Beijing Institute of Technology; Beijing Emergency Medical Center B
20 When Thinking Backfires: Mechanistic Insights into Reason-induced Misalignment
Mechanistic Interpretability Mechanistic Interpretability Explainability
2026 ICLR KAUST; King's College London; King's College London, University of London; Kings B
21 When Safety Alignment Fails to Generalize: Probing with Language Game Jailbreaks. 2026 ACL Chinese Academy of Sciences; University of Chinese Academy of Sciences B
22 When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models
Mechanistic Interpretability Mechanistic Interpretability Attribution
2026 ICLR Carnegie Mellon University; Columbia University; Pennsylvania State University; B
23 When Random Saliency Looks Trained: Architectural Center Bias in CNN Interpretability
Explanation Evaluation Saliency Map Interpretability
2026 ICML University of California, Berkeley; University of North Carolina at Chapel Hill B
24 When More is Less: Understanding Chain-of-Thought Length in LLMs
Other Explainability
2026 ICLR Google DeepMind; MIT; Peking University; Technische Universität München B
25 When Machine Learning Gets Personal: Evaluating Prediction and Explanation
Explanation Evaluation Attribution Explainability
2026 ICLR University California Santa Barbara; University of California, Santa Barbara B
26 When Does Sparsity Mitigate the Curse of Depth in LLMs 2026 Max Planck Institute for Intelligent Systems C
27 When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression
Mechanistic Interpretability Explainability
2026 ICML Michigan State University; University of Michigan - Ann Arbor; calte B
28 When Do Diffusion Models Learn to Generate Multiple Objects?
Other Attribution
2026 ICML KAIST & Cortiq; TU Darmstadt; Technische Universität Darmstadt; University of Ox B
29 What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
Mechanistic Interpretability Sparse Autoencoder Black-box Interpretability
2026 ICLR Facebook; Meta FAIR; University of California, Berkeley B
30 What is Missing? Explaining Neurons Activated by Absent Concepts
Explanation Methods Attribution Interpretability Explainability
2026 ICML ETH Zurich; Johannes-Gutenberg Universität Mainz; MPI for Informatics; Max-Planc B
31 What Makes Effective Supervision in Latent Chain-of-Thought: An Information-Theoretic Analysis
Mechanistic Interpretability Probing
2026 ICML Eastern Institute of Technology, Ningbo; Hong Kong Polytechnic University; Natio B
32 What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input Beliefs
Explanation Evaluation Attribution Post-hoc Explanation Explainability
2026 ICML Bar-Ilan University; New York University B
33 What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom
Probing Analysis Probing
2026 ICML Fudan University; Northwest Polytechnical University Xi'an; Peking University; S B
34 What Do Large Language Models Know About Opinions?
Probing Analysis Sparse Autoencoder Probing
2026 ICLR Hong Kong University of Science and Technology (Guangzhou); University of Califo B
35 What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
Other Probing
2026 ICML Meta; New York University / Meta FAIR; New York University and AmiLabs; Universi B
36 What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial Nudge. 2026 ACL Rochester Institute of Technology B
37 Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
Mechanistic Interpretability Superposition Rule/Symbolic Interpretability
2026 ICML Independent Researcher; Mila, University of Montreal; Peking University; Peking B
38 Weight-sparse transformers have interpretable circuits
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML Massachusetts Institute of Technology; OpenAI; Stanford University; University o BC
39 WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics. 2026 ACL Peking University; University of , USA; Beijing Tongming Lake Information Techno B
40 Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
Mechanistic Interpretability Interpretability
2026 ICLR Carnegie Mellon University B
41 WARP: Weight-Space Analysis for Recovering Training Data Portfolios 2026 University of Wisconsin–Madison C
42 VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
Intrinsically Interpretable Models Interpretability Explainability
2026 ICLR Arizona State University; Facebook; Massachusetts Institute of Technology; Meta B
43 Visual Persuasion: What Influences Decisions of Vision-Language Models?
Probing Analysis Interpretability
2026 ICML BITS Pilani; Dartmouth College; MIT Media Lab; Massachusetts Institute of Techno B
44 VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization
Intrinsically Interpretable Models Interpretability
2026 ICLR The Hong Kong University of Science and Technology; Xiamen University B
45 Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering. 2026 ACL The Hong Kong University of Science and Technology (Guangzhou); Hong Kong Univer B
46 VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
Explanation Methods Interpretability Explainability
2026 ICLR , University of Texas at Austin; Microsoft; Microsoft Research Asia; University B
47 Verifying Chain-of-Thought Reasoning via Its Computational Graph
Mechanistic Interpretability Circuit Analysis Attribution Black-box
2026 ICLR FAIR; Independent; UCL | FAIR (Meta); University of California, Santa Barbara; U BC
48 Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing. 2026 ACL University of Hong Kong; Sun Yat-sen University B
49 Verified SHAP: Provable Bounds for Exact Shapley Values of Neural Networks
Attribution Methods Shapley/SHAP Explainability
2026 ICML Hebrew University of Jerusalem; University of Konstanz B
50 Verification of the Implicit World Model in a Generative Model via Adversarial Sequences
Other Probing
2026 ICLR University of Szeged B
51 Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning. 2026 ACL University of Illinois Urbana-Champaign; Fudan University; Tsinghua University; B
52 Verbalizable Representations Form a Global Workspace in Language Models
monitoring representations
2026 Lab post (Anthropic) Anthropic AC
53 VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck. 2026 ACL Fudan University; Shanghai Artificial Intelligence Laboratory B
54 Unveiling the Visual Counting Bottleneck in Vision-Language Models
Probing Analysis Probing
2026 ICML Department of Computer Science, ETHZ - ETH Zurich; ETH Zurich; ETHZ - ETH Zurich B
55 Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection
Explanation Evaluation Counterfactual Interpretability Explainability
2026 ICLR Australian National University; SUN YAT-SEN UNIVERSITY; Sun Yat-sen University; B
56 Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
Probing Analysis Probing
2026 ICLR INRIA; Max Planck Institute of Biochemistry; Universiteit Gent; University of Ka B
57 Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
Attribution Methods Attribution Interpretability Explainability
2026 ICLR Nanyang Technological University; Umea University; Umeå University B
58 Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention. 2026 ACL Shanghai Jiao Tong University B
59 Universal Redundancies in Time Series Foundation Models
Mechanistic Interpretability Mechanistic Interpretability Attribution Interpretability
2026 ICML University of Texas at Austin B
60 Union-of-Experts: Neurons in Mixture-of-Experts are Secretly Routers. 2026 ACL Renmin University of China B
61 Unifying Low Dimensional Spectra in Deep Learning
Other Explainability
2026 ICML University of Oxford B
62 Unifying Formal Explanations: A Complexity-Theoretic Perspective
Explanation Methods Interpretability Explainability
2026 ICLR Hebrew University of Jerusalem; Nanyang Technological University B
63 Unified Time Series Explanations via Amortized Optimization and Instance-level Multi-Expert Knowledge Distillation
Explanation Methods Attribution Faithfulness Post-hoc Explanation Explainability
2026 ICML Forschungszentrum Juelich GmbH; North China University of Water Resources and El B
64 Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks 2026 CUHK / Carnegie Mellon University / Institute of Science Tokyo / UIUC / UC Santa Barba C
65 Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and Limitations
Mechanistic Interpretability Saliency Map
2026 ICLR Ohio State University, Columbus; Xi'an Jiaotong University B
66 Understanding Reasoning Collapse in LLM Agent Reinforcement Learning
Other Explainability
2026 ICML Apple; City University of Hong Kong; Computer Science Department, Stanford Unive B
67 Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation. 2026 ACL Nanjing University; Huawei Translation Services Center, Beijing, China B
68 Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
Probing Analysis Probing
2026 ICLR ByteDance; Department of Computer Science, University of Wisconsin - Madison; Un B
69 Understanding Emergent Misalignment via Feature Superposition Geometry. 2026 ACL The University of Tokyo; Google DeepMind (United Kingdom) B
70 Understanding Cross-layer Contributions to Mixture-of-Experts Routing in LLMs
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICLR Institute of Science Tokyo; RIKEN B
71 Uncovering the Latent Potential of Deep Intermediate Representations
Probing Analysis Interpretability
2026 ICML Indraprastha Institute of Information Technology, Delhi B
72 Uncovering Sentiment Analysis Circuit in Large Language Model. 2026 ACL Soochow University B
73 Uncovering Hidden Triggers: Backdoor Attribution in Language Models
Mechanistic Interpretability Mechanistic Interpretability Probing Attribution Black-box Interpretability
2026 ICML Abel AI; Nanyang Technological University; National University of Singapore; Squ B
74 Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR Harvard University; Stanford University; University of Michigan B
75 Uncovering Competency Gaps in Large Language Models and Their Benchmarks
Mechanistic Interpretability Sparse Autoencoder Concept-based
2026 ICML Google; Google DeepMdin; Google DeepMind; Google Inc; Stanford University B
76 Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
Probing Analysis Interpretability
2026 ICLR Amazon; Capital One; CapitalOne; University of California, Irvine; University of B
77 UCoder: Unsupervised Code Generation by Internal Probing of Large Language Models. 2026 ACL Beihang University; Huawei B
78 Tversky Neural Networks: Psychologically Plausible Deep Learning with Differentiable Tversky Similarity
Intrinsically Interpretable Models Interpretability
2026 ICLR Computer Science Department, Stanford University; Stanford University B
79 Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution 2026 Anthropic C
80 Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
Attribution Methods Attribution Faithfulness Explainability
2026 ICLR Uni Edinburgh / Uni Amsterdam; University of Amsterdam B
81 Trustworthy and Explainable Causal Representation Learning in Transformers. 2026 ACL Huazhong Agricultural University; Adelaide University; Hainan University B
82 TrustTable: A Neuro-Symbolic Auditing Framework for Faithful Table QA. 2026 ACL Nanjing University of Posts and Telecommunications; The State Key Laboratory of B
83 TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs. 2026 ACL Adelaide University B
84 TreeGrad-Ranker: Feature Ranking via $O(L)$-Time Gradients for Decision Trees
Attribution Methods Shapley/SHAP
2026 ICLR National University of Singapore; University of Waterloo Vector Institute B
85 Tree-of-Evidence: Efficient "System 2" Search for Faithful Multimodal Grounding. 2026 ACL Georgia Institute of Technology B
86 Tree-CoT-RT: An Explainable Multi-Path Tree-Guided Chain-of-Thought and Reinforcement Learning Framework for Aspect Sentiment Quad Prediction. 2026 ACL XU Exponential University of Applied Sciences, Germany; Universiti Sains Malaysi B
87 TravelBehaviorQA: A Benchmark Dataset for Behavioral Interpretation of GPS Trajectories. 2026 ACL University of Maryland, College Park B
88 Translation Heads: Disentangling meaning from language in LLM-based machine translation
Mechanistic Interpretability Mechanistic Interpretability Activation Steering Interpretability
2026 ICML INRIA; Inria; Inria Paris B
89 Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
Explanation Methods Explainability
2026 ICLR Harvard University; Skywork AI; Tencent; Tsinghua Univ.; Tsinghua University; Ts B
90 Transformers learn factored representations
Mechanistic Interpretability Interpretability Explainability
2026 ICML Astera Institute; Astera Institute, Simplex; Beyond Institute for Theoretical Sc B
91 Transformers are inherently succint :ICLR 2026 Outstanding 2026 ETH,RTPU C
92 Transformer Circuits Can Realize Clustering Algorithms
Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML IBM Research; Massachusetts Institute of Technology B
93 Training-free Counterfactual Explanation for Temporal Graph Model Inference
Explanation Methods Counterfactual Explanation Counterfactual Post-hoc Explanation Explainability
2026 ICLR Arizona State University; Case Western Reserve University; University of Massach B
94 Training large language models on narrow tasks can lead to broad misalignment 2026 Joint team: Truthful AI (Berlin) / Center on Long-Term Risk (London) / Warsaw University of Technology / University of Toronto / Stanford University / UC Berkeley et al. C
95 Tracking Equivalent Mechanistic Interpretations Across Neural Networks
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICLR Carnegie Mellon University B
96 Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
Probing Analysis Interpretability
2026 ICLR Emory University; Goethe University Frankfurt; Microsoft Research; NVIDIA B
97 Tracing the Persona Circuit: How Large Language Models Encode and Express Character Traits
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Attribution
2026 ICML Tencent AI Lab; University of Science and Technology of China B
98 Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
Mechanistic Interpretability
2026 ICML Beijing Jiaotong University; Nanyang Technological University; Peking University B
99 Tracing Logit Trajectories Across Layer Depth: Dataset-Level Explainability for Language Models. 2026 ACL Korea Advanced Institute of Science and Technology; Chungnam National University B
100 Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Method
Explanation Evaluation Explainability
2026 ICLR ByteDance; ByteDance Inc.; Institute of Automation, Chinese Academy of Sciences; B
101 TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
Mechanistic Interpretability Sparse Autoencoder Circuit Analysis
2026 ICML Nanjing University; Nanjing University of Science and Technology; Nanjing univer B
102 ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection. 2026 ACL East China Normal University B
103 ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway. 2026 ACL Korea University; Myongji University; University of Texas Health Science; AIGEN B
104 Towards the Explainability of Temporal Graph Networks via Memory Backtracking and Topological Attribution
Explanation Methods Attribution Faithfulness Explainability
2026 ICML Beijing University of Posts and Telecommunications; HKUST(GZ); Hong Kong Univers B
105 Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures. 2026 ACL Peking University; Beijing Academy of Artificial Intelligence; Tsinghua Universi B
106 Towards Understanding the Shape of Representations in Protein Language Models
Probing Analysis Faithfulness Interpretability
2026 ICLR Department of Physics, University of Oslo; University of Oslo B
107 Towards Understanding the Robustness of Sparse Autoencoders. 2026 ACL University of Virginia; Independent Age B
108 Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
Mechanistic Interpretability Sparse Autoencoder Circuit Analysis Superposition Interpretability
2026 ICLR Fudan University; Shanghai Innovation Institute B
109 Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
Mechanistic Interpretability Mechanistic Interpretability
2026 ICLR University of California Berkeley; University of Freiburg / MATS; University of B
110 Towards Understanding Massive Activations in Attention Sink Mechanism
Mechanistic Interpretability Mechanistic Interpretability Explainability
2026 ICML The Chinese University of Hong Kong B
111 Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions
Mechanistic Interpretability Activation Steering Post-hoc Explanation
2026 ICML Ant Group; Zhejiang University; antgroup B
112 Towards Spectroscopy: Susceptibility Clusters in Language Models
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML Resolution; The University of Melbourne; Timaeus B
113 Towards Proactive Information Probing: Customer Service Chatbots Harvesting Value from Conversation. 2026 ACL National University of Singapore; Sichuan University; Engineering Research Cente B
114 Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMs
Attribution Methods Attribution Faithfulness Interpretability Explainability
2026 ICML City University of Hong Kong; Harbin Institute of Technology BC
115 Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures. 2026 ACL Peking University; Beijing Academy of Artificial Intelligence; Nanjing Universit BC
116 Towards Interpretable Visual Decoding with Attention to Brain Representations
Explanation Methods Probing Interpretability Transparency
2026 ICLR Columbia University B
117 Towards Interpretable Tabular Reasoning: Enhancing LLM Reasoning on Tabular Data with Pre-Constructed Logic Graph. 2026 ACL Zhejiang University; MYbank, Ant Group B
118 Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation
Probing Analysis Attribution
2026 ICML Tsinghua University B
119 Towards Explainable Diagnosis: A Self-learned Explanatory Knowledge Base Approach. 2026 ACL Chinese Academy of Sciences B
120 Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
Intrinsically Interpretable Models Faithfulness Interpretability
2026 ICLR Carnegie Mellon University; Duke University; Indian Institute of Technology, Del B
121 Towards Atoms of Large Language Models
Mechanistic Interpretability Sparse Autoencoder Faithfulness
2026 ICML Institute of Automation, Chinese Academy of Sciences; Institute of automation, C B
122 Toward Safe Quantization-Aware Fine-tuning: Understanding and Mitigating Safety Alignment Degradation
Probing Analysis Interpretability
2026 ICML Institute of Intelligent Computing, University of Electronic Science and Technol B
123 Toward Identifiable Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder
2026 ICML Achira Inc; IST Austria; ISTA B
124 Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Faithfulness Post-hoc Explanation Interpretability
2026 ICLR University of Virginia; University of Virginia, Charlottesville B
125 Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts
Mechanistic Interpretability Probing
2026 ICML National University of Singapore; University of Science and Technology of China B
126 TimeSeg: An Information-Theoretic Segment-Wise Explainer for Time-Series Predictions
Explanation Methods Black-box Interpretability Explainability
2026 ICLR Chung-Ang University; Korea University B
127 TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
Mechanistic Interpretability Sparse Autoencoder Faithfulness Black-box Interpretability Explainability
2026 ICML TUM & Télécom Paris; Technical University of Denmark; Technical University of Mu B
128 Time series saliency maps: Explaining models across multiple domains
Attribution Methods Mechanistic Interpretability Saliency Map Attribution Faithfulness Interpretability
2026 ICML EPFL; Swiss Federal Institute of Technology Lausanne (EPFL) B
129 Thought-Action Graph Reasoning: Faithful and Efficient Reasoning of Large Language Models via Reusing Past Experience. 2026 ACL Tsinghua University; Beijing University of Posts and Telecommunications; Zhonggu B
130 Thought Branches: Interpreting LLM Reasoning Requires Resampling
Mechanistic Interpretability Mechanistic Interpretability Counterfactual Faithfulness
2026 ICLR Anthropic Fellows; DeepMind; Duke University; Google DeepMind B
131 This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-Critic
Intrinsically Interpretable Models Post-hoc Explanation Black-box Interpretability Explainability
2026 ICML EPITA Lyon; Sony AI; University of Roma "La Sapienza", IRISA B
132 ThinkPersona: Thinking with Persona Graphs for Faithful Individualized Role-Playing. 2026 ACL Zhejiang University B
133 Think in Latent, Explain in Language: Self-Explainable Latent Reasoning
Explanation Methods Self-explaining Post-hoc Explanation Black-box Interpretability Explainability
2026 ICML Peking University; University of Illinois Champaign Urbana; University of Illino B
134 There Was Never a Bottleneck in Concept Bottleneck Models
Concept Models Concept Bottleneck Interpretability
2026 ICLR Universidad de Zaragoza; University of Cambridge B
135 Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
Probing Analysis Probing
2026 ICLR Cornell University; Department of Computer Science; Department of Computer Scien B
136 The Value of Information in Human-AI Decision-making
Explanation Methods Shapley/SHAP Explainability
2026 ICLR Microsoft; Northwestern University; Northwestern University, Northwestern Univer B
137 The Tutor-Pupil Augmentation: Enhancing Learning and Interpretability via Input Corrections
Other Interpretability
2026 ICLR Eindhoven University of Technology; University of Minnesota; University of Minne B
138 The Tell-Tale Norm: $\ell_2$ Magnitude as a Signal for Reasoning Dynamics in Large Language Models
Mechanistic Interpretability Sparse Autoencoder Probing Faithfulness
2026 ICML Peking University; Zhejiang University B
139 The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity
Mechanistic Interpretability Mechanistic Interpretability Explainability
2026 ICML Huawei Noah's Ark Lab; National University of Singapore; Princeton University; T B
140 The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
Probing Analysis Interpretability
2026 ICLR Imperial College London; Queen Mary University of London; Queen Mary University B
141 The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models
Mechanistic Interpretability Probing
2026 ICML Nanjing University; nanjing university B
142 The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics
Other Interpretability
2026 ICLR Apple; Apple Inc.; Imperial College London B
143 The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Probing Analysis Interpretability
2026 ICML California Institute of Technology; Caltech and NVIDIA; Carnegie Mellon Universi B
144 The Perception–Physics Paradox: Probing Scientific Alignment with TC-Bench
Probing Analysis Probing Interpretability
2026 ICML ISTA; Institute of Science and Technology Austria B
145 The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
Probing Analysis Probing
2026 ICML FAR.AI; Google DeepMind B
146 The Mechanistic Emergence of Symbol Grounding in Language Models
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML Department of Computer Science, University of North Carolina at Chapel Hill; Uni B
147 The Mechanics of Interference: Defusing Distractors in RAG via Sparse Autoencoder Interventions. 2026 ACL Sapienza University of Rome; Universitas Multi Data Palembang; ISTI-CNR; Univers B
148 The Learnability of Model-Theoretic Interpretation Functions in Artificial Neural Networks. 2026 ACL University of California, Santa Cruz B
149 The Lattice Representation Hypothesis of Large Language Models
Mechanistic Interpretability
2026 ICLR Stanford University B
150 The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
Probing Analysis Explainability
2026 ICML TUM; TUM & Télécom Paris; Technical University of Munich; Technical University o B
151 The Information Geometry of Softmax: Probing and Steering
Probing Analysis Probing
2026 ICML DeepMind; University of Chicago; INSEAD; University of Chicago; University of Ch B
152 The Impact of Off-Policy Training Data on Probe Generalisation. 2026 ACL King's College London; University of Cambridge B
153 The Geometry of Representational Failures in Vision Language Models
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML CENTAI; Intesa Sanpaolo AI Research; Northeastern University London; Polytechnic B
154 The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution
Probing Analysis
2026 ICML Center for Information and Language Processing; Huawei Technologies Ltd.; LMU Mu B
155 The Geometry of Reasoning: Flowing Logics in Representation Space
Mechanistic Interpretability Interpretability
2026 ICLR Duke University; Facebook B
156 The Geometry of Narrow Fine-Tuning Degradation: Trajectory Lock-in and Spectral Bifurcation
Mechanistic Interpretability Probing
2026 ICML Huazhong University of Science and Technology; South China University of Technol B
157 The Geometric Origin of Grokking: Accelerating Generalization via Active Structural Reorganization
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML Hong Kong Baptist University; National University of Defense Technology B
158 The First Impression Problem: Internal Bias Triggers Overthinking in Reasoning Models
Probing Analysis Counterfactual Interpretability
2026 ICLR Nanjing University B
159 The Extra Tokens Matter: Disentangled Representation Learning with Vision Transformers
Concept Models Probing
2026 ICML University of Tennessee B
160 The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
Mechanistic Interpretability Probing Interpretability
2026 ICML University of Hamburg B
161 The Deleuzian Representation Hypothesis
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR CEA; CEA LIST B
162 The Cylindrical Representation Hypothesis for Language Model Steering
Mechanistic Interpretability Concept-based Activation Steering
2026 ICML Indian Institute of Technology Patna; MBZUAI; Mohamed bin Zayed University of Ar B
163 The Consciousness Cluster: Emergent Preferences of Models that Claim to be Conscious
representations safety/alignment
2026 Lab post (Anthropic) Anthropic A
164 The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
Probing Analysis Black-box Interpretability
2026 ICLR Carnegie Mellon University; Cornell University; KAIST; KAIST AI; Korea Advanced B
165 The Assistant Axis: Situating and Stabilizing the Character of Large Language Models
activations persona representations
2026 Lab post (Anthropic) Anthropic A
166 The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
Other Interpretability
2026 ICLR Harvard University; Imperial College London; Rocana Venture Partners B
167 The Achilles’ Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities
Mechanistic Interpretability Interpretability
2026 ICLR Beihang University; Beijing University of Aeronautics and Astronautics; Renmin U B
168 The Abstraction Gap in Vision-Language Causal Reasoning
Explanation Evaluation Probing Faithfulness Explainability
2026 ICML University of Nebraska-Lincoln B
169 The Price of Amortized inference in Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Interpretability
2026 ICLR King Abdullah University of Science and Technology; MBZUAI B
170 Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models. 2026 ACL University of Southern California B
171 Temporal superposition and feature geometry of RNNs under memory demands
Mechanistic Interpretability Mechanistic Interpretability Superposition Explainability
2026 ICLR G-Research; Imperial College; Imperial College London; Imperial College London, B
172 Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR Harvard; Harvard University; School of Engineering and Applied Sciences, Harvard B
173 Temporal Geometry of Deep Networks: Hyperbolic Representations of Training Dynamics for Intrinsic Explainability
Other Explainability
2026 ICLR Einndoven University, Tilburg University B
174 Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICML Carnegie Mellon University; EarthDynamics.ai; Max Planck Institute for Software B
175 Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability
Probing Analysis Probing
2026 ICML Virginia Polytechnic Institute and State University B
176 Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insights
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis
2026 ICLR Japan Advanced Institute of Science and Technology; Northwestern University; RIK B
177 Targeted Neuron Modulation via Contrastive Pair Search 2026 Nous Research C
178 Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
Mechanistic Interpretability Black-box Interpretability
2026 ICML ByteDance; ByteDance Inc.; UC Santa Cruz; University of California Santa Cruz; U B
179 Taming Polysemanticity in LLMs: Theory-Grounded Feature Recovery via Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder
2026 ICLR Shanghai Jiaotong University; University of California, San Diego; Yale; Yale Un B
180 Talent or Luck? Evaluating Attribution Bias in Large Language Models. 2026 ACL George Mason University; Bronx High School of Science; University of Washington B
181 Tackling the XAI Disagreement Problem with Adaptive Feature Grouping
Explanation Methods Shapley/SHAP LIME Faithfulness Post-hoc Explanation Explainability
2026 ICLR CortAIx Lab, Thales; Thales Digital Solutions B
182 TabReX: Tabular Referenceless eXplainable Evaluation. 2026 ACL Adobe Systems (United States); Adobe Gastroenterology B
183 TT-Sparse: Learning Sparse Rule Models with Differentiable Truth Tables
Intrinsically Interpretable Models Interpretability Transparency
2026 ICML Continental Automotive Singapore; Nanyang Technological University B
184 TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG. 2026 ACL University of Technology Sydney B
185 TN-SHAP-G: Graph-Structured Tensor Network Surrogates for Shapley Values and Interactions
Attribution Methods Shapley/SHAP Black-box
2026 ICML Montreal Institute for Learning Algorithms, University of Montreal, University o B
186 TIMESLIVER : SYMBOLIC-LINEAR DECOMPOSITION FOR EXPLAINABLE TIME SERIES CLASSIFICATION
Attribution Methods Attribution Faithfulness Post-hoc Explanation Interpretability Explainability
2026 ICLR Northwestern University; Northwestern University, Northwestern University B
187 TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts. 2026 ACL Tianjin University of Science and Technology; Tianjin University of Technology B
188 TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning
Probing Analysis Attention Visualization
2026 ICML Shandong University B
189 Synthesising Counterfactual Explanations via Label-Conditional Gaussian Mixture Variational Autoencoders
Explanation Methods Counterfactual Explanation Counterfactual Explainability
2026 ICLR Imperial College London; JPMorganChase; King's College London B
190 Syntax vs. Semantics: How Transformers Learn Deep Dependencies
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML Beijing University of Posts and Telecommunications B
191 Synergizing Stylometrics with Semantics: Dual-Path Framework for LLM Detection and Attribution. 2026 ACL Chongqing Key Laboratory of Image Cognition; School of Artificial Intelligence a B
192 Symmetry Reveals the In-Context Classifier: Transformers Implement Mean-Shift Dynamics
Mechanistic Interpretability Probing Interpretability
2026 ICML Boston University; Boston University, Google Research B
193 Symmetries in language statistics shape the geometry of model representations
Probing Analysis Probing
2026 ICML EPFL - EPF Lausanne; Johns Hopkins University; UC Berkeley B
194 SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models
Attribution Methods Shapley/SHAP Attribution
2026 ICML Department of Computer Science, University of Washington; University of Washingt B
195 SuperMAN: Interpretable and Expressive Networks over Temporally Sparse Heterogeneous Data
Intrinsically Interpretable Models Interpretability
2026 ICLR Aalborg University; Imperial College London; Meta; National Center of Excellence B
196 Summaries as Centroids for Interpretable and Scalable Text Clustering
Intrinsically Interpretable Models Interpretability
2026 ICLR York University B
197 Structural Inference: Interpreting Small Language Models with Susceptibilities
Mechanistic Interpretability Attribution Interpretability
2026 ICLR Timaeus; University of Melbourne B
198 Stretching Beyond the Obvious: A Gradient-Free Framework to Unveil the Hidden Landscape of Visual Invariance
Probing Analysis Probing Interpretability
2026 ICLR EPFL; Harvard Medical School; Harvard Medical School and MIT; International High B
199 Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
Probing Analysis Probing Activation Steering
2026 ICLR ELLIS Institute & MPI Intelligent Systems, Tübingen AI Center; EPFL; ETH Zurich; B
200 Stop Hardening Everything: A Training-Free Neuron-Level Defense for Neural Ranking Models. 2026 ACL State Key Laboratory of AI Safety; Chinese Academy of Sciences; University of Ch B
201 Step-Resolved Data Attribution for Looped Transformers
Attribution Methods Data Attribution Attribution Interpretability
2026 ICML Google; Harvard; Hasso Plattner Institute & Alphabet; Technical University of Mu B
202 Step-Level Sparse Autoencoder for Reasoning Process Interpretation
Mechanistic Interpretability Sparse Autoencoder Probing Interpretability
2026 ICML City University of Hong Kong; Li Auto Inc.; University of Science and Technology B
203 Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
Explanation Methods Interpretability
2026 ICML Anthropic; Google DeepMind; Harvard University, Harvard University; Independent; B
204 Steering LLM Thinking with Budget Guidance. 2026 ACL University of Massachusetts Amherst; Zhejiang University; MIT-IBM Watson AI Lab B
205 Steering Evaluation-Aware Language Models To Act Like They Are Deployed
Mechanistic Interpretability Activation Steering
2026 ICLR Astra Fellow; DeepMind; Jump Trading; Northeastern University B
206 Steering Away from Refusal: A Black-box Jailbreak Method Based on First-Token Distribution. 2026 ACL State Key Laboratory of AI Safety, Institute of Computing Technology; Chinese Ac B
207 Steering Autoregressive Music Generation with Recursive Feature Machines
Mechanistic Interpretability Probing Interpretability
2026 ICLR UC San Diego; University of California, San Diego B
208 Steer Like the LLM: Activation Steering that Mimics Prompting
Mechanistic Interpretability Activation Steering Interpretability
2026 ICML Nokia Bell Labs BC
209 State-Dependent Safety Failures in Multi-Turn Language Model Interaction
Mechanistic Interpretability Mechanistic Interpretability Probing
2026 ICML A*STAR; Beijing Electronic Science and Technology Institute; Nanyang Technologic B
210 Stable and Explainable Personality Trait Evaluation in Large Language Models with Internal Activations. 2026 ACL Georgia Institute of Technology; South China University of Technology B
211 Stabilizing Equation Learning via Zero-Point Constraints
Intrinsically Interpretable Models Rule/Symbolic Interpretability
2026 ICML Central China Normal University; Wuhan University of Technology B
212 Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML Alibaba Group; East China Normal University; Linköping University; The Universit B
213 Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities
Other Mechanistic Interpretability Black-box Interpretability
2026 ICML Fundação Getúlio Vargas (FGV); Independent; MATS; MRM Investments; Saarland Univ B
214 Splat Regression Models
Intrinsically Interpretable Models Interpretability
2026 ICLR Massachusetts Institute of Technology B
215 Spilled Energy in Large Language Models
Other Probing
2026 ICLR Sapienza University of Rome; University of Roma "La Sapienza" B
216 SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation. 2026 ACL Nankai University; Microsoft (Finland) B
217 Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation. 2026 ACL Jilin University; The University of Tokyo; Key Laboratory of Ancient Chinese Scr B
218 Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
Probing Analysis Sparse Autoencoder Explainability
2026 ICLR Department of Computer Science, ETHZ - ETH Zurich; ETH Zurich, Stanford; ETH Zür B
219 Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Probing Interpretability
2026 ICML Independent; University College London, University of London B
220 Sparse and Faithful Local Explanations with Piecewise Linear Surrogates
Explanation Methods LIME Faithfulness Post-hoc Explanation Black-box Interpretability
2026 ICML Sichuan University B
221 Sparse Relaxed-Lasso Steering: Automatic Sparse Autoencoder Feature Selection for Precise Image Editing
Mechanistic Interpretability Sparse Autoencoder Faithfulness Interpretability
2026 ICML Chinese Academy of Sciences, Chinese Academy of Sciences; Institute of Software, B
222 Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
Intrinsically Interpretable Models Sparse Autoencoder Post-hoc Explanation Interpretability
2026 ICLR Facebook; Meta; Stanford; University of Oxford B
223 Sparse Bayesian Deep Functional Learning with Structured Region Selection
Intrinsically Interpretable Models Interpretability
2026 ICML Shanghai University of Finance and Economics; Yale University B
224 Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
Mechanistic Interpretability Sparse Autoencoder Activation Steering Interpretability
2026 ICML College of William & Mary; College of William and Mary; William & Mary B
225 Sparse Autoencoders are Topic Models
Mechanistic Interpretability Sparse Autoencoder Explainability
2026 ICML TUM; Technical University of Munich B
226 Sparse Autoencoders Trained on the Same Data Learn Different Features
Explanation Evaluation Sparse Autoencoder Interpretability
2026 ICLR EleutherAI B
227 Sparling: End-to-End Spatial Concept Learning via Extremely Sparse Activations
Intrinsically Interpretable Models Concept-based
2026 ICLR Massachusetts Institute of Technology; Massachussets Institute of Technology; Un B
228 Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability
Mechanistic Interpretability Mechanistic Interpretability Attribution Interpretability
2026 ICLR Charles University; FAR.AI; Imperial College London; MATS Research; independent/ B
229 Singular Vectors of Attention Heads Align with Features
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICML Boston University; Boston University, Boston University B
230 Simul-COMET: A Quality Metric for Simultaneous Interpretation in Distant Language Pair Considering Word Order Difference. 2026 ACL Seikei University; Nara Institute of Science and Technology B
231 Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
Mechanistic Interpretability Sparse Autoencoder Black-box
2026 ICLR Berkeley; Independent; University of Chicago B
232 Shortcut-Resistant CAM Distillation for Long-Tailed Recognition
Attribution Methods Explainability
2026 ICML Huazhong University of Science and Technology; Nanjing University of Aeronautics B
233 Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Attribution Interpretability
2026 ICML Yonsei University B
234 Shapley Neuron Values for Continual Learning: Which Neurons Matter Most?
Attribution Methods Shapley/SHAP
2026 ICML Aarhus University B
235 Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
Explanation Evaluation Explainability
2026 ICLR Beijing Jiaotong University; Microsoft Research Asia; Shenzhen University B
236 Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
Mechanistic Interpretability Interpretability
2026 ICLR Apple; Apple, Carnegie Mellon University; Computer Science and Artificial Intell B
237 SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping. 2026 ACL Association for Computational Linguistics B
238 SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
Other Faithfulness Transparency
2026 ICLR Apple; Eberhard-Karls-Universität Tübingen; University of Tübingen B
239 Self-Explaining Hate Speech Detection with Moral Rationales. 2026 ACL Portland State University B
240 Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS. 2026 ACL Queen Mary University of London; Chongqing Key Laboratory of Big Data Intelligen B
241 Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNs
Intrinsically Interpretable Models Faithfulness Interpretability Explainability Transparency
2026 ICLR Iowa State University; University of Electronic Science and Technology of China B
242 Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection. 2026 ACL VNU University of Science, Vietnam; EDF LAB Singapore Pte Ltd (Singapore) B
243 Selective Concept Bottleneck Models Without Predefined Concepts
Concept Models Concept Bottleneck Concept-based Black-box Interpretability
2026 ICML University Freiburg; University of Freiburg; University of Freiburg, Anthropic A B
244 Segment-Level Attribution for Selective Learning of Long Reasoning Traces
Attribution Methods Attribution
2026 ICLR University of Southern California B
245 Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICML Centro Nacional de Inteligencia Artificial; Pontifical Catholic University of Ch B
246 Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems. 2026 ACL State Key Laboratory of Complex System Modeling and Simulation Technology; Chine B
247 Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
Probing Analysis Probing
2026 ICLR Amazon; Arizona State University; ByteDance Inc.; Microsoft; Pennsylvania State B
248 See the Emotion: A Facial Emoji Proxy Modeling for EEG Emotion Recognition
Explanation Methods Attribution Faithfulness Black-box Interpretability Explainability
2026 ICML Anhui University; Hefei University of Technology; MBZUAI; University of Electron B
249 Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories
Probing Analysis Activation Steering
2026 ICML ByteDance; Peking University B
250 SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity. 2026 ACL Vrije Universiteit Amsterdam; Wageningen University & Research; Association for B
251 Scalable and Interpretable Representation Alignment with Ordinal Similarity
Probing Analysis Interpretability
2026 ICML Helmholtz AI, Technical University of Munich; Helmholtz Munich / TUM; Helmholtz B
252 SafeSeek: Universal Attribution of Safety Circuits in Language Models
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Attribution Interpretability
2026 ICML Hong Kong Polytechnic University; Intelligent Science & Technology Academy of CA B
253 SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering. 2026 ACL Kathmandu University B
254 Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework. 2026 ACL Alibaba Group (China); The State Key Laboratory of Blockchain and Data Security, B
255 SVD as a Fast Interpretability Method for Transformers
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Faithfulness Post-hoc Explanation Interpretability
2026 ICML Heidelberg University B
256 STEM: Scaling Transformers with Embedding Modules (Authors) 2026 Meta AI,Carnegie Mellon University C
257 ST-TGExplainer: Disentangling Stability and Transition Patterns for Temporal GNN Interpretability
Explanation Methods Self-explaining Faithfulness Interpretability Explainability
2026 ICML Griffith University; Hangzhou Dianzi University; Royal Melbourne Institute of Te B
258 SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models. 2026 ACL Northeastern University, Shenyang, China; CAS Key Laboratory of Behavioral Scien B
259 SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks. 2026 ACL Oracle B
260 SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language Models. 2026 ACL Zhejiang Lab; Zhejiang University of Technology B
261 SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models. 2026 ACL Xidian University; University of Science and Technology of China; Xi'an Jiaotong B
262 SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models. 2026 ACL University of Maryland, College Park B
263 SLASH the Sink: Sharpening Structural Attention Inside LLMs
Mechanistic Interpretability Attention Visualization
2026 ICML IGSNRR, Chinese Academy of Sciences, Beijing, China; Shanghai Jiao Tong Universi B
264 SIM-CoT: Supervised Implicit Chain-of-Thought
Other Interpretability
2026 ICLR Fudan University; Microsoft; Nanyang Technological University; Shanghai AI Labor B
265 SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning
Attribution Methods Shapley/SHAP Attribution Interpretability Explainability
2026 ICML Edith Cowan University; Huazhong University of Science and Technology; Zhejiang B
266 SCOUT: Selective Coupling via Optimal Unbalanced Transport for Interpretable Text Classification. 2026 ACL Accessible Space; Zhejiang University; Hangzhou Pujian Medical Technology Co., L B
267 SAVOIR: Learning Social Savoir-Faire via Shapley-based Reward Attribution. 2026 ACL University of Hong Kong; Harbin Institute of Technology; Association for Computa B
268 SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability
2026 ICLR Alibaba Group; National University of Singapore; University of Science and Techn B
269 SAEs-BrainMap: Unveiling the Emergence of Specialized Concepts in Deep Models via Brain Alignment
Mechanistic Interpretability Sparse Autoencoder Probing Interpretability
2026 ICML BIT; Beijing Institute of Technology; Westlake University B
270 SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Concept-based Interpretability
2026 ICML CENTAI; Intesa Sanpaolo AI Research; Polytechnic Institute of Turin; University B
271 SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection. 2026 ACL Georgia Institute of Technology; New Jersey Institute of Technology; Chinese Uni B
272 SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training
Mechanistic Interpretability Sparse Autoencoder Black-box Interpretability
2026 ICLR MIT; Meituan; Peking University BC
273 RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
Mechanistic Interpretability Sparse Autoencoder Interpretability Explainability
2026 ICML Brown University; TU Wien B
274 Role-Sensitive Neurons: A Neuron-Level Gain Control Mechanism for Confidence Steering. 2026 ACL National Taiwan University B
275 Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models
Mechanistic Interpretability Mechanistic Interpretability Attribution
2026 ICML Beijing University of Posts and Telecommunications; Shandong University of Scien B
276 Robust Equation Structure Learning with Adaptive Refinement
Intrinsically Interpretable Models Rule/Symbolic
2026 ICLR Department of Computer Science and Engineering, The Chinese University of Hong K B
277 Rhetorical Questions in LLM Representations: A Linear Probing Study. 2026 ACL Independent Age; University of Cincinnati; Amazon B
278 Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models. 2026 ACL Ministry of Education B
279 Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
Mechanistic Interpretability Mechanistic Interpretability Post-hoc Explanation Interpretability
2026 ICML Centrale Supélec; CentraleSupélec; IRT Saint Exupery & Mila; IRT Saint Exupéry & B
280 Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis
Intrinsically Interpretable Models Black-box Interpretability Explainability
2026 ICML Wuhan University B
281 Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICLR CENIA; Centro Nacional de Inteligencia Artificial; Pontificia Universidad Catoli B
282 Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
Probing Analysis Probing Interpretability
2026 ICLR PAII Inc.; Ping An Technology; Pingan Group; Pingan Technology; University of Co BC
283 Rethinking LLM Reasoning: From Explicit Trajectories to Latent Representations 2026 Harbin Institute of Technology C
284 Responsible Text-to-Image Diffusion: Interpretable and Linearly Controllable Semantics for Fair and Safe Generation
Concept Models Interpretability
2026 ICML Kyungpook National University; Queen's University B
285 Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
Other Explainability
2026 ICLR University of Hong Kong B
286 Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal Transport
Probing Analysis Interpretability
2026 ICLR University of California San Diego; University of California, San Diego B
287 RepoShapley: Shapley-Enhanced Context Filtering for Repository-Level Code Completion. 2026 ACL Chinese University of Hong Kong, Shenzhen; ♯Shenzhen Future Network of Intellige B
288 Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics
Probing Analysis Probing Transparency
2026 ICML Johns Hopkins University; University of Amsterdam B
289 Reinforcement Learning Towards Broadly and Persistently Beneficial Models 2026 OpenAI C
290 Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs
Mechanistic Interpretability Circuit Analysis Probing Attribution
2026 ICLR Electronic Engineering, Tsinghua University, Tsinghua University; Tsinghua Unive B
291 Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies. 2026 ACL Chinese University of Hong Kong; Nara Institute of Science B
292 Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
Other Attribution
2026 ICLR State University of New York at Stony Brook; Stony Brook; Stony Brook University B
293 Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
Probing Analysis Probing
2026 ICML Cornell University; Goodfire; Goodfire AI; Harvard University; Harvard Universit B
294 Reasoning Fails Where Step Flow Breaks 2026 Shanghai Jiao Tong University / Fuzhou University / Jilin University C
295 Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing. 2026 ACL Shanghai Jiao Tong University; The University of Sydney; Shenzhen Campus of Sun B
296 Real-Time Visual Attribution Streaming in Thinking Model
Attribution Methods Attribution Faithfulness Attention Visualization
2026 ICML Amazon; Yonsei University; Yonsei University, Stanford University B
297 Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting
Probing Analysis Mechanistic Interpretability Probing
2026 ICML Fraunhofer FIT; Ludwig-Maximilians-Universität München; University of Maryland, B
298 Rashomon Sets of Falling Trees
Intrinsically Interpretable Models Interpretability
2026 ICML Department of Computer Science, Duke University; Duke; Duke University; Universi B
299 RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
Probing Analysis Probing
2026 ICLR Allen Institute for AI; Berkeley; Independent; University of California, Berkele B
300 RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering. 2026 ACL Tongji University; IGI Global Scientific Publishing (United States) B
301 RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
Explanation Evaluation Probing Counterfactual Faithfulness
2026 ICLR Seoul National University B
302 REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control. 2026 ACL National University of Singapore B
303 RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited Data
Explanation Methods Counterfactual Explanation Counterfactual Black-box Explainability
2026 ICML Forschungszentrum Juelich GmbH; Forschungszentrum Jülich, LMU Munich, MCML B
304 RAP-ID: Mechanistic Prompt Injection Detection via Impostor Behavior Analysis. 2026 ACL SK Innovation (South Korea) B
305 RAG-KT: Cross-platform Explainable Knowledge Tracing with Multi-view Fusion Retrieval Generation. 2026 ACL Inner Mongolia University B
306 R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling. 2026 ACL Shanghai Key Laboratory of Trustworthy Computing; East China Normal University; B
307 Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects
Mechanistic Interpretability Sparse Autoencoder Faithfulness Interpretability
2026 ICML Hanyang University; Hanyang Universty B
308 Query Circuits: Explaining How Language Models Answer User Prompts
Mechanistic Interpretability Sparse Autoencoder Circuit Analysis Faithfulness Explainability
2026 ICML University of Oxford; University of Oxford / Martian B
309 Quasi-Monte Carlo Methods Enable Extremely Low-Dimensional Deep Generative Models
Intrinsically Interpretable Models Post-hoc Explanation Interpretability Transparency
2026 ICLR Duke University; New York University B
310 Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML INRIA; MILA, Quebec, Canada; McGill University; McGill University, McGill Univer B
311 Quantifying Cross-Attention Interaction in Transformers for Interpreting TCR-pMHC Binding
Explanation Methods Mechanistic Interpretability Post-hoc Explanation Black-box Interpretability Explainability
2026 ICLR Tulane University; Tulane University School of Medicine B
312 Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness. 2026 ACL University of Southern California; University of California, Los Angeles B
313 Provably Explaining Neural Additive Models
Explanation Methods Post-hoc Explanation Interpretability Explainability
2026 ICLR Hebrew University of Jerusalem; Masaryk University; Technical University of Muni B
314 Prototype-Grounded Concept Models for Verifiable Concept Alignment
Concept Models Concept Bottleneck Concept-based Interpretability Transparency
2026 ICML IBM Research; KU Leuven B
315 Prototype Transformer: Towards Language Model Architectures Interpretable by Design
Intrinsically Interpretable Models Interpretability
2026 ICML Relational Intelligence; TU Wien; TU Wien and University of Oxford; Technische U B
316 ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting
Intrinsically Interpretable Models Interpretability Explainability Transparency
2026 ICLR Alibaba Group; Renmin University of China B
317 Protein Circuit Tracing via Cross-layer Transcoders
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML Georgia Institute of Technology; Georgia Tech B
318 Propaganda AI: An Analysis of Semantic Divergence in Large Language Models
Other Black-box
2026 ICLR Singapore Management University B
319 Prompt Injection as Role Confusion
Probing Analysis Probing
2026 ICML Massachusetts Institute of Technology; NBC News B
320 Profiling the Irrational Agent: Cognitive Modeling of LLM Behaviors in Sequential Jailbreaks
Probing Analysis Counterfactual Interpretability
2026 ICML Chinese Academy of Science; Chinese Academy of Sciences; Institute of Informatio B
321 Probing the Safety Robustness of LLMs in Latent Space. 2026 ACL Tsinghua University; Shanghai Artificial Intelligence Laboratory; Fudan Universi B
322 Probing the Plasticity and Correlation of LLM Value Systems: LLM Value Rankings are Not Stable. 2026 ACL Hong Kong University of Science and Technology (Guangzhou) B
323 Probing the Inductive Bias of Neural Networks through Learning Random Cellular Automata
Mechanistic Interpretability Circuit Analysis Probing
2026 ICML Johannes-Gutenberg Universität Mainz; University of Mainz B
324 Probing the Geometry of Diffusion Models with the String Method
Probing Analysis Probing
2026 ICML CFM; Capital Fund Management; New York University B
325 Probing for Reading Times. 2026 ACL ETH Zurich; Toyota Technological Institute; University College London B
326 Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups. 2026 ACL Politecnico di Milano; University of Science and Technology of China B
327 Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing. 2026 ACL University of Macau B
328 Probing Rotary Position Embeddings through Frequency Entropy
Probing Analysis Probing Explainability
2026 ICLR Ehime University; NTT; NTT Human Informatics Laboratories; NTT corporation B
329 Probing RLVR Training Instability through the Lens of Objective-Level Hacking
Probing Analysis Mechanistic Interpretability Probing Explainability
2026 ICML Alibaba; Alibaba Group; Peking University; Renmin University of China; Tsinghua B
330 Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation. 2026 ACL Johns Hopkins University; Chinese University of Hong Kong; University of Chicago B
331 Probing Cross-modal Information Hubs in Audio-Visual LLMs
Probing Analysis Probing
2026 ICML Chung-Ang University; KAIST; KAIST, Korea Advanced Institute of Science & Techno B
332 Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio. 2026 ACL Chinese University of Hong Kong; Chinese University of Hong Kong, Shenzhen; Stan B
333 ProbeLLM: Automating Principled Diagnosis of LLM Failures
Probing Analysis Probing Interpretability
2026 ICML IBM Research; LMU Munich; LMU Munich, MCML; MIT; Massachusetts Institute of Tech B
334 Probabilistically-routed Bayesian Additive Spanning Trees for Learning on Constrained Domains
Intrinsically Interpretable Models Interpretability
2026 ICML Eli Lilly and Company; Florida State University; Medical University of South Car B
335 ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs. 2026 ACL Peking University; Key Laboratory of High Confidence Software Technologies, Mini B
336 ProConMV: Provenance-Enabled Conceptual Framework for Interpretable Multi-View Diabetic Retinopathy Diagnosis
Concept Models Concept-based Interpretability
2026 ICML Harbin Institute of Technology; Shenzhen University; The University of Nottingha B
337 Priors in time: Missing inductive biases for language model interpretability
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR Bau Lab; Boston University; Harvard University; Kempner Institute, Harvard Unive B
338 Priority-Aware Shapley Value
Attribution Methods Shapley/SHAP Attribution Faithfulness
2026 ICML CMU, Carnegie Mellon University; Carnegie Mellon University; Ohio State Universi B
339 Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization. 2026 ACL McGill University; Mila - Quebec AI Institute; Mohamed bin Zayed University of A B
340 Precise and Interpretable Editing of Code Knowledge in Large Language Models
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICLR BIFOLD & TU Berlin; Heidelberg College; Heidelberg University; Ruprecht-Karls-Un B
341 Pre-training Limited Memory Language Models with Internal and External Knowledge
Other Black-box
2026 ICLR Cornell University; Department of Computer Science, Cornell University BC
342 Position: Your VLM May Not Be Thinking with Interleaved Images
Other Mechanistic Interpretability Transparency
2026 ICML Fudan University B
343 Position: When AI Decides Who Gets an Organ: Multi-Agentic AI Systems in Transplant Medicine Risk Amplifying Disparities Without Targeted Explainability and Deployment Strategies
Other Counterfactual Explainability
2026 ICML University Health Network; University of California, Irvine; University of Toron B
344 Position: Use Sparse Autoencoders to Discover Unknowns
Mechanistic Interpretability Sparse Autoencoder Interpretability Explainability
2026 ICML Cornell; Cornell Tech; Cornell University; UC Berkeley B
345 Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
Explanation Evaluation Interpretability
2026 ICML Arizona State University; Yale B
346 Position: Multi-Agent Explainability Needs Contracts Before Methods
Other Post-hoc Explanation Explainability
2026 ICML Dartmouth; Dartmouth College B
347 Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
Other Probing
2026 ICML Technical University of Munich; University of Exeter; University of Florida; Uni B
348 Position: Interpretability in Deep Time Series Models Demands Semantic Alignment
Other Black-box Interpretability
2026 ICML IBM Research; KU Leuven; SUPSI - University of Applied Sciences Southern Switzer B
349 Position: Interpretability Can Be Actionable
Explanation Evaluation Interpretability
2026 ICML Google DeepMind; Harvard University; Kempner Institute, Harvard University; Mila B
350 Position: In Defense of Information Leakage in Concept-based Models
Concept Models Concept-based Interpretability
2026 ICML University of Oxford B
351 Position: Genomic Model Research Must Move Beyond Anecdotal Evaluation of Interpretability Methods
Explanation Evaluation Faithfulness Interpretability Explainability
2026 ICML Living Systems Institute, University of Exeter; University of Electronic Science B
352 Position: Explanation Stability Is a Property of the Model–Method Pair, Not the Model
Explanation Evaluation Attribution Explainability
2026 ICML Duke-NUS medical School; Singapore Health Services (SingHealth) B
353 Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods
Other Sparse Autoencoder Attribution Explainability
2026 ICML Bosch Research North America; Duke; Google Research; Harvard; Harvard University B
354 Position: Don't Just "Fix it in Post'': A Science of AI Must Study Learning Dynamics
Other Mechanistic Interpretability Post-hoc Explanation Interpretability
2026 ICML EleutherAI; Humans&/CMU; Kempner Institute, Harvard University; Max Planck Insti B
355 Position: Causality is Key for Interpretability Claims to Generalise
Mechanistic Interpretability Counterfactual Interpretability
2026 ICML Boston University; Cold Spring Harbor Laboratory; Max Planck Institute for Intel B
356 Position: Behavioral Systems Require Behavioral Tests
Other Probing
2026 ICML Dartmouth College; MIT Media Lab; Massachusetts Institute of Technology (MIT) B
357 Position: Accountable Deployment of Agentic AI Demands Layered, System-Level Interpretability
Other Interpretability
2026 ICML Dalhousie University; Ontario Tech University; Vector Institute; Vector Institut B
358 PolySHAP: Extending KernelSHAP with Interaction-Informed Polynomial Regression
Attribution Methods Shapley/SHAP Explainability
2026 ICLR Bielefeld University; Claremont McKenna College; New York University B
359 PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
Mechanistic Interpretability Sparse Autoencoder Probing Interpretability
2026 ICML The Cyprus Institute; University of Athens & Archimedes AI, Athena RC; Universit B
360 Persona Features Control Emergent Misalignment
Mechanistic Interpretability Sparse Autoencoder
2026 ICLR Independent Researcher; Massachusetts Institute of Technology; OpenAI; Stanford B
361 PerfCoder: Large Language Models for Interpretable Code Performance Optimization. 2026 ACL University of Alberta; University of Victoria; Huawei Technologies Ltd., Toronto B
362 Patterning: The Dual of Interpretability
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML The University of Melbourne; Timaeus B
363 Patronus: Interpretable Diffusion Models with Prototypes
Intrinsically Interpretable Models Faithfulness Black-box Interpretability
2026 ICLR Technical University of Denmark B
364 PathwayLLM: Explainable Clinical Trajectory Modeling with Structured Pathways for Sepsis Prediction
Intrinsically Interpretable Models Interpretability Explainability
2026 ICML The Second Affiliated Hospital of Zhejiang Chinese Medical University; Xiamen Un B
365 Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
Mechanistic Interpretability Mechanistic Interpretability
2026 ICLR FAR AI; FAR.AI B
366 Partial Soft-Matching Distance For Neural Representational Comparison With Partial Unit Correspondence
Other Interpretability
2026 ICLR New York University; University of California, San Diego B
367 Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation. 2026 ACL Technische Universität Berlin; German Research Centre for Artificial Intelligenc B
368 Paradigm Shift of GNN Explainer from Label Space to Prototypical Representation Space
Explanation Methods Post-hoc Explanation Explainability
2026 ICLR Central South University; Griffith University; Hong Kong Polytechnic University; B
369 PV-SQL: Synergizing Database Probing and Rule-based Verification for Text-to-SQL Agents. 2026 ACL Purdue University West Lafayette B
370 PROBE: PROcess-Based BEnchmark for Hallucination Detection. 2026 ACL NVIDIA; Chinese University of Hong Kong B
371 PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues. 2026 ACL Indian Institute of Technology Patna; Indian Institute of Technology Jodhpur B
372 PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations. 2026 ACL Hong Kong University of Science and Technology (Guangzhou); NYU Shanghai; Dongbe B
373 PR-XAI: PageRank-Based Feature Attribution for Transformers. 2026 ACL Simon Fraser University; University of California, Berkeley B
374 PINNfluence: Interpreting PINNs through Influence Functions
Attribution Methods Data Attribution Attribution Influence Function Interpretability
2026 ICML Fraunhofer HHI; Fraunhofer HHI, Berlin; Fraunhofer HHI, Einsteinufer 37, 10587 B B
375 PGRF-Net: A Prototype-Guided Relational Fusion Network for Diagnostic Multivariate Time-Series Anomaly Detection
Intrinsically Interpretable Models Interpretability Transparency
2026 ICLR Yonsei University B
376 PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes. 2026 ACL Beijing Institute of Technology; Beihang University; Baidu (China) B
377 PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
Explanation Methods Explainability
2026 ICLR Auburn University; Carnegie Mellon University; None B
378 PAR: Training-Free Positional Perturbation and Attention Recycling for Faithful OCR. 2026 ACL Shanghai Jiao Tong University; Higher Education Commission; University of Hong K B
379 Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
Mechanistic Interpretability Faithfulness Black-box
2026 ICML Amazon; MATS; Redwood Research B
380 Origo: Interpretable Multi-physics PDE Foundation Model through Neural Operator Splitting
Intrinsically Interpretable Models Interpretability
2026 ICML Beijing University of Post and Telecommunications; NCEPU; UIC; University of Ken B
381 Optimal Transport Group Counterfactual Explanations
Explanation Methods Counterfactual Explanation Counterfactual Explainability
2026 ICML LMU; Technical University of Madrid; Universidad Politécnica de Madrid; Universi B
382 Open Sourcing Monitorability Evaluations
chain-of-thought monitoring safety/alignment
2026 Lab post (OpenAI) OpenAI A
383 One Probe Won’t Catch Them All: Towards Targeted Deception Detection
Probing Analysis Probing Post-hoc Explanation
2026 ICML Decode Research; Equivariant labs; LASR Labs; Lambda; UK AISI B
384 One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
Mechanistic Interpretability Mechanistic Interpretability Post-hoc Explanation
2026 ICML Stanford University B
385 One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework. 2026 ACL Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University; Jili B
386 Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models. 2026 ACL Qatar Computing Research Institute, HBKU, Doha, Qatar; Association for Computati B
387 On the Variability of Concept Activation Vectors
Explanation Evaluation Concept-based Feature Importance Explainability Transparency
2026 ICML Julius-Maximilians-Universität Würzburg - CAIDAS; University of Würzburg B
388 On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Superposition Interpretability
2026 ICML Columbia University; Stanford; Stanford University B
389 On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR MIT; Peking University B
390 On the Accuracy of Newton Step and Influence Function Data Attributions
Attribution Methods Data Attribution Attribution Influence Function Interpretability
2026 ICML Massachusetts Institute of Technology B
391 On The Geometry and Topology of Representations: the Manifolds of Modular Addition
Mechanistic Interpretability Circuit Analysis
2026 ICLR Leiden University, Dept. of Mathematics, Leiden University; McGill University; M B
392 On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
Other Faithfulness
2026 ICML Apple; Apple Inc.; Harvard University, Apple B
393 On Predictability of Reinforcement Learning Dynamics for Large Language Models
Other Interpretability
2026 ICLR The Hong Kong University of Science and Technology; Tsinghua University, Tsinghu B
394 On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM Agents 2026 CUHK (James Cheng's Group) C
395 OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation. 2026 ACL Northeastern University; University of Southern California; Stony Brook Universi B
396 ODASim: Ordered, Distinctive and Absolute Semantic Similarity for Code Explanation Evaluation. 2026 ACL IBM Research - India B
397 OCP: Outlier-Centric Probing for Dynamic Structured Pruning of LLMs. 2026 ACL The Hong Kong University of Science and Technology (Guangzhou); National Univers B
398 Nonparametric Data Attribution for Diffusion Models
Attribution Methods Data Attribution Attribution Interpretability
2026 ICML National University of Singapore; SMU, Singapore; Sea AI Lab; Tencent B
399 Neuronal Insights into LLM Attacks: Targeted Neuron Tuning for Precise and Robust Vulnerability Patching. 2026 ACL Tianjin University of Science and Technology; Tianjin University of Technology B
400 Neuron-Level Analysis of Cultural Understanding in Large Language Models
Mechanistic Interpretability
2026 ICLR The University of Tokyo; The University of Tokyo / Riken; University of Liverpoo B
401 Neuron-Aware Active Few-Shot Learning for LLMs. 2026 ACL University of Pittsburgh B
402 Neuro-Fuzzy Concept Learning for Interpretable Large Multimodal Models
Concept Models Concept-based Interpretability
2026 ICML Indian Institute of Technology Indore; Indian Institute of Technology, Indore B
403 Neural–Evolutionary Symbolic Regression with Global Constraints: Constraint-Aware Decoding and Reward Shaping
Intrinsically Interpretable Models Rule/Symbolic Interpretability
2026 ICML Beihang University; Beijing University of Aeronautics and Astronautics; Columbia B
404 Neural+Symbolic Approaches for Interpretable Actor-Critic Reinforcement Learning
Intrinsically Interpretable Models Black-box Interpretability Explainability Transparency
2026 ICLR Maincode; Monash University B
405 Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
Concept Models Concept-based Interpretability
2026 ICML German Research Center for AI; Max Planck Institute for Informatics; Technische B
406 NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-Neurons. 2026 ACL State Key Laboratory of General Artificial Intelligence; Peking University B
407 NeuRAG: End-to-End Neural Knowledge Augmentation via Hyper-Neurons. 2026 ACL Nankai University B
408 Negative Pre-activations Differentiate Syntax
Mechanistic Interpretability Interpretability
2026 ICLR Computer Science and Artificial Intelligence Laboratory, Electrical Engineering B
409 Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations (Authors) 2026 Anthropic C
410 Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
activations explainability
2026 Lab post (Anthropic) Anthropic A
411 Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICLR DeepMind; ENS Paris-Saclay; EPFL; Harvard University; Massachusetts Institute of B
412 NSF-CoT: Neuro-Symbolic Formal Verification of Chain-of-Thought Faithfulness in Contextual Question Answering. 2026 ACL The University of Texas at San Antonio B
413 NIMO: a Nonlinear Interpretable MOdel
Intrinsically Interpretable Models Faithfulness Post-hoc Explanation Interpretability Explainability
2026 ICLR University of Basel B
414 NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
Mechanistic Interpretability Sparse Autoencoder Post-hoc Explanation Interpretability
2026 ICML National University of Singapore; Tsinghua University; University of Science and B
415 NEAT: Neuron-Based Early Exit for Large Reasoning Models. 2026 ACL Northeastern University B
416 Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
Intrinsically Interpretable Models Transparency
2026 ICLR City University of Hong Kong; Huawei Technologies Ltd. B
417 Multilingual Routing in Mixture-of-Experts
Mechanistic Interpretability Interpretability
2026 ICLR Fudan University; Google & University of California, Los Angeles; University of B
418 Multi-scale Explainer for Graph Neural Networks
Explanation Methods Post-hoc Explanation Explainability Transparency
2026 ICML Shanxi University; Southeast University; Suzhou University B
419 Multi-ReduNet: Interpretable Class-Wise Decomposition of ReduNet
Intrinsically Interpretable Models Interpretability
2026 ICLR National University of Singapore B
420 Motion Attribution for Video Generation
Attribution Methods Data Attribution Attribution
2026 ICML MIT; NVIDIA; NVIDIA; UMich; Princeton University; Princeton University NVIDIA; U B
421 More Than What Was Chosen: LLM-based Explainable Recommendation Beyond Noisy User Preferences
Other Faithfulness Explainability
2026 ICLR KAIST; SK Telecom; SK Telecom, KAIST B
422 More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing
Probing Analysis Black-box
2026 ICML City University of Hong Kong; University of Science and Technology of China B
423 Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
Explanation Evaluation Mechanistic Interpretability Faithfulness Transparency
2026 ICML Harvard; Harvard University B
424 Modeling Hierarchical Thinking in Large Reasoning Models
Probing Analysis Activation Steering Interpretability
2026 ICML Sharif University of Technology; University of California, Riverside B
425 Mixture of Concept Bottleneck Experts
Concept Models Concept Bottleneck Rule/Symbolic Interpretability
2026 ICML Department of Information Engineering and Mathematical Sciences, University of S B
426 Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization
Intrinsically Interpretable Models Interpretability
2026 ICLR EPFL; EPFL - EPF Lausanne; Massachusetts Institute of Technology B
427 Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling
Probing Analysis Counterfactual Interpretability
2026 ICML KAIST; KAIST AI; KRAFTON; Korea Advanced Institute of Science & Technology B
428 Missingness Bias Calibration in Feature Attribution Explanations
Attribution Methods Probing Attribution Feature Importance Post-hoc Explanation Explainability
2026 ICLR University of Pennsylvania; University of Texas at Austin B
429 MidSteer: Optimal Affine Framework for Steering Generative Models
Mechanistic Interpretability Concept-based
2026 ICML Huawei London; MVP Lab; Quantum Light; Queen Mary University of London; Universi B
430 MicroC-KT: Modeling Community Effect via Learning Micro-Environment for Evidence-Grounded Explainable Knowledge Tracing. 2026 ACL Inner Mongolia University; Jilin University B
431 MetaOthello: A Controlled Study of Multiple World Models in Transformers
Mechanistic Interpretability Mechanistic Interpretability Probing Interpretability
2026 ICML University of Michigan - Ann Arbor; University of Vermont B
432 MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
Mechanistic Interpretability Interpretability
2026 ICML Apple; University of Texas at Austin B
433 Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMs. 2026 ACL Kyoto University; University of Tokyo; National Institute of Informatics; RIKEN B
434 Medical Interpretability and Knowledge Maps of Large Language Models
Mechanistic Interpretability Causal Intervention Saliency Map Interpretability
2026 ICLR Lumos AI B
435 MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning. 2026 ACL National University of Singapore; Chinese University of Hong Kong; Hunan Univers B
436 MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis. 2026 ACL Stanford University; Shenzhen University; Renmin University of China; Xi'an Jiao B
437 Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Interpretability
2026 ICML Wilfrid Laurier University B
438 Mechanistic Interpretability of Text-to-Image Diffusion Models via Cross-Attention Interventions. 2026 ACL School of Computer Science; University of Oklahoma B
439 Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy. 2026 ACL Sharif University of Technology B
440 Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Attribution Interpretability
2026 ICML CNRS; LIG / Université Grenoble Alpes; Université Grenoble Alpes B
441 Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders. 2026 ACL Stanford University B
442 Mechanistic Insights into Deferred Semantic Drift in LLMs. 2026 ACL Dalian University of Technology; Key Laboratory of Social Computing and Cognitiv B
443 Mechanistic Detection and Mitigation of Hallucination in Large Reasoning Models
Mechanistic Interpretability Mechanistic Interpretability
2026 ICLR Renmin University of China; Renmin University of China, Gaoling School of Artifi B
444 Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units;:ICML 2026 2026 Peking University C
445 Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Data Attribution Attribution Influence Function
2026 ICML Peking University B
446 Mechanistic Anomaly Detection via Functional Attribution
Mechanistic Interpretability Mechanistic Interpretability Attribution Influence Function
2026 ICML University of Melbourne B
447 Mechanisms of Introspective Awareness
Mechanistic Interpretability Circuit Analysis Activation Steering Interpretability
2026 ICML Anthropic; Constellation Research Center; Massachusetts Institute of Technology AB
448 Measuring and Mitigating Post-Hoc Rationalization in Reverse Chain-of-Thought Generation
Explanation Evaluation Post-hoc Explanation Explainability
2026 ICML BOSS Zhipin; BOSS Zhipin Career Science Lab; Peking University; University of El B
449 Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos. 2026 ACL Harbin Institute of Technology B
450 MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows. 2026 ACL MoE Key Lab of Artificial Intelligence, X-LANCE Lab; Shanghai Jiao Tong Universi B
451 Maximum Likelihood Reinforcement Learning 2026 Carnegie Mellon University C
452 Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
Other Faithfulness
2026 ICLR Australian Institute for Machine Learning (AIML); Nanjing University of Science B
453 Mask-to-Correct⁺: Leveraging Retriever Diversity for Masking-guided Faithful Fact Correction. 2026 ACL Indian Association for the Cultivation of Science B
454 Markovian Transformers for Informative Language Modeling
Explanation Methods Faithfulness
2026 ICLR Stanford University B
455 Mapping Semantic & Syntactic Relationships with Geometric Rotation
Mechanistic Interpretability Interpretability
2026 ICLR Fuel iX; TELUS Digital B
456 Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICLR NAVER AI Lab; Seoul National University B
457 Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
Attribution Methods Attribution Faithfulness Explainability
2026 ICML INEEJI; KAIST; KAIST/INEEJI; Korea Advanced Institute of Science and Technology B
458 Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing. 2026 ACL University of Delaware B
459 MOD-SR: Unifying Multimodal Learning and Direct Optimization with Gradient-Guided Diffusion Model for Symbolic Regression
Intrinsically Interpretable Models Rule/Symbolic Interpretability
2026 ICML Shanghai Jiao Tong University; Shanghai Jiaotong University B
460 MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models. 2026 ACL University of Science and Technology of China; Beijing Institute for General Art B
461 MICLIP: Learning to Interpret Representation in Vision Models
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Interpretability
2026 ICLR ShanghaiTech University B
462 MCLE-Mol: Empowering LLM with Molecular Comprehension and Low-Cost Continual Evolution for Interpretable Property Prediction. 2026 ACL East China University of Science and Technology B
463 MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
Explanation Methods Explainability
2026 ICML Peking University B
464 MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference. 2026 ACL Chung-Ang University; Adobe Research, USA; Adobe Research B
465 M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding
Explanation Evaluation Interpretability Transparency
2026 ICLR East China Normal University; Hangzhou Medical College; National University of S B
466 LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrieval. 2026 ACL University of Colorado System; University of Colorado Boulder; Loyola University B
467 Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration
Probing Analysis Attention Visualization
2026 ICML Shanghai Jiao Tong University; Shanghai Jiaotong University; Shanghai artificial B
468 Logit Distance Bounds Representational Similarity
Other Probing Interpretability
2026 ICML Helmholtz AI, Technical University of Munich; IT University of Copenhagen; Max P B
469 LogicXGNN: Grounded Logical Rules for Explaining Graph Neural Networks
Explanation Methods Faithfulness Post-hoc Explanation Interpretability Explainability
2026 ICLR McGill University; McGill University, McGill University; University of Toronto; B
470 Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models. 2026 ACL University of Hong Kong; Fudan University; LMU Munich; Tsinghua University; Tech B
471 Locate then Correct: Debiasing Attention Heads in CLIP
Mechanistic Interpretability Mechanistic Interpretability Post-hoc Explanation
2026 ICML A*STAR; Deakin University; Nanyang Technological University; School of Computer B
472 Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
Mechanistic Interpretability Mechanistic Interpretability Attribution Interpretability
2026 ICLR Japan Advanced Institute of Science and Technology; RIKEN / Tohoku Univ.; Univer B
473 Localizing Memorized Regions in Diffusion Models via Coordinate-Wise Curvature Differences
Probing Analysis Explainability
2026 ICML Hanyang University B
474 Local Mechanisms of Compositional Generalization
Probing Analysis
2026 ICML Apple B
475 Linguistic Properties and Model Scale in Brain Encoding: From Small to Compressed Language Models
Probing Analysis Mechanistic Interpretability Probing
2026 ICML Amazon; GE HealthCare; Indian Institute of Technology, Delhi; International Inst B
476 LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?
Probing Analysis Probing Interpretability
2026 ICLR Amazon; Georgia Institute of Technology B
477 Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
Mechanistic Interpretability Interpretability
2026 ICLR California Institute of Technology; Caltech; Facebook B
478 Lightweight and Interpretable Transformer via Unrolling of Mixed Graph Algorithms for Traffic Forecast
Intrinsically Interpretable Models Black-box Interpretability
2026 ICML Tsinghua University; Tsinghua University, Tsinghua University; York University B
479 Lightweight and Faithful Visual Condition Checking in Behavior Trees via Expert-Regularized Reinforcement Learning. 2026 ACL University of Toronto B
480 Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling
Intrinsically Interpretable Models Interpretability
2026 ICLR Autodesk; New York University; New York University Tandon School of Engineering; B
481 Lending Eyesight to Language Models: Modeling and Probing Human scanpath through Transformer Decoder. 2026 ACL Hong Kong Polytechnic University; Mohamed bin Zayed University of Artificial Int B
482 Left–Right Symmetry Breaking in CLIP-Style Vision-Language Models Trained on Synthetic Spatial-Relation Data
Probing Analysis Mechanistic Interpretability Probing
2026 ICML Toyota Motor Corporation B
483 Learning to Weight Parameters for Training Data Attribution
Attribution Methods Data Attribution Attribution
2026 ICLR EPFL; EPFL - EPF Lausanne; State University of New York at Stony Brook B
484 Learning to Interpret Weight Differences in Language Models
Mechanistic Interpretability Interpretability
2026 ICLR MIT; Massachusetts Institute of Technology B
485 Learning multimodal dictionary decompositions with group-sparse autoencoders
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR Dolby; Dolby Labs; Georgia Institute of Technology B
486 Learning for Highly Faithful Explainability
Explanation Methods Faithfulness Explainability
2026 ICLR Beijing Institute of Technology; Microsoft B
487 Learning a Generative Meta-Model of LLM Activations
Mechanistic Interpretability Sparse Autoencoder Probing Interpretability
2026 ICML Anthropic; Electrical Engineering & Computer Science Department; UC Berkeley; Un B
488 Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling. 2026 ACL Zhejiang University; State Key Laboratory of Transvascular Implantation Devices B
489 Learning Through Dialogue: Engagement and Efficacy Matter More Than Explanations. 2026 ACL Centre for Tactile Internet with Human-in-the-Loop; National University of Singa B
490 Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML AE Studio; Agency Enterprise Studio; Independent; Princeton University B
491 Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
Mechanistic Interpretability Interpretability
2026 ICLR University of Maryland, College Park B
492 Learning Protein Structure-Function Relationships through Knowledge-guided Representation Decomposition
Concept Models Interpretability
2026 ICML Fudan University; ICT; Shanghai Smart Logic Technology Co., Ltd.; Tsinghua Unive B
493 Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies
Explanation Methods Explainability
2026 ICLR ELLIS Institute; MPI for Intelligent Systems; Max Planck Institute for Intellige B
494 Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding. 2026 ACL Nanyang Technological University B
495 Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement Learning
Intrinsically Interpretable Models Interpretability
2026 ICML Huawei Technologies; National University of Singapore; Zhejiang University B
496 Learning Explicit Single-Cell Dynamics Using ODE Representations
Intrinsically Interpretable Models Mechanistic Interpretability Interpretability
2026 ICLR Chan Zuckerberg Initiative; IST Austria CZI; ISTA; Institute of Science and Tech B
497 Learning Efficient and Interpretable Multi-Agent Communication
Intrinsically Interpretable Models Interpretability
2026 ICLR BIGAI; Shandong University B
498 Learning Concept Bottleneck Models from Mechanistic Explanations
Concept Models Sparse Autoencoder Mechanistic Interpretability Concept Bottleneck Black-box Interpretability
2026 ICLR Computer Science and Artificial Intelligence Laboratory, Electrical Engineering B
499 Learning Coherent Representations: A Topological Approach to Interpretability
Intrinsically Interpretable Models Interpretability
2026 ICML NTNU; Norwegian University of Science and Technology B
500 Learning AND–OR Templates for Compositional Representation in Art and Design
Intrinsically Interpretable Models Interpretability Explainability
2026 ICLR Beijing Electronic Science and Technology Institute; Beijing Institute for Gener B
501 Learn from A Rationalist: Distilling Intermediate Interpretable Rationales
Intrinsically Interpretable Models Black-box Interpretability
2026 ICML University of Alberta B
502 LatentQA: Teaching LLMs to Decode Activations Into Natural Language
Mechanistic Interpretability Probing Transparency
2026 ICLR Transluce, UC Berkeley; UC Berkeley; University of California, Berkeley B
503 LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
Explanation Methods Interpretability
2026 ICML McGill University; Meta; Mila (Montreal)/McGill University; Mila - Quebec AI Ins B
504 Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
Probing Analysis Interpretability
2026 ICLR Ohio State University, Columbus; Xi'an Jiaotong University B
505 Latent Planning Emerges with Scale
Mechanistic Interpretability Mechanistic Interpretability
2026 ICLR Anthropic; University of Amsterdam B
506 Latent Concept Disentanglement in Transformer-based Language Models
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICLR Google; Google Research; USC; University of Oxford; University of Southern Calif B
507 LassoFlexNet: a Flexible Neural Architecture for Tabular Data
Intrinsically Interpretable Models Interpretability
2026 ICML Borealis AI; NACE.AI; Qube Research and Technologies; RBC Borealis AI B
508 Large Vision-Language Models Get Lost in Attention
Mechanistic Interpretability Attribution
2026 ICML Alibaba Group; Beijing Academy of Artificial Intelligence(BAAl) ; Beijing Univer B
509 Large Language Model Agents Are Not Always Faithful Self-Evolvers
Other Faithfulness
2026 ICML Harbin Institute of Technology; Singapore Management University B
510 Language Models are Injective and Hence Invertible
Other Interpretability Transparency
2026 ICLR EPFL; EPFL - EPF Lausanne; National and Kapodistrian University of Athens; Sapie BC
511 Language Models Represent and Transform Concepts with Shared Geometry 2026 Georgia Tech C
512 Language Model Circuits Are Sparse in the Neuron Basis
Mechanistic Interpretability Sparse Autoencoder Circuit Analysis Attribution Interpretability
2026 ICML Google; Stanford University; Transluce / MIT; University of California, Berkeley B
513 Label-Free Mitigation of Spurious Correlations in VLMs using Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR State University of New York at Buffalo; State University of New York, Buffalo; B
514 Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference. 2026 ACL UCLouvain; F.R.S.-FNRS B
515 LaVCa: LLM-assisted Visual Cortex Captioning
Probing Analysis Black-box
2026 ICLR Fukui Computer Holdings,Inc; Nagoya Institute of Technology; University of Osaka B
516 LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
Probing Analysis Probing
2026 ICLR Baidu; Baidu Inc; Institute of automation, Chinese academy of science, Chinese A B
517 LLMs Process Lists With General Filter Heads
Mechanistic Interpretability Interpretability
2026 ICLR Northeastern University B
518 LLMs Lean on Priors, Not Programming Language Semantics
Probing Analysis Probing
2026 ICML Cisco; Google; University of Texas at Austin B
519 LLM-induced Rationales for More Compact Explainable Style Classification Models. 2026 ACL Bentley University; Metropolitan University B
520 LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines. 2026 ACL Stanford University; Independent Age; University of California, Irvine; Chinese B
521 LLM Self-Recognition: Steering and Retrieving Activation Signatures
Mechanistic Interpretability Attribution Interpretability
2026 ICML FU Berlin; Freie Universitat Berlin; Freie Universität Berlin B
522 LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals. 2026 ACL Technion – Israel Institute of Technology B
523 LEAF: Towards Lightweight Explainable Hateful Video Detection via Self-Grounding CoT Guided Stage-Wise Distillation. 2026 ACL University of Electronic Science and Technology of China; Sichuan Youjianzhihui B
524 LDEDE: LRP-Driven Efficient Detection and Editing Framework for LLM Privacy Neurons. 2026 ACL Information Engineering University, Zhengzhou, 450001, Henan, China; State Key L B
525 LAFaCT: Attribution-based Localization and Focused Sequential Analysis of Fact-Critical Tokens for Hallucination Detection. 2026 ACL University of Science and Technology of China B
526 Knowing the Unknown: Interpretable Open-World Object Detection via Concept Decomposition Model
Concept Models Interpretability
2026 ICML Northwest Polytechnical University; Northwest Polytechnical University Xi'an; No B
527 Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
Mechanistic Interpretability Attribution
2026 ICML George Mason University B
528 KODA: Contrastive Representation Comparison and Alignment for Vision-Language Foundation Models
Probing Analysis Interpretability
2026 ICML Chinese University of Hong Kong (CUHK); The Chinese University of Hong Kong B
529 KANO: Kolmogorov-Arnold Neural Operator
Intrinsically Interpretable Models Interpretability
2026 ICLR Caltech; MIT; Massachusetts Institute of Technology; UC Santa Barbara; Universit B
530 Joint Distribution–Informed Shapley Values for Sparse Counterfactual Explanations
Explanation Methods Shapley/SHAP Counterfactual Explanation Counterfactual Attribution Post-hoc Explanation
2026 ICLR King/Microsoft; Technical University of Denmark; University of Copenhagen B
531 JX4MEI: Multimodal Semantically-Enhanced LLM for Joint Multimodal Emotion-Intent Explanation and Classification. 2026 ACL Northeastern University B
532 Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility
Probing Analysis Mechanistic Interpretability Interpretability
2026 ICLR Brown University; DeepMind; Johns Hopkins University B
533 Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML TU Dortmund University / Lamarr Institute; University of Tübingen, TU Dortmund B
534 Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training 2026 University of Minnesota / Peking University / Amazon C
535 Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers. 2026 ACL Department of Computer Science University of Texas at Dallas B
536 Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization. 2026 ACL University of North Carolina at Chapel Hill; University of North Carolina Health B
537 Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective. 2026 ACL National Key Laboratory for Multimedia Information Processing; Peking University B
538 Investigating Counterfactual Unfairness in LLMs towards Identities through Humor. 2026 ACL Yonsei University; Korea Advanced Institute of Science and Technology; Seoul Nat B
539 Introspection Adapters: Training LLMs to Report Their Learned Behaviors
Other Faithfulness
2026 ICML Anthropic; Anthropic Fellows B
540 Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
Mechanistic Interpretability Interpretability
2026 ICLR Brown University; Harvard University; Stanford; University of Michigan; York Uni B
541 Interpreting and Steering State-Space Models via Activation Subspace Bottlenecks
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICML AppViewX; Indian Institute of Technology Hyderabad; Microsoft Research B
542 Interpreting Physics in Video World Models
Mechanistic Interpretability Mechanistic Interpretability Probing Interpretability
2026 ICML AMI Labs; Advanced Machine Intelligence; Brown University; FAR.AI; Goodfire AI; B
543 Interpreting Genomic Language Models using Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML University of Pennsylvania; University of Pennsylvania, University of Pennsylvan B
544 Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation. 2026 ACL Arizona State University B
545 Interpretable Semantic Gradients in SSD: A PCA Sweep Approach and a Case Study on AI Discourse. 2026 ACL Warsaw University of Technology B
546 Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation
Explanation Methods Black-box Interpretability Explainability Transparency
2026 ICML Technical University of Munich; Technische Universität München B
547 Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation. 2026 ACL Beijing University of Posts and Telecommunications B
548 Interpretable Neural ODEs for Gene Regulatory Network Discovery under Perturbations
Intrinsically Interpretable Models Interpretability
2026 ICML Columbia University; Columbia University & New York Genome Center; Genentech / C B
549 Interpretable Functional Koopman Learning with Non-Markovian Closure for Spatiotemporal Systems
Intrinsically Interpretable Models Interpretability
2026 ICML Fudan University B
550 Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML Google; Google DeepMind; Massachusetts Institute of Technology; Stanford Univers BC
551 Interpretable Coreference Resolution Evaluation Using Explicit Semantics. 2026 ACL Sapienza University of Rome; Babelscape B
552 Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
Concept Models Concept-based Interpretability Explainability
2026 ICLR MPI Informatics; Max Planck Institute for Informatics; Saarland Informatics Camp B
553 Interpretability from the Ground Up: Stakeholder-Centric Design of Automated Scoring in Educational Assessments. 2026 ACL Stanford University B
554 Interpretability and Generalization Bounds for Learning Spatial Physics
Mechanistic Interpretability Mechanistic Interpretability Black-box Interpretability
2026 ICML Google; OpenAI; Sandia National Laboratories B
555 Interpretability Transfer from Language to Vision via Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML IIT Kanpur; Lambda; Samsung; University of Bath B
556 Interpretability Driven Evolutionary Approach for the Design of Biological Sequences
Explanation Methods Interpretability Explainability
2026 ICML Northwestern University B
557 Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
Probing Analysis Interpretability
2026 ICLR Carnegie Mellon University B
558 InteracSPARQL : An Interactive System for SPARQL Query Refinement Using Natural Language Explanations. 2026 ACL University of Waterloo B
559 Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Circuit Analysis Probing Interpretability
2026 ICML University of Delaware; University of Virginia B
560 Information Flow Reveals When to Trust Language Models
Probing Analysis Interpretability
2026 ICML HKUST(GZ); Jilin University; The Hong Kong University of Science and Technology B
561 InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition :ICML 2026 2026 ByteDance C
562 Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback
Intrinsically Interpretable Models Rule/Symbolic
2026 ICML AstraZeneca; Google DeepMind / University of Cambridge; University of Cambridge; B
563 Influence Dynamics and Stagewise Data Attribution
Attribution Methods Data Attribution Attribution
2026 ICLR Independent; Timaeus; University College London, University of London; Universit B
564 Inferring the Invisible: Neuro-Symbolic Rule Discovery for Missing Value Imputation
Intrinsically Interpretable Models Interpretability
2026 ICLR Texas A&M University; The Chinese University of Hong Kong; The Chinese Universit B
565 Induction Heads Interpolate N-Grams
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML EPFL; Indian Institute of Technology, Madras B
566 In-context learning of representations can be explained by induction circuits
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Explainability
2026 ICLR Northeastern University B
567 In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
Probing Analysis Transparency
2026 ICLR MPI-SWS; Max Planck Institute for Software Systems; Max Planck Institute for Sof B
568 Improving Adversarial Robustness of Attribution via Implicit Regularization
Attribution Methods Attribution Explainability
2026 ICML Brown University; KTH; KTH Royal Institute of Technology B
569 Imagination Helps Visual Reasoning, But Not Yet in Latent Space
Mechanistic Interpretability Probing
2026 ICML Beijing Jiaotong University; Institute of automation, Chinese academy of science B
570 ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
Explanation Evaluation Attribution Explainability
2026 ICLR Academia Sinica; Beever AI; CCHUML; Center for Intelligent Multidimensional Data B
571 INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs. 2026 ACL State Key Laboratory of AI Safety; Chinese Academy of Sciences; School of Advanc B
572 IDEA: An Interpretable and Editable Decision-Making Framework for LLMs via Verbal-to-Numeric Calibration. 2026 ACL The Hong Kong University of Science and Technology (Guangzhou); Huawei Technolog B
573 ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical Coding. 2026 ACL Pennsylvania State University; Stony Brook University B
574 I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR Adelaide University; The University of Adelaide; The University of Melbourne; Un B
575 How do LLMs Compute Verbal Confidence?
Probing Analysis Probing Activation Steering Post-hoc Explanation Black-box
2026 ICML Google; Google DeepMind; Google DeepMind / University of Cambridge B
576 How can embedding models bind concepts?
Probing Analysis Probing
2026 ICML KAIST & Cortiq; University of Tübingen B
577 How Transformers Represent Hierarchies: A Local-to-Global Mechanism
Mechanistic Interpretability Mechanistic Interpretability Probing
2026 ICML Yale University B
578 How Transformers Learn Causal Structures In-Context: Explainable Mechanism Meets Theoretical Guarantee
Mechanistic Interpretability Explainability
2026 ICLR Yale; Yale University B
579 How To Open the Black Box: Modern Models for Mechanistic Interpretability
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Circuit Analysis Probing Superposition
2026 ICLR Patsnap; University of British Columbia B
580 How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
Probing Analysis Faithfulness
2026 ICML University of California, San Diego B
581 How Language Models Process Negation
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICML USC; USC Information Sciences Institute; University of Southern California B
582 How Hard Is Science?
Intrinsically Interpretable Models Rule/Symbolic Interpretability
2026 ICML University of Cambridge B
583 How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
Mechanistic Interpretability Mechanistic Interpretability Superposition Explainability
2026 ICML Saarland University; Universität des Saarlandes B
584 How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning
Probing Analysis Probing
2026 ICML IBM Research; WeChat AI, Tencent; WeChat AI, Tencent Inc. B
585 How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
Mechanistic Interpretability Mechanistic Interpretability Interpretability
2026 ICLR Department of Computer Sciences, University of Wisconsin - Madison; University o B
586 How Do Language Models Speak Languages? A Case Study on Unintended Code-Switching
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML Alibaba Cloud; Alibaba Group; Beijing Automobile Works; UniTTEC Co. Ltd.; Univer B
587 How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study. 2026 ACL Beijing Institute of Technology; Key Laboratory of Computing Power Network and I B
588 Hinge Regression Tree: A Newton Method for Oblique Regression Tree Splitting
Intrinsically Interpretable Models Transparency
2026 ICLR Harbin Institute of Technology; Harbin Institute of Technology, Shenzhen B
589 Hierarchical Concept-based Interpretable Models
Concept Models Concept-based Interpretability Explainability
2026 ICLR University of Cambridge B
590 Hierarchical Causal Abduction: A Foundation Framework for Explainable Model Predictive Control
Explanation Methods LIME Faithfulness Interpretability Explainability
2026 ICML Technische Universität Chemnitz B
591 Hide&Seek: Learning to Explain in an End-to-End Differentiable Network
Attribution Methods Black-box
2026 ICML University of Technology Sydney B
592 Hidden in Plain Sight -- Class Competition Focuses Attribution Maps
Attribution Methods Attribution Transparency
2026 ICML CISPA Helmholtz; CISPA Helmholtz Center; Max Planck Institute for Informatics B
593 Hidden Breakthroughs in Language Model Training
Other Interpretability
2026 ICLR Google Research; Harvard University; Kempner Institute, Harvard University B
594 HiPPO Zoo: Explicit Memory Mechanisms for Interpretable State Space Models
Intrinsically Interpretable Models Interpretability
2026 ICML Duke University B
595 Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs
Attribution Methods Attribution Faithfulness Interpretability
2026 ICLR United States Military Academy; University of Florida; University of Oklahoma B
596 Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video Detection
Explanation Methods Interpretability Explainability
2026 ICML HKUST(GZ), CUHK; Shanghai Jiao Tong University; The Chinese University of Hong K B
597 Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
Mechanistic Interpretability Mechanistic Interpretability Probing Interpretability
2026 ICLR UMass Amherst; University of Massachusetts at Amherst; University of Massachuset B
598 Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation Shaping
Probing Analysis Counterfactual
2026 ICML Nanyang Technological University; Sichuan University; Zhejiang University B
599 Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization
Attribution Methods Attribution Interpretability Explainability
2026 ICLR Sapienza University of Rome; University of Genoa; University of Genoa sAIfer L B
600 Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing
Other Explainability
2026 ICML Columbia University; Northwestern University B
601 Hallucination Reduction with CASAL: Contrastive Activation Steering for Amortized Learning
Mechanistic Interpretability Activation Steering Interpretability
2026 ICLR Facebook; Independent; Meta; Meta FAIR; New York University; University of Cambr B
602 Hallucination Begins Where Saliency Drops
Attribution Methods Saliency Map Interpretability
2026 ICLR Alibaba Group; Dalian Martime University; Institute of automation, Chinese acade B
603 Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders 2026 Tsinghua University C
604 Guaranteed Optimal Compositional Explanations for Neurons
Explanation Methods Explainability
2026 ICML MIT; University of California, Santa Cruz B
605 Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation
Other Counterfactual Interpretability
2026 ICLR DFKI; German Research Center for AI; Saarland University, Universität des Saarla B
606 Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
Mechanistic Interpretability Mechanistic Interpretability Attribution Faithfulness
2026 ICLR MBZUAI; University of Maryland, College Park B
607 Graph Signal Processing Meets Mamba2: Adaptive Filter Bank via Delta Modulation
Intrinsically Interpretable Models Interpretability
2026 ICLR KAIST; Korea Advanced Institute of Science & Technology; Korea Advanced Institut B
608 Graph Explorer: Training Faithful KG Agents with Visibility-Grounded Supervision. 2026 ACL Boston University; University of Washington B
609 GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs. 2026 ACL University of North Carolina at Chapel Hill; The University of B
610 Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency 2026 KAUST C
611 Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
Probing Analysis Circuit Analysis Probing
2026 ICML Devoteam B
612 Geometric Collapse: When Vision Models Fail to Verify Physical Causality
Other Counterfactual
2026 ICML CUHK; Department of Computer Science and Engineering, The Chinese University of B
613 Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
Mechanistic Interpretability Interpretability
2026 ICML Northwestern; Northwestern University; Northwestern University, Northwestern Uni B
614 Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline. 2026 ACL Xi'an Jiaotong University; Hefei University of Technology; CSIRO; Northwestern P B
615 GenSR: Symbolic regression based on equation generative space
Intrinsically Interpretable Models Rule/Symbolic
2026 ICLR Eastern Institute of Technology, Ningbo; Imperial College London; Shanghai Jiaot B
616 Gauge-invariant representation holonomy
Probing Analysis Probing
2026 ICLR Athena RC; Athena Research Center B
617 GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
Attribution Methods Counterfactual Data Attribution Attribution
2026 ICML Sony AI; Sony Group Corporation; Stanford University B
618 GRASP: Awakening Latent Spatial Reasoning in LVLMs via Training-free Geometric Rectification
Probing Analysis Probing Counterfactual
2026 ICML Henan Univeristy; Soochow University; Suzhou University B
619 GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
Intrinsically Interpretable Models Black-box Interpretability Explainability
2026 ICLR Apple; Apple ML Research; Apple MLR and Mila; Meta B
620 GNN Explanations that do not Explain and How to find Them
Explanation Evaluation Self-explaining Faithfulness Explainability
2026 ICLR Fondazione Bruno Kessler; TU Wien; University of Trento B
621 GAVEL: Towards Rule-Based Safety through Activation Monitoring
Mechanistic Interpretability Interpretability Transparency
2026 ICLR Amrita Vishwa Vidyapeetham (Deemed University); Ben Gurion University of the Neg B
622 GARLIC: Graph Attention-based Relational Learning of Multivariate Time Series in Intensive Care
Intrinsically Interpretable Models Attribution Interpretability Explainability Transparency
2026 ICLR Department of Informatics, University of Zurich, University of Zurich; ETH Zuric B
623 GALAX: Graph-Augmented Language Model for Explainable Reinforcement-Guided Subgraph Reasoning in Precision Medicine
Explanation Methods Mechanistic Interpretability Interpretability Explainability
2026 ICLR Washington University; Washington University in Saint Louis; Washington Universi B
624 Functional Decomposition and Shapley Interactions for Interpreting Survival Models
Attribution Methods Shapley/SHAP Interpretability Explainability
2026 ICML Bielefeld University; Leibniz Institute for Prevention Research and Epidemiology B
625 Functional Attention: From Pairwise Affinities to Functional Correspondences : ICML 2026 2026 Technical University of Munich + University of Oxford + UT Austin C
626 Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
Mechanistic Interpretability Circuit Analysis Counterfactual Interpretability
2026 ICLR Salesforce AI Research; University of Southern California B
627 From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution. 2026 ACL Technical University of Applied Sciences Würzburg-Schweinfurt; JetBrains Researc B
628 From Weights to Activations: Is Steering the Next Frontier of Adaptation? 2026 ACL Saarland University; German Research Centre for Artificial Intelligence; Centre B
629 From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment. 2026 ACL Technical University of Munich; Munich Center for Machine Learning; Lund Univers B
630 From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets
Intrinsically Interpretable Models Interpretability
2026 ICML Department of Computer Science, Duke University; Duke; Duke University; Universi B
631 From RAG to Agentic RAG for Faithful Islamic Question Answering. 2026 ACL Qatar Computing Research Institute, HBKU, Qatar; Hamad Bin Khalifa University; A B
632 From Nodes to Narratives: Explaining Graph Neural Networks with LLMs and Graph Context. 2026 ACL University of Illinois Chicago; Indian Institute of Technology Kharagpur B
633 From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts? 2026 ACL Boston University; Harvard University; Mila – Quebec AI Institute; University of B
634 From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models. 2026 ACL Tokyo University of Science B
635 From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models. 2026 ACL Tianjin University; Alibaba Group (China) B
636 From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models. 2026 ACL Tongji University; University of Wisconsin–Madison B
637 From Growing to Looping: A Unified View of Iterative Computation in LLMs
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML Google; Helmholtz AI | Technical University of Munich; Helmholtz Munich; KTH Sto B
638 From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations. 2026 ACL Tata Consultancy Services Research; Indian Institute of Technology Patna; Univer B
639 From Data Statistics to Feature Geometry: How Correlations Shape Superposition
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Superposition Interpretability
2026 ICLR Imperial College; Imperial College London B
640 From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
Mechanistic Interpretability Attribution Interpretability
2026 ICLR Meta AI; New York University B
641 From Basis to Basis: Gaussian Particle Representation for Interpretable PDE Operators
Intrinsically Interpretable Models Interpretability
2026 ICML The Hong Kong University of Science and Technology; The Hong Kong University of B
642 From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards. 2026 ACL William & Mary; Anytime AI B
643 Fresh in memory: Training-order recency is linearly encoded in language model activations
Probing Analysis Probing
2026 ICLR University of Cambridge B
644 Formalizing the Binding Problem
Probing Analysis Probing
2026 ICML Carnegie Mellon University; Mila - Quebec AI Institute, University of Pennsylvan B
645 Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICLR Hebrew University of Jerusalem B
646 Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based Learning
Concept Models Concept-based Interpretability
2026 ICML Amazon; IIT Hyderabad; Indian Institute of Technology, Hyderabad B
647 Forget-It-All: Multi-Concept Machine Unlearning via Concept-Aware Neuron Masking
Mechanistic Interpretability Saliency Map
2026 ICML Clemson University; University of Arizona; University of Georgia; University of B
648 Forest Before Trees: Latent Superposition for Efficient Visual Reasoning. 2026 ACL Mohamed bin Zayed University of Artificial Intelligence; Fudan University; Renmi B
649 ForensicConcept: Transferable Forensic Concepts for AIGI Detection
Concept Models Attribution Black-box Explainability
2026 ICML Tencent Youtu Lab; Westlake University; Xiamen University B
650 Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs. 2026 ACL AI Singapore B
651 Focus and Dilution: The Multi-stage Learning Process of Attention
Other Explainability
2026 ICML Shanghai Jiao Tong University; Shanghai Jiaotong University B
652 FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models
Probing Analysis Post-hoc Explanation
2026 ICLR Delineate Inc.; Microsoft; Stevens Institute of Technology; University of Centra B
653 Flow-Disentangled Feature Importance
Attribution Methods Attribution Feature Importance Black-box Interpretability
2026 ICLR SUN YAT-SEN UNIVERSITY; St. Jude Children's Research Hospital; University of Hon B
654 Fix the Mind, Not the Move: Interpretable AI Assistance via Knowledge-Gap Localization
Other Interpretability
2026 ICML University of Southern California B
655 First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
Attribution Methods Influence Function
2026 ICLR University of South Florida B
656 FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models. 2026 ACL University of California, Los Angeles B
657 Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
Attribution Methods Attribution
2026 ICML Carnegie Mellon University; Goethe University Frankfurt; Sony AI B
658 Fiction Flows: A Replication and Reinterpretation of Narrative Sequentiality. 2026 ACL Cornell University; McGill University B
659 Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
Mechanistic Interpretability Sparse Autoencoder Superposition Interpretability
2026 ICLR Stony Brook University; University of Cambridge; University of Cambridge & Googl B
660 Feature segregation by signed weights in artificial vision systems and biological models
Probing Analysis
2026 ICLR Harvard Medical School; Harvard Medical School, Harvard University B
661 Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
Mechanistic Interpretability Attribution
2026 ICML The Chinese University of Hong Kong B
662 Fast Retrieval and Slow Reasoning for Explainable Multimodal Sentiment Analysis. 2026 ACL Anhui Province Key Laboratory of Affective Computing and Advanced Intelligent Ma B
663 Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML Google; Google DeepMind; Google Deepmind; Google Inc.; UT Austin; University of B
664 False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize. 2026 ACL National University of Singapore; Peking University; University of California, D B
665 FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning
Explanation Methods Black-box Interpretability Explainability
2026 ICLR Ant Group; AntGroup; Shanghai Jiao Tong University; Shanghai Jiaotong University B
666 Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation. 2026 ACL ETH Zurich; Mohamed bin Zayed University of Artificial Intelligence B
667 Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence. 2026 ACL The University of Texas at Austin; Universidad del Noreste; Scripps MD Anderson B
668 Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
faithfulness chain-of-thought monitoring
2026 Lab post (Anthropic) Anthropic A
669 Faithfulness Under the Distribution: A New Look at Attribution Evaluation
Explanation Evaluation Attribution Faithfulness
2026 ICLR MI2.AI - WarsawTech; Suzhou Yierqi; University of Technology Sydney; University B
670 Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs. 2026 ACL Shanghai Jiao Tong University; Hong Kong University of Science and Technology; A B
671 Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance. 2026 ACL Tel Aviv University B
672 Faithful Persona Steering under Incongruity via Dual-Stream Refinement. 2026 ACL National Yang Ming Chiao Tung University B
673 Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
Mechanistic Interpretability Causal Intervention Counterfactual Faithfulness Interpretability
2026 ICLR Ant Group; National Certification Technology (Hangzhou) Co., Ltd; Zhejiang Norma B
674 FaithLens: Detecting and Explaining Faithfulness Hallucination. 2026 ACL Fudan University; Tsinghua University; University of Illinois Urbana-Champaign; B
675 FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
Explanation Evaluation Counterfactual Faithfulness Interpretability Explainability Transparency
2026 ICLR Arizona State University; Department of Computer Science, University of North Ca B
676 Factorized Scheduling Principle: Learning Interpretable and Transferable Policies via Structured Additive Functions
Intrinsically Interpretable Models Interpretability
2026 ICML Sejong University B
677 FLAIR: Steering LLM Mathematical Problem Solving based on A Fuzzy-Logic-AssIsted Reasoner. 2026 ACL Central China Normal University; University of Wollongong; University of Wiscons B
678 FIPN: Forward Self-Organizing Interpretable Polynomial Networks for Time Series Forecasting
Intrinsically Interpretable Models Interpretability Explainability Transparency
2026 ICML College of Information Communication Technology University of Suwon; Linyi Unive B
679 FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation. 2026 ACL University of Oregon B
680 FAME: Formal Abstract Minimal Explanation for Neural Networks
Explanation Methods Explainability
2026 ICLR Airbus; Airbus SAS, IRT Saint Exupéry; Hebrew University of Jerusalem B
681 Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks
Explanation Evaluation Attribution Interpretability Explainability
2026 ICML Emory University; Griffith University (Australia); Michigan State University; Pe B
682 Exploring and Distilling Multi-Dimensional Clues for Interpretable Social Bot Detection. 2026 ACL Singapore Management University; Beijing Language and Culture University B
683 Exploring Layer Activation Dynamic of CoT via Knowledge Probe. 2026 ACL Southeast University B
684 Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
Concept Models Interpretability Explainability
2026 ICLR Microsoft Research; Microsoft Research Asia; Shanghai Artificial Intelligence La B
685 Exploring Accurate and Transparent Domain Adaptation in Predictive Healthcare via Concept-Grounded Orthogonal Inference
Intrinsically Interpretable Models Black-box Explainability Transparency
2026 ICML Stevens Institute of Technology; University of Massachusetts Chan Medical School B
686 Explanations are a Means to an End: Decision Theoretic Explanation Evaluation
Explanation Evaluation Mechanistic Interpretability Interpretability Explainability
2026 ICML Northwestern University; UCSD B
687 Explanation Quality Assessment as Ranking with Listwise Rewards. 2026 ACL Université d'Artois; AIL Research (United States) B
688 Explaining Sources of Uncertainty in Automated Fact-Checking. 2026 ACL University of Copenhagen B
689 Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
Other Explainability
2026 ICLR The University of Tokyo; Tokyo University, Tokyo Institute of Technology B
690 Explaining Concept Shift with Interpretable Feature Attribution
Attribution Methods Attribution Interpretability
2026 ICML CMU, Carnegie Mellon University; Carnegie Mellon University; School of Computer B
691 Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection. 2026 ACL Griffith University B
692 Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
Attribution Methods Explainability
2026 ICLR Alibaba Group; Nanyang Technological University; Zhejiang University; Zhejiang U B
693 Explainable Quantum Program Repair with Verifiable Proof Traces. 2026 ACL Singapore Management University B
694 Explainable Mixture Models through Differentiable Rule Learning
Intrinsically Interpretable Models Interpretability Explainability Transparency
2026 ICLR CISPA Helmholtz; CISPA Helmholtz Center for Information Security; Universität de B
695 Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
Explanation Methods Interpretability Explainability
2026 ICML Independent Researcher; Nanjing University; Nanjing University of Aeronautics an B
696 Explainable Federated Learning via Global–Local Attribution Alignment
Explanation Methods Attribution Faithfulness Interpretability Explainability
2026 ICML DEVCOM Army Research Laboratory; Virginia Polytechnic Institute and State Univer B
697 Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI. 2026 ACL University of Oregon; Adobe Research, CA, USA B
698 Explainable $ K $-means Neural Networks for Multi-view Clustering
Intrinsically Interpretable Models Explainability
2026 ICLR Fudan University; Shanghai University B
699 Explain the Synth: Interpretable Evaluation of LLM Data Synthesis. 2026 ACL Australian Regenerative Medicine Institute; Monash University; Centre de Recherc B
700 ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns
Mechanistic Interpretability
2026 ICML Amazon; ByteDance; Bytedance Seed & PSU; Fudan University; Soochow University; T B
701 Expert Heads: Robust Evidence Identification for Large Language Models
Mechanistic Interpretability Interpretability
2026 ICLR Jilin University; Renmin University of China; Soochow University B
702 Expand Neurons, Not Parameters
Mechanistic Interpretability Superposition Interpretability
2026 ICML Computer Science and Artificial Intelligence Laboratory, Electrical Engineering B
703 Exactly Computing do-Shapley Values
Attribution Methods Shapley/SHAP
2026 ICML Barcelona Supercomputing Center; Bielefeld University; Claremont McKenna College B
704 Exact Functional ANOVA Decomposition for Categorical Inputs Models
Attribution Methods Shapley/SHAP Interpretability Explainability
2026 ICML EDF & Sorbonne Université; Université Toulouse Paul Sabatier Institut de Mathéma B
705 Exact Functional ANOVA Decomposition for Categorical Inputs
Attribution Methods Shapley/SHAP Interpretability Explainability
2026 ICML EDF & Sorbonne Université; Université Toulouse Paul Sabatier Institut de Mathéma B
706 ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability. 2026 ACL Mohamed bin Zayed University of Artificial Intelligence; Tokyo Institute of Tech B
707 ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection
Explanation Methods Interpretability Explainability
2026 ICLR Alibaba Group; Tencent; University of Bath; University of Cambridge; University B
708 ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
Attribution Methods Attribution Post-hoc Explanation Interpretability Explainability
2026 ICML LMU Munich; Ludwig-Maximilians-Universität München; Saarland University B
709 Evolution of Concepts in Language Model Pre-Training
Mechanistic Interpretability Sparse Autoencoder Attribution Black-box Interpretability
2026 ICLR Fudan University; Shanghai Artificial Intelligence Laboratory B
710 Evidential Reasoning Advances Interpretable Real-World Disease Screening
Intrinsically Interpretable Models Saliency Map Post-hoc Explanation Interpretability Transparency
2026 ICML Hong Kong Polytechnic University; The Hong Kong Polytechnic University; Tsinghua B
711 Evidential Copula Concept Embedding Models
Concept Models Concept Bottleneck Concept-based Interpretability
2026 ICML Macquarie University; Shanghai University; Tongji University B
712 Evian: Towards Explainable Visual Instruction-tuning Data Auditing. 2026 ACL The Hong Kong University of Science and Technology (Guangzhou); ByteDance B
713 Evaluating and Steering Modality Preferences in Multi-modal LLMs
Probing Analysis Probing
2026 ICML Harbin Institute of Technology; Harbin Institute of Technology (Shenzhen); Harbi B
714 Evaluating SAE interpretability without generating explanations
Explanation Evaluation Sparse Autoencoder Interpretability Explainability
2026 ICLR EleutherAI B
715 Evaluating Data Influence in Meta Learning
Attribution Methods Data Attribution Attribution Influence Function
2026 ICLR KAUST; King Abdullah University of Science and Technology; MBZUAI; Shanghai Arti B
716 Escaping Low-Rank Traps: Interpretable Visual Concept Learning via Implicit Vector Quantization
Concept Models Probing Concept Bottleneck Concept-based Interpretability
2026 ICLR Fudan University; Fudan University; Shanghai AI Lab; Fudan university; Hong Kong B
717 Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
Mechanistic Interpretability Attribution Faithfulness
2026 ICLR MPI-SP; Max Planck Institute; Max Planck Institute for Security and Privacy; Seo B
718 Ensembling Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML University of Washington B
719 EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace Method
Attribution Methods Shapley/SHAP LIME Attribution Faithfulness Explainability
2026 ICLR Penn State B
720 Endogenous Resistance to Activation Steering in Language Models
Mechanistic Interpretability Sparse Autoencoder Circuit Analysis Activation Steering Transparency
2026 ICML AE Studio; Agency Enterprise Studio; Independent; Princeton University B
721 Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models
Probing Analysis Probing Interpretability
2026 ICLR Georgia Institute of Technology B
722 EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
Explanation Methods Interpretability Explainability
2026 ICLR Chinese University of Hong Kong, The Chinese University of Hong Kong; Microsoft; B
723 Emotion Concepts and their Function in a Large Language Model
steering concepts sycophancy representations reward-hacking
2026 Lab post (Anthropic) Anthropic A
724 EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness. 2026 ACL South China University of Technology B
725 Emergent Analogical Reasoning in Transformers
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML The University of Tokyo; The University of Tokyo, The University of Tokyo; Unive B
726 Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
Mechanistic Interpretability Superposition
2026 ICLR EECS, UC Berkeley; Meta AI Research; UC San Diego; University of California - Be B
727 Emergence and Localisation of Semantic Role Circuits in LLMs. 2026 ACL University of Manchester B
728 Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models. 2026 ACL Yonsei University B
729 Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
Attribution Methods Attribution Faithfulness Interpretability Explainability
2026 ICML University of Birmingham; University of Leicester B
730 Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models. 2026 ACL ETH Zurich; Mohamed bin Zayed University of Artificial Intelligence; University B
731 Efficient Hallucination Detection for LLMs Using Uncertainty-Aware Attention Heads
Mechanistic Interpretability
2026 ICML FusionBrain; FusionBrain Lab; Independent researcher; MBZUAI; Mohamed bin Zayed B
732 Efficient Estimation of Kernel Surrogate Models for Task Attribution
Attribution Methods Attribution Influence Function
2026 ICLR Northeastern University B
733 Effective Reasoning Chains Reduce Intrinsic Dimensionality
Other Explainability
2026 ICML Google; Google DeepMind; University of North Carolina at Chapel Hill; University B
734 EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI. 2026 ACL LMU Klinikum; Ludwig-Maximilians-Universität München B
735 EGG-SR: Embedding Symbolic Equivalence into Symbolic Regression via Equality Graph
Intrinsically Interpretable Models Rule/Symbolic
2026 ICLR Purdue University; University of Texas at El Paso B
736 EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions. 2026 ACL Georgia Institute of Technology; Virginia Tech B
737 ECSEL: Explainable Classification via Signomial Equation Learning
Intrinsically Interpretable Models Counterfactual Attribution Rule/Symbolic Interpretability Explainability
2026 ICML University of Amsterdam B
738 Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis
2026 ICLR Fraunhofer HHI; Oxford & Fraunhofer HHI; Technical University of Berlin; Univers B
739 Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure
Mechanistic Interpretability Probing
2026 ICML Alibaba Group; Harbin Institute of Technology; Harbin Institute of Technology (S B
740 Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers
Mechanistic Interpretability Causal Intervention Interpretability
2026 ICLR University of California, Berkeley; University of Chicago B
741 Dynamic Reflections: Probing Video Representations with Text Alignment
Probing Analysis Probing
2026 ICLR DeepMind; Google; Google DeepMind; Princeton University | Google DeepMind; Stanf B
742 Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
Mechanistic Interpretability Activation Steering
2026 ICLR East China Normal University B
743 Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML Seoul National University B
744 Domain Restriction via SAE Multi-Layer Transitions
Mechanistic Interpretability Sparse Autoencoder Black-box Interpretability
2026 ICML Technion - Israel Institute of Technology; Technion - Israel Institute of Techno B
745 Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking
Probing Analysis Circuit Analysis Probing Attribution
2026 ICML Corning Inc.; The University of Sydney; University of Central Florida; Universit B
746 Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders
Explanation Evaluation Sparse Autoencoder Interpretability
2026 ICLR Shenzhen Research Institute of Big Data; The Chinese University of Hong Kong; Un B
747 Do Transformers Grok Succinct Algorithms? Mechanistic Evidence for Counting Circuits. 2026 ACL Nanjing University B
748 Do Sparse Autoencoders Identify Reasoning Features in Language Models?
Mechanistic Interpretability Sparse Autoencoder Probing
2026 ICML UC Berkeley; University of California, Berkeley B
749 Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models. 2026 ACL Laboratory for Social and Neural Systems Research; National Center for Theoretic B
750 Do Neural Operators Forget Geometry? The Forgetting Hypothesis in Deep Operator Learning
Probing Analysis Probing
2026 ICML Tsinghua University B
751 Do Language Models Track Entities Across State Changes?
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML Boston University; Monash University; University of Vienna B
752 Do LLMs “Feel”? Emotion Circuits Discovery and Control
Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICML MBZUAI; Mohamed bin Zayed University of Artificial Intelligence; NYU Shanghai & BC
753 Do LLMs Signal When They’re Right? Evidence from Neuron Agreement
Probing Analysis
2026 ICML Fudan University; Harbin Institute of Technology; University of Hong Kong B
754 Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
Probing Analysis Probing
2026 ICML University of Southern California B
755 Do Activation Verbalization Methods Convey Privileged Information?
Explanation Evaluation Interpretability
2026 ICML Kempner Institute, Harvard University; Northeastern; Northeastern University B
756 Distributionally Robust Causal Abstractions
Mechanistic Interpretability
2026 ICML University of Warwick; University of Warwick & The Alan Turing Institute B
757 Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
Intrinsically Interpretable Models Black-box Interpretability Transparency
2026 ICML Peking University B
758 Dissecting the Safety Circuit: Neuronal Intervention for Transferable Adversarial Attacks on VLMs
Mechanistic Interpretability Circuit Analysis Probing Black-box
2026 ICML Chongqing University; Huawei Technologies Ltd.; Nanyang Technological University B
759 Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis
2026 ICML Beijing University of Posts and Telecommunications; Heidelberg University, Unive B
760 Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding
Mechanistic Interpretability Probing Faithfulness
2026 ICML University of Auckland; University of Electronic Science and Technology of China B
761 Disentangling Latent Risk Pathways via Bayesian Hypergraph Inference
Intrinsically Interpretable Models Black-box Interpretability
2026 ICML Yale University B
762 Disentangling Geometry, Performance, and Training in Language Models
Mechanistic Interpretability Interpretability
2026 ICML Carnegie Mellon University; USC; University of California, Los Angeles B
763 Disentangled Representation Learning for Parametric Partial Differential Equations
Probing Analysis Black-box Interpretability
2026 ICLR Amazon; International Business Machines; Lehigh University B
764 Discovering and Steering Interpretable Concepts in Large Generative Music Models
Mechanistic Interpretability Sparse Autoencoder Interpretability Transparency
2026 ICLR Dartmouth College; Massachusetts Institute of Technology B
765 Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models. 2026 ACL Johns Hopkins University; University of Edinburgh B
766 Discovering a Shared Logical Subspace: Steering LLM Logical Reasoning via Alignment of Natural-Language and Symbolic Views. 2026 ACL University of Illinois Urbana-Champaign B
767 Discovering Interpretable Algorithms by Decompiling Transformers to RASP
Mechanistic Interpretability Faithfulness Interpretability
2026 ICML NYU / AI2; Saarland University; University of Oxford; Universität des Saarlandes B
768 Discovering Implicit Large Language Model Alignment Objectives
Explanation Methods Interpretability Transparency
2026 ICML Stanford University; Stanford University & Apple; Stanford University // Virtue B
769 Directly Optimizing Natural Language Explanations for Behavioral Faithfulness: Simulatability and Recoverability
Explanation Methods Faithfulness Post-hoc Explanation Explainability
2026 ICML International Institute of Information Technology - Hyderabad; University of Nor B
770 Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
Mechanistic Interpretability Sparse Autoencoder
2026 ICML Fudan University; Shanghai Innovation Institute B
771 Diffusion-CAM: Faithful Visual Explanations for dMLLMs. 2026 ACL Shanghai Jiao Tong University; Sun Yat-sen University; Northwestern University; B
772 Dialectical Structured Reasoning for Explainable Multimodal Fake News Detection. 2026 ACL University of Science and Technology Beijing; National University of Singapore; B
773 Dialectic-Med: Mitigating Diagnostic Hallucinations via Counterfactual Adversarial Multi-Agent Debate. 2026 ACL Xi’an Jiaotong-Liverpool University B
774 Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
Probing Analysis Probing
2026 ICML Korea University; Seoul National University B
775 Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
Other Attribution Black-box
2026 ICML Arizona State University; Google DeepMind; Johns Hopkins University; University B
776 Diagnosing Generalization Failures from Representational Geometry Markers
Probing Analysis Mechanistic Interpretability Circuit Analysis Probing Interpretability
2026 ICLR Center of Computational Neuroscience, Flatiron Institute; Google DeepMind; Harva B
777 Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation
Attribution Methods Data Attribution Attribution
2026 ICML University of Illinois Urbana-Champaign; University of Southern California; Yale B
778 Detecting What Queries Seek: Steering LLM Safety with FFN Output Activation Monitoring. 2026 ACL Northeastern University B
779 Detecting Invariant Manifolds in ReLU-Based RNNs
Other Explainability
2026 ICLR Central Institute of Mental Health; Dept. Theoretical Neuroscience, Central Inst B
780 Density-Guided Robust Counterfactual Explanations on Tabular Data under Model Multiplicity
Explanation Methods Counterfactual Explanation Counterfactual Black-box Explainability
2026 ICML Central South University B
781 Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
Probing Analysis Probing
2026 ICML Harvard University; Yale University B
782 Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success
Other Interpretability
2026 ICML Sapienza University of Rome; University of California, San Diego B
783 Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring
Explanation Methods Faithfulness Explainability
2026 ICLR AITRICS; Korea Advanced Institute of Science & Technology; Sung Kyun Kwan Univer B
784 Deliberate Evolution: Agentic Reasoning for Sample-Efficient Symbolic Regression with LLMs
Intrinsically Interpretable Models Rule/Symbolic
2026 ICML HKBU / RIKEN; HKBU / Stanford; Hong Kong Baptist University; Lenovo Group Limite B
785 Deep neural networks divide and conquer dihedral multiplication
Mechanistic Interpretability Interpretability
2026 ICML Leiden University, Dept. of Mathematics, Leiden University; McGill University; M B
786 Deep networks learn to parse uniform-depth context-free languages from local statistics
Probing Analysis Post-hoc Explanation
2026 ICML EPFL; International Higher School for Advanced Studies Trieste B
787 Deep Single-Index Fréchet Regression
Intrinsically Interpretable Models Interpretability
2026 ICML University of California, Davis; Waymo B
788 Deconstructing Positional Information: From Attention Logits to Training Biases
Probing Analysis Probing
2026 ICLR Chinese Academy of Sciences; Institute of Information Engineering; Shanghai Jiao B
789 Deconstructing Guidance: A Semantic Hierarchy for Precise Diffusion Model Editing
Mechanistic Interpretability Interpretability
2026 ICLR Korea University B
790 Decomposition of Concept-Level Rules in Visual Scenes
Concept Models Interpretability Explainability
2026 ICLR Fudan University B
791 Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICLR Saarland University; Universität des Saarlandes B
792 Decomposing Query-Key Feature Interactions Using Contrastive Covariances
Mechanistic Interpretability Interpretability
2026 ICML Google; Harvard University; Technion, Technion B
793 Decomposing LLM Computation with Jets
Mechanistic Interpretability Interpretability
2026 ICLR AWS; Massachusetts Institute of Technology; University College London; Universit B
794 DecodeShare: Tracing the Shared Pathways of LLM Decode-Time Decisions
Mechanistic Interpretability Activation Steering
2026 ICML Department of Computer Science, Duke University; Duke University; Duke Universit B
795 Deciphering Cultural Representations in Large Language Models via Sparse Autoencoders. 2026 ACL Mohamed bin Zayed University of Artificial Intelligence; University of Toronto B
796 Debugging Concept Bottleneck Models through Removal and Retraining
Concept Models Concept Bottleneck Post-hoc Explanation Interpretability Explainability
2026 ICLR Cornell University B
797 De-Anonymization at Scale via Tournament-Style Attribution. 2026 ACL Peking University; Beihang University B
798 Data-Aware and Scalable Sensitivity Analysis for Decision Tree Ensembles
Other Interpretability
2026 ICLR IIT Bombay; Indian Institute of Technology Bombay, Mumbai, India; Indian Institu B
799 DVI-DTM: Dual-View Representation Learning for Interpretable Short Text Dynamic Topic Modeling. 2026 ACL University of Chinese Academy of Sciences; Beijing University of Posts and Telec B
800 DRIV-EX: Counterfactual Explanations for Driving LLMs. 2026 ACL Institut des langues et cultures d'Europe, Amérique, Afrique, Asie et Australie; B
801 DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole Slide Image Survival Prediction
Intrinsically Interpretable Models Interpretability Transparency
2026 ICML IIAI; Imperial colleges London; National University of Singapore; Peking Union M B
802 DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models. 2026 ACL Southeast University; Shanghai Jiao Tong University; East China Normal Universit B
803 DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Faithfulness Interpretability
2026 ICML Alibaba Group; The University of Hong Kong; University of Hong Kong; the Univers B
804 DISSOLVR: An Interpretable and Fast Framework for Aqueous and Organic Solubility Prediction
Intrinsically Interpretable Models Post-hoc Explanation Interpretability Explainability Transparency
2026 ICML IIT Delhi; Indian Institute of Technology, Delhi B
805 DEFT: Demystifying VLN Failures via a Unified Dual-View Explainability Framework for LLM-based Agents. 2026 ACL Institute of Software; Chinese Academy of Sciences B
806 DB-KSVD: Scalable Alternating Optimization for Disentangling High-Dimensional Embedding Spaces
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Interpretability
2026 ICML Google DeepMind; Stanford University; Stanford University/ ddx.inc B
807 DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
Attribution Methods Attribution Explainability
2026 ICML Jagiellonian University; Jagiellonian University in Krakow; Jagiellonian Univers B
808 Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings
Mechanistic Interpretability Sparse Autoencoder Probing Interpretability
2026 ICLR CNRS; ENS Paris Saclay; Harvard University; IRT Saint-Exupery B
809 CrisPrune: Combining Contextual Relevance and Intrinsic Saliency for Efficient Visual Token Pruning in MLLMs. 2026 ACL Tongji University; Ant Group B
810 Credal Concept Bottleneck Models for Epistemic-Aleatoric Uncertainty Decomposition. 2026 ACL Université d'Artois; AIL Research (United States) B
811 Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs. 2026 ACL University of Notre Dame B
812 Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System. 2026 ACL Bangalore , India B
813 Counterfactual Explanations on Robust Perceptual Geodesics
Explanation Methods Counterfactual Explanation Counterfactual Explainability
2026 ICLR QIMR Berghofer Medical Research Institute; Queensland University of Technology; B
814 Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs. 2026 ACL Yunnan University; National University of Singapore B
815 CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML Department of Computer Science, University College London, University of London; B
816 Controlled LLM Training on Spectral Sphere 2026 MSRA (Baining Guo et al.) C
817 Controllable and Explainable Personality Sliders for LLMs at Inference Time
Mechanistic Interpretability Probing Activation Steering Explainability
2026 ICML University College Dublin (UCD); University of Cambridge; University of Cambridg B
818 Controllable Molecule Generation via Sparse Representation Editing: An Interpretability-Driven Perspective
Mechanistic Interpretability Mechanistic Interpretability Post-hoc Explanation Interpretability
2026 ICML Department of Computing, The Hong Kong Polytechnic University; Hong Kong Polytec B
819 Controllable LLM Reasoning via Sparse Autoencoder-Based Steering. 2026 ACL University of Science and Technology of China; Alibaba Group (United States) BC
820 Contribution Weights: A Geometrical Analysis of Self-Attention Transformers
Mechanistic Interpretability Mechanistic Interpretability Faithfulness
2026 ICML Imperial College London; University College London, University of London B
821 Contrastive Symbolic Regression: Aligned Representations, Adaptive Prediction, and Diverse Ensembles
Intrinsically Interpretable Models Faithfulness Rule/Symbolic Interpretability
2026 ICML Michigan State University; Victoria University of Wellington B
822 Continuous Interpretive Steering for Scalar Diversity. 2026 ACL Sungkyunkwan University B
823 ContextCheck: Sentence-Level Faithfulness Verification with Context-Aware Disambiguation. 2026 ACL Microsoft (United States); University of Michigan; The University of Texas at Au B
824 Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding. 2026 ACL Tencent; McGill University; Wuhan University; Université de Montréal; Tsinghua U B
825 Context Attribution with Multi-Armed Bandit Optimization. 2026 ACL University of Notre Dame B
826 Constructing Interpretable Features from Compositional Neuron Groups. 2026 ACL Tel Aviv University; ESI Group (France) B
827 Concepts' Information Bottleneck Models
Concept Models Concept Bottleneck Faithfulness Interpretability
2026 ICLR University of Amsterdam; University of Hull; University of Oslo; University of t B
828 Concept-TRAK: Understanding how diffusion models learn concepts through concept attribution
Attribution Methods Attribution Influence Function Transparency
2026 ICLR Sony; Sony AI; Sony Group Corporation; University of Pennsylvania B
829 Concept Concentration for Faithful Representation Intervention
Mechanistic Interpretability Faithfulness
2026 ICML CUHK; Carnegie Mellon University & MBZUAI; HKBU / RIKEN; Johns Hopkins Universit B
830 ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution
Explanation Methods Saliency Map Concept-based Attribution Faithfulness Interpretability
2026 ICML Tel Aviv University; The Open University of Israel B
831 Compressed Sensing for Capability Localization in Large Language Models
Mechanistic Interpretability Interpretability
2026 ICML Carnegie Mellon University; Thinking Machines Lab B
832 Compositional Steering of Large Language Models with Steering Tokens. 2026 ACL University of Edinburgh; iMinds B
833 Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
Probing Analysis
2026 ICML Helmholtz AI, Technical University of Munich; KAIST & Cortiq; University of Tübi B
834 Composable Sparse Subnetworks via Maximum-Entropy Principle
Mechanistic Interpretability Interpretability
2026 ICLR Polytechnic of Turin; Sapienza University of Rome; University of Roma "La Sapien B
835 Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors. 2026 ACL Zhejiang University; Palo Alto Networks; OPPO Research Institute; Qilu Universit B
836 Comparing the learning dynamics of in-context learning and fine-tuning in language models
Probing Analysis Mechanistic Interpretability
2026 ICLR Harvard University; University College London; University College London, Univer B
837 Comparing Human and Large Language Model Interpretation of Implicit Information. 2026 ACL Politecnico di Milano B
838 Compact Example-Based Explanations for Language Models. 2026 ACL University of Vienna B
839 CombinationTS: A Modular Framework for Understanding Time-Series Forecasting Models
Probing Analysis Attribution
2026 ICML , Chinese Academy of Sciences; Beijing Normal-Hong Kong Baptist University; Comp B
840 Colorful Talks with Graphs: Human-Interpretable Graph Encodings for Large Language Models. 2026 ACL University of Illinois Chicago B
841 Cognitive models can reveal interpretable value trade-offs in language models
Other Probing Interpretability
2026 ICLR Google DeepMind; Harvard University; Johns Hopkins University; Kempner Institute B
842 Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement
Probing Analysis Interpretability
2026 ICML Artificial Intelligence Institute of South Carolina; Indian Institute of Technol B
843 CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
Mechanistic Interpretability Interpretability
2026 ICML Institute of Automation, Chinese Academy of Sciences; Institute of Information E B
844 CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs
Mechanistic Interpretability Probing
2026 ICLR Monash University; Southeast University; southeast university B
845 CoSToM: Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models. 2026 ACL National University of Singapore B
846 CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment. 2026 ACL Queen's University; Cornell University B
847 Clustered Influence Functions
Attribution Methods Influence Function
2026 ICML Eötvös Lorand University; Eötvös Loránd University B
848 CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation. 2026 ACL University of Waterloo; Hong Kong University of Science and Technology; William B
849 Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
Attribution Methods Probing Attribution
2026 ICLR Duke University; Duke University / Apple; Google; Simon Fraser University B
850 CircuitSynth: Reliable Synthetic Data Generation. 2026 ACL University of Oxford B
851 CircuitPrint: Mechanistic Circuit Fingerprints for Large Language Models
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis
2026 ICML Hunan University B
852 Circuit Insights: Towards Interpretability Beyond Activations
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Attribution Interpretability Explainability
2026 ICLR Fraunhofer HHI; Fraunhofer HHI, Fraunhofer IAIS; Fraunhofer Heinrich Hertz Insti B
853 CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization. 2026 ACL The Hong Kong University of Science and Technology (Guangzhou); Chinese Universi B
854 Characterizing, Evaluating, and Optimizing Complex Reasoning 2026 Shanghai AI Lab + SJTU + CUHK C
855 Characterizing the Expressivity of Local Attention in Transformers :arXiv 2026.05(ACL 2026 Best Paper) 2026 ETH Zürich C
856 Characterizing Pattern Matching and Its Limits on Compositional Task Structures
Other Interpretability
2026 ICLR Amazon; KAIST; KAIST AI; Korea Advanced Institute of Science & Technology; LG AI B
857 Characteristic Root Analysis and Regularization for Linear Time Series Forecasting
Other Interpretability
2026 ICLR Bosch; Bosch (China) Investment Co., Ltd.; Bosch Cooperate Research; Robert Bosc B
858 Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Explanation Evaluation Faithfulness Post-hoc Explanation
2026 ICML Google DeepMind; MATS(Neel/Nanda); ML Alignment & Theory Scholars; Poseidon Rese B
859 Chain-of-Relations: Faithful and Efficient LLM Reasoning over Knowledge Graphs via Relation-Centric Exploration. 2026 ACL Sun Yat-sen University B
860 Certified Evaluation of Model-Level Explanations for Graph Neural Networks
Explanation Evaluation Explainability
2026 ICLR Indian Statistical Institute B
861 Certified Circuits: Stability Guarantees for Mechanistic Circuits
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Black-box Interpretability Explainability
2026 ICML CISPA Helmholtz Center; CISPA Helmholtz Center for Infomation Security; MPI for B
862 CausalXRL: Explainable Reinforcement Learning through Causal Graph Reasoning
Explanation Methods Faithfulness Interpretability Explainability Transparency
2026 ICML State University of New York at Stony Brook; Stony Brook University B
863 CausalX: A Unified and Causally-Interpretable Plug-and-Play Model for Multi-modal Spatio-Temporal Forecasting
Intrinsically Interpretable Models Attribution Interpretability
2026 ICML Shandong University; Zhejiang University of Technology B
864 CausalGaze: Unveiling Hallucinations via Counterfactual Graph Intervention in Large Language Models. 2026 ACL National University of Defense Technology; Anhui Province Key Laboratory of Cybe B
865 CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
Attribution Methods Attribution Explainability
2026 ICML Google; Google DeepMind; SK hynix Research; Seoul National University B
866 Causal-Steer: Disentangled Continuous Style Control without Parallel Corpora
Mechanistic Interpretability Activation Steering
2026 ICLR Zhejiang University B
867 Causal Interpretation of Neural Network Computations with Contribution Decomposition
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Interpretability
2026 ICLR Stanford University B
868 Capturing Visual Environment Structure Correlates with Control Performance
Probing Analysis Probing
2026 ICLR Department of Computer Science, University of Illinois at Urbana-Champaign; Toyo B
869 Capacity without Access: Reinterpreting the Mid-Depth Spectral Plateau in LLMs
Probing Analysis Probing
2026 ICML Chung-Ang University; Chung-Ang university; TTA B
870 Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
Mechanistic Interpretability Sparse Autoencoder
2026 ICLR Northeastern University B
871 Can LLMs Reason Soundly in Law? Auditing Inference Patterns for Legal Judgment
Attribution Methods Faithfulness Explainability
2026 ICLR Beijing Institute for General Artificial Intelligence; Shanghai Artificial Intel B
872 Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs 2026 Yale University C
873 CSI: An Investigative Multi-Agent Framework for Explainable Short Video Fake News Detection. 2026 ACL Inner Mongolia University B
874 CRISP: Persistent Concept Unlearning via Sparse Autoencoders. 2026 ACL Technion – Israel Institute of Technology; University of Zagreb B
875 CRISP: Compressing Redundancy in Chain-of-Thought via Intrinsic Saliency Pruning. 2026 ACL College of Artificial Intelligence; Nanjing University of Aeronautics and Astron B
876 COCOGEC: Counterfactual Generation for Robust Grammatical Error Correction. 2026 ACL East China Normal University B
877 CNN Interpretability with Multivector Tucker Saliency Maps for Self-Supervised Models
Attribution Methods Saliency Map Interpretability
2026 ICLR ENS Ulm; Ecole normale supérieure B
878 CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark. 2026 ACL Saarland University; German Research Centre for Artificial Intelligence; Centre B
879 CLUE: Conflict-guided Localization for LLM Unlearning Framework
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICLR Department of Computer Science and Engineering, The Chinese University of Hong K B
880 CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
Probing Analysis Probing
2026 ICLR University of Tübingen; University of Tübingen, Max Planck Institute for Intelli B
881 CLARITree: Cholesky and Lookahead Accelerations for Regression with Interpretable Piecewise Linear Trees
Intrinsically Interpretable Models Interpretability
2026 ICML Duke; Duke University; University of British Columbia B
882 CB-SLICE: Concept-Based Interpretable Error Slice Discovery
Concept Models Concept Bottleneck Concept-based Faithfulness Interpretability Explainability
2026 ICML University of Cambridge; University of Oxford B
883 CAMP: Coherent Alignment of Multimodal Prototypes for Explainable Complementary Learning
Intrinsically Interpretable Models Interpretability Explainability
2026 ICML J.P. Morgan Chase; JPMorgan Chase; Lancaster University; University College Dubl B
884 C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICML College of Computer Science, Chongqing University; Renmin University of China; U B
885 Budget Alignment: Making Models Reason in the User's Language
Other Faithfulness Interpretability
2026 ICLR Department of Computer Science; ETH Zürich University of Gro B
886 Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
Probing Analysis Faithfulness Interpretability Explainability
2026 ICML Seoul National University B
887 Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers 2026 Microsoft Research / ETH Zürich / KRAFTON C
888 Bridging Radiology and Pathology Foundation Models via Concept-Based Multimodal Co-Adaptation
Concept Models Concept-based
2026 ICLR Hong Kong Polytechnic University; Stanford University; University of Hong Kong; B
889 Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM Generation. 2026 ACL Beijing Normal University B
890 Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
Explanation Methods Black-box Explainability
2026 ICLR Saarland University; Saarland University, Saarland University; Saarland Universi B
891 Bridging Explainability and Embeddings: BEE Aware of Spuriousness
Probing Analysis Probing Explainability Transparency
2026 ICLR Bitdefender; Bitdefender, Bucharest, Romania; Mila; University of Bucharest B
892 Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression
Intrinsically Interpretable Models Rule/Symbolic Interpretability
2026 ICML Heidelberg University; IWR Heidelberg B
893 Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability Interpretability
2026 ICML National University of Singapore; Shanghai Jiaotong University; Shenzhen Institu B
894 Block Recurrent Dynamics in Vision Transformers
Mechanistic Interpretability Mechanistic Interpretability Probing Interpretability
2026 ICLR Harvard University; Universität Osnabrück B
895 Biases in the Blind Spot: Detecting What LLMs Fail to Mention
Other Black-box
2026 ICML Independent; Poseidon Research; UCL; University College London, University of Lo B
896 Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying Prompts
Attribution Methods Attribution
2026 ICLR Ocean University of China; Xiamen University; vivo B
897 Beyond Static Personas: Situational Personality Steering for Large Language Models. 2026 ACL University of Science and Technology of China; Singapore Management University B
898 Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme Understanding. 2026 ACL National University of Defense Technology B
899 Beyond Self-Report: Bridging the Intention-Behavior Gap in Critical Thinking Assessment via Interpretable Multi-Agent System. 2026 ACL Tsinghua University B
900 Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering. 2026 ACL Xi'an Jiaotong University; National University of Singapore; Sichuan University; B
901 Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization. 2026 ACL Zhejiang Normal University; Huawei Technologies; GS1 Hong Kong; Sun Yat-sen Univ B
902 Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
Probing Analysis Probing Black-box Interpretability
2026 ICLR Queen Mary University of London; University of Oxford; University of Oxford / Ma B
903 Beyond Fixed Biases: Decoding the Role of Reasoning Uncertainty in MLLM Modality Conflicts
Probing Analysis Probing
2026 ICML KAUST; MBZUAI; Peking University; South China University of Technology; Universi B
904 Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
Other Transparency
2026 ICML MBZUAI; Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong Universi B
905 Beyond Evidence: Belief-Chain Conditioning for Persuasive Misinformation Debunking Explanation. 2026 ACL Academia Sinica; National Tsing Hua University; Pennsylvania State University B
906 Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
Other Mechanistic Interpretability
2026 ICLR Westlake University; Zhejiang University B
907 Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks. 2026 ACL Université de Reims Champagne-Ardenne; Chochoy Conseil B
908 Beyond Black-Box Interventions: Latent Probing for Faithful Retrieval-Augmented Generation. 2026 ACL Xiamen University; Alibaba Group (China); Jilin University; University of Chines B
909 Beyond Additive Decompositions: Interpretability Through Separability
Intrinsically Interpretable Models Shapley/SHAP Faithfulness Black-box Interpretability Explainability
2026 ICML University of Copenhagen B
910 Beyond Accuracy and Complexity: The Effective Information Criterion for Structurally Stable Symbolic Regression
Intrinsically Interpretable Models Rule/Symbolic Interpretability
2026 ICML Tsinghua University; Tsinghua University, Tsinghua University B
911 Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations. 2026 ACL University of Southern California; California Southern University B
912 Behavior Learning (BL)
Intrinsically Interpretable Models Interpretability Transparency
2026 ICLR Eberhard-Karls-Universität Tübingen; Xi'an Jiaotong University; Xiamen Universit B
913 Bayesian Neural Networks for Functional ANOVA Model
Intrinsically Interpretable Models Interpretability
2026 ICLR Seoul National University; University of Twente B
914 Bayesian Influence Functions for Hessian-Free Data Attribution
Attribution Methods Data Attribution Attribution Influence Function
2026 ICLR Timaeus; Timaeus Research; University of Colorado at Boulder; University of Melb B
915 Bayesian Gated Non-Negative Contrastive Learning
Intrinsically Interpretable Models Interpretability
2026 ICML MBZUAI; Mohamed bin Zayed University of Artificial Intelligence B
916 Base Models Know How to Reason, Thinking Models Learn When
Mechanistic Interpretability Sparse Autoencoder Activation Steering Interpretability
2026 ICML Google DeepMind; Oxford; Poseidon Research; University of Oxford BC
917 BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation. 2026 ACL Independent University, Bangladesh B
918 BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML Carnegie Mellon University B
919 Awakening Dormant Experts: Counterfactual Routing to Mitigate MoE Hallucinations. 2026 ACL Xi'an Jiaotong University; China Telecom; Beijing Foreign Studies University B
920 Automatically Finding Reward Model Biases
Other Interpretability
2026 ICML Google DeepMind; Massachusetts Institute of Technology; Poseidon Research B
921 Automatic Image-Level Morphological Trait Annotation for Organismal Images
Mechanistic Interpretability Sparse Autoencoder Interpretability
2026 ICLR Ohio State University; The Ohio State University; The Ohio State University, Col B
922 Automatic Construction of Clinical Scoring Systems with LLM Agents
Intrinsically Interpretable Models Interpretability
2026 ICML University of Cambridge; University of Cambridge and UCLA B
923 Automated Knowledge Component Generation and Interpretable Knowledge Tracing in Coding Problems. 2026 ACL University of Pittsburgh; University of Massachusetts Amherst B
924 Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
Explanation Evaluation Sparse Autoencoder Mechanistic Interpretability Interpretability Explainability
2026 ICLR University of Bristol B
925 AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning. 2026 ACL University of Notre Dame; Uniphore B
926 Authorship Attribution in Multilingual Machine-Generated Texts. 2026 ACL University of Calabria; Kempelen Institute of Intelligent Technologies B
927 Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions
Attribution Methods Attribution
2026 ICML Centre for Credible AI, University of Warsaw; Children Clinical Hospital, Medica B
928 AudioStealer: Extracting Audio Prompts via Shapley Value-Guided Query Search. 2026 ACL Hong Kong Polytechnic University; New York University B
929 Attribution-Guided Decoding
Attribution Methods Attribution Interpretability
2026 ICLR Fraunhofer HHI; Fraunhofer Heinrich Hertz Institut; Fraunhofer Heinrich Hertz In B
930 Attribution-Based Analysis and Optimization of Modular Agentic Workflows. 2026 ACL Shanghai Jiao Tong University B
931 Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models. 2026 ACL Center for Scalable Data Analytics and Artificial Intelligence; Technische Unive B
932 Attributing Response to Context: A Jensen–Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
Attribution Methods Mechanistic Interpretability Attribution
2026 ICLR Department of Computer Science, University College London; Nanyang Technological B
933 Attentive Multi-Layer Fusion for Vision Transformers
Probing Analysis Probing
2026 ICML Aignostics; Google DeepMind, TU Munich; TU Berlin; Technical University of Munic B
934 Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
Probing Analysis Attention Visualization
2026 ICML Wroclaw University of Science and Technology; Wroclaw University of Science and B
935 At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization
Mechanistic Interpretability Sparse Autoencoder Mechanistic Interpretability
2026 ICML INRIA; McGill University; McGill University, McGill University; Meta; Sapient In B
936 Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility. 2026 ACL University of Maryland, College Park B
937 Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
Probing Analysis Probing
2026 ICML Tiktok; University of California, San Diego; University of Illinois at Urbana-Ch B
938 Are Reasoning LLMs Robust to Interventions on their Chain-of-Thought?
Other Transparency
2026 ICLR Helmholtz AI; Technische Universität München B
939 Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering. 2026 ACL Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry B
940 Arboreal Neural Network
Intrinsically Interpretable Models Faithfulness Interpretability Transparency
2026 ICML Beijing University of Technology; Du Xiaoman Technology(BeiJing); DuXiaoman Tech B
941 Analysing the Safety Pitfalls of Steering Vectors. 2026 ACL Technical University of Munich B
942 An Odd Estimator for Shapley Values
Attribution Methods Shapley/SHAP Attribution Feature Importance
2026 ICML Bielefeld University; Claremont McKenna College; UC Berkeley; UC Berkeley, Snork B
943 An Information-Theoretic Parameter-Free Bayesian Framework for Probing Labeled Dependency Trees from Attention Score
Probing Analysis Probing Explainability Transparency
2026 ICLR Beijing University of Post and Telecommunication; Beijing University of Posts an B
944 All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection. 2026 ACL University of Manchester; Stevens Institute of Technology; Columbia University; B
945 All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Superposition Faithfulness Explainability
2026 ICML Northwestern University; Peking University; Rutgers; Rutgers University; TU Darm B
946 Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation 2026 Stanford University C
947 Aligning What LLMs Do and Say: Towards Self-Consistent Explanations. 2026 ACL Technion – Israel Institute of Technology; University of Edinburgh B
948 Aligned, Orthogonal or In-conflict: When Can We Safely Optimize Chain-of-Thought?
chain-of-thought monitoring
2026 Lab post (Google DeepMind) Google DeepMind A
949 Algorithmic Recourse of In-Context Learning for Tabular Data
Explanation Methods Post-hoc Explanation Black-box
2026 ICML HKUST(GZ); KAUST; King Abdullah University of Science and Technology (KAUST); MB B
950 Aitchison Embeddings for Learning Compositional Graph Representations
Intrinsically Interpretable Models Probing Post-hoc Explanation Interpretability Explainability
2026 ICML Natera; University of Peloponnese; Yale University; École Normale Supérieure (EN B
951 Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations. 2026 ACL Ludwig-Maximilians-Universität München; UCLouvain; University of Vienna; Georget B
952 Aggregate Models, Not Explanations: Improving Feature Importance Estimation
Attribution Methods Feature Importance Explainability
2026 ICML Inria; Roche Pharma Research & Early Development (pRED); Institut de Mathémathiq B
953 AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction
Other Black-box Interpretability
2026 ICML Fudan University; Shanghai Jiao Tong University; Shanghai Jiaotong University; T B
954 AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
Attribution Methods Counterfactual Attribution
2026 ICLR Beijing University of Post and Telecommunications; Guangdong OPPO Mobile Telecom B
955 Adversarial Vulnerability from Interference Between Features in Superposition
Mechanistic Interpretability Superposition Explainability
2026 ICML Bogazici University; Imperial College London B
956 Addressing divergent representations from causal interventions on neural networks
Mechanistic Interpretability Mechanistic Interpretability Counterfactual Faithfulness Interpretability Explainability
2026 ICLR Stanford University B
957 AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations. 2026 ACL Zhejiang University; University of Illinois Chicago; Chinese University of Hong B
958 Adaptive Node Feature Selection for Graph Neural Networks
Attribution Methods Feature Importance
2026 ICML Rice University B
959 Adaptive Concept Discovery for Interpretable Few-Shot Text Classification
Concept Models Concept Bottleneck Interpretability
2026 ICLR HKUST; Hong Kong University of Science and Technology B
960 Adalina: Adaptive Linear Approximation for the Shapley Value and Beyond
Attribution Methods Shapley/SHAP Attribution
2026 ICML National University of Singapore; University of Waterloo B
961 ActivationReasoning: Logical Reasoning in Latent Activation Spaces
Mechanistic Interpretability Sparse Autoencoder Interpretability Transparency
2026 ICLR Adobe, hessian.AI; CS Department, TU Darmstadt; German Research Center for AI; M B
962 Activation Steering with a Feedback Controller
Mechanistic Interpretability Activation Steering Interpretability
2026 ICLR Hanoi University of Science and Technology; National University of Singapore B
963 Activation Steering for Chain-of-Thought Compression. 2026 ACL University of Southern California; California Southern University B
964 Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
Mechanistic Interpretability Black-box
2026 ICML Anthropic; ENS Paris-Saclay / EPFL; EPFL - EPF Lausanne; Independent; MIT Tegmar BC
965 Activation Decomposition and Steering for LLM Backdoor Remediation. 2026 ACL Macquarie University B
966 Accumulating Context Changes the Beliefs of Language Models 2026 Carnegie Mellon University & Princeton University & Standfor C
967 AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
Mechanistic Interpretability Sparse Autoencoder Probing Interpretability
2026 ICLR Ohio State University, Columbus B
968 AURA: Visually Interpretable Affective Understanding via Robust Archetypes
Intrinsically Interpretable Models Interpretability Transparency
2026 ICML Queen Mary University of London; Queen Mary University of London & Xi'an Jiaoton B
969 ATEX-CF: Attack-Informed Counterfactual Explanations for Graph Neural Networks
Explanation Methods Counterfactual Explanation Counterfactual Faithfulness Explainability
2026 ICLR AI Institute - University of Central Florida; Aalborg University; Bowling Green B
970 ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
Mechanistic Interpretability Mechanistic Interpretability Circuit Analysis Interpretability
2026 ICLR Korea University B
971 APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation. 2026 ACL Tianjin University; Zhejiang University; North University of China; Institute of B
972 AIR: Post-training Data Selection for Reasoning via Attention Head Influence
Attribution Methods Mechanistic Interpretability
2026 ICML Alibaba Group; Beihang University; Beihang University; Beijing University of Ae B
973 AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives. 2026 ACL Arizona State University; Clemson University; Washington University in St. Louis B
974 ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
Mechanistic Interpretability Mechanistic Interpretability Attribution Interpretability
2026 ICLR Beijing University of Aeronautics and Astronautics; HKUST (GZ); The Hong Kong Un B
975 A Syntactic and Semantic Probe into Language Evolution based on Large Language Models. 2026 ACL Dalian University of Technology B
976 A Probabilistic Hard Concept Bottleneck for Steerable Generative Models
Concept Models Concept Bottleneck Interpretability
2026 ICLR Saarland University, Saarland University; Saarland University, Universität des S B
977 A Positive Case for Faithfulness: Explanations Help Predict Model Behavior
Explanation Evaluation Counterfactual Faithfulness Explainability
2026 ICML Arcadia Impact; Google DeepMind; UC Berkeley; University of Oxford B
978 A Narrowing Geometry in Contaminated Reasoning
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML Institute of automation, Chinese academy of science; Institute of automation, Ch B
979 A Mechanistic Perspective and Difficulty Metric for Unlearning. 2026 ACL University of Massachusetts Lowell B
980 A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies
Probing Analysis Mechanistic Interpretability
2026 ICML Amazon; The University of Texas at Austin; University of Texas at Austin; Univer B
981 A Lightweight Explainable Guardrail for Prompt Safety. 2026 ACL University of Arizona B
982 A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
Other Probing
2026 ICLR Dartmouth College; Massachusetts Institute of Technology; University of Californ B
983 A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insights
Mechanistic Interpretability Mechanistic Interpretability
2026 ICML International Higher School for Advanced Studies Trieste; SISSA (IT); SISSA (Int B
984 A Factorized Low-Rank RNN Framework for Uncovering Independent Neural Latent Dynamics and Connectivity
Intrinsically Interpretable Models Interpretability
2026 ICML Emory University; Georgia Institute of Technology; Georgia Institute of Technolo B
985 A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
Mechanistic Interpretability Mechanistic Interpretability Faithfulness Interpretability
2026 ICML Fudan University; Xi'an Jiaotong University B
986 A Counterfactual Explanation Framework for Retrieval Models. 2026 ACL University of California San Diego; University of Liverpool B
987 A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
Probing Analysis Attribution
2026 ICLR Microsoft Research; The University of Tokyo B
988 A Capacity-Based Rationale for Multi-Head Attention
Mechanistic Interpretability Superposition
2026 ICML Computer Science and Artificial Intelligence Laboratory, Electrical Engineering B
989 A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
Probing Analysis Probing Interpretability
2026 ICML CapitalOne; Fraunhofer HHI; Indiana University at Bloomington; Northeastern Univ B
990 A 'Diff' Tool for AI: Finding Behavioral Differences in New Models 2026 Lab post (Anthropic) Anthropic A
991 $\texttt{ShaplEIG}$: Bayesian Experimental Design for Shapley Value Estimation
Attribution Methods Shapley/SHAP Attribution Feature Importance Interpretability
2026 ICML Bielefeld University; LMU; LMU Munich; LMU Munich, MCML; Lamarr Institute, TU Do B
992 Can Training Only a Single Transformer Block Match or Even Surpass Full-Parameter RL? 2025 C
993 [Survey] Interpreting Language Models Through Concept Descriptions: A Survey 2025 C
994 [MASK]ED - Language Modeling for Explainable Classification and Disentangling of Socially Unacceptable Discourse. 2025 EMNLP CY Cergy Paris Université B
995 Zero-shot protein stability prediction by inverse folding models: a free energy interpretation. 2025 NeurIPS Technical University of Denmark; University of B
996 Zero-Shot Natural Language Explanations. 2025 ICLR Vrije Universiteit Brussel, Brussels, Belgium B
997 XDAC: XAI-Driven Detection and Attribution of LLM-Generated News Comments in Korean. 2025 ACL National Security Research Institute; Korea Advanced Institute of Science and Te B
998 XAIguiFormer: explainable artificial intelligence guided transformer for brain disorder identification. 2025 ICLR B
999 X2-DFD: A framework for explainable and extendable Deepfake Detection. 2025 NeurIPS School of Data Science, The Chinese University of Hong Kong, Shenzhen, Guangdong B
1000 X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning. 2025 EMNLP Rochester Institute of Technology; DEVCOM Army Research Laboratory; United State B
1001 Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA. 2025 ACL University of California, Santa Cruz; Carnegie Mellon University B
1002 Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers. 2025 ICLR FZI Research Center for Information Technology; Karlsruhe Institute of Technolog B
1003 Who's the Author? How Explanations Impact User Reliance in AI-Assisted Authorship Attribution. 2025 EMNLP University of Maryland B
1004 Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. 2025 ICML Pennsylvania State University; Duke University; University of Washington; Nanyan B
1005 Where Did That Come From? Sentence-Level Error-Tolerant Attribution. 2025 EMNLP Mila - Quebec Artificial Intelligence Institute; McGill University; Bar-Ilan Uni B
1006 When Models Manipulate Manifolds: The Geometry of a Counting Task
circuits representations
2025 Lab post (Anthropic) Anthropic A
1007 When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations. 2025 ACL Columbia University; Nanyang Technological University; Meta; Rutgers, The State B
1008 What’s the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns 2025 Center for Information and Language Processing, LMU Munich C
1009 What should a neuron aim for? Designing local objective functions based on information theory. 2025 ICLR Faculty of Physics, Institute for the Dynamics of Complex Systems, University of B
1010 What makes an Ensemble (Un) Interpretable? 2025 ICML Hebrew University of Jerusalem B
1011 What Makes a Reward Model a Good Teacher? An Optimization Perspective 2025 Princeton University C
1012 What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models. 2025 ICML Dana Foundation; Harvard University B
1013 Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model 2025 Shanghai Jiao Tong University C
1014 Wasserstein Distances, Neuronal Entanglement, and Sparsity. 2025 ICLR Institute of Science and Technology Austria; Neural Magic; Red Hat B
1015 Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations. 2025 ICLR Microsoft Research B
1016 WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification. 2025 EMNLP Monash University B
1017 WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data. 2025 ACL National University of Singapore; Institute for Infocomm Research; AI Singapore; B
1018 Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting. 2025 NeurIPS Toyota Technological Institute at Chicago; University of California, Berkeley B
1019 Verbosity-Aware Rationale Reduction: Sentence-Level Rationale Reduction for Efficient and Effective Reasoning. 2025 ACL University of Illinois Urbana-Champaign; Kyungpook National University B
1020 Variational Counterfactual Intervention Planning to Achieve Target Outcomes. 2025 ICML University of Science and Technology of China; Nanyang Technological University B
1021 Validating Mechanistic Interpretations: An Axiomatic Approach. 2025 ICML University of Wisconsin–Madison; Colorado State University; Scale AI; Carnegie M B
1022 VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models. 2025 NeurIPS Nanchang University B
1023 VISA: Retrieval Augmented Generation with Visual Source Attribution. 2025 ACL University of Waterloo; CSIRO B
1024 VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware Tree. 2025 NeurIPS School of Software, Xi’an Jiaotong University; College of Computer Science and T B
1025 V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models. 2025 EMNLP Tongji University; University of Wisconsin–Madison B
1026 Using Shapley interactions to understand how models use structure. 2025 ACL Harvard University Press B
1027 Unveiling the Magic of Code Reasoning through Hypothesis Decomposition and Amendment 2025 USTC (ICLR 2025) C
1028 Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiency. 2025 EMNLP Tianjin University; GMT Technology (Shenzhen) Co., Ltd. (China) B
1029 Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders. 2025 ACL University of Science and Technology of China; Sichuan University B
1030 Unstructured Evidence Attribution for Long Context Query Focused Summarization. 2025 EMNLP University of Michigan B
1031 Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection. 2025 ICML Jinan University; University of Macau; Nanyang Technological University B
1032 Unlearning-based Neural Interpretations. 2025 ICLR Massachusetts Institute of Technology; University of Oxford; Pioneer Centre for B
1033 Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment. 2025 ICML York University; State Research Center of Virology and Biotechnology VECTOR; Har B
1034 Understanding and Mitigating Hallucination in Large Vision-Language Models via Modular Attribution and Intervention. 2025 ICLR B
1035 Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs. 2025 EMNLP Beijing Institute for General Artificial Intelligence; Wuhan University BC
1036 Understanding and Improving Adversarial Robustness of Neural Probabilistic Circuits. 2025 NeurIPS University of Illinois Urbana-Champaign B
1037 Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron. 2025 ICLR Mila - Quebec AI Institute, Montreal, QC, Canada; University of Montreal, QC, Ca B
1038 Understanding Refusal in Language Models with Sparse Autoencoders. 2025 EMNLP Nanyang Technological University; Singapore University of Technology and Design; B
1039 Understanding Neural Networks Through Sparse Circuits
mechanistic circuits neurons concepts interpretability
2025 Lab post (OpenAI) OpenAI A
1040 Understanding How Value Neurons Shape the Generation of Specified Values in LLMs. 2025 EMNLP Provable Responsible AI and Data Analytics (PRADA) Lab; King Abdullah University B
1041 Uncovering Gaps in How Humans and LLMs Interpret Subjective Language. 2025 ICLR UC Berkeley B
1042 UTILITY: Utilizing Explainable Reinforcement Learning to Improve Reinforcement Learning. 2025 ICLR B
1043 UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective 2025 HKU Team C
1044 Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training. 2025 NeurIPS Zhejiang University; Shandong University B
1045 Turning Logic Against Itself: Probing Model Defenses Through Contrastive Questions. 2025 EMNLP Ubiquitous Knowledge Processing Lab (UKP Lab); hessian.AI; Technische Universitä B
1046 Tree-of-Quote Prompting Improves Factuality and Attribution in Multi-Hop and Medical Reasoning. 2025 EMNLP University of Oxford; Saarland University; All India Institute of Medical Scienc B
1047 Treble Counterfactual VLMs: A Causal Approach to Hallucination. 2025 EMNLP University of Southern California; National University of Singapore; University B
1048 Transformer-Based Spatial-Temporal Counterfactual Outcomes Estimation. 2025 ICML National University of Defense Technology; PLA Academy of Military Science B
1049 Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders. 2025 NeurIPS Tohoku University; RIKEN BC
1050 Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language Models. 2025 EMNLP State Key Lab of AI Safety, Institute of Computing Technology, CAS; Chinese Acad B
1051 Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy. 2025 EMNLP T-Tech; HSE University B
1052 Tracing the Thoughts of a Large Language Model
circuits
2025 Lab post (Anthropic) Anthropic AC
1053 Tracing the Roots: Leveraging Temporal Dynamics in Diffusion Trajectories for Origin Attribution. 2025 NeurIPS Imperial College London; Apple B
1054 Towards the Pedagogical Steering of Large Language Models for Tutoring: A Case Study with Modeling Productive Failure. 2025 ACL ETH Zurich; École Polytechnique; ETH AI Center; Association for Computational Li B
1055 Towards counterfactual fairness through auxiliary variables. 2025 ICLR University of Maryland, College Park, ^2Clemson University B
1056 Towards an Explainable Comparison and Alignment of Feature Embeddings. 2025 ICML Chinese University of Hong Kong B
1057 Towards a Mechanistic Explanation of Diffusion Model Generalization. 2025 ICML University of British Columbia; Alberta Machine Intelligence Institute B
1058 Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures. 2025 ICLR OpenMOSS Team, School of Computer Science, Fudan University B
1059 Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization. 2025 ICLR B
1060 Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis. 2025 ICML University of Hong Kong; Chinese University of Hong Kong, Shenzhen B
1061 Towards Synergistic Path-based Explanations for Knowledge Graph Completion: Exploration and Evaluation. 2025 ICLR Massachusetts Institute of Technology, Cambridge, MA, USA B
1062 Towards Robustness and Explainability of Automatic Algorithm Selection. 2025 ICML Hong Kong Polytechnic University; Chongqing University B
1063 Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration. 2025 ICML Shanghai University; Tencent Youtu Lab B
1064 Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control. 2025 ICLR Independent Researchers; Google DeepMind B
1065 Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few. 2025 NeurIPS School of Artificial Intelligence, Beijing University of Posts and Telecommunica BC
1066 Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders. 2025 NeurIPS University of Wisconsin–Madison; Queen Mary University of London; University of B
1067 Towards Global-level Mechanistic Interpretability: A Perspective of Modular Circuits of Large Language Models. 2025 ICML University of Virginia; Florida State University B
1068 Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models. 2025 EMNLP Nanyang Technological University; Institute of High Performance Computing; Agenc B
1069 Towards Explaining the Power of Constant-depth Graph Neural Networks for Structured Linear Programming. 2025 ICLR B
1070 Towards Explainable Temporal Reasoning in Large Language Models: A Structure-Aware Generative Framework. 2025 ACL Wuhan University; The Hong Kong University of Science and Technology (Guangzhou) B
1071 Towards Explainable Hate Speech Detection. 2025 ACL Universitätsmedizin Greifswald; Universität Greifswald B
1072 Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning. 2025 ICML Nanyang Technological University B
1073 Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales. 2025 EMNLP Harbin Institute of Technology; Pengcheng Laboratory, Shenzhen, China; Shaoguan B
1074 Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness. 2025 ACL Chinese Academy of Sciences B
1075 Towards Automated Knowledge Integration From Human-Interpretable Representations. 2025 ICLR University of Cambridge B
1076 Towards Attributions of Input Variables in a Coalition. 2025 ICML Sun Yat-sen University B
1077 Towards Achieving Concept Completeness for Textual Concept Bottleneck Models. 2025 EMNLP Sorbonne Université; Ekimetrics B
1078 Toward Real-world Text Image Forgery Localization: Structured and Interpretable Data Synthesis. 2025 NeurIPS Sun Yat-sen University; Beihang University; Guangzhou University; Peng Cheng Lab B
1079 Toward Inclusive Language Models: Sparsity-Driven Calibration for Systematic and Interpretable Mitigation of Social Biases in LLMs. 2025 EMNLP George Mason University B
1080 Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language Models. 2025 EMNLP University of Illinois Urbana-Champaign, IL, USA B
1081 TopInG: Topologically Interpretable Graph Learning via Persistent Rationale Filtration. 2025 ICML Rutgers, The State University of New Jersey B
1082 TokenShapley: Token Level Context Attribution with Shapley Value. 2025 ACL TikTok; Princeton University B
1083 To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts. 2025 ICLR Duke University B
1084 To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models. 2025 ICML Work done during an . Morgan AI Research; J.P. Morgan B
1085 To See a World in a Spark of Neuron: Disentangling Multi-Task Interference for Training-Free Model Merging. 2025 EMNLP Xiamen University Malaysia B
1086 TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research. 2025 EMNLP Martian; Apart Research; NVIDIA; Thoughtworks; Hubei Polytechnic University B
1087 TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop. 2025 NeurIPS School of Computing, University of Connecticut; System Security Department, NEC B
1088 ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations. 2025 EMNLP The University of Sydney B
1089 Thought Anchors: Which LLM Reasoning Steps Matter? 2025 DeepSeek C
1090 ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models. 2025 EMNLP Universidad Católica Santo Domingo B
1091 Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation. 2025 ICLR Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, B
1092 TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding. 2025 ACL University of Waterloo; State Research Center of Virology and Biotechnology VECT B
1093 The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It. 2025 EMNLP University of Trento; EU Business School, Munich; Munich Center for Machine Lear B
1094 The Two Paradigms of LLM Detection: Authorship Attribution vs Authorship Verification. 2025 ACL Leipzig University B
1095 The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs. 2025 EMNLP Japan Advanced Institute of Science and Technology; RIKEN B
1096 The Superposition of Diffusion Models Using the Itô Density Estimator. 2025 ICLR University of Toronto; Vector Institute; University of Oxford; Mila - Quebec AI B
1097 The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas. 2025 EMNLP Chinese Academy of Sciences; University of Chinese Academy of Sciences; Zhonggua B
1098 The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability? 2025 NeurIPS ETH Zurich; École Polytechnique Fédérale de Lausanne B
1099 The Logical Implication Steering Method for Conditional Interventions on Transformer Generation. 2025 ICML Salesforce B
1100 The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination 2025 UIUC / Columbia University / Northwestern University / Stanford University C
1101 The Knowledge Microscope: Features as Better Analytical Lenses than Neurons. 2025 ACL Institute for Complex Systems; Chinese Academy of Sciences; University of Chines B
1102 The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information Steering. 2025 ICML Rutgers, The State University of New Jersey; Stanford University; Google DeepMin B
1103 The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness. 2025 NeurIPS Microsoft; ELLIS Institute Tübingen; MPI for Intelligent Systems; Tübingen AI Ce B
1104 The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization. 2025 NeurIPS Michigan State University B
1105 The Emergence of Abstract Thought in Large Language Models Beyond Any Language 2025 National University of Singapore / Peking University C
1106 The Coverage Principle: How Pre-Training Enables Post-Training 2025 Microsoft Research/Princeton C
1107 The Computational Complexity of Circuit Discovery for Inner Interpretability. 2025 ICLR University of Bristol; Goethe University Frankfurt; Memorial University of Newfo B
1108 The Anatomy of Evidence: An Investigation Into Explainable ICD Coding. 2025 ACL Fraunhofer Institute for Applied and Integrated Security; Lamarr Institute for M B
1109 Test-Time Steering for Lossless Text Compression via Weighted Product of Experts. 2025 EMNLP University of British Columbia; State Research Center of Virology and Biotechnol B
1110 Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models. 2025 NeurIPS Rutgers University B
1111 Temporal Misalignment in ANN-SNN Conversion and its Mitigation via Probabilistic Spiking Neurons. 2025 ICML Department of ML, MBZUAI, Abu Dhabi, UAE; City University of Macau; Jilin Univer B
1112 Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations. 2025 NeurIPS α X-LANCE Lab, Department of Computer Science and Engineering; MoE Key Lab of Ar BC
1113 Taming Hyperparameter Sensitivity in Data Attribution: Practical Selection Without Costly Retraining. 2025 NeurIPS University of Michigan Ann Arbor; University of Illinois Urbana-Champaign B
1114 Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers. 2025 EMNLP National Institute of Informatics; The University of Tokyo; National Taiwan Univ B
1115 TS-LIF: A Temporal Segment Spiking Neuron Network for Time Series Forecasting. 2025 ICLR Nanyang Technological University; University of Chinese Academy of Sciences; Ten B
1116 TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection. 2025 EMNLP National University of Singapore B
1117 TIMING: Temporality-Aware Integrated Gradients for Time Series Explanation. 2025 ICML AITRICS; Korea Advanced Institute of Science and Technology B
1118 TAO: Using test-time compute to train efficient LLMs without labeled data 2025 databricks C
1119 SynC-LLM: Generation of Large-Scale Synthetic Circuit Code with Hierarchical Language Models. 2025 EMNLP Hong Kong University of Science and Technology B
1120 Supervised and Unsupervised Probing of Shortcut Learning: Case Study on the Emergence and Evolution of Syntactic Heuristics in BERT. 2025 ACL KU Leuven B
1121 Superposition Yields Robust Neural Scaling. 2025 NeurIPS Massachusetts Institute of Technology B
1122 SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines 2025 ByteDance C
1123 Steering off Course: Reliability Challenges in Steering Language Models. 2025 ACL The Ohio State University; University of Washington; Fastino AI; Allen Institute B
1124 Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models. 2025 ACL Georgia Institute of Technology; Apple (United States) B
1125 Steering Protein Language Models. 2025 ICML Tencent AI Lab B
1126 Steering Protein Family Design through Profile Bayesian Flow. 2025 ICLR Institute of AI Industry Research (AIR), Tsinghua University; School of Pharmace B
1127 Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction. 2025 ICLR Université de Montréal, ^ 2 Mila, ^ 3 Dreamfold, ^ 4 Duke University, ^ 5 McGill B
1128 Steering Large Language Models between Code Execution and Textual Reasoning. 2025 ICLR Microsoft B
1129 Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect. 2025 EMNLP University of Mannheim; Ho Technical University B
1130 Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation. 2025 EMNLP Southeast University; Key Laboratory of New Generation Artificial Intelligence T B
1131 Steering LLM Reasoning Through Bias-Only Adaptation. 2025 EMNLP T-Tech; Central University BC
1132 Steering Information Utility in Key-Value Memory for Language Model Post-Training. 2025 NeurIPS Rice University BC
1133 Steering Generative Models with Experimental Data for Protein Fitness Optimization. 2025 NeurIPS California Institute of Technology; Mathematical Sciences; Microsoft Corporation B
1134 SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models. 2025 EMNLP Virginia Tech B
1135 Start Smart: Leveraging Gradients For Enhancing Mask-based XAI Methods. 2025 ICLR B
1136 Spurious Forgetting in Continual Learning of Language Models 2025 South China University of Technology C
1137 Splitting & Integrating: Out-of-Distribution Detection via Adversarial Gradient Attribution. 2025 ICML Suzhou University of Technology; University of Malaya; University of Technology B
1138 SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking. 2025 ICLR School of Artificial Intelligence, University of Chinese Academy of Sciences; In B
1139 Speech Act Patterns for Improving Generalizability of Explainable Politeness Detection Models. 2025 ACL Bentley University B
1140 SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder 2025 USTC + Huawei C
1141 SparseMVC: Probing Cross-view Sparsity Variations for Multi-view Clustering. 2025 NeurIPS China University of Geosciences; The Hong Kong University of Science and Technol B
1142 Sparse autoencoders reveal selective remapping of visual concepts during adaptation. 2025 ICLR Institute of Computational Biology, Computational Health Center, Helmholtz Munic B
1143 Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs. 2025 EMNLP Brown University; Drexel University B
1144 Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. 2025 ICLR Northeastern University BC
1145 Sparse Autoencoders, Again? 2025 ICML Fudan University; Amazon Web Services B
1146 Sparse Autoencoders for Hypothesis Generation. 2025 ICML Cornell University B
1147 Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models. 2025 ICLR Institute for Human Centered Design; Max Planck Institute for Biological Cyberne B
1148 Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models. 2025 NeurIPS Technical University of Munich; Munich Center for Machine Learning; Munich Data B
1149 Sparse Autoencoders Do Not Find Canonical Units of Analysis. 2025 ICLR Durham University; Decode Research; Apollo Research B
1150 Sparse Autoencoder Features for Classifications and Transferability. 2025 EMNLP Harvard University; Mass General Brigham; Boston Children's Hospital; Johns Hopk B
1151 Sound Logical Explanations for Mean Aggregation Graph Neural Networks. 2025 NeurIPS University of Oxford B
1152 Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment. 2025 EMNLP Indian Institute of Technology Kharagpur; Eindhoven University of Technology B
1153 Solving Satisfiability Modulo Counting Exactly with Probabilistic Circuits. 2025 ICML Purdue University B
1154 Small Changes, Big Impact: How Manipulating a Few Neurons Can Drastically Alter LLM Aggression. 2025 ACL Konkuk University; Electronics and Telecommunications Research Institute B
1155 Since Faithfulness Fails: The Performance Limits of Neural Causal Discovery. 2025 ICML University of Warsaw; Poznan University of Technology, Poznan, Poland; Polish Ac B
1156 Should I Trust You? Detecting Deception in Negotiations using Counterfactual RL. 2025 ACL University of Maryland, College Park; Northwestern University; The University of B
1157 Shedding Light on Time Series Classification using Interpretability Gated Networks. 2025 ICLR B
1158 Shapley-Guided Utility Learning for Effective Graph Inference Data Valuation. 2025 ICLR Rensselaer Polytechnic Institute, Troy, NY, United States B
1159 Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents. 2025 NeurIPS Antai College of Economics and Management, Shanghai Jiao Tong University, Shangh B
1160 ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference 2025 ByteDance Seed C
1161 Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers. 2025 ACL ENS Paris-Saclay; Université Paris-Saclay; Northeastern University; Princeton Un B
1162 Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors. 2025 ICLR School of Informatics, University of Edinburgh; Huawei Technologies Co., Ltd; Sc B
1163 SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models. 2025 ICML Massachusetts Institute of Technology; FAIR B
1164 Self-Supervised Discovery of Neural Circuits in Spatially Patterned Neural Responses with Graph Neural Networks. 2025 NeurIPS Department of Artificial Intelligence; Hanyang University B
1165 Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models. 2025 ACL Chinese Academy of Sciences B
1166 Self-Critique and Refinement for Faithful Natural Language Explanations. 2025 EMNLP University of Copenhagen B
1167 Scaling up Test-Time Compute with Latent Reasoning 2025 ELLIS Institute Tübingen C
1168 Scaling and evaluating sparse autoencoders. 2025 ICLR OpenAI B
1169 Scaling Sparse Feature Circuits For Studying In-Context Learning. 2025 ICML ETH Zurich; Georgia Institute of Technology B
1170 Scaling Probabilistic Circuits via Monarch Matrices. 2025 ICML University of California, Los Angeles B
1171 Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow Matching. 2025 NeurIPS ♣ The Leiden Institute of Advanced Computer Science (LIACS), Leiden University; B
1172 Scalable Mechanistic Neural Networks. 2025 ICLR Institute of Science and Technology Austria B
1173 Salvage: Shapley-distribution Approximation Learning Via Attribution Guided Exploration for Explainable Image Classification. 2025 ICLR B
1174 Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning. 2025 EMNLP Hanyang University B
1175 SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning. 2025 EMNLP Shanghai Jiao Tong University; Alibaba Group (China) B
1176 SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior. 2025 ICML Allen Institute for Artificial Intelligence B
1177 Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety. 2025 ACL Carnegie Mellon University B
1178 SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations. 2025 ICLR University of Chicago, ^ 2 Virtue AI, ^ 3 University of Illinois, Urbana-Champai B
1179 SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals. 2025 EMNLP University of Illinois Urbana-Champaign B
1180 STSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex. 2025 NeurIPS Neuroscience Interdepartmental Program, Stanford University, Stanford, CA; Depar B
1181 STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment. 2025 EMNLP Nanyang Technological University; Harbin Institute of Technology B
1182 SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning. 2025 NeurIPS Zhejiang University B
1183 SPEX: Scaling Feature Interaction Explanations for LLMs. 2025 ICML University of California, Berkeley B
1184 SHARP: Steering Hallucination in LVLMs via Representation Engineering. 2025 EMNLP State Key Laboratory of Pattern Recognition; State Key Laboratory of Multimodal B
1185 SHAP zero Explains Biological Sequence Models with Near-zero Marginal Cost for Future Queries. 2025 NeurIPS Georgia Institute of Technology B
1186 SEK: Self-Explained Keywords Empower Large Language Models for Code Generation. 2025 ACL Zhejiang University B
1187 SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling. 2025 EMNLP EPFL B
1188 SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs. 2025 NeurIPS University of Electronic Science and Technology of China (UESTC); Shenzhen Insti B
1189 SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation 2025 Southeast University C
1190 SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text Generation. 2025 ICLR Sorbonne Université, CNRS, ISIR, F-75005 Paris, France; Miles Team, LAMSADE, Uni B
1191 SCDTour: Embedding Axis Ordering and Merging for Interpretable Semantic Change Detection. 2025 EMNLP University of Liverpool B
1192 SCBench: A KV Cache-Centric Analysis of Long-Context Methods 2025 Microsoft C
1193 SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders. 2025 ICML Warsaw University of Technology; IDEAS Research Institute B
1194 SAKE: Steering Activations for Knowledge Editing. 2025 ACL ÅF (Switzerland); AXA; Sorbonne Université; Polish Academy of Sciences B
1195 SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs. 2025 EMNLP Texas A&M University; University of Milano-Bicocca; Hamad bin Khalifa University B
1196 SAEs Are Good for Steering - If You Select the Right Features. 2025 EMNLP Technion – Israel Institute of Technology; Boston University B
1197 SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders. 2025 NeurIPS Peking University BC
1198 SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. 2025 ICML Decode Research; University College London; MATS Research; Anthropic (United Sta BC
1199 SAE-V: Interpreting Multimodal Models for Enhanced Alignment. 2025 ICML Peking University B
1200 SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models. 2025 EMNLP New Jersey Institute of Technology; Rutgers, The State University of New Jersey; B
1201 Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset. 2025 ACL University of Cambridge; SB Intuitions; Tohoku University; RIKEN; NetMind.AI; Un B
1202 Route Sparse Autoencoder to Interpret Large Language Models. 2025 EMNLP University of Science and Technology of China; Douyin Co., Ltd B
1203 Robustly identifying concepts introduced during chat fine-tuning using crosscoders 2025 EPFL、ETHZ、Ecole Normale Supérieure Paris-Saclay、Université P C
1204 Rhetorical Device-Aware Sarcasm Detection with Counterfactual Data Augmentation. 2025 ACL Tongji University; Ministry of Education, Shanghai 201804, China; DC Arts & Huma B
1205 Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability. 2025 NeurIPS Tel-Aviv University B
1206 Revisiting LLM Value Probing Strategies: Are They Robust and Expressive? 2025 EMNLP University of Michigan; LG AI Research B
1207 Revisiting In-context Learning Inference Circuit in Large Language Models. 2025 ICLR Japan Advanced Institute of Science and Technology; RIKEN B
1208 Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing. 2025 ACL University of Chinese Academy of Sciences; Institute for Complex Systems; Chines B
1209 Retrieve to Explain: Evidence-driven Predictions for Explainable Drug Target Identification. 2025 ACL BenevolentAI (United Kingdom) B
1210 RetrievalAttention: ACCELERATING LONG-CONTEXT LLM INFERENCE VIA VECTOR RETRIEVAL 2025 Microsoft C
1211 Retrieval Head Mechanistically Explains Long-Context Factuality. 2025 ICLR Peking University; University of Washington; University of Edinburgh BC
1212 Rethinking and Improving Autoformalization: Towards a Faithful Metric and a Dependency Retrieval-based Approach. 2025 ICLR B
1213 Rethinking Visual Counterfactual Explanations Through Region Constraint. 2025 ICLR University of Warsaw; Warsaw University of Technology; University of Warsaw, War B
1214 Rethinking Shapley Value for Negative Interactions in Non-convex Games. 2025 ICLR B
1215 Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words. 2025 ICLR The University of Tokyo B
1216 Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective 2025 Zhejiang University / Tencent C
1217 Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates. 2025 NeurIPS School of Computer Science and Technology; Xi’an Jiaotong University; School of B
1218 Restyling Unsupervised Concept Based Interpretable Networks with Generative Models. 2025 ICLR Sorbonne Université; Institut Systèmes Intelligents et de Robotique; Helmholtz Z B
1219 Residualized Similarity for Faithfully Explainable Authorship Verification. 2025 EMNLP Department of Computer Science; Department of Linguistics; Institute for Advance B
1220 Rescorla-Wagner Steering of LLMs for Undesired Behaviors over Disproportionate Inappropriate Context. 2025 EMNLP University of Illinois Urbana-Champaign B
1221 Rescaled Influence Functions: Accurate Data Attribution in High Dimension. 2025 NeurIPS EECS, MIT, Cambridge, MA B
1222 Representational Similarity via Interpretable Visual Concepts. 2025 ICLR University of Edinburgh B
1223 Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition. 2025 NeurIPS Nanjing University of Aeronautics and Astronautics, Nanjing, China B
1224 Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering 2025 Xiaomi C
1225 Reinforced Learning Explicit Circuit Representations for Quantum State Characterization from Local Measurements. 2025 ICML University of Hong Kong B
1226 Regional Explanations: Bridging Local and Global Variable Importance. 2025 NeurIPS J.P. Morgan AI Research; LaMME, ENSIIE, University Paris Saclay B
1227 Refining Attention for Explainable and Noise-Robust Fact-Checking with Transformers. 2025 EMNLP EURECOM B
1228 Redundancy Undermines the Trustworthiness of Self-Interpretable GNNs. 2025 ICML University of Electronic Science and Technology of China; Iowa State University B
1229 Reducing Hallucinations in Large Vision-Language Models via Latent Space Steering. 2025 ICLR Stanford University B
1230 Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation. 2025 NeurIPS Indian Institute of Information Technology Dharwad; Indian Institute of Technolo B
1231 Reconsidering Faithfulness in Regular, Self-Explainable and Domain Invariant GNNs. 2025 ICLR University of Trento B
1232 Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation. 2025 ACL Westlake University; Hangzhou Dianzi University B
1233 Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought. 2025 NeurIPS Meta AI B
1234 Reasoning Models Don't Always Say What They Think
faithfulness chain-of-thought
2025 Lab post (Anthropic) Anthropic AC
1235 Reasoning Elicitation in Language Models via Counterfactual Feedback. 2025 ICLR Harvard University; Microsoft Research Cambridge; Cornell Tech BC
1236 Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference. 2025 ACL Idiap Research Institute B
1237 ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning 2025 Baichuan AI / Tongji University / University of Edinburgh / Zhejiang University C
1238 ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability. 2025 ICLR Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, B
1239 ReCoVeR the Target Language: Language Steering without Sacrificing Task Performance. 2025 EMNLP University of Cambridge; University of Würzburg B
1240 Rationalize and Align: Enhancing Writing Assistance with Rationale via Self-Training for Improved Alignment. 2025 ACL National University of Singapore B
1241 Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability. 2025 ACL University of Science and Technology of China; Beijing University of Posts and T B
1242 RankSHAP: Shapley Value Based Feature Attributions for Learning to Rank. 2025 ICLR Manning College of Information and Computer Sciences; University of Massachusett B
1243 RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models 2025 Houmo AI C
1244 RISE: Radius of Influence based Subgraph Extraction for 3D Molecular Graph Explanation. 2025 ICML Stony Brook University; New Jersey Institute of Technology; Arizona State Univer B
1245 RD-MCSA: A Multi-Class Sentiment Analysis Approach Integrating In-Context Classification Rationales and Demonstrations. 2025 EMNLP Beijing Institute of Mathematical Sciences and Applications; Renmin University o B
1246 RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals. 2025 ICML University of Chicago B
1247 R.I.P.: Better Models by Survival of the Fittest Prompts 2025 Meta + New York University + UC Berkeley C
1248 Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models. 2025 Tsinghua University C
1249 Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization 2025 The University of Tokyo C
1250 Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering. 2025 EMNLP Iowa State University B
1251 Quantifying Misattribution Unfairness in Authorship Attribution. 2025 ACL Stony Brook University; University of Pennsylvania B
1252 QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models 2025 New York University C
1253 QPM: Discrete Optimization for Globally Interpretable Image Classification. 2025 ICLR Institute for Information Processing (tnt); Leibniz University Hannover; Intel L B
1254 QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models. 2025 EMNLP Inspired Spine; University of Science and Technology of China; ByteDance; Fudan B
1255 Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment 2025 Seoul National University C
1256 PychoAgent: Psychology-driven LLM Agents for Explainable Panic Prediction on Social Media during Sudden Disaster Events. 2025 EMNLP National University of Defense Technology; Tsinghua University; Department of Ps B
1257 Provably Robust Explainable Graph Neural Networks against Graph Perturbation Attacks. 2025 ICLR Cranberry-Lemon University; University of the Witwatersrand; Illinois Institute B
1258 Provably Accurate Shapley Value Estimation via Leverage Score Sampling. 2025 ICLR New York University B
1259 ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering. 2025 EMNLP Shandong University; Harvard University B
1260 ProtoPairNet: Interpretable Regression through Prototypical Pair Reasoning. 2025 NeurIPS University of Maine; University of New Hampshire at Manchester; Dartmouth Colleg B
1261 ProtoLens: Advancing Prototype Learning for Fine-Grained Interpretability in Text Classification. 2025 ACL University University B
1262 PropXplain: Can LLMs Enable Explainable Propaganda Detection? 2025 EMNLP University of Toronto B
1263 Programming Refusal with Conditional Activation Steering. 2025 ICLR University of Pennsylvania; IBM Research B
1264 Probing the Latent Hierarchical Structure of Data via Diffusion Models. 2025 ICLR Institute of Physics, EPFL; EPFL B
1265 Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks. 2025 ACL Zhejiang University; Zhejiang Normal University B
1266 Probing for Arithmetic Errors in Language Models. 2025 EMNLP ETH Zrich B
1267 Probing and Boosting Large Language Models Capabilities via Attention Heads. 2025 EMNLP Harbin Institute of Technology; Peng Cheng Laboratory B
1268 Probing Visual Language Priors in VLMs. 2025 ICML University of Michigan B
1269 Probing Subphonemes in Morphology Models. 2025 ACL Ben-Gurion University of the Negev B
1270 Probing Semantic Routing in Large Mixture-of-Expert Models. 2025 EMNLP Intel Labs; Intel (United Arab Emirates); Oracle B
1271 Probing Relative Interaction and Dynamic Calibration in Multi-modal Entity Alignment. 2025 ACL Northeastern University B
1272 Probing Political Ideology in Large Language Models: How Latent Political Representations Generalize Across Tasks. 2025 EMNLP University of Chicago B
1273 Probing Neural Combinatorial Optimization Models. 2025 NeurIPS Singapore Management University; Massachusetts Institute of; Technology B
1274 Probing Narrative Morals: A New Character-Focused MFT Framework for Use with Large Language Models. 2025 EMNLP McGill University B
1275 Probing Logical Reasoning of MLLMs in Scientific Diagrams. 2025 EMNLP University of Pittsburgh B
1276 Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set. 2025 ACL EU Business School, Munich; Munich Center for Machine Learning B
1277 Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding. 2025 EMNLP University of Wisconsin–Madison B
1278 Probing Equivariance and Symmetry Breaking in Convolutional Networks. 2025 NeurIPS AMLab, University of Amsterdam; QurAI, University of Amsterdam; Independent Rese B
1279 Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models. 2025 ICLR College of Cyber Science, Nankai University; Independent Researcher; Alibaba Gro B
1280 Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing. 2025 ICLR University of Minnesota; University of North Carolina at Charlotte B
1281 PrimeX: A Dataset of Worldview, Opinion, and Explanation. 2025 EMNLP Apple (United States); University of Southern California B
1282 Prediction via Shapley Value Regression. 2025 ICML KTH Royal Institute of Technology B
1283 Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMs. 2025 ICLR University of Science and Technology of China; Department of Computer Science an B
1284 Practical do-Shapley Explanations with Estimand-Agnostic Causal Inference. 2025 NeurIPS Barcelona Supercomputing Center B
1285 Practical Guide for Model Selection for Real‑World Use Cases (Blog) 2025 Openai C
1286 Position-aware Automatic Circuit Discovery. 2025 ACL Technion – Israel Institute of Technology; Northeastern University BC
1287 Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs. 2025 EMNLP Computer Science Lab, SRI; Johns Hopkins University B
1288 PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration 2025 Renmin University of China C
1289 Pixel-level Certified Explanations via Randomized Smoothing. 2025 ICML Helmholtz Center for Information Security B
1290 Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis. 2025 EMNLP The Hong Kong University of Science and Technology (Guangzhou); Hong Kong Univer B
1291 Personalized Text Generation with Contrastive Activation Steering. 2025 ACL Chinese Academy of Sciences; University of Chinese Academy of Sciences; Northeas B
1292 Personality Alignment of Large Language Models 2025 Zhejiang University C
1293 Persona Vectors: Monitoring and Controlling Character Traits in Language Models
activations monitoring persona hallucination sycophancy
2025 Lab post (Anthropic) Anthropic A
1294 Performative Validity of Recourse Explanations. 2025 NeurIPS Tübingen AI Center, ^2University of Tübingen; Korteweg-de Vries Institute for Ma B
1295 ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization 2025 Meta AI C
1296 ParamΔ for Direct Weight Mixing: Post-Train Large Language Model at Zero Cost 2025 Meta C
1297 ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation. 2025 NeurIPS School of Computer Science and Engineering, Northeastern University, China; Depa B
1298 PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-Encoder. 2025 ACL National University of Singapore; Harbin Institute of Technology B
1299 POISONING ATTACKS ON LLMS REQUIRE A NEAR-CONSTANT NUMBER OF POISON SAMPLES 2025 Anthropic C
1300 Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning. 2025 NeurIPS École Normale Supérieure Paris-Saclay; École Normale Supérieure B
1301 Optimal Information Retention for Time-Series Explanations. 2025 ICML Beijing Jiaotong University; Beijing Emergency Medical Center B
1302 OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution. 2025 EMNLP University of Calabria B
1303 Open-World Authorship Attribution. 2025 ACL National University of Singapore B
1304 Open-Sourcing Circuit-Tracing Tools
circuits attribution-graph attribution neurons
2025 Lab post (Anthropic) Anthropic A
1305 Open Problems in Mechanistic Interpretability
mechanistic concepts interpretability
2025 Lab post (Google DeepMind) Google DeepMind A
1306 One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models. 2025 NeurIPS École Polytechnique Fédérale de Lausanne; Northeastern University B
1307 One Wave To Explain Them All: A Unifying Perspective On Feature Attribution. 2025 ICML Électricité de France (France); Centre Observation, Impacts, Énergie; Hôpital Sa B
1308 One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs. 2025 NeurIPS Beijing University of Posts and Telecommunications; Shihezi University B
1309 On the Versatility of Sparse Autoencoders for In-Context Learning. 2025 EMNLP University of Southern California B
1310 On the Biology of a Large Language Model
circuits attribution-graph attribution hallucination
2025 Lab post (Anthropic) Anthropic A
1311 On Synthesizing Data for Context Attribution in Question Answering. 2025 ACL NEC Laboratories Europe; NEC Laboratories America; NUC Corporation (Japan); Univ B
1312 On Relation-Specific Neurons in Large Language Models. 2025 EMNLP Ludwig-Maximilians-Universität München; Technical University of Munich; Google D B
1313 On Minimizing Adversarial Counterfactual Error in Adversarial Reinforcement Learning. 2025 ICLR Singapore Management University; Rutgers, The State University of New Jersey B
1314 On Logic-based Self-Explainable Graph Neural Networks. 2025 NeurIPS INSA Lyon; EPITA B
1315 On Explaining Equivariant Graph Networks via Improved Relevance Propagation. 2025 ICML Texas A&M University; University of Houston B
1316 OmniKV: Dynamic Context Selection for Efficient Long-Context LLMs 2025 Ant Group C
1317 OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature. 2025 EMNLP University of Massachusetts Amherst; University of Maryland, College Park B
1318 ON THE ROLE OF ATTENTION HEADS IN LARGE LANGUAGE MODEL SAFETY 2025 Tongyi Lab (ICLR Oral) C
1319 Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning. 2025 ICLR The University of Texas at Austin; University of California San Diego; Universit B
1320 Not Lost After All: How Cross-Encoder Attribution Challenges Position Bias Assumptions in LLM Summarization. 2025 EMNLP Dalhousie University; State Research Center of Virology and Biotechnology VECTOR B
1321 Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human Diversity. 2025 EMNLP University of Illinois Urbana-Champaign B
1322 Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions Explainability. 2025 ACL Corti; University of Copenhagen; IT University of Copenhagen; LUT University B
1323 No Need for Explanations: LLMs can implicitly learn from mistakes in-context. 2025 EMNLP Imperial College London B
1324 No Black Boxes: Interpretable and Interactable Predictive Healthcare with Knowledge-Enhanced Agentic Causal Discovery. 2025 EMNLP Stevens Institute of Technology B
1325 Neurons as Detectors of Coherent Sets in Sensory Dynamics. 2025 NeurIPS Center for Computational Neuroscience, Flatiron Institute, Simons Foundation, Ne B
1326 NeuronTune: Towards Self-Guided Spurious Bias Mitigation. 2025 ICML University of Virginia B
1327 NeuronMerge: Merging Models via Functional Neuron Groups. 2025 ACL Zhejiang University; Alibaba Group (China); Zhejiang Gongshang University B
1328 Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models. 2025 ICLR University of Southern California, CA, USA; University of California, Riverside, B
1329 Neuron-Level Sequential Editing for Large Language Models. 2025 ACL University of Science and Technology of China; National University of Singapore; B
1330 Neuron-Level Differentiation of Memorization and Generalization in Large Language Models. 2025 EMNLP National Taiwan University; National Tsing Hua University B
1331 Neuron based Personality Trait Induction in Large Language Models. 2025 ICLR Gaoling School of Artificial Intelligence, Renmin University of China; Tongyi La B
1332 Neuron Platonic Intrinsic Representation From Dynamics Using Contrastive Learning. 2025 ICLR Peking University; University of Georgia; Chinese Academy of Sciences; Institute B
1333 Neuron Empirical Gradient: Discovering and Quantifying Neurons' Global Linear Controllability. 2025 ACL Advanced Institute of Industrial Technology; The University of Tokyo B
1334 Neuron Activation Modulation for Text Style Transfer: Guiding Large Language Models. 2025 ACL Beijing University of Posts and Telecommunications B
1335 NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning. 2025 EMNLP Institute for Bulgarian Language B
1336 Neural Interpretable PDEs: Harmonizing Fourier Insights with Attention for Scalable and Interpretable Physics Discovery. 2025 ICML Lehigh University; Global Embedded Technologies (United States) B
1337 Neural Causal Graph for Interpretable and Intervenable Classification. 2025 ICLR B
1338 NeurFlow: Interpreting Neural Networks through Neuron Groups and Functional Interactions. 2025 ICLR Institute for Learning Innovation; Hanoi University of Science and Technology; U B
1339 NetFormer: An interpretable model for recovering dynamical connectivity in neuronal population dynamics. 2025 ICLR B
1340 Negative Results for Sparse Autoencoders on Downstream Tasks and Deprioritising SAE Research
SAE
2025 Lab post (Google DeepMind) Google DeepMind A
1341 NavBench: Probing Multimodal Large Language Models for Embodied Navigation. 2025 NeurIPS The University of Adelaide; The University of Queensland; Mohamed bin Zayed Univ B
1342 Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability. 2025 ICLR University of Technology Sydney; University of Sydney B
1343 NarratEX Dataset: Explaining the Dominant Narratives in News Texts. 2025 EMNLP Universidade do Porto; Athens University of Economics and Business B
1344 NarGINA: Towards Accurate and Interpretable Children's Narrative Ability Assessment via Narrative Graphs. 2025 ACL Nanjing Normal University B
1345 NOBLE - Neural Operator with Biologically-informed Latent Embeddings to Capture Experimental Variability in Biological Neuron Models. 2025 NeurIPS ETH Zürich; California Institute of Technology; University of Alberta; Alberta M B
1346 Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems. 2025 EMNLP Technische Universität Berlin; German Research Centre for Artificial Intelligenc B
1347 Multi-Level Explanations for Generative Language Models. 2025 ACL Harvard University; IBM Research; Merck Research Labs B
1348 Multi-Domain Explainability of Preferences. 2025 EMNLP Decision Sciences International Corporation (United States); IIBM Research B
1349 Multi-Attribute Steering of Language Models via Targeted Intervention. 2025 ACL University of North Carolina at Chapel Hill; University of North Carolina Health B
1350 Monet: Mixture of Monosemantic Experts for Transformers. 2025 ICLR Korea University, ^2KAIST, ^3AIGEN Sciences B
1351 MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correction. 2025 EMNLP Carnegie Mellon University; Hong Kong University of Science and Technology B
1352 Models of Heavy-Tailed Mechanistic Universality. 2025 ICML The University of Melbourne; University of California, Berkeley; International C B
1353 Model Unlearning via Sparse Autoencoder Subspace Guided Projections. 2025 EMNLP University of Hong Kong; Chinese University of Hong Kong, Shenzhen B
1354 Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws. 2025 ICML Texas A&M University; Oracle; Indiana University; Google; University of Florida B
1355 Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models. 2025 ACL University of Notre Dame; University of Pennsylvania; Georgia Institute of Techn B
1356 MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines. 2025 ACL Charles University; Sorbonne Université B
1357 Mixture of Experts Made Intrinsically Interpretable. 2025 ICML University of Oxford; National University of Singapore B
1358 Mitigating Spurious Correlations via Counterfactual Contrastive Learning. 2025 EMNLP University of Amsterdam; Peking University B
1359 Mitigating Overthinking in Large Reasoning Models via Manifold Steering. 2025 NeurIPS Institute of Artificial Intelligence, Beihang University, Beijing 100191, China; B
1360 Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values? 2025 University of Washington C
1361 MicroEdit: Neuron-level Knowledge Disentanglement and Localization in Lifelong Model Editing. 2025 EMNLP Jilin University; Engineering Research Center of Knowledge-Driven Human-Machine B
1362 MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning 2025 University of Illinois Urbana-Champaign C
1363 Metric-Driven Attributions for Vision Transformers. 2025 ICLR B
1364 MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs. 2025 EMNLP Yale University; Google (United States); University of Toronto B
1365 MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social Media. 2025 EMNLP Beijing University of Technology; Wuhan University; Pitié-Salpêtrière Hospital B
1366 MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models. 2025 EMNLP Chinese University of Hong Kong; University of International Relations; Tencent B
1367 MemeIntel: Explainable Detection of Propagandistic and Hateful Memes. 2025 EMNLP Qatar Computing Research Institute, Qatar; Blackbird.AI; APAVI.AI B
1368 MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models. 2025 ICLR Department of Electrical and Computer Engineering, University of Southern Califo B
1369 Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization. 2025 ICML University of Maryland; Georgia Institute of Technology; University of Bristol; B
1370 Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models. 2025 EMNLP EU Business School, Munich; Munich Center for Machine Learning B
1371 Mechanistic Permutability: Match Features Across Layers. 2025 ICLR T-Tech, ^2 Moscow Institute of Physics and Technologies, ^3 HSE University B
1372 Mechanistic PDE Networks for Discovery of Governing Equations. 2025 ICML Institute of Science and Technology, Klosterneuburg, Austria; University of Amst B
1373 Mechanistic Interpretability of RNNs emulating Hidden Markov Models. 2025 NeurIPS Institute of Neuroinformatics, University of Zurich; ETH Zurich B
1374 Mechanistic Interpretability of Emotion Inference in Large Language Models. 2025 ACL University of Southern California; University of California, Los Angeles B
1375 Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations. 2025 EMNLP Stanford University B
1376 Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models. 2025 NeurIPS Zidi Xiong, Harvard University; Harvard University B
1377 Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks. 2025 EMNLP University of Zurich B
1378 Measuring and Guiding Monosemanticity. 2025 NeurIPS Technische Universität Darmstadt; German Research Center for AI (DFKI) B
1379 Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse. 2025 ICLR Singapore University of Technology and Design; DSO National Laboratories B
1380 Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps. 2025 EMNLP Technion – Israel Institute of Technology; University of Utah BC
1381 Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning. 2025 NeurIPS Hong Kong University of Science and Technology; Mila - Québec AI Institute; Univ B
1382 MarathiEmoExplain: A Dataset for Sentiment, Emotion, and Explanation in Low-Resource Marathi. 2025 EMNLP Indian Institute of Technology Jammu B
1383 Make Information Diffusion Explainable: LLM-based Causal Framework for Diffusion Prediction. 2025 NeurIPS Tianjin University B
1384 MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Model. 2025 EMNLP Seoul National University; AIGENDRUG Co., Ltd. (South Korea) B
1385 MPRF: Interpretable Stance Detection through Multi-Path Reasoning Framework. 2025 EMNLP University of Chinese Academy of Sciences; State Key Laboratory of AI Safety; Ch B
1386 MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
reward-hacking
2025 Lab post (Google DeepMind) Google DeepMind A
1387 MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models. 2025 NeurIPS University of Rochester; Purdue University; NVIDIA B
1388 MMDEND: Dendrite-Inspired Multi-Branch Multi-Compartment Parallel Spiking Neuron for Sequence Modeling. 2025 ACL Institute for Complex Systems; Chinese Academy of Sciences; University of Chines B
1389 MLIP Arena: Advancing Fairness and Transparency in Machine Learning Interatomic Potentials via an Open, Accessible Benchmark Platform. 2025 NeurIPS University of California, Berkeley; Lawrence Berkeley National Laboratory; Imper B
1390 MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models. 2025 ICLR School of Artificial Intelligence, University of Chinese Academy of Sciences; Th B
1391 MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion Prediction. 2025 NeurIPS Renmin University of China; University of Southern California; Amazon Web Servic B
1392 MIB: A Mechanistic Interpretability Benchmark. 2025 ICML University of Maryland; Brown University; Technion – Israel Institute of Technol B
1393 MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation. 2025 EMNLP Southern States University B
1394 MELODI: Exploring Memory Compression for Long Contexts 2025 DeepMind C
1395 MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis. 2025 ACL University of California, Santa Barbara; NEC Laboratories Europe, Heidelberg, Ge B
1396 MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data. 2025 ACL Maastricht University B
1397 MAPLE: Enhancing Review Generation with Multi-Aspect Prompt LEarning in Explainable Recommendation. 2025 ACL Department of Computer Science and Information Engineering; National Cheng Kung B
1398 MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset. 2025 ACL University of Colorado Boulder; University of Chicago; Johannes Gutenberg Univer B
1399 MAGE: Model-Level Graph Neural Networks Explanations via Motif-based Graph Generation. 2025 ICLR Iowa State University B
1400 LucidPPN: Unambiguous Prototypical Parts Network for User-centric Interpretable Computer Vision. 2025 ICLR Jagiellonian University; Institute of Applied Psychology B
1401 Low-Rank Adapting Models for Sparse Autoencoders. 2025 ICML Massachusetts Institute of Technology B
1402 Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models 2025 C
1403 LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data. 2025 ACL Digital Development Communications International (United States); The Hong Kong B
1404 Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs. 2025 EMNLP National Centre for Atmospheric Science B
1405 Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs. 2025 ACL University of Toronto; Technische Universität Darmstadt; Microsoft Research B
1406 LittleBit: Ultra Low-Bit Quantization via Latent Factorization 2025 Samsung Research C
1407 Linguistic Neuron Overlap Patterns to Facilitate Cross-lingual Transfer on Low-resource Languages. 2025 EMNLP Beijing Foreign Studies University; King's College London B
1408 LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss. 2025 NeurIPS Apple, ^2Sapienza, ^ B
1409 LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference. 2025 EMNLP EU Business School, Munich; Munich Center for Machine Learning; Faculty of Compu B
1410 Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language Models. 2025 ACL Centre for Argument Technology; University of Dundee B
1411 Leveraging Variation Theory in Counterfactual Data Augmentation for Optimized Active Learning. 2025 ACL University of Notre Dame B
1412 Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility. 2025 ACL University of Massachusetts Amherst; Northwestern University B
1413 Less is More: Explainable and Efficient ICD Code Prediction with Clinical Entities. 2025 ACL The University of Sydney; CSIRO Data61; Beamtree B
1414 Learning to cluster neuronal function. 2025 NeurIPS Institute of Computer Science and Campus Institute Data Science, University Gött B
1415 Learning to Steer: Input-dependent Steering for Multimodal LLMs. 2025 NeurIPS ISIR, Sorbonne Université, Paris, France B
1416 Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention. 2025 ACL Hong Kong Polytechnic University B
1417 Learning and aligning single-neuron invariance manifolds in visual cortex. 2025 ICLR University of Göttingen, Germany; University of Tübingen, Institute for Neurobio B
1418 Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning. 2025 ACL Adobe B
1419 Learning Multi-Level Features with Matryoshka Sparse Autoencoders. 2025 ICML Independent Researchers (MATS); Google DeepMind BC
1420 Learning Interpretable Hierarchical Dynamical Systems Models from Time Series Data. 2025 ICLR Central Institute of Mental Health (CIMH); Interdisciplinary Center for Scientif B
1421 Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression 2025 Nanyang Technological University C
1422 Learning Counterfactual Outcomes Under Rank Preservation. 2025 NeurIPS Beijing Technology and Business University; Peking University; Zhejiang Universi B
1423 LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching. 2025 NeurIPS Munich Center for Machine Learning (MCML), LMU Munich, Germany; Aarhus Universit B
1424 LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language Models. 2025 NeurIPS Chinese Academy of Sciences; University of Chinese Academy of Sciences B
1425 Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations. 2025 EMNLP University of the District of Columbia; University of Missouri; Columbia Univers B
1426 Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion. 2025 NeurIPS Seoul National University B
1427 Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability. 2025 EMNLP Northwestern University; New Jersey Institute of Technology; University of Brist B
1428 Large Language Models are Interpretable Learners. 2025 ICLR UCLA; Google Research B
1429 Large Language Models Are Cross-Lingual Knowledge-Free Reasoners 2025 Nanjing University C
1430 LaMAGIC2: Advanced Circuit Formulations for Language Model-Based Analog Topology Generation. 2025 ICML Duke University; University of California, Los Angeles; IBM T. J. Watson Researc B
1431 LORE: Continual Logit Rewriting Fosters Faithful Generation. 2025 EMNLP William & Mary B
1432 LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering. 2025 ACL Ludwig-Maximilians-Universität München; Munich Research Center, Huawei Technolog B
1433 LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing. 2025 ACL Polytechnic University of Bari; Sapienza University of Rome B
1434 LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations. 2025 EMNLP University of Oxford; Trinity College Dublin B
1435 LLM Interpretability with Identifiable Temporal-Instantaneous Representation. 2025 NeurIPS Carnegie Mellon University; Mohamed bin Zayed University of Artificial Intellige B
1436 LIMEFLDL: A Local Interpretable Model-Agnostic Explanations Approach for Label Distribution Learning. 2025 ICML Nanjing University of Science and Technology; Hong Kong Polytechnic University; B
1437 LIME: Less Is More for MLLM Evaluation. 2025 ACL University of Manchester B
1438 LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement Learning. 2025 ICLR Carnegie Mellon University B
1439 LDIR: Low-Dimensional Dense and Interpretable Text Embeddings with Relative Representations. 2025 ACL Shenzhen University B
1440 LAQuer: Localized Attribution Queries in Content-grounded Generation. 2025 ACL Bar-Ilan University; University of North Carolina at Chapel Hill B
1441 Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models. 2025 ACL National University of Singapore B
1442 Know Thyself by Knowing Others: Learning Neuron Identity from Population Context. 2025 NeurIPS University of Pennsylvania, ^ 2 Columbia University, ^ 3 McGill University, ^ 4 B
1443 Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents 2025 Tsinghua University / Moonshot AI C
1444 Kernel Density Steering: Inference-Time Scaling via Mode Seeking for Image Restoration. 2025 NeurIPS Google, ^2Washington University in St. Louis B
1445 KLay: Accelerating Arithmetic Circuits for Neurosymbolic AI. 2025 ICLR Centre for Applied Autonomous Sensor Systems; Örebro University B
1446 Just a Scratch: Enhancing LLM Capabilities for Self-harm Detection through Intent Differentiation and Emoji Interpretation. 2025 ACL Fondazione Bruno Kessler; Indian Institute of Technology Patna; Indian Institute B
1447 JoPA: Explaining Large Language Model's Generation via Joint Prompt Attribution. 2025 ACL Pennsylvania State University B
1448 Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations. 2025 ICML University of Bristol B
1449 J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning 2025 Meta C
1450 Iterative Vectors: In-Context Gradient Steering without Backpropagation. 2025 ICML Peking University, School of Intelligence Science and Technology, State Key Labo B
1451 It Helps to Take a Second Opinion: Teaching Smaller LLMs To Deliberate Mutually via Selective Rationale Optimisation. 2025 ICLR Media and Data Science Research Lab, Adobe B
1452 Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness. 2025 ICLR Institute of Computing Technology; University of Chinese Academy of Sciences; Un B
1453 Investigating the Impact of Conceptual Metaphors on LLM-based NLI through Shapley Interactions. 2025 EMNLP Leibniz University Hannover; EU Business School, Munich; Bielefeld University; S B
1454 Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution. 2025 ACL Indian Institute of Technology Delhi; Independent Researcher; University of Tart B
1455 Investigating Pattern Neurons in Urban Time Series Forecasting. 2025 ICLR National University of Singapore B
1456 Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors. 2025 EMNLP Nara Institute of Science and Technology; Mohamed bin Zayed University of Artifi B
1457 Investigating Context Faithfulness in Large Language Models: The Roles of Memory Strength and Evidence Style. 2025 ACL Department of Computer Science; Iowa State University B
1458 Introducing Nested Learning: A new ML paradigm for continual learning 2025 Google Research / University of Southern California (USC) C
1459 Intrinsic User-Centric Interpretability through Global Mixture of Experts. 2025 ICLR EPFL B
1460 Interpreting the Second-Order Effects of Neurons in CLIP. 2025 ICLR University of California, Berkeley B
1461 Interpreting Language Reward Models via Contrastive Explanations. 2025 ICLR Imperial College London; J.P. Morgan AI Research B
1462 Interpreting CLIP with Hierarchical Sparse Autoencoders. 2025 ICML University of Warsaw B
1463 Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety. 2025 EMNLP Georgia Institute of Technology B
1464 Interpretable and Parameter Efficient Graph Neural Additive Models with Random Fourier Features. 2025 NeurIPS Fujitsu Research of India B
1465 Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational Pathology. 2025 ICLR University of Electronic Science and Technology of China B
1466 Interpretable Unsupervised Joint Denoising and Enhancement for Real-World low-light Scenarios. 2025 ICLR Tsinghua University B
1467 Interpretable Text Embeddings and Text Similarity Explanation: A Survey. 2025 EMNLP University of Zurich; University of Stuttgart B
1468 Interpretable Next-token Prediction via the Generalized Induction Head. 2025 NeurIPS Microsoft Research; Seoul National University; Stanford University B
1469 Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximization. 2025 EMNLP University of Massachusetts Amherst B
1470 Interpretable Causal Representation Learning for Biological Data in the Pathway Space. 2025 ICLR CIMA University of Navarra, CCUN, IdiSNA, Pamplona, Spain; TECNUN, University of B
1471 Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks. 2025 ICLR The Hong Kong University of Science and Technology; Sun Yat-Sen Memorial Hospita B
1472 Interpretability Analysis of Arithmetic In-Context Learning in Large Language Models. 2025 EMNLP University of Tübingen B
1473 Interpret and Improve In-Context Learning via the Lens of Input-Label Mappings. 2025 ACL Inspired Spine; University of Science and Technology of China; Independent Age; B
1474 Interpolating Neural Network-Tensor Decomposition (INN-TD): a scalable and interpretable approach for large-scale physics-based problems. 2025 ICML Northwestern University; Institute of Computational Modeling B
1475 Internal Value Alignment in LLM through Controlled Value Vector Activation 2025 C
1476 InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized Rationales. 2025 ICLR University of Virginia B
1477 InstaSHAP: Interpretable Additive Models Explain Shapley Values Instantly. 2025 ICLR University of Southern California B
1478 InfoCons: Identifying Interpretable Critical Concepts in Point Clouds via Information Theory. 2025 ICML School of Computer Science, Fudan University, China B
1479 Influence Functions for Scalable Data Attribution in Diffusion Models. 2025 ICLR University of Cambridge; University of Toronto; Max Planck Institute for Intelli B
1480 Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models 2025 Tsinghua University + CMU (ICLR 2025) C
1481 Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning. 2025 ICML Imperial College London B
1482 Inducing Argument Facets for Faithful Opinion Summarization. 2025 EMNLP Shandong University B
1483 InCoDe: Interpretable Compressed Descriptions For Image Generation. 2025 ICLR B
1484 In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention. 2025 ICML Yale University; Nanjing University B
1485 Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanations. 2025 EMNLP Peking University; Fuzhou University B
1486 Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templates. 2025 EMNLP University of Notre Dame; Amazon B
1487 Improving LLM Reasoning through Interpretable Role-Playing Steering. 2025 EMNLP EU Business School, Munich; Northwestern University; Munich Center for Machine L B
1488 Improving Instruction-Following in Language Models through Activation Steering. 2025 ICLR ETH Zurich; Microsoft Research (India) B
1489 Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced Optimization. 2025 ACL Harbin Institute of Technology; Peng Cheng Laboratory; Northeastern University, B
1490 Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE. 2025 ACL NASK National Research Institute; Vrije Universiteit Amsterdam B
1491 Improved Representation Steering for Language Models. 2025 NeurIPS Stanford University B
1492 Identifying interactions across brain areas while accounting for individual-neuron dynamics with a Transformer-based variational autoencoder. 2025 NeurIPS Carnegie Mellon University B
1493 Identifying and Answering Questions with False Assumptions: An Interpretable Approach. 2025 EMNLP Department of Computer Science; University of Arizona B
1494 Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework. 2025 EMNLP Hong Kong University of Science and Technology B
1495 Identification of Multiple Logical Interpretations in Counter-Arguments. 2025 EMNLP Tohoku University; RIKEN; Beyond Reason; Ricoh; Japan Advanced Institute of Scie B
1496 IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response Theory. 2025 ACL University of Science and Technology of China; Institute of Artificial Intellige B
1497 IRIS: Interpretable Retrieval-Augmented Classification for Long Interspersed Document Sequences. 2025 ACL Duke University B
1498 IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector. 2025 NeurIPS Computer Science and Engineering, Hong Kong University of Science and Technology B
1499 ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs. 2025 ACL Peking University B
1500 ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests. 2025 NeurIPS Gaoling School of Artificial Intelligence, Renmin University of China B
1501 IBCircuit: Towards Holistic Circuit Discovery with Information Bottleneck. 2025 ICML Chinese University of Hong Kong; Alibaba Group (China); The Hong Kong University B
1502 I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts. 2025 ICML University of Pennsylvania; University of North Texas; University of Science and B
1503 I2AM: Interpreting Image-to-Image Latent Diffusion Models via Bi-Attribution Maps. 2025 ICLR Dongguk University B
1504 I-GUARD: Interpretability-Guided Parameter Optimization for Adversarial Defense. 2025 EMNLP King's College London B
1505 HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks. 2025 ICLR Stanford University; ♠ Confirm Labs; ♣ Ghent University B
1506 Hybrid Re-matching for Continual Learning with Parameter-Efficient Tuning 2025 Nankai University & Tsinghua University C
1507 How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations. 2025 ICLR Max Planck Institute for Informatics, Saarland Informatics Campus, Germany, ^2In B
1508 How to Generalize the Detection of AI-Generated Text: Confounding Neurons. 2025 EMNLP Italian institute for Genomic Medicine B
1509 How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders. 2025 EMNLP RIKEN; Tohoku University B
1510 How Programming Concepts and Neurons Are Shared in Code Language Models. 2025 ACL Munich Center for Machine Learning; Sorbonne Université B
1511 How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation. 2025 EMNLP Fudan University; University of Southern California B
1512 How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis. 2025 EMNLP University of Oxford; Jagiellonian University; Harvard University B
1513 How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training. 2025 ACL Zhejiang University; National University of Singapore; ♢Zhejiang Key Laboratory BC
1514 High-order Interactions Modeling for Interpretable Multi-Agent Q-Learning. 2025 NeurIPS School of Management and Engineering; Nanjing University; School of Information B
1515 High-dimensional neuronal activity from low-dimensional latent dynamics: a solvable model. 2025 NeurIPS University College London; École Polytechnique Fédérale de Lausanne; Shanghai Ji B
1516 High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity. 2025 ICLR Dalian University of Technology; vivo Mobile Communication Co., Ltd B
1517 Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion Trajectory. 2025 NeurIPS Fudan University; The University of Hong Kong B
1518 HiBug2: Efficient and Interpretable Error Slice Discovery for Comprehensive Model Debugging. 2025 ICLR The Chinese University of Hong Kong B
1519 HeadMap: Locating and Enhancing Knowledge Circuits in LLMs. 2025 ICLR B
1520 Head Pursuit: Probing Attention Specialization in Multimodal Transformers. 2025 NeurIPS Sapienza University of Rome, Italy; Institute of Science and Technology, Austria B
1521 Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models. 2025 EMNLP University of Science, VNU-HCM; Indiana University B
1522 HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models. 2025 ICLR Georgia Institute of Technology; University of Oregon; University of Illinois Ur B
1523 H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs 2025 Tsinghua University C
1524 Gumbel Counterfactual Generation From Language Models. 2025 ICLR Bar-Ilan University; Allen Institute for Artificial Intelligence B
1525 Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups. 2025 EMNLP London School of Economics and Political Science B
1526 GraphNarrator: Generating Textual Explanations for Graph Neural Networks. 2025 ACL Emory University B
1527 Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models. 2025 ICML Monash University; Nanjing University of Science and Technology; Shanghai Jiao T B
1528 Graph-Guided Textual Explanation Generation Framework. 2025 EMNLP University of Copenhagen; University of Mannheim; University of Technology Nurem B
1529 Graph Inverse Style Transfer for Counterfactual Explainability. 2025 ICML Sapienza University of Rome B
1530 Gradient-based Explanations for Deep Learning Survival Models. 2025 ICML Leibniz Institute for Prevention Research and Epidemiology - BIPS; University of B
1531 GradPS: Resolving Futile Neurons in Parameter Sharing Network for Multi-Agent Reinforcement Learning. 2025 ICML Xiamen University; Key Laboratory of Multimedia Trusted Perception and Efficient B
1532 Going Beyond Static: Understanding Shifts with Time-Series Attribution. 2025 ICLR Tsinghua University, Department of Computer Science and Technology, Beijing, Chi B
1533 Gnothi Seauton: Empowering Faithful Self-Interpretability in Black-Box Transformers. 2025 ICLR School of Artificial Intelligence, Shanghai Jiao Tong University; Efficient and B
1534 Generating Likely Counterfactuals Using Sum-Product Networks. 2025 ICLR Faculty of Electrical Engineering, Czech Technical University B
1535 Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution. 2025 ACL Chinese University of Hong Kong; Massachusetts Institute of Technology; Associat B
1536 Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow. 2025 ACL Simon Fraser University; University of California, Berkeley B
1537 Gemma Scope 2
SAE transcoder faithfulness chain-of-thought interpretability
2025 Lab post (Google DeepMind) Google DeepMind A
1538 Gaussian Mixture Counterfactual Generator. 2025 ICLR B
1539 GSE: Group-wise Sparse and Explainable Adversarial Attacks. 2025 ICLR Department for AI in Society, Science, and Technology, Zuse Institute Berlin, Ge B
1540 GEMS: Generation-Based Event Argument Extraction via Multi-perspective Prompts and Ontology Steering. 2025 ACL University of Electronic Science and Technology of China B
1541 GEFA: A General Feature Attribution Framework Using Proxy Gradient Estimation. 2025 ICML Freie Universität Berlin B
1542 From Synapses to Dynamics: Obtaining Function from Structure in a Connectome Constrained Model of the Head Direction Circuit. 2025 NeurIPS Institute of Cognitive and Brain Sciences; Massachusetts Institute of Technology B
1543 From Shortcuts to Balance: Attribution Analysis of Speech-Text Feature Utilization in Distinguishing Original from Machine-Translated Texts. 2025 EMNLP Center for Language Studies; University of Groningen; Universitat d’Alacant; Uni B
1544 From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers 2025 The University of Tokyo & RIKEN AIP C
1545 From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models. 2025 EMNLP Microsoft B
1546 From Probability to Counterfactuals: the Increasing Complexity of Satisfiability in Pearl's Causal Hierarchy. 2025 ICLR Saarland University, Germany; University of Lübeck, Germany B
1547 From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature Selection. 2025 ICML ETH Zurich B
1548 From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment. 2025 ACL Xiamen University B
1549 From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language Models. 2025 ICML Columbia University; Ginkgo Bioworks, Inc. (United States) B
1550 From Imitation to Introspection: Probing Self-Consciousness in Language Models. 2025 ACL Shanghai Artificial Intelligence Laboratory; Tongji University; Fudan University B
1551 From GNNs to Trees: Multi-Granular Interpretability for Graph Neural Networks. 2025 ICLR Zhejiang University; Nanyang Technological University B
1552 From Black-box to Causal-box: Towards Building More Interpretable Models. 2025 NeurIPS Causal Artificial Intelligence Lab; Columbia University B
1553 Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages. 2025 ICML Moody Foundation; University of Notre Dame; MIT-IBM Watson AI Lab, IBM Research B
1554 Forward Knows Efficient Backward Path: Saliency-Guided Memory-Efficient Fine-tuning of Large Language Models. 2025 ACL Korea University B
1555 Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents. 2025 EMNLP University of Maryland; Adobe Research B
1556 Focus On This, Not That! Steering LLMs with Adaptive Feature Specification. 2025 ICML University of Oxford; University of Illinois Urbana-Champaign; University of Chi B
1557 FlowSearch: Advancing deep research with dynamic structured knowledge flow 2025 Shanghai AI Lab C
1558 FlowMixer: A Depth-Agnostic Neural Architecture for Interpretable Spatiotemporal Forecasting. 2025 NeurIPS New York University in Abu Dhabi; New York University B
1559 FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models. 2025 EMNLP School of Intelligent Systems Engineering, Sun Yat-Sen University; DP Technology B
1560 Fix False Transparency by Noise Guided Splatting. 2025 NeurIPS Case Western Reserve University B
1561 FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation. 2025 ACL Technische Universität Berlin; German Research Centre for Artificial Intelligenc BC
1562 First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training 2025 Shanghai Jiao Tong University & Zhongguancun Academy C
1563 Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State Tracking. 2025 ACL Key Laboratory of High Confidence Software Technology (PKU), MOE, China; Peking B
1564 Final-Model-Only Data Attribution with a Unifying View of Gradient-Based Methods. 2025 NeurIPS IBM Research; Merck Research Labs B
1565 FinGrAct: A Framework for FINe-GRrained Evaluation of ACTionability in Explainable Automatic Fact-Checking. 2025 EMNLP Université de Sherbrooke B
1566 FiDeLiS: Faithful Reasoning in Large Language Models for Knowledge Graph Question Answering. 2025 ACL National University of Singapore; University of Science and Technology of China B
1567 Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations. 2025 NeurIPS University of Maryland, College Park B
1568 Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders. 2025 ACL Skolkovo Institute of Science and Technology; AI Foundation and Algorithm Lab; A B
1569 Feature Responsiveness Scores: Model-Agnostic Explanations for Recourse. 2025 ICLR Haverford College B
1570 Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models. 2025 EMNLP New Jersey Institute of Technology; University of California, Santa Barbara; Geo B
1571 FastCAV: Efficient Computation of Concept Activation Vectors for Explaining Deep Neural Networks. 2025 ICML Institute of Data Science, German Aerospace Center, Jena, Germany; Friedrich Sch B
1572 Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models. 2025 NeurIPS University of Oxford; Polytechnique Montréal B
1573 FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models. 2025 ICLR School of Electronic and Computer Engineering, Peking University; Peking Univers B
1574 FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation. 2025 ACL Xiamen University; Hong Kong Polytechnic University; Migu Meland Co., Ltd; Sooch B
1575 Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations. 2025 ACL University of Manchester; Idiap Research Institute; University of Sheffield B
1576 Faithful Group Shapley Value. 2025 NeurIPS The Ohio State University; Carnegie Mellon University B
1577 Faithful Dynamic Imitation Learning from Human Intervention with Dynamic Regret Minimization. 2025 NeurIPS Southeast University; Nanjing University of Science and Technology B
1578 FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge. 2025 EMNLP Seoul National University; Adobe Research; LG AI Research B
1579 FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows". 2025 ICLR Salesforce AI Research; University of Texas at Austin B
1580 Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs 2025 C
1581 Fairness on Principal Stratum: A New Perspective on Counterfactual Fairness. 2025 ICML Peking University; Carnegie Mellon University; Sun Yatsen University; Renmin Uni B
1582 FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering. 2025 ACL Zhejiang University; Zhejiang Key Laboratory of Medical Imaging Artificial Intel B
1583 Factor Graph-based Interpretable Neural Networks. 2025 ICLR Dalian University of Technology, ^ 2 Jilin University, ^ 3 RMIT University B
1584 Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning. 2025 NeurIPS MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, USTC; B
1585 Fact Recall, Heuristics or Pure Guesswork? Precise Interpretations of Language Models for Fact Completion. 2025 ACL Chalmers University of Technology; University of Gothenburg; Linköping Universit B
1586 FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models. 2025 EMNLP Zhongguancun Laboratory; Renmin University of China; The Hong Kong University of B
1587 FaVe: Factored and Verified Search Rationale for Long-form Answer. 2025 ACL Seoul National University B
1588 FaCT: Faithful Concept Traces for Explaining Neural Network Decisions. 2025 NeurIPS Max Planck Institute for Informatics, Saarland Informatics Campus, Saarbrücken, B
1589 FLARE: Faithful Logic-Aided Reasoning and Exploration. 2025 EMNLP University of Copenhagen; University of Edinburgh; Miniml.AI; Cohere (Canada); N B
1590 F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. 2025 ACL MoE Key Lab of Artificial Intelligence, X-LANCE Lab, School of Computer Science; B
1591 F-Fidelity: A Robust Framework for Faithfulness Evaluation of Explainable AI. 2025 ICLR NEC Laboratories America, Princeton, United States; University of California, Sa B
1592 Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward Pass. 2025 EMNLP TU Dresden; ScaDS.AI Dresden/Leipzig, Germany B
1593 Exploring the Translation Mechanism of Large Language Models 2025 Harbin Institute of Technology (Shenzhen) / Pengcheng Laboratory C
1594 Exploring Explanations Improves the Robustness of In-Context Learning. 2025 ACL CyberAgent (Japan) B
1595 Explanations of GNN on Evolving Graphs via Axiomatic Layer edges. 2025 ICLR B
1596 Explaining, Fast and Slow: Abstraction and Refinement of Provable Explanations. 2025 ICML Hebrew University of Jerusalem B
1597 Explaining the role of Intrinsic Dimensionality in Adversarial Training. 2025 ICML Qatar Computing Research Institute, HBKU, Doha, Qatar; Dalhousie University B
1598 Explaining novel senses using definition generation with open language models. 2025 EMNLP University of Oslo; KU Leuven B
1599 Explaining Puzzle Solutions in Natural Language: An Exploratory Study on 6x6 Sudoku. 2025 ACL University of Colorado Boulder; University of Colorado System B
1600 Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation. 2025 ICLR The Blavatnik School of Computer Science, Tel Aviv University B
1601 Explaining Matters: Leveraging Definitions and Semantic Expansion for Sexism Detection. 2025 ACL University of Warwick; University of Leeds B
1602 Explaining Length Bias in LLM-Based Preference Evaluations. 2025 EMNLP University of Southern California B
1603 Explaining Differences Between Model Pairs in Natural Language through Sample Learning. 2025 EMNLP University of North Carolina at Chapel Hill B
1604 Explainably Safe Reinforcement Learning. 2025 NeurIPS Masaryk University; Graz University of Technology; Technical University of Munic B
1605 Explainable Text Classification with LLMs: Enhancing Performance through Dialectical Prompting and Explanation-Guided Training. 2025 EMNLP School of Computing and Artificial Intelligence; Southwestern University of Fina B
1606 Explainable Reinforcement Learning from Human Feedback to Improve Alignment. 2025 NeurIPS Department of Electrical Engineering, Pennsylvania State University; Department B
1607 Explainable Hallucination through Natural Language Inference Mapping. 2025 ACL University of Bonn; University of Sheffield; Lamarr Institute for Machine Learni B
1608 Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation. 2025 ACL King's College London; Southeast University; The Alan Turing Institute B
1609 Explainable Concept Generation through Vision-Language Preference Learning for Understanding Neural Networks' Internal Representations. 2025 ICML School of Computing and Augmented Intelligence, Arizona; State University, Tempe B
1610 Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics. 2025 EMNLP University of California San Diego; Adobe Research B
1611 Explainability and Interpretability of Multilingual Large Language Models: A Survey. 2025 EMNLP University of Cambridge B
1612 Explain then Rank: Scale Calibration of Neural Rankers Using Natural Language Explanations from LLMs. 2025 ACL Snowflake Inc. (United States); Dataminr (United States) B
1613 Explain Yourself, Briefly! Self-Explaining Neural Networks with Concise Sufficient Reasons. 2025 ICLR IBM Research; The Hebrew University of Jerusalem; Bar-Ilan University B
1614 ExpProof : Operationalizing Explanations for Confidential Models with ZKPs. 2025 ICML UC San Diego; Stanford University B
1615 Exogenous Isomorphism for Counterfactual Identifiability. 2025 ICML East China Normal University B
1616 Exact Computation of Any-Order Shapley Interactions for Graph Neural Networks. 2025 ICLR Universitat de València; University of Padua B
1617 ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation. 2025 ACL University of Massachusetts Amherst B
1618 ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning. 2025 NeurIPS University of Texas at Austin B
1619 Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language Models. 2025 ICML Shenzhen Campus of Sun Yat-Sen University, School of Cyber Science and Technolog B
1620 Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable? 2025 ICLR Laboratoire d'Informatique de Grenoble; GIPSA-Lab B
1621 Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition. 2025 ICML University of Michigan B
1622 Evaluation of Attribution Bias in Generator-Aware Retrieval-Augmented Large Language Models. 2025 ACL Leiden University; University of Strathclyde; eBay; University of Amsterdam B
1623 Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration. 2025 ACL Korea Advanced Institute of Science and Technology; KT Corporation; Sogang Unive B
1624 Evaluating Neuron Explanations: A Unified Framework with Sanity Checks. 2025 ICML CSE, UC San Diego, CA, USA; HDSI, UC San Diego B
1625 Evaluating Chain-of-Thought Monitorability
faithfulness chain-of-thought monitoring
2025 Lab post (OpenAI) OpenAI A
1626 Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis. 2025 ACL University of Chinese Academy of Sciences; Chinese Academy of Sciences; Tsinghua B
1627 Enhancing the Comprehensibility of Text Explanations via Unsupervised Concept Discovery. 2025 ACL Chinese Academy of Sciences; University of Chinese Academy of Sciences B
1628 Enhancing Uncertainty Estimation and Interpretability with Bayesian Non-negative Decision Layer. 2025 ICLR Xidian University; Xi'an Jiaotong University; Asia School of Business B
1629 Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering Perspective. 2025 ICML The University of Queensland; The University of Melbourne; Mohamed bin Zayed Uni B
1630 Enhancing Training Data Attribution with Representational Optimization. 2025 NeurIPS Carnegie Mellon University; University of Toronto; Vector Institute B
1631 Enhancing Recommendation Explanations through User-Centric Refinement. 2025 EMNLP Renmin University of China; Huawei B
1632 Enhancing Pre-trained Representation Classifiability can Boost its Interpretability. 2025 ICLR Key Lab of Intell. Info. Process., Inst. of Comput. Tech., CAS; University of Ch B
1633 Enhancing Performance of Explainable AI Models with Constrained Concept Refinement. 2025 ICML University of Michigan; Princeton University B
1634 Enhancing Interpretable Image Classification Through LLM Agents and Conditional Concept Bottleneck Models. 2025 ACL Monash University B
1635 Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategy. 2025 EMNLP Nara Institute of Science and Technology; University of the Philippines Diliman B
1636 Enhancing Graph Of Thought: Enhancing Prompts with LLM Rationales and Dynamic Temperature Control. 2025 ICLR B
1637 Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data. 2025 ICLR School of Computing, University of Georgia; Department of Radiology, Massachuset B
1638 Enhancing Chain-of-Thought Reasoning via Neuron Activation Differential Analysis. 2025 EMNLP Renmin University of China; iFLYTEK AI Research B
1639 Enhancing Automated Interpretability with Output-Centric Feature Descriptions. 2025 ACL Tel Aviv University; R Group BC
1640 Enhanced Noun-Noun Compound Interpretation through Textual Enrichment. 2025 EMNLP Brandeis University B
1641 EnSToM: Enhancing Dialogue Systems with Entropy-Scaled Steering Vectors for Topic Maintenance. 2025 ACL Pohang University of Science and Technology B
1642 Emergent Introspective Awareness in Large Language Models
activations concepts introspection
2025 Lab post (Anthropic) Anthropic AC
1643 Emergence and Evolution of Interpretable Concepts in Diffusion Models. 2025 NeurIPS University of Southern California B
1644 Eliminating Position Bias of Language Models: A Mechanistic Approach. 2025 ICLR Texas A&M University B
1645 Efficiently Verifiable Proofs of Data Attribution. 2025 NeurIPS Morgan Stanley Machine Learning Research and Harvard Business School; Google Res B
1646 Efficient and Generalizable Mixed-Precision Quantization via Topological Entropy 2025 Shanxi University & Northeastern University C
1647 Efficient and Accurate Explanation Estimation with Distribution Compression. 2025 ICLR University of Warsaw; Munich Center for Machine Learning; Warsaw University of T B
1648 Efficient Neuron Segmentation in Electron Microscopy by Affinity-Guided Queries. 2025 ICLR B
1649 Efficient Dictionary Learning with Switch Sparse Autoencoders. 2025 ICLR Massachusetts Institute of Technology; University of Oxford BC
1650 Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition. 2025 ICLR CSAIL, MIT; Center for Computational Biology B
1651 Effective and Efficient Time-Varying Counterfactual Prediction with State-Space Models. 2025 ICLR B
1652 Editable Concept Bottleneck Models. 2025 ICML King Abdullah University of Science and Technology; Shanghai Jiao Tong Universit B
1653 Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs. 2025 NeurIPS Key Lab of Intell. Info. Process., Inst. of Comput. Tech., CAS; University of Ch B
1654 EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations. 2025 ACL Seoul National University; Coxwave B
1655 ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations. 2025 ACL University of Southern California; Stanford University B
1656 EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification. 2025 NeurIPS King Abdullah University of Science and Technology; Harbin Institute of Technolo B
1657 E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel Time. 2025 ICML Dartmouth College B
1658 Dynamic Steering With Episodic Memory For Large Language Models. 2025 ACL Deakin University; Meta B
1659 Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models 2025 East China Normal University C
1660 Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models. 2025 ACL Ping An (China); University of Electronic Science and Technology of China B
1661 Dually Self-Improved Counterfactual Data Augmentation Using Large Language Model. 2025 ACL Beijing Institute of Technology; Harbin Institute of Technology B
1662 Dual-Path Counterfactual Integration for Multimodal Aspect-Based Sentiment Classification. 2025 EMNLP Chinese Academy of Sciences; China Mobile (China); ByteDance; GMT Technology (Sh B
1663 Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs. 2025 EMNLP Zhejiang University; Shanghai Innovation Institute; Southeast University; Shangh B
1664 Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic Inference. 2025 ACL King's College London B
1665 Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains. 2025 ACL Peking University B
1666 Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation. 2025 ACL Korea Advanced Institute of Science and Technology B
1667 Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models. 2025 EMNLP Hanyang University B
1668 DocVXQA: Context-Aware Visual Explanations for Document Question Answering. 2025 ICML Universitat Autònoma de Barcelona; UiT The Arctic University of Norway; Inria B
1669 Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing. 2025 EMNLP Munich Center for Machine Learning B
1670 Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations? 2025 ICLR Heidelberg University B
1671 Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs 2025 C
1672 Do Large Language Models Have "Emotion Neurons"? Investigating the Existence and Role. 2025 ACL Konkuk University B
1673 Do LLMs Behave as Claimed? Investigating How LLMs Follow Their Own Claims using Counterfactual Questions. 2025 EMNLP Harbin Institute of Technology B
1674 Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs. 2025 EMNLP Kyoto University B
1675 Dissecting Persona-Driven Reasoning in Language Models via Activation Patching. 2025 EMNLP Independent Researchers B
1676 Disentangling Superpositions: Interpretable Brain Encoding Model with Sparse Concept Atoms. 2025 NeurIPS University of California, Berkeley B
1677 Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition. 2025 NeurIPS Kyung Hee University; Korea University B
1678 Discursive Circuits: How Do Language Models Understand Discourse Relations? 2025 EMNLP National University of Singapore B
1679 Discovering Influential Neuron Path in Vision Transformers. 2025 ICLR ShanghaiTech University, ^2Tencent PCG B
1680 Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing. 2025 ACL Language Science (South Korea); University of Edinburgh B
1681 Diffusion Counterfactual Generation with Semantic Abduction. 2025 ICML Imperial College London B
1682 Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion Models. 2025 ICLR The University of Sydney, ^ 2 City University of Hong Kong B
1683 Differentially Private Steering for Large Language Model Alignment. 2025 ICLR Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science an B
1684 Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework. 2025 ACL Tencent; Tsinghua University; Chinese Academy of Sciences; Huazhong University o B
1685 Detecting Misbehavior in Frontier Reasoning Models
chain-of-thought monitoring reward-hacking
2025 Lab post (OpenAI) OpenAI A
1686 Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation. 2025 EMNLP Massachusetts Institute of Technology; Nanyang Technological University B
1687 Dense SAE Latents Are Features, Not Bugs. 2025 NeurIPS ETH Zürich; University of Sheffield B
1688 Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling. 2025 NeurIPS University of Electronic Science and Technology of China; The Chinese University B
1689 Dementia Through Different Eyes: Explainable Modeling of Human and LLM Perceptions for Early Awareness. 2025 EMNLP Decision Sciences International Corporation (United States); University of Washi B
1690 DeepSeek-V3.2 and DeepSeekMath-V2 2025 DeepSeek C
1691 DeepLayout: Learning Neural Representations of Circuit Placement Layout. 2025 ICML Peking University; Business Innovation Centre (Czechia); Wuhan University; Chine B
1692 DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale. 2025 ICLR The Chinese University of Hong Kong; Shanghai Jiao Tong University B
1693 Deep Linear Probe Generators for Weight Space Learning. 2025 ICLR School of Computer Science and Engineering; The Hebrew University of Jerusalem, B
1694 Deep Learning Alternatives Of The Kolmogorov Superposition Theorem. 2025 ICLR University of Pennsylvania B
1695 Deep Bayesian Filter for Bayes-Faithful Data Assimilation. 2025 ICML Preferred Networks (Japan) B
1696 Decoupling Memories, Muting Neurons: Towards Practical Machine Unlearning for Large Language Models. 2025 ACL Huazhong University of Science and Technology B
1697 Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains. 2025 ACL Chinese University of Hong Kong; Massachusetts Institute of Technology B
1698 Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis. 2025 ACL The Hong Kong University of Science and Technology (Guangzhou); Hong Kong Univer B
1699 Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval. 2025 EMNLP Sungkyunkwan University B
1700 Decision Information Meets Large Language Models: The Future of Explainable Operations Research. 2025 ICLR Department of Computer Science, City University of Hong Kong; Huawei Noah’s Ark B
1701 DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction. 2025 ACL Graduate School of Artificial Intelligence, POSTECH, Republic of Korea; Republic B
1702 Data-centric Prediction Explanation via Kernelized Stein Discrepancy. 2025 ICLR Dalhousie University B
1703 Data Shapley in One Training Run. 2025 ICLR Princeton University B
1704 DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models. 2025 EMNLP Fudan University; Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Ton B
1705 DO I KNOW THIS ENTITY? KNOWLEDGE AWARENESS AND HALLUCINATIONS IN LANGUAGE MODELS 2025 Universitat Politècnica de Catalunya C
1706 DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models. 2025 NeurIPS University of Catania; University of Central Florida B
1707 DCBM: Data-Efficient Visual Concept Bottleneck Models. 2025 ICML University of Mannheim; Clausthal University of Technology; Max Planck Institute B
1708 DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models. 2025 NeurIPS Carnegie Mellon University; University of Michigan B
1709 DAPO: An Open-Source LLM Reinforcement Learning System at Scale 2025 Tsinghua University / ByteDance Seed C
1710 DAPI: Domain Adaptive Toxicity Probe Vector Intervention, for Fine-Grained Detoxification. 2025 ACL Sungkyunkwan University B
1711 D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models 2025 The Ohio State University / USTC C
1712 Cyclic Counterfactuals under Shift-Scale Interventions. 2025 NeurIPS Indian Statistical Institute B
1713 Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models. 2025 ACL Mohamed bin Zayed University of Artificial Intelligence; University of Notre Dam B
1714 Cross-Lingual Generalization and Compression: From Language-Specific to Shared Neurons. 2025 ACL Heidelberg University; Association for Computational Linguistics B
1715 Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction. 2025 EMNLP Key Laboratory of Aerospace Information Security and Trusted Computing; Ministry B
1716 Cracking Factual Knowledge: A Comprehensive Analysis of Degenerate Knowledge Neurons in Large Language Models. 2025 ACL Chinese Academy of Sciences B
1717 Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models. 2025 ACL Artificial Intelligence in Medicine (Canada) B
1718 Counterfactual reasoning: an analysis of in-context emergence. 2025 NeurIPS Max Planck Institute for Intelligent Systems; ETH Zurich; University of Cambridg B
1719 Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation. 2025 ICML University of Illinois Chicago; University of Michigan; LG AI Research (South Ko B
1720 Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models. 2025 NeurIPS Renmin University of China; Microsoft Research Asia B
1721 Counterfactual Realizability. 2025 ICLR Causal Artificial Intelligence Lab; Columbia University B
1722 Counterfactual Identifiability via Dynamic Optimal Transport. 2025 NeurIPS Imperial College London, UK B
1723 Counterfactual Graphical Models: Constraints and Inference. 2025 ICML Universidad Autonoma de Manizales B
1724 Counterfactual Generative Modeling with Variational Causal Inference. 2025 ICLR University of California, Berkeley B
1725 Counterfactual Evolution of Multimodal Datasets via Visual Programming. 2025 NeurIPS Zhejiang University; National University of Singapore; Nanyang Technological Uni B
1726 Counterfactual Effect Decomposition in Multi-Agent Sequential Decision Making. 2025 ICML Max Planck Institute for Software Systems, Germany B
1727 Counterfactual Contrastive Learning with Normalizing Flows for Robust Treatment Effect Estimation. 2025 ICML Shanxi University; Agency for Science, Technology and Research B
1728 Counterfactual Concept Bottleneck Models. 2025 ICLR Università della Svizzera italiana; IBM Research; Work conducted while employed B
1729 Correcting on Graph: Faithful Semantic Parsing over Knowledge Graphs with Large Language Models. 2025 ACL Huazhong University of Science and Technology B
1730 CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models. 2025 ICML Duke University B
1731 Contrastive Prompting Enhances Sentence Embeddings in LLMs through Inference-Time Steering. 2025 ACL Nanjing University B
1732 Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models. 2025 EMNLP Xi'an Jiaotong University B
1733 Continued Pretraining and Interpretability-Based Evaluation for Low-Resource Languages: A Galician Case Study. 2025 ACL Citrix (Switzerland) B
1734 Context-DPO: Aligning Language Models for Context-Faithfulness. 2025 ACL University of Chinese Academy of Sciences; Microsoft (Finland); University of Ca B
1735 Context Steering: Controllable Personalization at Inference Time. 2025 ICLR UC Berkeley BC
1736 Context Copying Modulation: The Role of Entropy Neurons in Managing Parametric and Contextual Knowledge Conflicts. 2025 EMNLP Sorbonne Université; Institut Systèmes Intelligents et de Robotique; BNP Paribas B
1737 Constrain Alignment with Sparse Autoencoders. 2025 ICML King's College London B
1738 Constitutional AI: Harmlessness from AI Feedback 2025 Anthropic C
1739 Connectome Mapping: Shape-Memory Network via Interpretation of Contextual Semantic Information. 2025 ICLR B
1740 ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning. 2025 ICLR University of Edinburgh, ^2Samsung AI Research Centre, Cambridge B
1741 ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features. 2025 ICML Virginia Tech; Georgia Institute of Technology B
1742 Concept-Centric Token Interpretation for Vector-Quantized Generative Models. 2025 ICML University of Georgia; New Jersey Institute of Technology; New York University B
1743 Concept Bottleneck Large Language Models. 2025 ICLR University of California San Diego B
1744 Concept Bottleneck Language Models For Protein Design. 2025 ICLR University of California San Diego; Guide Labs; Department of Computer Science, B
1745 ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability. 2025 ACL Université Toulouse III - Paul Sabatier; Université Toulouse-I-Capitole; IRT M2P B
1746 Computing Circuits Optimization via Model-Based Circuit Genetic Evolution. 2025 ICLR B
1747 Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders. 2025 ICML Australian National University; Nazarbayev University; Cold Spring Harbor Labora B
1748 Compositional Generalisation for Explainable Hate Speech Detection. 2025 EMNLP University of Edinburgh; Cohere (Canada) B
1749 Colloquial Singaporean English Style Transfer with Fine-Grained Explainable Control. 2025 ACL Singapore Management University; DSO National Laboratories; Australian National B
1750 Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs 2025 C
1751 CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models. 2025 ACL The Hong Kong University of Science and Technology (Guangzhou); Tencent; Beijing B
1752 CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models. 2025 ACL The University of Sydney B
1753 CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark. 2025 ICLR Tokyo Institute of Technology; School of Engineering Westlake Univeristy B
1754 CoDy: Counterfactual Explainers for Dynamic Graphs. 2025 ICML Karlsruhe Institute of Technology; University of Amsterdam B
1755 CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis. 2025 ACL Shenzhen Research Institute of Big Data; Chinese University of Hong Kong, Shenzh B
1756 CoCoA: A Minimum Bayes Risk Framework Bridging Confidence and Consistency for Uncertainty Quantification in LLMs 2025 MBZUAI C
1757 CiteEval: Principle-Driven Citation Evaluation for Source Attribution. 2025 ACL AWS AI Labs; Orby AI; Google B
1758 Circuits Updates (monthly series, Jan-Dec 2025)
mechanistic SAE transcoder circuits faithfulness persona
2025 Lab post (Anthropic) Anthropic A
1759 CircuitFusion: Multimodal Circuit Representation Learning for Agile Chip Design. 2025 ICLR The Hong Kong University of Science and Technology B
1760 Circuit Transformer: A Transformer That Preserves Logical Equivalence. 2025 ICLR University College London B
1761 Circuit Tracing: Revealing Computational Graphs in Language Models
transcoder circuits attribution-graph attribution interpretability
2025 Lab post (Anthropic) Anthropic AC
1762 Circuit Stability Characterizes Language Model Generalization. 2025 ACL Carnegie Mellon University B
1763 Circuit Representation Learning with Masked Gate Modeling and Verilog-AIG Alignment. 2025 ICLR The Chinese University of Hong Kong; Shanghai Artificial Intelligence Laboratory B
1764 Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models. 2025 ACL EU Business School, Munich; Munich Center for Machine Learning; University of Os B
1765 Circuit Complexity Bounds for RoPE-based Transformer Architecture. 2025 EMNLP Middle Tennessee State University; Stevens Institute of Technology; The Universi B
1766 CheXalign: Preference fine-tuning in chest X-ray interpretation models without human feedback. 2025 ACL University of Oxford; Stanford University B
1767 ChatGPT and the art of post-training 2025 Former OpenAI researchers C
1768 ChartLens: Fine-grained Visual Attribution in Charts. 2025 ACL Adobe Systems (United States); University of Maryland, College Park B
1769 Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models. 2025 ICLR Department of Computer Science, Northwestern University, Evanston, IL 60208, USA B
1770 Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
chain-of-thought monitoring safety/alignment
2025 Lab post (Joint (OpenAI + DeepMind + Anthropic + others)) Joint (OpenAI + DeepMind + Anthropic + others) AC
1771 Certifying Counterfactual Bias in LLMs. 2025 ICLR UIUC, ^2 Amazon, ^3 Oracle Health B
1772 Causally Reliable Concept Bottleneck Models. 2025 NeurIPS Università della Svizzera Italiana; University of Liechtenstein; IBM Research; S B
1773 Causality Meets the Table: Debiasing LLMs for Faithful TableQA via Front-Door Intervention. 2025 NeurIPS Anhui University B
1774 Causal Logistic Bandits with Counterfactual Fairness Constraints. 2025 ICML Iowa State University; Mohamed bin Zayed University of Artificial Intelligence B
1775 Causal Explanation-Guided Learning for Organ Allocation. 2025 NeurIPS Vrije Universiteit Brussel; King's College London B
1776 Causal Attribution Analysis for Continuous Outcomes. 2025 ICML Beijing Technology and Business University; Alibaba Group (China) B
1777 Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework. 2025 NeurIPS Technische Universität Berlin, Germany; UMI Lab, ATB Potsdam, Germany; Fraunhofe B
1778 CapeX: Category-Agnostic Pose Estimation from Textual Point Explanation. 2025 ICLR Tel Aviv University B
1779 Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs. 2025 EMNLP Osnabrück University B
1780 Can LLMs Explain Themselves Counterfactually? 2025 EMNLP Ruhr University Bochum; University Alliance Ruhr Research Center for Trustworthy B
1781 Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs. 2025 ACL Southeast University B
1782 Can Input Attributions Explain Inductive Reasoning in In-Context Learning? 2025 ACL Tohoku University; Mohamed bin Zayed University of Artificial Intelligence; RIKE B
1783 Calibrating LLM Confidence by Probing Perturbed Representation Stability. 2025 EMNLP Michigan State University; Independent Researcher; JPMorgan AI Research; Henry F B
1784 CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners. 2025 EMNLP National University of Singapore; University of California, Los Angeles B
1785 CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic. 2025 NeurIPS College of Computer Science and Technology, Zhejiang University, China; Research B
1786 COSDA: Counterfactual-based Susceptibility Risk Framework for Open-Set Domain Adaptation. 2025 ICML Ocean University of China; National University of Defense Technology; University B
1787 CONDA: Adaptive Concept Bottleneck for Foundation Models Under Distribution Shifts. 2025 ICLR University of Wisconsin-Madison, Madison, USA B
1788 COMRECGC: Global Graph Counterfactual Explainer through Common Recourse. 2025 ICML University of Illinois Chicago B
1789 CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection. 2025 ACL Yunnan University; National University of Singapore B
1790 CLIX: Cross-Lingual Explanations of Idiomatic Expressions. 2025 ACL University of Colorado Boulder B
1791 CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction. 2025 ACL Tsinghua University; Huazhong University of Science and Technology; ByteDance; P B
1792 CEAES: Bidirectional Reinforcement Learning Optimization for Consistent and Explainable Essay Assessment. 2025 ACL Guangdong University of Foreign Studies B
1793 CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environments. 2025 EMNLP École Polytechnique Fédérale de Lausanne B
1794 CAIR: Counterfactual-based Agent Influence Ranker for Agentic AI Workflows. 2025 EMNLP Fujitsu (United Kingdom) B
1795 Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency. 2025 EMNLP Abu Dhabi University; Abu Dhabi Health Services B
1796 Bridging Relevance and Reasoning: Rationale Distillation in Retrieval-Augmented Generation. 2025 ACL City University of Hong Kong; University of Science and Technology of China; Hua B
1797 Bridging Brains and Concepts: Interpretable Visual Decoding from fMRI with Semantic Bottlenecks. 2025 NeurIPS Department of Biomedicine and Prevention; University of Rome Tor Vergata; Viale B
1798 Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization. 2025 ICLR School of Computer Science and Technology, HUST; Faculty of Artificial Intellige B
1799 Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders. 2025 EMNLP Siebel School of Computing and Data Science; University of Illinois Urbana-Champ B
1800 Bounds on the computational complexity of neurons due to dendritic morphology. 2025 NeurIPS University of Washington; Allen Institute B
1801 BottleHumor: Self-Informed Humor Explanation using the Information Bottleneck Principle. 2025 ACL University of British Columbia; State Research Center of Virology and Biotechnol B
1802 Boosting the visual interpretability of CLIP via adversarial fine-tuning. 2025 ICLR B
1803 Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation. 2025 ACL University of Chinese Academy of Sciences; Chinese Academy of Sciences; Departme B
1804 Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing. 2025 NeurIPS Department of Electronic Engineering, Tsinghua University; School of Computer Sc B
1805 BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments 2025 Fudan University / Shanghai AI Lab C
1806 Bilinear MLPs enable weight-based mechanistic interpretability. 2025 ICLR University of Antwerp; University of Antwerp, sqIRL/IDLab; Apollo Research B
1807 Beyond single neurons: population response geometry in digital twins of mouse visual cortex. 2025 ICLR B
1808 Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels. 2025 EMNLP University of Washington; Université Paris Cité B
1809 Beyond Topological Self-Explainable GNNs: A Formal Explainability Perspective. 2025 ICML University of Trento B
1810 Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing. 2025 EMNLP Peking University B
1811 Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms. 2025 ACL Zhejiang University; Tencent; National University of Singapore B
1812 Beyond Linear Steering: Unified Multi-Attribute Control for Language Models. 2025 EMNLP University of Oxford B
1813 Beyond Last-Click: An Optimal Mechanism for Ad Attribution. 2025 NeurIPS Gaoling School of Artificial Intelligence; Renmin University of China; School of B
1814 Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness. 2025 ICLR Peking University; MIT CSAIL; New York University; MIT EECS, CSAIL B
1815 Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders. 2025 EMNLP Northwestern University; University of Georgia; New Jersey Institute of Technolo B
1816 Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence. 2025 ICML The University of Tokyo B
1817 Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits. 2025 NeurIPS Indian Institute of Technology Kanpur (IIT Kanpur) B
1818 Beyond Circuit Connections: A Non-Message Passing Graph Transformer Approach for Quantum Error Mitigation. 2025 ICLR B
1819 Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases. 2025 ACL Center for Data Science; Department of Linguistics; New York University B
1820 Better Training Data Attribution via Better Inverse Hessian-Vector Products. 2025 NeurIPS University of Toronto; Vector Institute for Artificial Intelligence; Tübingen AI B
1821 Beneath the Facade: Probing Safety Vulnerabilities in LLMs via Auto-Generated Jailbreak Prompts. 2025 EMNLP School of Computing, KAIST; Graduate School of Data Science, KAIST B
1822 Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks. 2025 EMNLP Korea University; Soongsil University B
1823 Bayesian Concept Bottleneck Models with LLM Priors. 2025 NeurIPS University of California, San Francisco; Microsoft Research; National University B
1824 BEDAA: Bayesian Enhanced DeBERTa for Uncertainty-Aware Authorship Attribution. 2025 ACL University of Manchester; Imperial College London B
1825 BANMIME : Misogyny Detection with Metaphor Explanation on Bangla Memes. 2025 EMNLP Chittagong University of Engineering & Technology; Dhaka International Universit B
1826 AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders. 2025 ICML Stanford University B
1827 Automating Steering for Safe Multimodal Large Language Models. 2025 EMNLP Zhejiang University; National University of Singapore B
1828 Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation. 2025 ACL King University; Peking University B
1829 Automatically Identifying Local and Global Circuits with Linear Computation Graphs 2025 Anthropic C
1830 AutoCT: Automating Interpretable Clinical Trial Prediction with LLM Agents. 2025 EMNLP University of Pennsylvania; Massachusetts Institute of Technology; Oracle; Santa B
1831 Auditing Language Models for Hidden Objectives
SAE interpretability safety/alignment
2025 Lab post (Anthropic) Anthropic A
1832 Attribution and Application of Multiple Neurons in Multimodal Large Language Models. 2025 EMNLP Beijing Language and Culture University B
1833 AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution. 2025 ICLR University of Toronto, Vector Institute B
1834 Attention with Dependency Parsing Augmentation for Fine-Grained Attribution. 2025 ACL Chinese Academy of Sciences; Institute of Computing Technology, CAS, Beijing 100 B
1835 Attention Consistency for LLMs Explanation. 2025 EMNLP HB Studio; INALCO; Sorbonne Université; University of Washington; VitaSight B
1836 Associative memory and dead neurons. 2025 ICLR AIRI; Skolkovo Institute of Science and Technology BC
1837 Artificial Kuramoto Oscillatory Neurons. 2025 ICLR Bernstein Center for Computational Neuroscience Tübingen; Tübingen AI Center B
1838 Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) 2025 OpenAI C
1839 Artifcial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) 2025 University of Washington C
1840 Around the World in 24 Hours: Probing LLM Knowledge of Time and Place. 2025 ACL Universität Hamburg; Bocconi University B
1841 Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. 2025 ICML Massachusetts Institute of Technology B
1842 Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation. 2025 EMNLP Shanghai Maritime University; École Polytechnique Fédérale de Lausanne; Xi’an Ji B
1843 Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice. 2025 ACL Università della Svizzera italiana; National Institute of Informatics; Brown Uni B
1844 Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models. 2025 ICML Harvard University; Kempner Institute B
1845 ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation [Technical Report] 2025 CUHK (Shenzhen) C
1846 Angular Steering: Behavior Control via Rotation in Activation Space. 2025 NeurIPS National University of Singapore B
1847 Analyze Feature Flow to Enhance Interpretation and Steering in Language Models. 2025 ICML T-Tech; FXM Research B
1848 Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation? 2025 EMNLP University of Sheffield B
1849 AnalogGenie: A Generative Engine for Automatic Discovery of Analog Circuit Topologies. 2025 ICLR Northeastern University, ^ 2; The George Washington University B
1850 AnalogGenie-Lite: Enhancing Scalability and Precision in Circuit Topology Discovery through Lightweight Graph Modeling. 2025 ICML Department of Electrical and Computer Engineering, Northeastern University, Bost B
1851 An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks. 2025 ICML Max Planck Institute for Intelligent Systems; ELLIS Institute Tübingen B
1852 An Approach to Technical AGI Safety and Security
monitoring interpretability safety/alignment
2025 Lab post (Google DeepMind) Google DeepMind A
1853 All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning 2025 CMU / Cornell University C
1854 Aligned at the Start: Conceptual Groupings in LLM Embeddings 2025 Virginia Tech C
1855 Africa Health Check: Probing Cultural Bias in Medical LLMs. 2025 EMNLP Georgia Institute of Technology; Google (United States) B
1856 Adversarial Attacks on Data Attribution. 2025 ICLR University of Michigan; University of Illinois Urbana-Champaign B
1857 Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization. 2025 ICML Universit´e de Montreal 2 Mila - Qu´ebec AI; Institute 3 HEC Montr´eal 4 Univers B
1858 Adaptive Transformer Programs: Bridging the Gap Between Performance and Interpretability in Transformers. 2025 ICLR B
1859 Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimates. 2025 EMNLP Northeastern University B
1860 Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search. 2025 NeurIPS Mohamed bin Zayed University of Artificial Intelligence (MBZUAI); University of B
1861 AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness. 2025 ACL Beijing University of Posts and Telecommunications; Hong Kong Baptist University B
1862 Active feature acquisition via explainability-driven ranking. 2025 ICML Boston University; Southern Illinois University School of Medicine; Worcester Po B
1863 Activation Steering Decoding: Mitigating Hallucination in Large Vision-Language Models through Bidirectional Hidden State Intervention. 2025 ACL Hong Kong Polytechnic University; State Key Laboratory of Pattern Recognition; S B
1864 Accelerating Training with Neuron Interaction and Nowcasting Networks. 2025 ICLR Samsung – SAIT AI Lab, Montreal; Concordia University; Université de Montréal; M B
1865 Abstract Counterfactuals for Language Model Agents. 2025 NeurIPS King’s College in London B
1866 AUTOCIRCUIT-RL: Reinforcement Learning-Driven LLM for Automated Circuit Topology Generation. 2025 ICML IBM Almaden Research Center; IBM Research - Thomas J. Watson Research Center B
1867 AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text. 2025 ICLR ♡University of Washington; ♠Allen Institute for Artificial Intelligence B
1868 ADIFF: Explaining audio difference using natural language. 2025 ICLR Carnegie Mellon University B
1869 A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders. 2025 NeurIPS LASR Labs; University College London; Tübingen AI Center, University of Tübingen B
1870 A data and task-constrained mechanistic model of the mouse outer retina shows robustness to contrast variations. 2025 NeurIPS University of Tübingen; University of Washington; Institute for Ophthalmic Resea B
1871 A Versatile Influence Function for Data Attribution with Non-Decomposable Loss. 2025 ICML University of Illinois Urbana-Champaign; Carnegie Mellon University B
1872 A Unified Framework for Provably Efficient Algorithms to Estimate Shapley Values. 2025 NeurIPS Global Technology Applied Research, JPMorganChase, New York, NY 10001, USA B
1873 A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models. 2025 EMNLP Northwestern University; University of Georgia; New Jersey Institute of Technolo B
1874 A Survey on Large Language Models with Multilingualism 2025 Beijing Jiaotong University + Université de Montréal C
1875 A Survey of Context Engineering for Large Language Models 2025 Chinese Academy of Sciences C
1876 A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMs. 2025 EMNLP Nvidia (United Kingdom); University of Groningen B
1877 A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression 2025 Renmin University of China C
1878 A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI. 2025 ACL EU Business School, Munich; Munich Center for Machine Learning; University of Ca B
1879 A Quantum Circuit-Based Compression Perspective for Parameter-Efficient Learning. 2025 ICLR Graduate Institute of Applied Physics, National Taiwan University, Taipei, Taiwa B
1880 A New Approach to Backtracking Counterfactual Explanations: A Unified Causal Framework for Efficient Model Interpretability. 2025 ICML Technical University of Munich; Munich Center for Machine Learning; École Polyte B
1881 A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations. 2025 EMNLP University of Maryland B
1882 A Lens into Interpretable Transformer Mistakes via Semantic Dependency. 2025 ICML The University of Sydney; Hong Kong Baptist University B
1883 A Hierarchy of Graphical Models for Counterfactual Inferences. 2025 NeurIPS Causal Artificial Intelligence Lab; Columbia University B
1884 A General Framework for Producing Interpretable Semantic Text Embeddings. 2025 ICLR National University of Singapore; Harbin Institute of Technology (Shenzhen) B
1885 A General Framework for Inference-time Scaling and Steering of Diffusion Models. 2025 ICML Department of Computer Science, New York University; Columbia University; Center B
1886 A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability. 2025 ACL King University B
1887 A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models. 2025 EMNLP Department of Computer Science; University of Maryland, College Park B
1888 A Causal Lens for Evaluating Faithfulness Metrics. 2025 EMNLP University of North Carolina at Chapel Hill; University of North Carolina Health B
1889 "I've Decided to Leak": Probing Internals Behind Prompt Leakage Intents. 2025 EMNLP Tsinghua University; Ant Group (China) B
1890 A Survey of Multilingual & Factual Recall Research 2024 C
1891 xTower: A Multilingual LLM for Explaining and Correcting Translation Errors. 2024 EMNLP Instituto de Telecomunicações; Instituto Superior Técnico, Universidade de Lisbo B
1892 xMIL: Insightful Explanations for Multiple Instance Learning in Histopathology. 2024 NeurIPS Berlin Institute for the Foundations of Learning and Data; Technische Universitä B
1893 shapiq: Shapley Interactions for Machine Learning. 2024 NeurIPS LMU Munich; University of Warsaw; Munich Center for Machine Learning; Warsaw Uni B
1894 dattri: A Library for Efficient Data Attribution. 2024 NeurIPS University of Illinois Urbana-Champaign B
1895 Yuan 2.0-M32: Mixture of Experts with Attention Router 2024 C
1896 XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs. 2024 EMNLP University of California, Santa Barbara; Nanyang Technological University B
1897 Xmodel-1.5: an 1B-Scale Multilingual LLM. Authors: Wang Qun, Liu Yang, Lin Qingquan, Jiang Ling, from XiaoduoAI. Summary: Xmodel-1.5 is a new 1B-parameter multilingual large model developed by the AI Lab of Xiaoduo Technology, pretrained on roughly 2 trillion tokens. The model shows strong performance across many languages, standing out especially on Thai, Arabic and French, while also performing well on Chinese and English. The authors additionally release a Thai evaluation dataset containing hundreds of questions annotated by students of the Integrated Innovation program at Chulalongkorn University, providing a valuable resource for future Thai NLP research. Although the results are encouraging, the authors acknowledge substantial room for improvement. They hope this work advances multilingual AI research and promotes better cross-lingual understanding across a range of NLP tasks. The model and code are publicly available on GitHub. 2024 C
1898 XRec: Large Language Models for Explainable Recommendation. 2024 EMNLP University of Hong Kong B
1899 XDetox: Text Detoxification with Token-Level Toxicity Explanations. 2024 EMNLP Hanyang University B
1900 X-ACE: Explainable and Multi-factor Audio Captioning Evaluation. 2024 ACL University of Science and Technology of China; Inspired Spine; University of Cal B
1901 What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation. 2024 ICML University College London; Google DeepMind (United Kingdom) B
1902 What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models. 2024 EMNLP Integrated Vision and Language Lab , KAIST B
1903 What does the Knowledge Neuron Thesis Have to do with Knowledge? 2024 ICLR University of Toronto, ^2University of Waterloo, ^3Stevens Institute of Technolo B
1904 What Would Gauss Say About Representations? Probing Pretrained Image Models using Synthetic Gaussian Benchmarks. 2024 ICML Massachusetts Institute of Technology; IBM Research B
1905 What Matters in Memorizing and Recalling Facts? Multifaceted Benchmarks for Knowledge Probing in Language Models. 2024 EMNLP Advanced Institute of Industrial Technology; The University of Tokyo B
1906 What Makes and Breaks Safety Fine-tuning? A Mechanistic Study. 2024 NeurIPS Five AI; University of Michigan; Harvard University; University of Oxford; Max P B
1907 What Does Parameter-free Probing Really Uncover? 2024 ACL University of Helsinki B
1908 What Do Language Models Hear? Probing for Auditory Representations in Language Models. 2024 ACL Massachusetts Institute of Technology B
1909 WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models. 2024 NeurIPS Michigan State University; IBM Research B
1910 Visual Pinwheel Centers Act as Geometric Saliency Detectors. 2024 NeurIPS Fudan University; Frontiers Center for Brain Science of the Ministry of Educatio B
1911 Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving. 2024 EMNLP University of Manchester; Idiap Research Institute B
1912 VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance. 2024 NeurIPS UC San Diego B
1913 VISIT: Visualizing and Interpreting the Semantic Information Flow of Transformers 2024 C
1914 VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation. 2024 ACL University of Waterloo; IN.AI Research♡; oiled science; content machine B
1915 VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models. 2024 ACL University of California, Los Angeles B
1916 Utilizing Human Behavior Modeling to Manipulate Explanations in AI-Assisted Decision Making: The Good, the Bad, and the Scary. 2024 NeurIPS Department of Computer Science; Purdue University B
1917 Using Natural Language Explanations to Improve Robustness of In-context Learning. 2024 ACL Weco AI; University College London; University of Edinburgh B
1918 Unveiling Factual Recall Behaviors of Large Language Models through Knowledge Neurons. 2024 EMNLP State Key Laboratory of Multimodal Artificial Intelligence Systems; Chinese Acad B
1919 Unsupervised Distractor Generation via Large Language Model Distilling and Counterfactual Contrastive Decoding. 2024 ACL King University; Peking University B
1920 Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models. 2024 EMNLP Institute for Complex Systems; Chinese Academy of Sciences; University of Chines B
1921 Unified Lexical Representation for Interpretable Visual-Language Alignment. 2024 NeurIPS Amazon Web Services; Fudan University B
1922 Unelicitable Backdoors via Cryptographic Transformer Circuits. 2024 NeurIPS Contramont Research; Institute of Mathematics and Computer Science, University o B
1923 Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective. 2024 NeurIPS The University of Tokyo B
1924 Understanding Faithfulness and Reasoning of Large Language Models on Plain Biomedical Summaries. 2024 EMNLP Data61 B
1925 Understanding Emergent Abilities of Language Models from the Loss Perspective 2024 Zhipu AI C
1926 Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense. 2024 NeurIPS Hong Kong University of Science and Technology; Pennsylvania State University B
1927 Unconditional stability of a recurrent neural circuit implementing divisive normalization. 2024 NeurIPS Courant Institute of Mathematical Sciences; Centre for Nano and Soft Matter Scie B
1928 UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models. 2024 ICLR Department of Artificial Intelligence; Sungkyunkwan University; Republic of Kore B
1929 Trust Regions for Explanations via Black-Box Probabilistic Certification. 2024 ICML IBM Research, Yorktown Heights, NY, USA; University of Tübingen B
1930 Tree-of-Counterfactual Prompting for Zero-Shot Stance Detection. 2024 ACL The University of Texas at Dallas; Human Language Technology Research Institute B
1931 Transferable and Efficient Non-Factual Content Detection via Probe Training with Offline Consistency Checking. 2024 ACL Renmin University of China; Tsinghua University B
1932 Transcoders find interpretable LLM feature circuits. 2024 NeurIPS Yale University; Columbia University B
1933 Training for Stable Explanation for Free. 2024 NeurIPS Harbin Institute of Technology; Key Laboratory of Trustworthy Distributed Comput B
1934 Training Data Attribution via Approximate Unrolling. 2024 NeurIPS University of Toronto; Vector Institute; NVIDIA B
1935 Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech. 2024 ACL Indian Institute of Technology Delhi; Association for Computational Linguistics B
1936 Towards a Greek Proverb Atlas: Computational Spatial Exploration and Attribution of Greek Proverbs. 2024 EMNLP Athens University of Economics and Business; Athena Research and Innovation Cent B
1937 Towards Verifiable Generation: A Benchmark for Knowledge-aware Language Model Attribution. 2024 ACL Nanyang Technological University; Fudan University; University of California, Sa B
1938 Towards Robust Fidelity for Evaluating Explainability of Graph Neural Networks. 2024 ICLR College Information Sciences and Technology, The Pennsylvania State University, B
1939 Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights. 2024 EMNLP Monash University B
1940 Towards Next-Generation Logic Synthesis: A Scalable Neural Circuit Generation Framework. 2024 NeurIPS Inspired Spine; University of Science and Technology of China; Noah’s Ark Lab, H B
1941 Towards Neuron Attributions in Multi-Modal Large Language Models. 2024 NeurIPS University of Science and Technology of China; National University of Singapore; B
1942 Towards Multi-dimensional Explanation Alignment for Medical Classification. 2024 NeurIPS Provable Responsible AI and Data Analytics (PRADA) Lab; King Abdullah University B
1943 Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models. 2024 EMNLP Apart Research; University of Oxford B
1944 Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation. 2024 ICML King Abdullah University of Science and Technology; Harbin Institute of Technolo B
1945 Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering. 2024 ACL University of Zurich; ETH Zurich; University of Regensburg; Swiss Finance Instit B
1946 Towards Faithful XAI Evaluation via Generalization-Limited Backdoor Watermark. 2024 ICLR Tsinghua University, Department of Mechanical Engineering, Beijing, China B
1947 Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering. 2024 EMNLP Harbin Institute of Technology; Queen Mary University of London; XtalPi (China) B
1948 Towards Faithful Explanations: Boosting Rationalization with Shortcuts Discovery. 2024 ICLR : State Key Laboratory of Cognitive Intelligence, University of Science and Tech B
1949 Towards Explainable Computerized Adaptive Testing with Large Language Model. 2024 EMNLP University of Science and Technology of China; Institute of Artificial Intellige B
1950 Towards Explainable Chinese Native Learner Essay Fluency Assessment: Dataset, Tasks, and Method. 2024 EMNLP East China Normal University; Microsoft Research (India) B
1951 Towards Characterizing Domain Counterfactuals for Invertible Latent Causal Models. 2024 ICLR Elmore Family School of Electrical and Computer Engineering; Purdue University B
1952 Towards Best Practices of Activation Patching in Language Models: Metrics and Methods. 2024 ICLR Google DeepMind B
1953 Towards Artwork Explanation in Large-scale Vision Language Models. 2024 ACL Nara Institute of Science and Technology; Google; EleutherAI; University of Cali B
1954 Towards 3D Molecule-Text Interpretation in Language Models. 2024 ICLR niversity of Science and Technology of China; ational University of Singapore; o B
1955 TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving Scenes. 2024 NeurIPS Institute of Computing Technology; Chinese Academy of Sciences; University of Ch B
1956 TimeX++: Learning Time-Series Explanations with Information Bottleneck. 2024 ICML Nanjing University; Microsoft Research (India); Pennsylvania State University; F B
1957 Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability. 2024 EMNLP Peking University B
1958 The motion planning neural circuit in goal-directed navigation as Lie group operator search. 2024 NeurIPS Center for Life Sciences; McGovern Institute for Brain Research; Peking Universi B
1959 The mechanistic basis of data dependence and abrupt learning in an in-context classification task. 2024 ICLR Informatics Labs, NTT Research Inc; Center for Brain Science, Harvard University B
1960 The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models. 2024 ACL University College London B
1961 The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and Analysis. 2024 EMNLP King's College London B
1962 The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI. 2024 EMNLP Technical University of Munich; University of Michigan B
1963 The Illusion of Competence: Evaluating the Effect of Explanations on Users' Mental Models of Visual Question Answering Systems. 2024 EMNLP Bielefeld University B
1964 The Geometry of Concepts: Sparse Autoencoder Feature Structure 2024 MIT C
1965 The Expressive Leaky Memory Neuron: an Efficient and Expressive Phenomenological Neuron Model Can Solve Long-Horizon Tasks. 2024 ICLR University of Tübingen, Germany; Max Planck Institute for Intelligent Systems, T B
1966 The Effect of Weight Precision on the Neuron Count in Deep ReLU Networks. 2024 ICML Rutgers, The State University of New Jersey B
1967 The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value Factorization. 2024 NeurIPS National University of Defense Technology; Istituto Nazionale di Fisica Nucleare B
1968 The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models. 2024 ICLR ◆ Chinese University of Hong Kong; Peking University; ▶ National University of S B
1969 The Bayesian sampling in a canonical recurrent circuit with a diversity of inhibitory interneurons. 2024 NeurIPS Southwestern Medical Center B
1970 TextGenSHAP: Scalable Post-Hoc Explanations in Text Generation with Long Documents. 2024 ACL University of Southern California; Google Cloud AI Research, Sunnyvale, CA B
1971 Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game. 2024 ICLR University of California, Berkeley; Carnegie Mellon University B
1972 Temporally Consistent Factuality Probing for Large Language Models. 2024 EMNLP Indian Institute of Technology Delhi; Wipro B
1973 Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs. 2024 ICLR Georgia Institute of Technology; University of California, Berkeley; ⋄Microsoft B
1974 Technical Report: Enhancing LLM Reasoning with Reward-guided Tree Search 2024 Jinhao Jiang, Zhipeng Chen, Yingqian Min, Jie Chen, Xiaoxue C
1975 Technical Report on Slow Thinking with LLMs: II - Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems 2024 Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng C
1976 Teaching Small Language Models Reasoning through Counterfactual Distillation. 2024 EMNLP Zhejiang University; Ant Group (China) B
1977 TaPERA: Enhancing Faithfulness and Interpretability in Long-Form Table QA by Content Planning and Execution-based Reasoning. 2024 ACL Yale University; Zhejiang University; Allen Institute for Artificial Intelligenc B
1978 TVE: Learning Meta-attribution for Transferable Vision Explainer. 2024 ICML Rice University; Wake Forest University; New Jersey Institute of Technology; Tex B
1979 TELLER: A Trustworthy Framework for Explainable, Generalizable and Controllable Fake News Detection. 2024 ACL City University of Hong Kong; Nanyang Technological University; University of El B
1980 SyntaxShap: Syntax-aware Explainability Method for Text Generation. 2024 ACL ETH Zurich B
1981 Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation. 2024 EMNLP University of California, Los Angeles B
1982 Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation. 2024 ICML Apple, Cupertino, CA, USA; Meta, Menlo xPark, CA, USA (*Work done ) B
1983 SummaCoz: A Dataset for Improving the Interpretability of Factual Consistency Detection for Summarization. 2024 EMNLP Iowa State University B
1984 Successor Heads: Recurring, Interpretable Attention Heads In The Wild. 2024 ICLR University of Cambridge B
1985 StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements. 2024 EMNLP University of Washington; Allen Institute for Artificial Intelligence B
1986 Style-Specific Neurons for Steering LLMs in Text Style Transfer. 2024 EMNLP Technical University of Munich; Munich Center for Machine Learning; EU Business B
1987 Structured Matrix Basis for Multivariate Time Series Forecasting with Interpretable Dynamics. 2024 NeurIPS Harbin Institute of Technology B
1988 Structure Your Data: Towards Semantic Graph Counterfactuals. 2024 ICML National Technical University of Athens B
1989 Stochastic Concept Bottleneck Models. 2024 NeurIPS Department of Computer Science; ETH Zurich B
1990 Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution. 2024 NeurIPS Stanford University; University of Washington B
1991 Steering Llama 2 via Contrastive Activation Addition. 2024 ACL Center for Human Genetics BC
1992 Standardized Interpretable Fairness Measures for Continuous Risk Scores. 2024 ICML SCHUFA Holding AG, Germany B
1993 Speaking Your Language: Spatial Relationships in Interpretable Emergent Communication. 2024 NeurIPS University of Southampton B
1994 SparseFit: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations. 2024 ACL ETH Zurich; University of Edinburgh; University College London B
1995 Sparse Autoencoders Find Highly Interpretable Features in Language Models. 2024 ICLR EleutherAI; MATS; Bristol AI Safety Centre; Apollo Research B
1996 Socratic Human Feedback (SoHF): Expert Steering Strategies for LLM Code Generation. 2024 EMNLP Amazon Web Services; Amazon B
1997 Social Bias Probing: Fairness Benchmarking for Language Models. 2024 EMNLP ETH Zurich; University of Copenhagen B
1998 Single-Model Attribution of Generative Models Through Final-Layer Inversion. 2024 ICML Ruhr University Bochum; Universität Hamburg B
1999 Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair. 2024 EMNLP Nara Institute of Science and Technology B
2000 Sign Gradient Descent-based Neuronal Dynamics: ANN-to-SNN Conversion Beyond ReLU Network. 2024 ICML Department of Computer Science & Engineering, Seoul Na-; tional University, Seou B
2001 ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors. 2024 EMNLP Tsinghua University; Peking University; Beihang University B
2002 Shaping the distribution of neural responses with interneurons in a recurrent circuit model. 2024 NeurIPS Flatiron Institute; New York University B
2003 Semantics or spelling? Probing contextual word embeddings with orthographic noise. 2024 ACL Cornell University B
2004 Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP. 2024 EMNLP Seoul National University; LG AI Research (South Korea) B
2005 SelfIE: Self-Interpretation of Large Language Model Embeddings. 2024 ICML Columbia University B
2006 Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models 2024 Alibaba Qwen Team C
2007 Self-Supervised Interpretable End-to-End Learning via Latent Functional Modularity. 2024 ICML Korea Advanced Institute of Science and Technology B
2008 Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations. 2024 EMNLP LFI - Learning, Fuzzy and Intelligent systems (France); Ekimetrics (36 Rue La Fa B
2009 Selective Explanations. 2024 NeurIPS Harvard University; IBM Research B
2010 Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability. 2024 EMNLP Hong Kong University of Science and Technology; Tencent AI Lab B
2011 SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation. 2024 NeurIPS University of Cambridge; University of Hong Kong; Google DeepMind (United Kingdo B
2012 Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation 2024 Jointly conducted by researchers from the MAPLE Lab at Westlake University, South China University of Technology, Peking University, and the Westlake Institute for Advanced Study C
2013 Scaling Tractable Probabilistic Circuits: A Systems Perspective. 2024 ICML National University of Singapore; University of California, Los Angeles B
2014 Scaling Synthetic Data Creation with 1,000,000,000 Personas 2024 C
2015 Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension 2024 University of Maryland + Microsoft C
2016 Scaling Continuous Latent Variable Models as Probabilistic Integral Circuits. 2024 NeurIPS Eindhoven University of Technology; University of Edinburgh B
2017 SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales. 2024 EMNLP Purdue University; University of Illinois Urbana-Champaign; University of Southe B
2018 Saliency-driven Experience Replay for Continual Learning. 2024 NeurIPS University of Catania; University of Modena and Reggio Emilia B
2019 Saliency strikes back: How filtering out high frequencies improves white-box explanations. 2024 ICML Harvard University; Brown University B
2020 SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation. 2024 ICLR Michigan State University, ^ ‡ University of Pennsylvania, ^§IBM Research B
2021 Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations. 2024 EMNLP Singapore University of Technology and Design; Indian Institute of Technology Kh B
2022 STORYSUMM: Evaluating Faithfulness in Story Summarization. 2024 EMNLP Columbia International University; University of Missouri; Columbia University B
2023 START: A Generalized State Space Model with Saliency-Driven Token-Aware Transformation. 2024 NeurIPS Southeast University B
2024 SPIN: Sparsifying and Integrating Internal Neurons in Large Language Models for Text Classification. 2024 ACL University of Toronto; Technical University of Munich B
2025 SOInter: A Novel Deep Energy-Based Interpretation Method for Explaining Structured Output Models. 2024 ICLR Sharif University of Technology B
2026 SIN: Selective and Interpretable Normalization for Long-Term Time Series Forecasting. 2024 ICML Nanjing University B
2027 SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning. 2024 NeurIPS University of Maryland; Clemson University; Rutgers, The State University of New B
2028 SEER: Facilitating Structured Reasoning and Explanation via Reinforcement Learning. 2024 ACL Shanghai Artificial Intelligence Laboratory; Chinese Academy of Sciences; Shangh B
2029 Reverse Thinking Makes LLMs Stronger Reasoners 2024 Google C
2030 Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods. 2024 ACL University of Copenhagen; Language Science (South Korea) B
2031 Revealing Personality Traits: A New Benchmark Dataset for Explainable Personality Recognition on Dialogues. 2024 EMNLP Renmin University of China; Independent Age B
2032 Retrieval-Guided Reinforcement Learning for Boolean Circuit Minimization. 2024 ICLR University of Calgary B
2033 Rethinking the symmetry-preserving circuits for constrained variational quantum algorithms. 2024 ICLR B
2034 Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits. 2024 ICML Princeton University; East China Normal University; Stanford University; Columbi B
2035 Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition. 2024 ICLR Seoul National University B
2036 Representing Molecules as Random Walks Over Interpretable Grammars. 2024 ICML MIT-IBM Watson AI Lab, IBM Research; Massachusetts Institute of Technology B
2037 Representation Surgery: Theory and Practice of Affine Steering. 2024 ICML Bar-Ilan University; Google B
2038 Relational Concept Bottleneck Models. 2024 NeurIPS Scuola Normale Superiore B
2039 RegExplainer: Generating Explanations for Graph Neural Networks in Regression Tasks. 2024 NeurIPS New Jersey Institute of Technology; Florida International University; Arizona St B
2040 Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning. 2024 ICLR Monash University; Griffith University B
2041 Rationales for Answers to Simple Math Word Problems Confuse Large Language Models. 2024 ACL Sichuan University; University of California, Berkeley; Anthropic; Robotics and B
2042 Rationale-Aware Answer Verification by Pairwise Self-Evaluation. 2024 EMNLP Asahi Shimbun Company (Japan) B
2043 RORA: Robust Free-Text Rationale Evaluation. 2024 ACL Johns Hopkins University B
2044 RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation. 2024 ICML Northwestern University B
2045 RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation. 2024 EMNLP Seoul National University B
2046 RDRec: Rationale Distillation for LLM-based Recommendation. 2024 ACL Graduate School of Engineering; Interdisciplinary Graduate School; University of B
2047 RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations. 2024 ACL Stanford University B
2048 RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering. 2024 ICLR B
2049 RA2FD: Distilling Faithfulness into Efficient Dialogue Systems. 2024 EMNLP Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory B
2050 Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing. 2024 NeurIPS Department of Computer Science; University of California San Diego B
2051 QORA: Zero-Shot Transfer via Interpretable Object-Relational Model Learning. 2024 ICML Texas A&M University B
2052 Q-Probe: A Lightweight Approach to Reward Maximization for Language Models. 2024 ICML New York University; Courant Institute of Mathematical Sciences B
2053 Putting Gale & Shapley to Work: Guaranteeing Stability Through Learning. 2024 NeurIPS Penn State University, USA; University of Leeds, UK B
2054 Provably Better Explanations with Optimized Aggregation of Feature Attributions. 2024 ICML University of Oxford B
2055 Prospector Heads: Generalized Feature Attribution for Large Models & Data. 2024 ICML Stanford University; Art Institute of Portland; University of Waterloo B
2056 Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks. 2024 ACL Rutgers, The State University of New Jersey; AWS AI Labs B
2057 Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal Features. 2024 ICLR University of California Berkeley, CA, USA B
2058 Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions. 2024 ICML JPMorganChase AI Research B
2059 Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models 2024 Laura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kama C
2060 Probing the Uniquely Identifiable Linguistic Patterns of Conversational AI Agents. 2024 ACL University of Manchester; Manchester University B
2061 Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games. 2024 ACL Apple (United States) B
2062 Probing the Emergence of Cross-lingual Alignment during LLM Training. 2024 ACL University of Edinburgh B
2063 Probing the Decision Boundaries of In-context Learning in Large Language Models. 2024 NeurIPS Department of Computer Science; University of California, Los Angeles B
2064 Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction. 2024 EMNLP Brandeis University B
2065 Probing Social Bias in Labor Market Text Generation by ChatGPT: A Masked Language Model Approach. 2024 NeurIPS University of Alberta; Lancaster University; Concordia University; Tsinghua Univ B
2066 Probing Language Models for Pre-training Data Detection. 2024 ACL Institute of Artificial Intelligence, School of Computer Science and Technology; B
2067 Probabilistic Generating Circuits - Demystified. 2024 ICML Saarland University B
2068 Probabilistic Constrained Reinforcement Learning with Formal Interpretability. 2024 ICML Imperial College London B
2069 Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models. 2024 ICML Rutgers, The State University of New Jersey B
2070 Presentations are not always linear! GNN meets LLM for Text Document-to-Presentation Transformation with Attribution. 2024 EMNLP Microsoft; Adobe Research B
2071 Predictive, scalable and interpretable knowledge tracing on structured domains. 2024 ICLR University of Tübingen, 2; Tübingen AI Center, 4 B
2072 Pre-trained Language Models Return Distinguishable Probability Distributions to Unfaithfully Hallucinated Texts. 2024 EMNLP Korea University B
2073 Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience. 2024 ICML Institute for Neuroscience; Goethe University Frankfurt; New York University; he BC
2074 Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs 2024 Liyi Chen, Panrong Tong, Zhongming Jin, Ying Sun, Jieping Ye C
2075 Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models. 2024 EMNLP KU Leuven; Sailplane AI B
2076 Piecewise Linear Parametrization of Policies: Towards Interpretable Deep Reinforcement Learning. 2024 ICLR McGill University, Montreal, Canada B
2077 Persuasiveness of Generated Free-Text Rationales in Subjective Decisions: A Case Study on Pairwise Argument Ranking. 2024 EMNLP University of Pittsburgh; Microsoft B
2078 Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization. 2024 NeurIPS Pennsylvania State University B
2079 Peering into the Mind of Language Models: An Approach for Attribution in Contextual Question Answering. 2024 ACL Adobe Systems (United States) B
2080 Paying More Attention to Source Context: Mitigating Unfaithful Translations from Large Language Model. 2024 ACL Harbin Institute of Technology; Peng Cheng Laboratory B
2081 Path Choice Matters for Clear Attributions in Path Methods. 2024 ICLR Tsinghua University B
2082 Partial observation can induce mechanistic mismatches in data-constrained models of neural dynamics. 2024 NeurIPS Harvard University B
2083 Parameter Competition Balancing for Model Merging 2024 Guodong Du, Junlin Lee, Jing Li, Runhua Jiang, Yifei Guo, Sh C
2084 PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning. 2024 ACL Shenzhen University; Nanyang Technological University B
2085 PRIME: Prioritizing Interpretability in Failure Mode Extraction. 2024 ICLR Department of Computer Science, University of Maryland B
2086 PREALIGN 2024 NLP Group, Nanjing University C
2087 PEDANTS: Cheap but Effective and Interpretable Answer Equivalence. 2024 EMNLP University of Maryland, College Park B
2088 PE: A Poincare Explanation Method for Fast Text Hierarchy Generation. 2024 EMNLP East China Normal University; NPPA Key Laboratory of Publishing Integration Deve B
2089 P-MMEval 2024 Qwen Team C
2090 Optimization Algorithm Design via Electric Circuits. 2024 NeurIPS Stanford University; Rice University B
2091 Optimal ablation for interpretability. 2024 NeurIPS Harvard University B
2092 Ontologically Faithful Generation of Non-Player Character Dialogues. 2024 EMNLP Johns Hopkins University; Microsoft B
2093 Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment 2024 Keming Lu, Bowen Yu, Fei Huang, Yang Fan, Runji Lin, Chang Z C
2094 On the Tractability of SHAP Explanations under Markovian Distributions. 2024 ICML Nantes Université B
2095 On the Similarity of Circuits across Languages: a Case Study on the Subject-verb Agreement Task. 2024 EMNLP Universitat Politècnica de Catalunya; Meta; Association for Computational Lingui B
2096 On the Feasibility of Single-Pass Full-Capacity Learning in Linear Threshold Neurons with Binary Input Vectors. 2024 ICML Syracuse University B
2097 On the Expressive Power of Tree-Structured Probabilistic Circuits. 2024 NeurIPS Department of Computer Science; University of Illinois Urbana-Champaign B
2098 On Mechanistic Knowledge Localization in Text-to-Image Generative Models. 2024 ICML University of Maryland; Adobe Research B
2099 On Measuring Faithfulness or Self-consistency of Natural Language Explanations. 2024 ACL Heidelberg University B
2100 On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey 2024 Zhejiang University + Harbin Institute of Technology — Survey on Synthetic Data Generation with LLMs C
2101 On Gradient-like Explanation under a Black-box Setting: When Black-box Explanations Become as Good as White-box. 2024 ICML Freie Universität Berlin B
2102 On Evaluating Explanation Utility for Human-AI Decision Making in NLP. 2024 EMNLP Kahlert School of Computing, University of Utah; University of Utah B
2103 O1 Replication Journey: A Strategic Progress Report 2024 Shanghai Jiao Tong University C
2104 Not All Language Model Features Are Linear 2024 MIT C
2105 Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics Discovery. 2024 NeurIPS Lehigh University; Global Engineering and Materials (United States); Johns Hopki B
2106 Non-asymptotic Approximation Error Bounds of Parameterized Quantum Circuits. 2024 NeurIPS Wuhan University; National University of Singapore; Hubei Key Laboratory of Comp B
2107 Neurons in Large Language Models: Dead, N-gram, Positional. 2024 ACL Meta; Universitat Politècnica de Catalunya B
2108 Neuronal Competition Groups with Supervised STDP for Spike-Based Classification. 2024 NeurIPS École Centrale de Lille; CRIStAL, University of Lille B
2109 Neuron-Level Knowledge Attribution in Large Language Models. 2024 EMNLP National Centre for Atmospheric Science B
2110 Neuron-Enhanced AutoEncoder Matrix Completion and Collaborative Filtering: Theory and Practice. 2024 ICLR The Chinese University of Hong Kong, Shenzhen, China; University of Texas at Arl B
2111 Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine Translation. 2024 EMNLP Language Technology Lab; University of Amsterdam B
2112 Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization. 2024 ICLR City University of Hong Kong; The University of Tokyo B
2113 NeurRev: Train Better Sparse Neural Network Practically via Neuron Revitalization. 2024 ICLR B
2114 Nearest Neighbor Speculative Decoding for LLM Generation and Attribution. 2024 NeurIPS Cohere (Canada); Meta; University of Chicago; Carnegie Mellon University; Univer B
2115 Navigating the OverKill in Large Language Models 2024 C
2116 Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics. 2024 NeurIPS German Cancer Research Center; ETH Zurich; Heidelberg University; University of B
2117 Natural Counterfactuals With Necessary Backtracking. 2024 NeurIPS Chinese University of Hong Kong; University of California San Diego; Rutgers, Th B
2118 NDOT: Neuronal Dynamics-based Online Training for Spiking Neural Networks. 2024 ICML Department of Machine Learning, MBZUAI, Abu Dhabi, UAE; Technology Innovation In B
2119 NALA: an Effective and Interpretable Entity Alignment Method. 2024 EMNLP Northeastern University B
2120 NAISR: A 3D Neural Additive Model for Interpretable Shape Representation. 2024 ICLR University of North Carolina at Chapel Hill, ^2Wake Forest School of Medicine, ^ B
2121 MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering. 2024 NeurIPS Tsinghua University; Pharmolix Inc B
2122 Multiply-Robust Causal Change Attribution. 2024 ICML Massachusetts Institute of Technology; Amazon B
2123 Multi-Aspect Controllable Text Generation with Disentangled Counterfactual Augmentation. 2024 ACL Nanjing University B
2124 MorphGrower: A Synchronized Layer-by-layer Growing Approach for Plausible Neuronal Morphology Generation. 2024 ICML School of Artificial Intelligence & De-; partment of Computer Science and Engine B
2125 Model Reconstruction Using Counterfactual Explanations: A Perspective From Polytope Theory. 2024 NeurIPS University of Maryland B
2126 Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation. 2024 EMNLP University of Groningen; Northeastern University B
2127 MoE-CT: A Novel Approach For Large Language Models Training With Resistance To Catastrophic Forgetting 2024 Alibaba C
2128 Mixture of a Million Experts 2024 DeepMind C
2129 Mitigating Privacy Seesaw in Large Language Models: Augmented Privacy Neuron Editing via Activation Patching. 2024 ACL Tianjin University B
2130 Mitigating Language Bias of LMMs in Social Intelligence Understanding with Virtual Counterfactual Calibration. 2024 EMNLP Monash University; Beihang University B
2131 Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination. 2024 ACL Seoul National University; LG AI Research; University of Michigan B
2132 MindMerger: Efficient Boosting LLM Reasoning in non-English Languages. Authors: Zixian Huang, Wenhao Zhu, Gong Cheng, Lei Li, Fei Yuan, respectively from the State Key Laboratory for Novel Software Technology at Nanjing University, Carnegie Mellon University, and Shanghai AI Laboratory. 2024 C
2133 Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models. 2024 NeurIPS Korea Advanced Institute of Science and Technology B
2134 MetaGPT: Merging Large Language Models Using Model Exclusive Task Arithmetic 2024 Meta C
2135 MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing. 2024 ACL Indraprastha Institute of Information Technology Delhi; Indian Institute of Tech B
2136 Mediator Interpretation and Faster Learning Algorithms for Linear Correlated Equilibria in General Sequential Games. 2024 ICLR Carnegie Mellon University, Pittsburgh, USA B
2137 Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks. 2024 ICLR University of Cambridge, UK; University College London, UK; EECS Department, Uni B
2138 Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations. 2024 EMNLP University of Toronto; McGill University; Mila – Québec AI Institute; University B
2139 Mechanistic Neural Networks for Scientific Machine Learning. 2024 ICML \addr Informatics Institute, University of Amsterdam; \addr Institute of Science B
2140 Mechanistic Design and Scaling of Hybrid Architectures. 2024 ICML Qure.ai B
2141 Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals. 2024 EMNLP Allen Institute for Artificial Intelligence; University of Washington; Universit B
2142 Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models. 2024 NeurIPS University of Mannheim; Harvard University; Northeastern University B
2143 Measuring Per-Unit Interpretability at Scale Without Humans. 2024 NeurIPS University of Tübingen B
2144 Mastering the game of Go without human knowledge 2024 Google C
2145 Many-Shot In-Context Learning 2024 Google DeepMind C
2146 Manifold Integrated Gradients: Riemannian Geometry for Feature Attribution. 2024 ICML ARC Training Centre for Information Resilience (CIRES), Brisbane, Australia; The B
2147 MambaLRP: Explaining Selective State Space Sequence Models. 2024 NeurIPS Technische Universität Berlin; Link (Germany); Berlin Institute for the Foundati B
2148 MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education. 2024 EMNLP Rice University B
2149 Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning. 2024 EMNLP École Polytechnique Fédérale de Lausanne B
2150 MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model. 2024 EMNLP The Hong Kong University of Science and Technology (Guangzhou); Hong Kong Univer B
2151 MIDGArD: Modular Interpretable Diffusion over Graphs for Articulated Designs. 2024 NeurIPS The Advanced Reality Lab B
2152 MG-Net: Learn to Customize QAOA with Circuit Depth Awareness. 2024 NeurIPS The University of Sydney; Wuhan University; Nanyang Technological University B
2153 MEQA: A Benchmark for Multi-hop Event-centric Question Answering with Explanations. 2024 NeurIPS Xi’an Jiaotong-Liverpool University; University of Liverpool B
2154 MARE: Multi-Aspect Rationale Extractor on Unsupervised Rationale Extraction. 2024 EMNLP Hunan Provincial Key Lab on Bioinformatics, School of Computer Science and Engin B
2155 Local vs. Global Interpretability: A Computational Complexity Perspective. 2024 ICML Hebrew University of Jerusalem B
2156 Local Feature Selection without Label or Feature Leakage for Interpretable Machine Learning Predictions. 2024 ICML Leibniz University Hannover B
2157 Linear Explanations for Individual Neurons. 2024 ICML CSE, UC San Diego, CA, USA; HDSI, UC San Diego B
2158 Lie Neurons: Adjoint-Equivariant Neural Networks for Semisimple Lie Algebras. 2024 ICML University of Michigan, Ann Arbor, MI B
2159 LiDAR: Sensing Linear Probing Performance in Joint Embedding SSL Architectures. 2024 ICLR Apple B
2160 Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding. 2024 NeurIPS University of Illinois Urbana-Champaign; Carnegie Mellon University B
2161 Leveraging Machine-Generated Rationales to Facilitate Social Meaning Detection in Conversations. 2024 ACL Carnegie Mellon University B
2162 Less is More: Fewer Interpretable Region via Submodular Subset Selection. 2024 ICLR Institute of Information Engineering, Chinese Academy of Sciences, Beijing 10009 B
2163 Legal Judgment Reimagined: PredEx and the Rise of Intelligent AI Interpretation in Indian Courts. 2024 ACL Indian Institute of Technology Kanpur; Indian Institute of Science Education and B
2164 Learning to Intervene on Concept Bottlenecks. 2024 ICML Technische Universität Darmstadt; German Research Center for AI (DFKI) B
2165 Learning the greatest common divisor: explaining transformer predictions. 2024 ICLR Meta AI B
2166 Learning interpretable control inputs and dynamics underlying animal locomotion. 2024 ICLR B
2167 Learning from Natural Language Explanations for Generalizable Entity Matching. 2024 EMNLP ♢Northeastern University; Amazon B
2168 Learning a Single Neuron Robustly to Distributional Shifts and Adversarial Label Noise. 2024 NeurIPS University of Wisconsin–Madison B
2169 Learning Interpretable Legal Case Retrieval via Knowledge-Guided Case Reformulation. 2024 EMNLP Renmin University of China B
2170 Learning Disentangled Semantic Spaces of Explanations via Invertible Neural Networks. 2024 ACL University of Manchester; Idiap Research Institute B
2171 Learnable Privacy Neurons Localization in Language Models. 2024 ACL Zhejiang University B
2172 Latent Logic Tree Extraction for Event Sequence Explanation from LLMs. 2024 ICML Nanyang Technological University; Chinese University of Hong Kong, Shenzhen B
2173 Latent Concept-based Explanation of NLP Models. 2024 EMNLP Dalhousie University; Hamad bin Khalifa University; Independent Age B
2174 Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment. 2024 ACL Alibaba Inc B
2175 Large Language Models Can Self-Improve in Long-context Reasoning 2024 Siheng Li, Cheng Yang, Zesen Cheng, Lemao Liu, Mo Yu, Yuyu Y C
2176 Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models. 2024 ACL Renmin University of China; Microsoft Research Asia (China) B
2177 Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication. 2024 NeurIPS University of Pittsburgh; Honda Research Institute, USA; Carnegie Mellon Univers B
2178 LaMAGIC: Language-Model-based Topology Generation for Analog Integrated Circuits. 2024 ICML IBM T. J. Watson Research Center; Duke University; MIT-IBM Watson AI Lab; New Je B
2179 LOFIT: Localized Fine-tuning on LLM Representations 2024 C
2180 LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages. Authors: Yinquan Lu, Wenhao Zhu, Lei Li, Yu Qiao, and Fei Yuan. Summary: LLaMAX aims to extend the linguistic boundaries of large language models (LLMs) by enhancing translation capabilities to cover more than 100 languages. Through extensive multilingual continual pretraining on the LLaMA series, the work achieves translation support for over 100 languages. The team developed LLaMAX via a comprehensive analysis of training strategies such as vocabulary extension and data augmentation. Without sacrificing generalization ability, LLaMAX attains significantly higher translation performance than existing open-source LLMs, and performs on par with the dedicated translation model M2M-100-12B on the Flores-101 benchmark. 2024 C
2181 LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study. 2024 EMNLP University of Marburg; University of Mannheim B
2182 LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression. 2024 ACL Tsinghua University; Microsoft (Finland) B
2183 LLMFactor: Extracting Profitable Factors through Prompts for Explainable Stock Movement Prediction. 2024 ACL Tokyo University of Agriculture B
2184 LLM Explainability via Attributive Masking Learning. 2024 EMNLP Tel Aviv University B
2185 LLM Circuit Analyses Are Consistent Across Training and Scale. 2024 NeurIPS EleutherAI; University of Amsterdam; Brown University B
2186 LILO: Learning Interpretable Libraries by Compressing and Documenting Code. 2024 ICLR MIT CSAIL; MIT Brain and Cognitive Sciences; Harvey Mudd College B
2187 LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation. 2024 ACL Tianjin University; Tsinghua University; Baidu (China) B
2188 Knowledge Mechanisms in Large Language Models: A Survey and Perspective. Authors: Mengru Wang, Yunzhi Yao, Ziwen Xu, Shuofei Qiao, Shumin Deng, Peng Wang, Xiang Chen, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang — respectively from Zhejiang University, the NUS-NLP Joint Lab at the National University of Singapore, the University of California, Los Angeles, and Alibaba Group. Summary: This survey examines knowledge mechanisms in large language models (LLMs), which are critical to advancing trustworthy AI. Starting from a novel taxonomy, it reviews the analysis of knowledge mechanisms along two axes: knowledge utilization and knowledge evolution. Knowledge utilization covers memorization, comprehension, application and creation mechanisms, while knowledge evolution focuses on the dynamic progression of knowledge within individual and group LLMs. The paper also discusses what knowledge LLMs have learned, the reasons for the fragility of parametric knowledge, and the potentially challenging "dark knowledge" hypothesis. The authors hope this work helps understand knowledge in LLMs and offers insights for future research. 2024 C
2189 Knowledge Circuits in Pretrained Transformers. 2024 NeurIPS Zhejiang University; National University of Singapore; Zhejiang Key Laboratory o BC
2190 KernelSHAP-IQ: Weighted Least Square Optimization for Shapley Interactions. 2024 ICML Bielefeld University B
2191 KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache 2024 Rice University,Texas A&M University,Stevens Institute of Te C
2192 KAN:Kolmogorov-Arnold Networks 2024 MIT C
2193 Iterative Search Attribution for Deep Neural Networks. 2024 ICML The University of Sydney; University of Malaya; CSIRO Data61; University of Woll B
2194 Iteration Head: A Mechanistic Study of Chain-of-Thought. 2024 NeurIPS Meta (United States); Laboratoire de Mathématiques d'Orsay; Centre Inria de Sacl B
2195 Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-Rationalization. 2024 NeurIPS School of Computer Science and Technology; Huazhong University of Science and Te B
2196 Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching. 2024 ICLR SERI MATS SERI MATS; OpenAI, 2023 B
2197 Is Epistemic Uncertainty Faithfully Represented by Evidential Deep Learning Methods? 2024 ICML Ghent University; Institute of Communications and Navigation, German Aerospace C B
2198 Investigating the Impact of Model Instability on Explanations and Uncertainty. 2024 ACL Machine Science B
2199 Intriguing Properties of Data Attribution on Diffusion Models. 2024 ICLR Singapore Management University; Sea AI Lab, Singapore B
2200 Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis. 2024 EMNLP National Centre for Atmospheric Science B
2201 Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models. 2024 ACL Microsoft (Finland); Purdue University B
2202 Interpretable Sparse System Identification: Beyond Recent Deep Learning Techniques on Time-Series Prediction. 2024 ICLR B
2203 Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts. 2024 EMNLP University of Illinois Urbana-Champaign; University of Wisconsin–Madison B
2204 Interpretable Meta-Learning of Physical Systems. 2024 ICLR École Normale Supérieure - PSL; Département d'Informatique B
2205 Interpretable Mesomorphic Networks for Tabular Data. 2024 NeurIPS Department of Representation Learning; University of Freiburg; Department of Mac B
2206 Interpretable Lightweight Transformer via Unrolling of Learned Graph Smoothness Priors. 2024 NeurIPS York University; Seattle University B
2207 Interpretable Image Classification with Adaptive Prototype-based Vision Transformers. 2024 NeurIPS Dartmouth College; Duke University; University of Maine B
2208 Interpretable Generalized Additive Models for Datasets with Missing Values. 2024 NeurIPS Department of Computer Science; Duke University; University of British Columbia B
2209 Interpretable Diffusion via Information Decomposition. 2024 ICLR University of California Riverside, ^2University of Southern California B
2210 Interpretable Deep Clustering for Tabular Data. 2024 ICML Technion – Israel Institute of Technology B
2211 Interpretable Concept-Based Memory Reasoning. 2024 NeurIPS Università della Svizzera italiana; University of Cambridge; Scuola Normale Supe B
2212 Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents. 2024 NeurIPS Technische Universität Darmstadt; Hessian Center for Artificial Intelligence; Ge B
2213 Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional Understanding. 2024 EMNLP MoE Key Laboratory of Brain-inspired; University of Science and Technology of Ch B
2214 Interpretability-based Tailored Knowledge Editing in Transformers. 2024 EMNLP The London College; University College London B
2215 Interpretability of Language Models via Task Spaces. 2024 ACL Universitat Pompeu Fabra B
2216 Interpretability Illusions in the Generalization of Simplified Models. 2024 ICML Princeton University B
2217 Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification. 2024 NeurIPS Xi’an Jiaotong-Liverpool University; University of Liverpool; Duke Kunshan Unive B
2218 InterpreTabNet: Distilling Predictive Signals from Tabular Data by Salient Feature Interpretation. 2024 ICML University of Toronto; State Research Center of Virology and Biotechnology VECTO B
2219 InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques. 2024 NeurIPS Universidad de Buenos Aires B
2220 IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning. 2024 EMNLP University of Maryland, College Park B
2221 Inherently Interpretable Time Series Classification via Multiple Instance Learning. 2024 ICLR Amazon Prime Video, UK B
2222 Inference to the Best Explanation in Large Language Models. 2024 ACL Idiap Research Institute B
2223 InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance 2024 Fudan University C
2224 Incorporating Information into Shapley Values: Reweighting via a Maximum Entropy Approach. 2024 ICML University of Minnesota B
2225 In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering. 2024 ICML New York University B
2226 Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders. 2024 NeurIPS Google DeepMind (United Kingdom) B
2227 Improving Quotation Attribution with Fictional Character Embeddings. 2024 EMNLP Laboratoire Lorrain de Recherche en Informatique et ses Applications B
2228 Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining. 2024 ICML Harvard University B
2229 Improving LLM Attributions with Randomized Path-Integration. 2024 EMNLP Tel Aviv University B
2230 Improving Interpretation Faithfulness for Vision Transformers. 2024 ICML King Abdullah University of Science and Technology; SDAIA-KAUST AI Center; Lehig B
2231 Improving Alignment and Robustness with Circuit Breakers. 2024 NeurIPS Gray Swan AI; Carnegie Mellon University; Center for AI Safety B
2232 Improve Mathematical Reasoning in Language Models by Automated Process Supervision 2024 Google DeepMind C
2233 Image Inpainting via Tractable Steering of Diffusion Models. 2024 ICLR Department of Computer Science, University of California, Los Angeles; Departmen B
2234 IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware Neurons. 2024 NeurIPS Tianjin University B
2235 IPO: Interpretable Prompt Optimization for Vision-Language Models. 2024 NeurIPS University of Amsterdam; University of Science and Technology of China B
2236 INViTE: INterpret and Control Vision-Language Models with Text Explanations. 2024 ICLR Columbia University, New York, NY, USA B
2237 ICLEF: In-Context Learning with Expert Feedback for Explainable Style Transfer. 2024 ACL Columbia University B
2238 Hypothesis Testing the Circuit Hypothesis in LLMs. 2024 NeurIPS Columbia University; University of Michigan; FAR AI B
2239 How do Large Language Models Handle Multilingualism? 2024 Renmin University of China C
2240 How connectivity structure shapes rich and lazy learning in neural circuits. 2024 ICLR University of Washington, Seattle, WA, USA; Allen Institute for Brain Science, S B
2241 How Interpretable Are Interpretable Graph Neural Networks? 2024 ICML Tencent AI Lab; Chinese University of Hong Kong; Hong Kong Baptist University B
2242 How Do Large Language Models Acquire Factual Knowledge During Pretraining? 2024 KAIST, UCL, KT C
2243 How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition 2024 Alibaba C
2244 Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model Predictions. 2024 ICML National University of Singapore; Institute for Infocomm Research B
2245 HelmFluid: Learning Helmholtz Dynamics for Interpretable Fluid Prediction. 2024 ICML School of Software, BNRist, Tsinghua; University B
2246 HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models. 2024 EMNLP University of Maryland; Microsoft Research B
2247 Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning. 2024 ICLR National University of Singapore, 2; Loyola Marymount University; Element, Inc., B
2248 Harder Tasks Need More Experts: Dynamic Routing in MoE Models 2024 Peking University C
2249 Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL. 2024 ACL Korea Advanced Institute of Science and Technology; Microsoft Research (India) B
2250 HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language Model. 2024 NeurIPS Carnegie Mellon University B
2251 Grounding Language Plans in Demonstrations Through Counterfactual Perturbations. 2024 ICLR Cranberry-Lemon University; University of the Witwatersrand B
2252 Grokking of Implicit Reasoning in Transformers: A Mechanistic Journey to the Edge of Generalization. 2024 NeurIPS The Ohio State University; Carnegie Mellon University B
2253 GraphTrail: Translating GNN Predictions into Human-Interpretable Logical Rules. 2024 NeurIPS Department of Computer Science & Engineering; Department of Computer Science; In B
2254 Graph Neural Networks and Arithmetic Circuits. 2024 NeurIPS Institute of Theoretical Computer Science; Leibniz University Hannover; School o B
2255 Graph Neural Network Explanations are Fragile. 2024 ICML Nanchang University; Illinois Institute of Technology; Milwaukee School of Engin B
2256 Gradient-based Visual Explanation for Transformer-based CLIP. 2024 ICML City University of Hong Kong; Sensetime (China) B
2257 Going Beyond Neural Network Feature Similarity: The Network Feature Complexity and Its Interpretation Using Category Theory. 2024 ICLR Department of Computer Science and Engineering and MoE Key Lab of Artificial Int B
2258 Generating and Evaluating Plausible Explanations for Knowledge Graph Completion. 2024 ACL NEC Laboratories Europe; University College London B
2259 Generating In-Distribution Proxy Graphs for Explaining Graph Neural Networks. 2024 ICML Florida International University; New Jersey Institute of Technology; University B
2260 Gender Identity in Pretrained Language Models: An Inclusive Approach to Data Creation and Probing. 2024 EMNLP University of Stuttgart B
2261 Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 2024 Google DeepMind C
2262 GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation. 2024 EMNLP Georgia Institute of Technology B
2263 GOAt: Explaining Graph Neural Networks via Graph Output Attribution. 2024 ICLR Department of Electrical and Computer Engineering, University of Alberta; Soluti B
2264 GNNBoundary: Towards Explaining Graph Neural Networks through the Lens of Decision Boundaries. 2024 ICLR Ohio State University, Columbus, USA B
2265 F²RL: Factuality and Faithfulness Reinforcement Learning Framework for Claim-Guided Evidence-Supported Counterspeech Generation. 2024 EMNLP National University of Defense Technology; PLA Academy of Military Science B
2266 From Neurons to Neutrons: A Case Study in Interpretability. 2024 ICML The NSF AI Institute for Artificial Intelligence and Fundamental Interactions; M B
2267 From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP. 2024 EMNLP Mila Quebec AI Institute; McGill University; Saarland University; Pontificia Uni B
2268 Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge 2024 Jiahuan Li, Yiqing Cao, Shujian Huang, Jiajun Chen, from the Dept. of Computer Science, Nanjing University C
2269 Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting. 2024 EMNLP University of Oxford; Vienna University of Technology; Rayscape; University Coll B
2270 Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion Perception. 2024 NeurIPS Nanyang Technological University; Agency for Science, Technology and Research; N B
2271 Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation. 2024 ACL City University of Hong Kong; Chinese University of Hong Kong; Shenzhen Universi B
2272 Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers. 2024 ACL University of Science and Technology of China; Fudan University; Tsinghua Univer B
2273 Finding Transformer Circuits With Edge Pruning. 2024 NeurIPS Princeton University B
2274 Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion Models. 2024 NeurIPS German Research Centre for Artificial Intelligence; Technische Universität Darms B
2275 Finding NEM-U: Explaining unsupervised representation learning through neural network generated explanation masks. 2024 ICML University of Copenhagen; UiT The Arctic University of Norway; Norwegian Computi B
2276 Finding Blind Spots in Evaluator LLMs with Interpretable Checklists. 2024 EMNLP Indian Institute of Technology Madras B
2277 FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents. 2024 EMNLP Yale-NUS College B
2278 Figuratively Speaking: Authorship Attribution via Multi-Task Figurative Language Modeling. 2024 ACL Department of Computer Science, *; Department of Cognitive Science; Rensselaer P B
2279 Federated Self-Explaining GNNs with Anti-shortcut Augmentations. 2024 ICML University of Science and Technology of China; Institute of Artificial Intellige B
2280 Federated Behavioural Planes: Explaining the Evolution of Client Behaviour in Federated Learning. 2024 NeurIPS Università della Svizzera italiana; Syneos Health B
2281 Feature Attribution with Necessity and Sufficiency via Dual-stage Perturbation Test for Causal Explanation. 2024 ICML Guangdong University of Technology; Guangzhou Laboratory; University of Cambridg B
2282 Faithfulness Measurable Masked Language Models. 2024 ICML Mila - Quebec AI Institute B
2283 Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach. 2024 EMNLP Poznan University of Technology, Faculty of Computing and Telecommunications, Po B
2284 Faithful and Efficient Explanations for Neural Networks via Neural Tangent Kernel Surrogate Models. 2024 ICLR Pacific Northwest National Laboratory; University of California; Courant Institu B
2285 Faithful Vision-Language Interpretation via Concept Bottleneck Models. 2024 ICLR B
2286 Faithful Rule Extraction for Differentiable Rule Learning Models. 2024 ICLR University of Oxford, UK B
2287 Faithful Persona-based Conversational Dataset Generation with Large Language Models. 2024 ACL Google (United States) B
2288 Faithful Logical Reasoning via Symbolic Chain-of-Thought. 2024 ACL National University of Singapore; University of California, Santa Barbara; Unive B
2289 Faithful Explanations of Black-box NLP Models Using LLM-generated Counterfactuals. 2024 ICLR Technion – Israel Institute of Technology; Columbia University; Google B
2290 Faithful Chart Summarization with ChaTS-Pi. 2024 ACL Google (United States) B
2291 FRVA: Fact-Retrieval and Verification Augmented Entailment Tree Generation for Explainable Question Answering. 2024 ACL Shanxi University B
2292 FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model. 2024 ACL Seoul National University B
2293 FFAM: Feature Factorization Activation Map for Explanation of 3D Detectors. 2024 NeurIPS Sun Yat-sen University B
2294 FANTAstic SEquences and Where to Find Them: Faithful and Efficient API Call Generation through State-tracked Constrained Decoding and Reranking. 2024 EMNLP Texas A&M University; Amazon B
2295 Exploring the trade-off between deep-learning and explainable models for brain-machine interfaces. 2024 NeurIPS Robotics Research (United States); ETH Zurich; BioSurfaces (United States); Eind B
2296 Explanations that reveal all through the definition of encoding. 2024 NeurIPS New York University B
2297 Explanation-aware Soft Ensemble Empowers Large Language Model In-context Learning. 2024 ACL Google (United States); Cornell University B
2298 Explaining and Improving Contrastive Decoding by Extrapolating the Probabilities of a Huge and Hypothetical LM. 2024 EMNLP University of Massachusetts Amherst; Amazon; University of Southern California B
2299 Explaining Time Series via Contrastive and Locally Sparse Perturbations. 2024 ICLR Nanjing University, 2; Pennsylvania State University, 4; Tsinghua University, 5; B
2300 Explaining Probabilistic Models with Distributional Values. 2024 ICML Amazon B
2301 Explaining Mixtures of Sources in News Articles. 2024 EMNLP University of Southern California; University of California, Los Angeles; Stanfo B
2302 Explaining Kernel Clustering via Decision Trees. 2024 ICLR Technical University of Munich; Amazon Research B
2303 Explaining Graph Neural Networks with Large Language Models: A Counterfactual Perspective on Molecule Graphs. 2024 EMNLP University of Virginia; Florida State University B
2304 Explaining Graph Neural Networks via Structure-aware Interaction Index. 2024 ICML Yale University; VinAI Research B
2305 Explaining Datasets in Words: Statistical Models with Natural Language Parameters. 2024 NeurIPS All authors affiliated with UC Berkeley B
2306 Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster. 2024 ACL University of Edinburgh; Snap Inc. B
2307 Exogenous Matching: Learning Good Proposals for Tractable Counterfactual Estimation. 2024 NeurIPS East China Normal University B
2308 Exact Soft Analytical Side-Channel Attacks using Tractable Circuits. 2024 ICML Graz University of Technology; Institute of Information and Communication Techno B
2309 Evaluating Readability and Faithfulness of Concept-based Explanations. 2024 EMNLP Renmin University of China; University of Science and Technology of China; Kuais B
2310 Evaluate then Cooperate: Shapley-based View Cooperation Enhancement for Multi-view Clustering. 2024 NeurIPS University of Defence; Academy of Military Science B
2311 Episodic Memory Retrieval from LLMs: A Neuromorphic Mechanism to Generate Commonsense Counterfactuals for Relation Extraction. 2024 ACL Wuhan University B
2312 Enhancing Training Data Attribution for Large Language Models with Fitting Error Consideration. 2024 EMNLP Chinese Academy of Sciences; Institute of Computing Technology, CAS; University B
2313 Enhancing Semantic Consistency of Large Language Models through Model Editing: An Interpretability-Oriented Approach. 2024 ACL Tianjin University; The Sense Innovation and Research Center B
2314 Enhancing Robustness of Graph Neural Networks on Social Media with Explainable Inverse Reinforcement Learning. 2024 NeurIPS Key Laboratory of Trustworthy Distributed Computing and Service (BUPT); Beijing B
2315 Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer Decomposition. 2024 EMNLP Adobe Systems (United States) B
2316 Enhancing Explainable Rating Prediction through Annotated Macro Concepts. 2024 ACL Hong Kong Polytechnic University B
2317 Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic Interpretations. 2024 ICLR The Hong Kong University of Science and Technology, ^2University of Washington, B
2318 End-to-End Neuro-Symbolic Reinforcement Learning with Textual Explanations. 2024 ICML Beijing Institute for General Artificial Intelligence B
2319 Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective. 2024 EMNLP King's College London; Carnegie Mellon University; Mohamed bin Zayed University B
2320 Embedded Named Entity Recognition using Probing Classifiers. 2024 EMNLP Karlsruhe Institute of Technology; Technische Universität Dresden B
2321 Eliciting Latent Predictions from Transformers with the Tuned Lens 2024 C
2322 EiG-Search: Generating Edge-Induced Subgraphs for GNN Explanation in Linear Time. 2024 ICML University of Alberta; Université de Montréal; Huawei B
2323 Efficient and Interpretable Grammatical Error Correction with Mixture of Experts. 2024 EMNLP National University of Singapore; Mohamed bin Zayed University of Artificial Int B
2324 Efficient Sketches for Training Data Attribution and Studying the Loss Landscape. 2024 NeurIPS Google DeepMind (United Kingdom); Imec the Netherlands B
2325 Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization. 2024 ICLR Center for Innovation in Data Engineering and Science, University of Pennsylvani B
2326 EX-FEVER: A Dataset for Multi-hop Explainable Fact Verification. 2024 ACL University of Chinese Academy of Sciences; State Key Laboratory of Pattern Recog B
2327 ESCoT: Towards Interpretable Emotional Support Dialogue Systems. 2024 ACL Renmin University of China; Independent Age B
2328 EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free Probing. 2024 NeurIPS State Key Lab of CAD&CG, Zhejiang University; China Mobile (Zhejiang) Research & B
2329 ELAD: Explanation-Guided Large Language Models Active Distillation. 2024 ACL Emory University; Emory and Henry College B
2330 Dynamic Multi-granularity Attribution Network for Aspect-based Sentiment Analysis. 2024 EMNLP State Key Laboratory of Cognitive Intelligence; University of Science and Techno B
2331 Dynamic Evaluation of Large Language Models by Meta Probing Agents. 2024 ICML Microsoft Research; University of Science and Technology of China B
2332 Dynamic Discounted Counterfactual Regret Minimization. 2024 ICLR B
2333 Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection. 2024 EMNLP Peking University B
2334 Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations. 2024 EMNLP Singapore Management University; National University of Singapore B
2335 Does Large Language Model Contain Task-Specific Neurons? 2024 EMNLP Faculty of Information Engineering and Automation; Kunming University of Science B
2336 Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations. 2024 ICML Columbia University; University of California, Berkeley; New York University B
2337 Do Llamas Work in English? On the Latent Language of Multilingual Transformers 2024 C
2338 Do Large Language Models Latently Perform Multi-Hop Reasoning? 2024 Google C
2339 Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates. 2024 ICML University of Wisconsin–Madison B
2340 Do LLMs Build World Representations? Probing Through the Lens of State Abstraction. 2024 NeurIPS McGill University B
2341 Do Counterfactually Fair Image Classifiers Satisfy Group Fairness? - A Theoretical and Empirical Study. 2024 NeurIPS Seoul National University B
2342 Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional Semantics. 2024 ACL Chinese University of Hong Kong; University of Cambridge B
2343 Dissect Black Box: Interpreting for Rule-Based Explanations in Unsupervised Anomaly Detection. 2024 NeurIPS Shanghai Artificial Intelligence Laboratory; Shenzhen University; Tsinghua Shenz B
2344 Disentangling Interpretable Factors with Supervised Independent Subspace Principal Component Analysis. 2024 NeurIPS BGI Genomics; Columbia University; New York Genome Center B
2345 Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models' Understanding of Discourse Relations. 2024 ACL National University of Singapore; Sichuan University; Institute for Infocomm Res B
2346 Discovering plasticity rules that organize and maintain neural circuits. 2024 NeurIPS Department of Physics; Department of Physiology and Biophysics; University of Wa B
2347 Discovering Biases in Information Retrieval Models Using Relevance Thesaurus as Global Explanation. 2024 EMNLP University of Massachusetts Amherst B
2348 Digital Socrates: Evaluating LLMs through Explanation Critiques. 2024 ACL Allen Institute for Artificial Intelligence B
2349 Diffusion-TS: Interpretable Diffusion for General Time Series Generation. 2024 ICLR Hefei University of Technology B
2350 Dial BeInfo for Faithfulness: Improving Factuality of Information-Seeking Dialogue via Behavioural Fine-Tuning. 2024 EMNLP LTL, University of Cambridge B
2351 DetoxLLM: A Framework for Detoxification with Explanations. 2024 EMNLP University of British Columbia B
2352 Detecting, Explaining, and Mitigating Memorization in Diffusion Models. 2024 ICLR University of Maryland; Zhejiang University; Sony Digital Audio Disc Corporation B
2353 Designs for Enabling Collaboration in Human-Machine Teaming via Interactive and Explainable Systems. 2024 NeurIPS MIT Lincoln Laboratory; The University of; Georgia Institute of Technology B
2354 Designing Decision Support Systems using Counterfactual Prediction Sets. 2024 ICML Max Planck Institute for Software Systems B
2355 Denoising Diffusion Path: Attribution Noise Reduction with An Auxiliary Diffusion Model. 2024 NeurIPS Shanghai Artificial Intelligence Laboratory; Fudan University; Institute of Scie B
2356 DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search 2024 DeepSeek C
2357 Decomposing Co-occurrence Matrices into Interpretable Components as Formal Concepts. 2024 ACL Japan Advanced Institute of Science and Technology; Tokyo Denki University; JSPS B
2358 Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning 2024 Akshara Prabhakar, Thomas L. Griffiths, R. Thomas McCoy, respectively from C
2359 Data-faithful Feature Attribution: Mitigating Unobservable Confounders via Instrumental Variables. 2024 NeurIPS University of Illinois Urbana-Champaign B
2360 Data-Centric Explainable Debiasing for Improving Fairness in Pre-trained Language Models. 2024 ACL Jilin University; New Jersey Institute of Technology; Key Laboratory of Symbolic B
2361 Data Debugging with Shapley Importance over Machine Learning Pipelines. 2024 ICLR Microsoft, USA; University of Amsterdam, The Netherlands B
2362 Data Attribution for Text-to-Image Models by Unlearning Synthesized Images. 2024 NeurIPS Carnegie Mellon University; Adobe Research; University of California, Berkeley B
2363 Dancing in Chains: Reconciling Instruction Following and Faithfulness in Language Models. 2024 EMNLP Stanford University; Samaya AI; Orby AI; AWS AI Labs; Google; NVIDIA; Denser.ai B
2364 DU-Shapley: A Shapley Value Proxy for Efficient Dataset Valuation. 2024 NeurIPS Inria; ENSAE Paris B
2365 DISCRET: Synthesizing Faithful Explanations For Treatment Effect Estimation. 2024 ICML Peking University; University of Pennsylvania; Harvard University B
2366 DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning. 2024 NeurIPS National University of Singapore; Institute for Infocomm Research; AI Singapore; B
2367 DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection. 2024 ACL School of Computer Science and Technology; Xi'an Jiaotong University; University B
2368 Credit Attribution and Stable Compression. 2024 NeurIPS Tel Aviv University; Technion and Google Research; Georgetown University and Goo B
2369 Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions. 2024 NeurIPS Berkeley College; University of California, Berkeley; Microsoft Research (United B
2370 Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training. 2024 ICML Nanjing University of Aeronautics and Astronautics; RIKEN Center for Advanced In B
2371 Counterfactual Metarules for Local and Global Recourse. 2024 ICML J.P. Morgan B
2372 Counterfactual Image Editing. 2024 ICML Department of Computer Science, Columbia Uni- B
2373 Counterfactual Fairness by Combining Factual and Counterfactual Predictions. 2024 NeurIPS Purdue University B
2374 Counterfactual Density Estimation using Kernel Stein Discrepancies. 2024 ICLR Carnegie Mellon University B
2375 CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples. 2024 ACL University of Wisconsin–Madison; Microsoft Research B
2376 Cosmopedia: how to create large-scale synthetic data for pre-training 2024 AI2 C
2377 Controlling Risk of Retrieval-augmented Generation: A Counterfactual Prompting Framework. 2024 EMNLP Institute of Computing Technology, Chinese Academy of Sciences; University of Ch B
2378 Controlling Counterfactual Harm in Decision Support Systems Based on Prediction Sets. 2024 NeurIPS Max Planck Institute for Software Systems B
2379 Consistent Document-level Relation Extraction via Counterfactuals. 2024 EMNLP EU Business School, Munich; Munich Center for Machine Learning B
2380 Confidence Regulation Neurons in Language Models. 2024 NeurIPS ETH Zurich; University of Sheffield B
2381 Concept Bottleneck Generative Models. 2024 ICLR Emory University, Atlanta, USA B
2382 Concept Alignment 2024 Princeton C
2383 Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks. 2024 ICML Computer and Information Science, University of Pennsylvania; Center for Brain S B
2384 Complex priors and flexible inference in recurrent circuits with dendritic nonlinearities. 2024 ICLR New York University B
2385 Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals. 2024 ACL University of Trieste; ETH Zurich; AREA Science Park; Max Planck Institute for I B
2386 Compact Proofs of Model Performance via Mechanistic Interpretability. 2024 NeurIPS Massachusetts Institute of Technology B
2387 Codebook Features: Sparse and Discrete Interpretability for Neural Networks. 2024 ICML Stanford University B
2388 CodePlan: Unlocking Reasoning Potential in Large Language Models by Scaling Code-form Planning 2024 Tsinghua University + Ant Group C
2389 Coarse-to-Fine Concept Bottleneck Models. 2024 NeurIPS Laboratoire d'Informatique, de Robotique et de Microélectronique de Montpellier; B
2390 CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems. 2024 EMNLP German Research Centre for Artificial Intelligence; Technische Universität Berli B
2391 CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity. 2024 EMNLP Intel (United States) B
2392 CoSy: Evaluating Textual Explanations of Neurons. 2024 NeurIPS Technische Universität Berlin; BIFOLD; UMI Lab, ATB Potsdam, Germany; Fraunhofer B
2393 Cluster-Norm for Unsupervised Probing of Knowledge. 2024 EMNLP Cadenza Labs; University of Cambridge; ENS Paris-Saclay; University of Californi B
2394 ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs. 2024 EMNLP University of Washington; University of California, Berkeley; Stanford Universit B
2395 CircuitNet 2.0: An Advanced Dataset for Promoting Machine Learning Innovations in Realistic Chip Design Environment. 2024 ICLR B
2396 Circuit Component Reuse Across Tasks in Transformer Language Models. 2024 ICLR Brown University; School of Medicine; University of Tübingen B
2397 ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline 2024 Zhipu AI C
2398 ChartCheck: Explainable Fact-Checking over Real-World Chart Images. 2024 ACL King's College London; University of Utah; University of Pennsylvania; TIB – Lei B
2399 Causality-Inspired Spatial-Temporal Explanations for Dynamic Graph Neural Networks. 2024 ICLR B
2400 CausalGym: Benchmarking causal interpretability methods on linguistic tasks. 2024 ACL Nielsen Engineering & Research (United States) B
2401 Causal Discovery from Event Sequences by Local Cause-Effect Attribution. 2024 NeurIPS CISPA Helmholtz Center; DER Security (United States); Institute for Structural P B
2402 Causal Contrastive Learning for Counterfactual Regression Over Time. 2024 NeurIPS Paris-Saclay University, CentraleSupélec, MICS Lab, Gif-sur-Yvette, France; Sain B
2403 Causal Action Influence Aware Counterfactual Data Augmentation. 2024 ICML ETH Zurich; Max Planck Institute for Intelligent Systems; University of Tübingen B
2404 Can Large Language Models Mine Interpretable Financial Factors More Effectively? A Neural-Symbolic Factor Mining Agent Model. 2024 ACL Renmin University of China; Faculty of Information Engineering and Automation; K B
2405 Can Large Language Models Interpret Noun-Noun Compounds? A Linguistically-Motivated Study on Lexicalized and Novel Compounds. 2024 ACL University of Bologna; Hong Kong Polytechnic University B
2406 Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words? 2024 EMNLP Google (United States) B
2407 Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales? 2024 NeurIPS Hong Kong Baptist University; Wuhan University B
2408 Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments. 2024 ACL Nanjing University; Microsoft B
2409 Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring. 2024 EMNLP King's College London; AQA; The Alan Turing Institute; University of Warwick B
2410 COLEP: Certifiably Robust Learning-Reasoning Conformal Prediction via Probabilistic Circuits. 2024 ICLR Delft University of Technology B
2411 CLOMO: Counterfactual Logical Modification with Large Language Models. 2024 ACL City University of Hong Kong; Tsinghua University; Seattle University; The Hong B
2412 CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural Networks. 2024 ICML The Hong Kong University of Science and Technology (Guangzhou); Beihang Universi B
2413 CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems. 2024 ACL Huawei Noah’s Ark Lab; Université de Montréal B
2414 CF-OPT: Counterfactual Explanations for Structured Prediction. 2024 ICML Polytechnique Montréal B
2415 CAUSE: Counterfactual Assessment of User Satisfaction Estimation in Task-Oriented Dialogue Systems. 2024 ACL Leiden University; University of Amsterdam B
2416 Bridging Word-Pair and Token-Level Metaphor Detection with Explainable Domain Mining. 2024 ACL Chinese Academy of Sciences B
2417 Born Differently Makes a Difference: Counterfactual Study of Bias in Biography Generation from a Data-to-Text Perspective. 2024 ACL Data61 B
2418 Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image Classification. 2024 ICLR University of Pennsylvania, PA, USA; Johns Hopkins University, Department of Bio B
2419 Binding in hippocampal-entorhinal circuits enables compositionality in cognitive maps. 2024 NeurIPS Neurosciences Institute; Université Paris-Saclay; University of California, Davi B
2420 BiasWipe: Mitigating Unintended Bias in Text Classifiers through Model Interpretability. 2024 EMNLP Indian Institute of Technology Patna B
2421 Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations. 2024 EMNLP Sichuan University; Singapore Management University; ♡Engineering Research Cente B
2422 Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning. 2024 EMNLP University of Illinois Urbana-Champaign; Vanderbilt University B
2423 Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS 2024 Jinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che, Zengqi We C
2424 Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics. 2024 EMNLP Unitelma Sapienza University; Sapienza University of Rome B
2425 Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable? 2024 NeurIPS ETH Zurich B
2426 Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals. 2024 EMNLP Beijing Normal University B
2427 Beyond Accuracy: Ensuring Correct Predictions With Correct Rationales. 2024 NeurIPS University of Delaware B
2428 Benchmarking the Attribution Quality of Vision Models. 2024 NeurIPS Technische Universität Darmstadt B
2429 Benchmarking Deletion Metrics with the Principled Explanations. 2024 ICML Elmore Family School of Electrical and Computer Engineering, Purdue University, B
2430 Benchmarking Counterfactual Image Generation. 2024 NeurIPS National and Kapodistrian University of Athens; Athena Research and Innovation C B
2431 Beam Enumeration: Probabilistic Explainability For Sample Efficient Self-conditioned Molecular Design. 2024 ICLR Laboratory of Artificial Chemical Intelligence (LIAC), Institut des Sciences et B
2432 Bayesian Power Steering: An Effective Approach for Domain Adaptation of Diffusion Models. 2024 ICML Department of Applied Mathematics, The Hong Kong; Polytechnic University, Hong K B
2433 Balancing Speciality and Versatility: a Coarse to Fine Framework for Supervised Fine-tuning Large Language Model 2024 EleutherAI C
2434 Balanced Resonate-and-Fire Neurons. 2024 ICML Institute of Robotics B
2435 Backward Lens: Projecting Language Model Gradients into the Vocabulary Space 2024 Shahar Katz, Yonatan Belinkov:Technion – Israel Institute of C
2436 BAN: Detecting Backdoors Activated by Adversarial Neuron Noise. 2024 NeurIPS Radboud University; Delft University of Technology; Vrije Universiteit Amsterdam B
2437 BAM!Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts 2024 Cohere C
2438 B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable. 2024 NeurIPS Max Planck Institute for Informatics; RTG Neuroexplicit Models of Language, Visi B
2439 AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments. 2024 EMNLP Princeton University B
2440 Autaptic Synaptic Circuit Enhances Spatio-temporal Predictive Learning of Spiking Neural Networks. 2024 ICML Peking University; Shenzhen Maternity and Child Healthcare Hospital B
2441 Auditing Local Explanations is Hard. 2024 NeurIPS University of Tübingen B
2442 AttributionBench: How Hard is Automatic Attribution Evaluation? 2024 ACL The Ohio State University B
2443 Attribute Based Interpretable Evaluation Metrics for Generative Models. 2024 ICML Yonsei University B
2444 Attention Meets Post-hoc Interpretability: A Mathematical Perspective. 2024 ICML Université Côte d'Azur; Institut de Biologie Valrose; Laboratoire Jean-Alexandre B
2445 AttEXplore: Attribution for Explanation with model parameters eXploration. 2024 ICLR B
2446 Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability. 2024 ACL Soongsil University; Adobe Research, USA B
2447 Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics 2024 Yaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan Belinkov C
2448 Are self-explanations from Large Language Models faithful? 2024 ACL Mila – Quebec AI Institute; Polytechnique Montréal; McGill University; CIFAR; Fa B
2449 Arcee’s MergeKit: A Toolkit for Merging Large Language Models 2024 Toolkit developed by the Arcee team (Florida, USA), including Charles Goddard, Shamane Sir C
2450 Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution. 2024 ACL Bocconi University B
2451 Analysing the Generalisation and Reliability of Steering Vectors. 2024 NeurIPS University College London; FAR AI; Athena Research and Innovation Center In Info B
2452 An interpretable error correction method for enhancing code-to-code translation. 2024 ICLR B
2453 An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare Records. 2024 EMNLP University of Copenhagen; Cork University Hospital B
2454 An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs. 2024 ACL Department of Computer Science; George Mason University BC
2455 An Interpretable Evaluation of Entropy-based Novelty of Generative Models. 2024 ICML Department of Computer Science and Engineering, The Chinese University of Hong K B
2456 An Empirical Examination of Balancing Strategy for Counterfactual Estimation on Time Series. 2024 ICML Jilin University; University of Southern California; University of California Sa B
2457 Almost-Linear RNNs Yield Highly Interpretable Symbolic Codes in Dynamical Systems Reconstruction. 2024 NeurIPS Central Institute of Mental Health; Heidelberg University B
2458 Advancing Large Language Model Attribution through Self-Improving. 2024 EMNLP Harbin Institute of Technology; Peng Cheng Laboratory; Northeastern University, B
2459 Adaptive Quantization Error Reconstruction for LLMs with Mixed Precision 2024 Alibaba Cloud C
2460 Activation Scaling for Steering and Interpreting Language Models. 2024 EMNLP Pioneer (United States); Pioneer Hi-Bred B
2461 Accelerating Nash Equilibrium Convergence in Monte Carlo Settings Through Counterfactual Value Based Fictitious Play. 2024 NeurIPS Huazhong University of Science and Technology; National Key Laboratory of Scienc B
2462 Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe Sampling. 2024 NeurIPS Université de Montréal B
2463 Abstracted Shapes as Tokens - A Generalizable and Interpretable Model for Time-series Classification. 2024 NeurIPS Rensselaer Polytechnic; Stony Brook University; Touro University California; IBM B
2464 AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation. 2024 ACL Department of Computer Science and Engineering, HKUST; Tencent AI Lab; Amazon; N B
2465 AXCEL: Automated eXplainable Consistency Evaluation using LLMs. 2024 EMNLP Amazon B
2466 AR-Pro: Counterfactual Explanations for Anomaly Repair with Formal Properties. 2024 NeurIPS Department of Computer and Information Science; University of Pennsylvania B
2467 AGR: Reinforced Causal Agent-Guided Self-explaining Rationalization. 2024 ACL Shanxi University B
2468 A theoretical design of concept sets: improving the predictability of concept bottleneck models. 2024 NeurIPS University of Cambridge B
2469 A hierarchical decomposition for explaining ML performance discrepancies. 2024 NeurIPS University of California, San Francisco; Center for Devices and Radiological Hea B
2470 A Survey on Natural Language Counterfactual Generation. 2024 EMNLP Nanyang Technological University; Shenzhen University B
2471 A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis. 2024 ICLR The Ohio State University; Amazon Alexa; Princeton University; Rensselaer Polyte B
2472 A Robust Dual-debiasing VQA Model based on Counterfactual Causal Effect. 2024 EMNLP Research & Development Institute; Northwestern Polytechnical University B
2473 A Neural Network Approach for Efficiently Answering Most Probable Explanation Queries in Probabilistic Models. 2024 NeurIPS The University of Texas at Dallas B
2474 A Multimodal Automated Interpretability Agent. 2024 ICML Massachusetts Institute of Technology B
2475 A Metalearned Neural Circuit for Nonparametric Bayesian Inference. 2024 NeurIPS Department of Computer Science; Princeton University; Department of Psychology B
2476 A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. 2024 ICML The University of Sydney B
2477 A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task. 2024 ACL University of Mannheim; Georgia Institute of Technology; Heinrich Heine Universi B
2478 A Linear Algebraic Framework for Counterfactual Generation. 2024 ICLR B
2479 A Hierarchical Adaptive Multi-Task Reinforcement Learning Framework for Multiplier Circuit Design. 2024 ICML University of Science and Technology of China; The Hong Kong University of Scien B
2480 A Geometric Explanation of the Likelihood OOD Detection Paradox. 2024 ICML University of Toronto B
2481 A General Protocol to Probe Large Vision Models for 3D Physical Understanding. 2024 NeurIPS University of Oxford; Shanghai Jiao Tong University B
2482 A Dual-module Framework for Counterfactual Estimation over Time. 2024 ICML University of Science and Technology of China; Nanyang Technological University; B
2483 A Concept-Based Explainability Framework for Large Multimodal Models. 2024 NeurIPS Sorbonne Université; Institut Polytechnique de Paris B
2484 A Compositional Atlas for Algebraic Circuits. 2024 NeurIPS University of California, Los Angeles; Universidade de São Paulo; Arizona State B
2485 A Circuit Domain Generalization Framework for Efficient Logic Synthesis in Chip Design. 2024 ICML Key Laboratory of Technology in GIP AS, University of Science and; Technology of B
2486 A Causal Approach for Counterfactual Reasoning in Narratives. 2024 ACL Hong Kong Polytechnic University B
2487 A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution. 2024 EMNLP University of Maryland, College Park B
2488 "What Data Benefits My Classifier?" Enhancing Model Performance and Interpretability through Influence-Based Data Selection. 2024 ICLR University of South Florida, Bellini College of Artificial Intelligence, Cyberse B
2489 "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations? 2024 EMNLP EU Business School, Munich; Munich Center for Machine Learning; University of Ca B
2490 Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning 2023 C
2491 Language Representation Projection: Can We Transfer Factual Knowledge across Languages in Multilingual Language Models? 2023 Tianjin University C
2492 Hyperpolyglot LLMs: Cross-Lingual Interpretability in Token Embeddings 2023 C
2493 Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models 2023 UNC Chapel Hill / Google Research C
2494 Context-DPO 2023 Chinese Academy of Sciences / Microsoft C
2495 The Geometry of Multilingual Language Model Representations 2022 C
2496 Discovering Low-rank Subspaces for Language-agnostic Multilingual Representations 2021 Shanghai Jiao Tong University C
2497 Paper: Enhancing Automated Interpretability with Output-Centric Feature Descriptions Yoav Gur-Arieh, Roy Mayan, Chen Agassy (Tel Aviv University) C
2498 Paper 2: FADE: Why Bad Descriptions Happen to Good Features Tianjin University C
2499 Scaling Still Matters Most in Model Training — An Interview with Baosong Yang, Head of Multilingual at Alibaba Tongyi Qwen Tongyi Lab C
2500 Multi-Turn Planning Techniques for LLM Agent RL Training: An Explainer Thread Covering Several Papers C
2501 An Overnight Reversal: The Brazilian LLM That "Broke Into the Top Tier" Turned Out to Be a Re-Skinned Chinese Model IT Company of the Rio de Janeiro City Government, Brazil C
2502 ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning Shanghai Jiao Tong University C
2503 mHC: Manifold-Constrained Hyper-Connections DeepSeek C
2504 ZHEN: Check which papers this paper cites, and its "" experiments C
2505 WorldPM Qwen Team C
2506 Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation Zhejiang University C
2507 When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers Rensselaer Polytechnic Institute, USA C
2508 When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Harvard + Amazon C
2509 When AI builds itself : Anthropic Blog Anthropic C
2510 What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs Signals and Interactive Systems Lab, University of Trento, Italy C
2511 Weight Patching: Toward Source-Level Mechanistic Localization in LLMs University of Science and Technology of China C
2512 Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning USTC C
2513 Unlearning Isn’t Deletion: Investigating Reversibility of Machine Unlearning in LLMs The Hong Kong Polytechnic University C
2514 Understanding the Dark Side of LLMs’ Intrinsic Self-Correction Oxford / Google DeepMind / Mila C
2515 Understanding and Enforcing Weight Disentanglement in Task Arithmetic Nanjing University C
2516 Trust-Region Adaptive Policy Optimization Tsinghua University / Ant Group C
2517 Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models Alibaba C
2518 TriAttention: Efficient Long Reasoning with Trigonometric KV Compression MIT / NVIDIA / Zhejiang University C
2519 Trans-Zero ByteDance C
2520 Training-Free Looped Transformers :arXiv University of Texas at Austin C
2521 Training Transformers for KV Cache Compressibility University of Oxford、Technion、AITHYRA、NVIDIA C
2522 Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning Harbin Institute of Technology C
2523 Topology of Reasoning The University of Tokyo / Google DeepMind C
2524 Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Renmin University of China / BAAI / Kuaishou C
2525 Token-Importance Guided Direct Preference Optimization Chinese Academy of Sciences / ByteDance C
2526 Titans: Learning to Memorize at Test Time Google Research C
2527 TileRT: Tile-Based Runtime for Ultra-Low-Latency LLM Inference TileRT C
2528 The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks :arXiv LeCun's Team C
2529 The Latent Space NUS Team C
2530 The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models Oxford & Anthropic C
2531 TTRL: Test-Time Reinforcement Learning Tsinghua University / Shanghai AI Laboratory C
2532 TOKEN ALIGNMENT HEADS: UNVEILING ATTENTION’S ROLE IN LLM MULTILINGUAL TRANSLATION ByteDance C
2533 Superposition, Memorization, and Double Descent anthropic C
2534 Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data Anthropic,Berkerly C
2535 Subliminal Learning Is Steering Vector Distillation Stanford University C
2536 Steering Language Models with Weight Arithmetic Anthropic C
2537 Spurious rewards: rethinking training signals in RLVR University of Washington / Allen Institute for AI / UC Berkeley C
2538 Sparse Crosscoders for Cross-Layer Features and Model Diffing Anthropic C
2539 Sparse Attention Post-Training for Mechanistic Interpretability Max Planck Institute for Intelligent Systems (MPI-IS) + Oxford + C
2540 Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance Meta C
2541 SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis Renmin University of China / BAAI / DataCanvas C
2542 Sigma-Moe-Tiny Technical Report Microsoft C
2543 Self-Distillation Enables Continual Learning MIT、ETH C
2544 Self-DC: When to Reason and When to Act Self Divide-and-Conquer for Compositional Questions The Chinese University of Hong Kong C
2545 Seed-Thinking-v1.5: Advancing Superb Reasoning Models with Reinforcement Learning ByteDance C
2546 Scaling sparse feature circuit finding for in-context learning ETH Zurich C
2547 Scaling and context steer LLMs along the same computational path as the human brain Meta C
2548 Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining University of Chicago C
2549 STRESS-TESTING MODEL SPECS REVEALS CHARACTER DIFFERENCES AMONG LANGUAGE MODELS Anthropic C
2550 SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Chinese Academy of Sciences / Meituan C
2551 SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model Training bilibili C
2552 SPARSE FEATURE CIRCUITS Northwestern University C
2553 SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs Zhejiang University / Ant Group C
2554 SE-BENCH: Benchmarking Self-Evolution with Knowledge Internalization Tsinghua University C
2555 Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging UC Berkeley C
2556 Rethinking Thinking Tokens: LLMs as Improvement Operators Meta C
2557 Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution Shanghai Jiao Tong University / Alibaba Tongyi Lab C
2558 Reinforcement Learning with Rubric Anchors Ant Group + Zhejiang University C
2559 Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection Tsinghua Shenzhen International Graduate School / Alibaba C
2560 Recursive Multi-Agent Systems :arXiv UIUC、Stanford University、NVIDIA、MIT C
2561 Reasoning Models Struggle to Control their Chains of Thought:arXiv OpenAI et al. C
2562 Reasoning Models Generate Societies of Thought Google / University of Chicago C
2563 ReFT: Reasoning with Reinforced Fine-Tuning ByteDance C
2564 R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning Renmin University of China / Beijing Institute of Technology / DataCanvas C
2565 Quantize What Counts: More for Keys, Less for Values Case Western Reserve University; Rice University; Meta C
2566 Pretraining with hierarchical memories: separating long-tail and common knowledge Apple C
2567 Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter Kimi (Moonshot AI) / Tsinghua University C
2568 PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models Fudan University / Soochow University / Shanghai AI Laboratory / Stony Brook University / CUHK C
2569 OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization Together AI、University of Sydney、UIUC C
2570 Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights MIT CSAIL C
2571 Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models ByteDance Seed / Peking University C
2572 Navigating the Accuracy-Size Trade-Off with Flexible Model Merging EPFL C
2573 Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training Xi'an Jiaotong University / CUHK / China Mobile / Nanyang Technological University C
2574 Monitoring Monitorability OpenAI C
2575 Model Merging in Pre-training of Large Language Models SEED C
2576 Mixture-of-Depths Attention Seed C
2577 Memory in the Age of Al Agents National University of Singapore, Renmin University of China C
2578 Memory in the Age of AI Agents National University of Singapore et al. C
2579 Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning Ruoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpe C
2580 Memento-Skills: Let Agents Design Agents Memento C
2581 Mapping Post-Training Forgetting in Language Models at Scale University of Tübingen C
2582 Mamba-3: Improved Sequence Modeling using State Space Principles:ICLR Carnegie Mellon University / Princeton University C
2583 MODEL MERGING WITH FUNCTIONAL DUAL ANCHORS CUHK / Westlake University C
2584 MLP Memory: A Retriever-Pretrained Memory for Large Language Models Shanghai Jiao Tong University C
2585 MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection Beihang University / IQuest Research / Shanghai Jiao Tong University / UBC / Langboat C
2586 MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging University of Electronic Science and Technology of China / Dalian University of Technology C
2587 MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs EPFL, Switzerland C
2588 MAmmoTH2: Scaling Instructions from the Web Carnegie Mellon University C
2589 Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression Northwestern University C
2590 Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Tencent Hunyuan C
2591 Linking forward-pass dynamics in Transformers and real-time human processing Harvard University C
2592 Linear representations in language models can change dramatically over a conversation Google Deepmind C
2593 Let’s Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model University of Macau / Tiger Research C
2594 Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs Meta C
2595 Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts. Beijing Jiaotong University C
2596 Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs :arXiv University of Georgia C
2597 Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs University of Georgia / University of California C
2598 Learning to Reason in 13 Parameters Meta C
2599 Learning is Forgetting: LLM Training As Lossy Compression Princeton University C
2600 Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning. Authors: Yun Qu, Yuhang Jiang, Boyuan Wang, Yixiu Mao et al., Prof. Xiangyang Ji's team, Tsinghua University C
2601 Language as a Latent Variable for Reasoning Optimization Tongyi, ZJU C
2602 Language Models' Factuality Depends on the Language of Inquiry. Harvard University C
2603 LLaMAX2: Your Translation-Enhanced Model also Performs Well in Reasoning Nanjing University / Shanghai AI Lab C
2604 KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning CMU C
2605 Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation Peking University C
2606 Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference Microsoft C
2607 KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction Seoul National University C
2608 KV Cache Transform Coding for Compact Storage in LLM Inference Nvidia C
2609 Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders Google Deepmind C
2610 Is In-Context Learning Learning? Microsoft C
2611 Interpreting Language Models Through Concept Descriptions: A Survey Technical University of Berlin C
2612 Internal bias in reasoning models leads to overthinking Nanjing University C
2613 Internal Value Alignment in Large Language Models throughControlled Value Vector Activation USTC C
2614 Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models Neel Nanda (Google DeepMind) / Durham University C
2615 Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures University of Chicago C
2616 Improving Continual Pre-training Through Seamless Data Packing Fudan University C
2617 Identifying indicators of consciousness in Al systems University of Oxford C
2618 ICA Lens: Interpreting LLM Embeddings and Activations with Independent Component Analysis EEEAI Lab C
2619 Hyperloop Transformers :arXiv MIT C
2620 How large language models encode theory-of-mind: a study on sparse parameter patterns Stanford University C
2621 How does Alignment Enhance LLMs’ Multilingual Capabilities? A Language Neurons Perspective Nanjing University / Microsoft Research Asia C
2622 How Do Multilingual Language Models Remember Facts? University of Copenhagen C
2623 Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders (SAERL) Tsinghua University C
2624 Generalist Reward Models Nanjing University C
2625 GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs Maxim Zhelnin, Viktor Moskvoretskii, Egor Shvetsov, Egor Ven C
2626 From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning Stanford University / LeCun C
2627 From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms Hong Kong Baptist University / South China Normal University C
2628 From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence CMU & NYU C
2629 FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling Alibaba C
2630 Feeling the Strength but Not the Source: Partial Introspection in LLMs Harvard University C
2631 FLEXOLMO: Open Language Models for Flexible Data Use Weijia Shi, Akshita Bhagia, Kevin Farhat, Niklas Muennighoff C
2632 Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Models Renmin University of China C
2633 Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking Huawei Noah's Ark Lab C
2634 Exclusive Self Attention :arXiv Apple C
2635 EvoWiki: Evaluating LLMs on Evolving Knowledge Fudan University C
2636 Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Google DeepMind C
2637 Episodic memories enable powerful algorithms for goal-directed decision making that are context-aware and explainable Graz University of Technology C
2638 End-to-End Test-Time Training for Long Context Stanford University / NVIDIA / UC Berkeley / UC San Diego / Astera Institute C
2639 Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning Google C
2640 ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer. google C
2641 Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space SEED C
2642 Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates Advanced Micro Devices C
2643 Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features Standford C
2644 Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference :arXiv Carnegie Mellon University、University of Maryland C
2645 Diverse Preference Learning for Capabilities and Alignment MIT C
2646 DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging The University of Tokyo C
2647 CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis Tongyi Lab C
2648 Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss ByteDance Seed C
2649 Could Thinking Multilingually Empower LLM Reasoning? (Close Reading) Nanjing University C
2650 Copyright-Protected Language Generation via Adaptive Model Fusion ETH Zurich C
2651 Conditional Memory via Scalable Lookup:A New Axis of Sparsity for Large Language Models DeepSeek-AI / Peking University C
2652 Compression is all you need: modeling mathematics Michael Freedman, Princeton C
2653 Code-Switching and Syntax: A Large-Scale Experiment University of Cambridge C
2654 Chain-of-Thought is Not Explainability Oxford / Google DeepMind / Mila C
2655 CL-bench: A Benchmark for Context Learning Tencent Hunyuan C
2656 CL-bench Life: Can Language Models Learn from Real-Life Context? Tencent Hunyuan LLM C
2657 Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies Institute of Automation, Chinese Academy of Sciences C
2658 Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Qwen C
2659 Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs Anthropic C
2660 BLEUBERI: BLEU is a surprisingly effective reward for instruction following University of Maryland C
2661 AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora HKUST C
2662 Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Fudan University C
2663 Attention Residuals Kimi C
2664 Anthropic plans Claude memory update with new Memory Files :Blog Antropic C
2665 Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels Fudan University C
2666 AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models USTC C
2667 Alignment Faking in Large Language Models Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Rog C
2668 Activation Steering via Generative Causal Mediation MIT + Stanford + Goodfire Research C
2669 ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality MIT / Stanford University / Google C
2670 AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents Harbin Institute of Technology C
2671 ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning Peking University C
2672 A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility Tübingen AI Center, University of Tübingen / University of Cambridge C
2673 A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce Salesforce AI Research, UIUC C
2674 A Mechanistic Analysis of Looped Reasoning Language Models University of Oxford C
2675 A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models University of Oregon,Adobe,Cisco C
2676 A Formal Comparison Between Chain of Thought and Latent Thought : ICML26 The University of Tokyo C
2677 A Comprehensive Survey of Reward Models: Taxonomy,Applications, Challenges, and Future Peking University / Fudan University C
2678 "RAG is an algorithm, not magic." — A First-Principles Analysis of Why RAG Works Unisound C