Deduplicated union of three source lists — A frontier-lab posts (OpenAI · DeepMind · Anthropic), B venue papers 2024–2026 (ACL/EMNLP/ICLR/ICML/NeurIPS), C team reading list. 2810 raw entries → 2678 unique after removing 132 duplicates.
| # | Title | Year | Venue / Field | Affiliation | Lists |
|---|---|---|---|---|---|
| 1 | xRFM: Accurate, scalable, and interpretable feature learning models for tabular data | 2026 | ICLR | Inria; Massachusetts Institute of Technology; UC San Diego; University of Califo | B |
| 2 | xHC: Expanded Hyper-Connections | 2026 | Shanghai Jiao Tong University / Xiaohongshu (RED) | C | |
| 3 | scCBGM: Single-Cell Editing via Concept Bottlenecks | 2026 | ICML | Genentech; Genentech Inc.; Guide Labs; New York University; Yale University | B |
| 4 | daVinci-Dev: Agent-native Mid-training for Software Engineering | 2026 | SII / Shanghai Jiao Tong University / GAIR | C | |
| 5 | Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language | 2026 | ICML | KAIST; Korea Advanced Institute of Science & Technology; POSTECH; Pohang Univers | B |
| 6 | Your VAR Model is Secretly an Efficient and Explainable Generative Classifier | 2026 | ICLR | Purdue University | B |
| 7 | Your Language Model Secretly Contains Personality Subnetworks | 2026 | ICLR | Northwestern University; Northwestern University, NVIDIA; University of Arizona; | B |
| 8 | Why These Documents? Explainable Generative Retrieval with Hierarchical Category Paths. | 2026 | ACL | Yonsei University; Korea University | B |
| 9 | Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics. | 2026 | ACL | Zhejiang University; Alibaba Group (China) | BC |
| 10 | Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention | 2026 | ICLR | Tsinghua University | B |
| 11 | Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention | 2026 | Stanford/Harvard Kempner/Anthropic | C | |
| 12 | Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations. | 2026 | ACL | University of Illinois Chicago; University of Illinois Urbana-Champaign | B |
| 13 | Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models. | 2026 | ACL | Tianjin University; German Research Centre for Artificial Intelligence; Saarland | B |
| 14 | Why Attention Patterns Exist: A Unifying Temporal Perspective Analysis | 2026 | ICLR | Huawei Technologies Ltd.; Tianjin University; University of Science and Technolo | B |
| 15 | Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons | 2026 | ICML | Harbin Institute of Technology; Nanjing University; Nanjing University of Scienc | B |
| 16 | Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents. | 2026 | ACL | Chinese Academy of Sciences; Sinoma Science & Technology Co., Ltd. (China); Stat | B |
| 17 | Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing | 2026 | ICLR | Singapore Management University | B |
| 18 | Where Concept Erasure Should Occur: Concept–Layer Alignment in Text-to-Video Diffusion Models | 2026 | ICML | Huazhong University of Science and Technology; University of Nevada Reno | B |
| 19 | Where CoT Reasoning Commits: Entropy Traces Identify Interpretable Attention Heads. | 2026 | ACL | Beijing Institute of Technology; Beijing Emergency Medical Center | B |
| 20 | When Thinking Backfires: Mechanistic Insights into Reason-induced Misalignment | 2026 | ICLR | KAUST; King's College London; King's College London, University of London; Kings | B |
| 21 | When Safety Alignment Fails to Generalize: Probing with Language Game Jailbreaks. | 2026 | ACL | Chinese Academy of Sciences; University of Chinese Academy of Sciences | B |
| 22 | When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models | 2026 | ICLR | Carnegie Mellon University; Columbia University; Pennsylvania State University; | B |
| 23 | When Random Saliency Looks Trained: Architectural Center Bias in CNN Interpretability | 2026 | ICML | University of California, Berkeley; University of North Carolina at Chapel Hill | B |
| 24 | When More is Less: Understanding Chain-of-Thought Length in LLMs | 2026 | ICLR | Google DeepMind; MIT; Peking University; Technische Universität München | B |
| 25 | When Machine Learning Gets Personal: Evaluating Prediction and Explanation | 2026 | ICLR | University California Santa Barbara; University of California, Santa Barbara | B |
| 26 | When Does Sparsity Mitigate the Curse of Depth in LLMs | 2026 | Max Planck Institute for Intelligent Systems | C | |
| 27 | When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression | 2026 | ICML | Michigan State University; University of Michigan - Ann Arbor; calte | B |
| 28 | When Do Diffusion Models Learn to Generate Multiple Objects? | 2026 | ICML | KAIST & Cortiq; TU Darmstadt; Technische Universität Darmstadt; University of Ox | B |
| 29 | What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data | 2026 | ICLR | Facebook; Meta FAIR; University of California, Berkeley | B |
| 30 | What is Missing? Explaining Neurons Activated by Absent Concepts | 2026 | ICML | ETH Zurich; Johannes-Gutenberg Universität Mainz; MPI for Informatics; Max-Planc | B |
| 31 | What Makes Effective Supervision in Latent Chain-of-Thought: An Information-Theoretic Analysis | 2026 | ICML | Eastern Institute of Technology, Ningbo; Hong Kong Polytechnic University; Natio | B |
| 32 | What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input Beliefs | 2026 | ICML | Bar-Ilan University; New York University | B |
| 33 | What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom | 2026 | ICML | Fudan University; Northwest Polytechnical University Xi'an; Peking University; S | B |
| 34 | What Do Large Language Models Know About Opinions? | 2026 | ICLR | Hong Kong University of Science and Technology (Guangzhou); University of Califo | B |
| 35 | What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT | 2026 | ICML | Meta; New York University / Meta FAIR; New York University and AmiLabs; Universi | B |
| 36 | What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial Nudge. | 2026 | ACL | Rochester Institute of Technology | B |
| 37 | Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer | 2026 | ICML | Independent Researcher; Mila, University of Montreal; Peking University; Peking | B |
| 38 | Weight-sparse transformers have interpretable circuits | 2026 | ICML | Massachusetts Institute of Technology; OpenAI; Stanford University; University o | BC |
| 39 | WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics. | 2026 | ACL | Peking University; University of , USA; Beijing Tongming Lake Information Techno | B |
| 40 | Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs | 2026 | ICLR | Carnegie Mellon University | B |
| 41 | WARP: Weight-Space Analysis for Recovering Training Data Portfolios | 2026 | University of Wisconsin–Madison | C | |
| 42 | VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation | 2026 | ICLR | Arizona State University; Facebook; Massachusetts Institute of Technology; Meta | B |
| 43 | Visual Persuasion: What Influences Decisions of Vision-Language Models? | 2026 | ICML | BITS Pilani; Dartmouth College; MIT Media Lab; Massachusetts Institute of Techno | B |
| 44 | VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization | 2026 | ICLR | The Hong Kong University of Science and Technology; Xiamen University | B |
| 45 | Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering. | 2026 | ACL | The Hong Kong University of Science and Technology (Guangzhou); Hong Kong Univer | B |
| 46 | VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL | 2026 | ICLR | , University of Texas at Austin; Microsoft; Microsoft Research Asia; University | B |
| 47 | Verifying Chain-of-Thought Reasoning via Its Computational Graph | 2026 | ICLR | FAIR; Independent; UCL | FAIR (Meta); University of California, Santa Barbara; U | BC |
| 48 | Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing. | 2026 | ACL | University of Hong Kong; Sun Yat-sen University | B |
| 49 | Verified SHAP: Provable Bounds for Exact Shapley Values of Neural Networks | 2026 | ICML | Hebrew University of Jerusalem; University of Konstanz | B |
| 50 | Verification of the Implicit World Model in a Generative Model via Adversarial Sequences | 2026 | ICLR | University of Szeged | B |
| 51 | Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning. | 2026 | ACL | University of Illinois Urbana-Champaign; Fudan University; Tsinghua University; | B |
| 52 | Verbalizable Representations Form a Global Workspace in Language Models | 2026 | Lab post (Anthropic) | Anthropic | AC |
| 53 | VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck. | 2026 | ACL | Fudan University; Shanghai Artificial Intelligence Laboratory | B |
| 54 | Unveiling the Visual Counting Bottleneck in Vision-Language Models | 2026 | ICML | Department of Computer Science, ETHZ - ETH Zurich; ETH Zurich; ETHZ - ETH Zurich | B |
| 55 | Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection | 2026 | ICLR | Australian National University; SUN YAT-SEN UNIVERSITY; Sun Yat-sen University; | B |
| 56 | Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification | 2026 | ICLR | INRIA; Max Planck Institute of Biochemistry; Universiteit Gent; University of Ka | B |
| 57 | Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models | 2026 | ICLR | Nanyang Technological University; Umea University; Umeå University | B |
| 58 | Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention. | 2026 | ACL | Shanghai Jiao Tong University | B |
| 59 | Universal Redundancies in Time Series Foundation Models | 2026 | ICML | University of Texas at Austin | B |
| 60 | Union-of-Experts: Neurons in Mixture-of-Experts are Secretly Routers. | 2026 | ACL | Renmin University of China | B |
| 61 | Unifying Low Dimensional Spectra in Deep Learning | 2026 | ICML | University of Oxford | B |
| 62 | Unifying Formal Explanations: A Complexity-Theoretic Perspective | 2026 | ICLR | Hebrew University of Jerusalem; Nanyang Technological University | B |
| 63 | Unified Time Series Explanations via Amortized Optimization and Instance-level Multi-Expert Knowledge Distillation | 2026 | ICML | Forschungszentrum Juelich GmbH; North China University of Water Resources and El | B |
| 64 | Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks | 2026 | CUHK / Carnegie Mellon University / Institute of Science Tokyo / UIUC / UC Santa Barba | C | |
| 65 | Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and Limitations | 2026 | ICLR | Ohio State University, Columbus; Xi'an Jiaotong University | B |
| 66 | Understanding Reasoning Collapse in LLM Agent Reinforcement Learning | 2026 | ICML | Apple; City University of Hong Kong; Computer Science Department, Stanford Unive | B |
| 67 | Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation. | 2026 | ACL | Nanjing University; Huawei Translation Services Center, Beijing, China | B |
| 68 | Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding | 2026 | ICLR | ByteDance; Department of Computer Science, University of Wisconsin - Madison; Un | B |
| 69 | Understanding Emergent Misalignment via Feature Superposition Geometry. | 2026 | ACL | The University of Tokyo; Google DeepMind (United Kingdom) | B |
| 70 | Understanding Cross-layer Contributions to Mixture-of-Experts Routing in LLMs | 2026 | ICLR | Institute of Science Tokyo; RIKEN | B |
| 71 | Uncovering the Latent Potential of Deep Intermediate Representations | 2026 | ICML | Indraprastha Institute of Information Technology, Delhi | B |
| 72 | Uncovering Sentiment Analysis Circuit in Large Language Model. | 2026 | ACL | Soochow University | B |
| 73 | Uncovering Hidden Triggers: Backdoor Attribution in Language Models | 2026 | ICML | Abel AI; Nanyang Technological University; National University of Singapore; Squ | B |
| 74 | Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders | 2026 | ICLR | Harvard University; Stanford University; University of Michigan | B |
| 75 | Uncovering Competency Gaps in Large Language Models and Their Benchmarks | 2026 | ICML | Google; Google DeepMdin; Google DeepMind; Google Inc; Stanford University | B |
| 76 | Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering | 2026 | ICLR | Amazon; Capital One; CapitalOne; University of California, Irvine; University of | B |
| 77 | UCoder: Unsupervised Code Generation by Internal Probing of Large Language Models. | 2026 | ACL | Beihang University; Huawei | B |
| 78 | Tversky Neural Networks: Psychologically Plausible Deep Learning with Differentiable Tversky Similarity | 2026 | ICLR | Computer Science Department, Stanford University; Stanford University | B |
| 79 | Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution | 2026 | Anthropic | C | |
| 80 | Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations | 2026 | ICLR | Uni Edinburgh / Uni Amsterdam; University of Amsterdam | B |
| 81 | Trustworthy and Explainable Causal Representation Learning in Transformers. | 2026 | ACL | Huazhong Agricultural University; Adelaide University; Hainan University | B |
| 82 | TrustTable: A Neuro-Symbolic Auditing Framework for Faithful Table QA. | 2026 | ACL | Nanjing University of Posts and Telecommunications; The State Key Laboratory of | B |
| 83 | TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs. | 2026 | ACL | Adelaide University | B |
| 84 | TreeGrad-Ranker: Feature Ranking via $O(L)$-Time Gradients for Decision Trees | 2026 | ICLR | National University of Singapore; University of Waterloo Vector Institute | B |
| 85 | Tree-of-Evidence: Efficient "System 2" Search for Faithful Multimodal Grounding. | 2026 | ACL | Georgia Institute of Technology | B |
| 86 | Tree-CoT-RT: An Explainable Multi-Path Tree-Guided Chain-of-Thought and Reinforcement Learning Framework for Aspect Sentiment Quad Prediction. | 2026 | ACL | XU Exponential University of Applied Sciences, Germany; Universiti Sains Malaysi | B |
| 87 | TravelBehaviorQA: A Benchmark Dataset for Behavioral Interpretation of GPS Trajectories. | 2026 | ACL | University of Maryland, College Park | B |
| 88 | Translation Heads: Disentangling meaning from language in LLM-based machine translation | 2026 | ICML | INRIA; Inria; Inria Paris | B |
| 89 | Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations | 2026 | ICLR | Harvard University; Skywork AI; Tencent; Tsinghua Univ.; Tsinghua University; Ts | B |
| 90 | Transformers learn factored representations | 2026 | ICML | Astera Institute; Astera Institute, Simplex; Beyond Institute for Theoretical Sc | B |
| 91 | Transformers are inherently succint :ICLR 2026 Outstanding | 2026 | ETH,RTPU | C | |
| 92 | Transformer Circuits Can Realize Clustering Algorithms | 2026 | ICML | IBM Research; Massachusetts Institute of Technology | B |
| 93 | Training-free Counterfactual Explanation for Temporal Graph Model Inference | 2026 | ICLR | Arizona State University; Case Western Reserve University; University of Massach | B |
| 94 | Training large language models on narrow tasks can lead to broad misalignment | 2026 | Joint team: Truthful AI (Berlin) / Center on Long-Term Risk (London) / Warsaw University of Technology / University of Toronto / Stanford University / UC Berkeley et al. | C | |
| 95 | Tracking Equivalent Mechanistic Interpretations Across Neural Networks | 2026 | ICLR | Carnegie Mellon University | B |
| 96 | Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning | 2026 | ICLR | Emory University; Goethe University Frankfurt; Microsoft Research; NVIDIA | B |
| 97 | Tracing the Persona Circuit: How Large Language Models Encode and Express Character Traits | 2026 | ICML | Tencent AI Lab; University of Science and Technology of China | B |
| 98 | Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection | 2026 | ICML | Beijing Jiaotong University; Nanyang Technological University; Peking University | B |
| 99 | Tracing Logit Trajectories Across Layer Depth: Dataset-Level Explainability for Language Models. | 2026 | ACL | Korea Advanced Institute of Science and Technology; Chungnam National University | B |
| 100 | Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Method | 2026 | ICLR | ByteDance; ByteDance Inc.; Institute of Automation, Chinese Academy of Sciences; | B |
| 101 | TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention | 2026 | ICML | Nanjing University; Nanjing University of Science and Technology; Nanjing univer | B |
| 102 | ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection. | 2026 | ACL | East China Normal University | B |
| 103 | ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway. | 2026 | ACL | Korea University; Myongji University; University of Texas Health Science; AIGEN | B |
| 104 | Towards the Explainability of Temporal Graph Networks via Memory Backtracking and Topological Attribution | 2026 | ICML | Beijing University of Posts and Telecommunications; HKUST(GZ); Hong Kong Univers | B |
| 105 | Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures. | 2026 | ACL | Peking University; Beijing Academy of Artificial Intelligence; Tsinghua Universi | B |
| 106 | Towards Understanding the Shape of Representations in Protein Language Models | 2026 | ICLR | Department of Physics, University of Oslo; University of Oslo | B |
| 107 | Towards Understanding the Robustness of Sparse Autoencoders. | 2026 | ACL | University of Virginia; Independent Age | B |
| 108 | Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition | 2026 | ICLR | Fudan University; Shanghai Innovation Institute | B |
| 109 | Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer | 2026 | ICLR | University of California Berkeley; University of Freiburg / MATS; University of | B |
| 110 | Towards Understanding Massive Activations in Attention Sink Mechanism | 2026 | ICML | The Chinese University of Hong Kong | B |
| 111 | Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions | 2026 | ICML | Ant Group; Zhejiang University; antgroup | B |
| 112 | Towards Spectroscopy: Susceptibility Clusters in Language Models | 2026 | ICML | Resolution; The University of Melbourne; Timaeus | B |
| 113 | Towards Proactive Information Probing: Customer Service Chatbots Harvesting Value from Conversation. | 2026 | ACL | National University of Singapore; Sichuan University; Engineering Research Cente | B |
| 114 | Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMs | 2026 | ICML | City University of Hong Kong; Harbin Institute of Technology | BC |
| 115 | Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures. | 2026 | ACL | Peking University; Beijing Academy of Artificial Intelligence; Nanjing Universit | BC |
| 116 | Towards Interpretable Visual Decoding with Attention to Brain Representations | 2026 | ICLR | Columbia University | B |
| 117 | Towards Interpretable Tabular Reasoning: Enhancing LLM Reasoning on Tabular Data with Pre-Constructed Logic Graph. | 2026 | ACL | Zhejiang University; MYbank, Ant Group | B |
| 118 | Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation | 2026 | ICML | Tsinghua University | B |
| 119 | Towards Explainable Diagnosis: A Self-learned Explanatory Knowledge Base Approach. | 2026 | ACL | Chinese Academy of Sciences | B |
| 120 | Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment | 2026 | ICLR | Carnegie Mellon University; Duke University; Indian Institute of Technology, Del | B |
| 121 | Towards Atoms of Large Language Models | 2026 | ICML | Institute of Automation, Chinese Academy of Sciences; Institute of automation, C | B |
| 122 | Toward Safe Quantization-Aware Fine-tuning: Understanding and Mitigating Safety Alignment Degradation | 2026 | ICML | Institute of Intelligent Computing, University of Electronic Science and Technol | B |
| 123 | Toward Identifiable Sparse Autoencoders | 2026 | ICML | Achira Inc; IST Austria; ISTA | B |
| 124 | Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders | 2026 | ICLR | University of Virginia; University of Virginia, Charlottesville | B |
| 125 | Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts | 2026 | ICML | National University of Singapore; University of Science and Technology of China | B |
| 126 | TimeSeg: An Information-Theoretic Segment-Wise Explainer for Time-Series Predictions | 2026 | ICLR | Chung-Ang University; Korea University | B |
| 127 | TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models | 2026 | ICML | TUM & Télécom Paris; Technical University of Denmark; Technical University of Mu | B |
| 128 | Time series saliency maps: Explaining models across multiple domains | 2026 | ICML | EPFL; Swiss Federal Institute of Technology Lausanne (EPFL) | B |
| 129 | Thought-Action Graph Reasoning: Faithful and Efficient Reasoning of Large Language Models via Reusing Past Experience. | 2026 | ACL | Tsinghua University; Beijing University of Posts and Telecommunications; Zhonggu | B |
| 130 | Thought Branches: Interpreting LLM Reasoning Requires Resampling | 2026 | ICLR | Anthropic Fellows; DeepMind; Duke University; Google DeepMind | B |
| 131 | This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-Critic | 2026 | ICML | EPITA Lyon; Sony AI; University of Roma "La Sapienza", IRISA | B |
| 132 | ThinkPersona: Thinking with Persona Graphs for Faithful Individualized Role-Playing. | 2026 | ACL | Zhejiang University | B |
| 133 | Think in Latent, Explain in Language: Self-Explainable Latent Reasoning | 2026 | ICML | Peking University; University of Illinois Champaign Urbana; University of Illino | B |
| 134 | There Was Never a Bottleneck in Concept Bottleneck Models | 2026 | ICLR | Universidad de Zaragoza; University of Cambridge | B |
| 135 | Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration? | 2026 | ICLR | Cornell University; Department of Computer Science; Department of Computer Scien | B |
| 136 | The Value of Information in Human-AI Decision-making | 2026 | ICLR | Microsoft; Northwestern University; Northwestern University, Northwestern Univer | B |
| 137 | The Tutor-Pupil Augmentation: Enhancing Learning and Interpretability via Input Corrections | 2026 | ICLR | Eindhoven University of Technology; University of Minnesota; University of Minne | B |
| 138 | The Tell-Tale Norm: $\ell_2$ Magnitude as a Signal for Reasoning Dynamics in Large Language Models | 2026 | ICML | Peking University; Zhejiang University | B |
| 139 | The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity | 2026 | ICML | Huawei Noah's Ark Lab; National University of Singapore; Princeton University; T | B |
| 140 | The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology | 2026 | ICLR | Imperial College London; Queen Mary University of London; Queen Mary University | B |
| 141 | The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models | 2026 | ICML | Nanjing University; nanjing university | B |
| 142 | The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics | 2026 | ICLR | Apple; Apple Inc.; Imperial College London | B |
| 143 | The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs | 2026 | ICML | California Institute of Technology; Caltech and NVIDIA; Carnegie Mellon Universi | B |
| 144 | The Perception–Physics Paradox: Probing Scientific Alignment with TC-Bench | 2026 | ICML | ISTA; Institute of Science and Technology Austria | B |
| 145 | The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes | 2026 | ICML | FAR.AI; Google DeepMind | B |
| 146 | The Mechanistic Emergence of Symbol Grounding in Language Models | 2026 | ICML | Department of Computer Science, University of North Carolina at Chapel Hill; Uni | B |
| 147 | The Mechanics of Interference: Defusing Distractors in RAG via Sparse Autoencoder Interventions. | 2026 | ACL | Sapienza University of Rome; Universitas Multi Data Palembang; ISTI-CNR; Univers | B |
| 148 | The Learnability of Model-Theoretic Interpretation Functions in Artificial Neural Networks. | 2026 | ACL | University of California, Santa Cruz | B |
| 149 | The Lattice Representation Hypothesis of Large Language Models | 2026 | ICLR | Stanford University | B |
| 150 | The Latent Color Subspace: Emergent Order in High-Dimensional Chaos | 2026 | ICML | TUM; TUM & Télécom Paris; Technical University of Munich; Technical University o | B |
| 151 | The Information Geometry of Softmax: Probing and Steering | 2026 | ICML | DeepMind; University of Chicago; INSEAD; University of Chicago; University of Ch | B |
| 152 | The Impact of Off-Policy Training Data on Probe Generalisation. | 2026 | ACL | King's College London; University of Cambridge | B |
| 153 | The Geometry of Representational Failures in Vision Language Models | 2026 | ICML | CENTAI; Intesa Sanpaolo AI Research; Northeastern University London; Polytechnic | B |
| 154 | The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution | 2026 | ICML | Center for Information and Language Processing; Huawei Technologies Ltd.; LMU Mu | B |
| 155 | The Geometry of Reasoning: Flowing Logics in Representation Space | 2026 | ICLR | Duke University; Facebook | B |
| 156 | The Geometry of Narrow Fine-Tuning Degradation: Trajectory Lock-in and Spectral Bifurcation | 2026 | ICML | Huazhong University of Science and Technology; South China University of Technol | B |
| 157 | The Geometric Origin of Grokking: Accelerating Generalization via Active Structural Reorganization | 2026 | ICML | Hong Kong Baptist University; National University of Defense Technology | B |
| 158 | The First Impression Problem: Internal Bias Triggers Overthinking in Reasoning Models | 2026 | ICLR | Nanjing University | B |
| 159 | The Extra Tokens Matter: Disentangled Representation Learning with Vision Transformers | 2026 | ICML | University of Tennessee | B |
| 160 | The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level | 2026 | ICML | University of Hamburg | B |
| 161 | The Deleuzian Representation Hypothesis | 2026 | ICLR | CEA; CEA LIST | B |
| 162 | The Cylindrical Representation Hypothesis for Language Model Steering | 2026 | ICML | Indian Institute of Technology Patna; MBZUAI; Mohamed bin Zayed University of Ar | B |
| 163 | The Consciousness Cluster: Emergent Preferences of Models that Claim to be Conscious | 2026 | Lab post (Anthropic) | Anthropic | A |
| 164 | The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think | 2026 | ICLR | Carnegie Mellon University; Cornell University; KAIST; KAIST AI; Korea Advanced | B |
| 165 | The Assistant Axis: Situating and Stabilizing the Character of Large Language Models | 2026 | Lab post (Anthropic) | Anthropic | A |
| 166 | The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives | 2026 | ICLR | Harvard University; Imperial College London; Rocana Venture Partners | B |
| 167 | The Achilles’ Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities | 2026 | ICLR | Beihang University; Beijing University of Aeronautics and Astronautics; Renmin U | B |
| 168 | The Abstraction Gap in Vision-Language Causal Reasoning | 2026 | ICML | University of Nebraska-Lincoln | B |
| 169 | The Price of Amortized inference in Sparse Autoencoders | 2026 | ICLR | King Abdullah University of Science and Technology; MBZUAI | B |
| 170 | Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models. | 2026 | ACL | University of Southern California | B |
| 171 | Temporal superposition and feature geometry of RNNs under memory demands | 2026 | ICLR | G-Research; Imperial College; Imperial College London; Imperial College London, | B |
| 172 | Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability | 2026 | ICLR | Harvard; Harvard University; School of Engineering and Applied Sciences, Harvard | B |
| 173 | Temporal Geometry of Deep Networks: Hyperbolic Representations of Training Dynamics for Intrinsic Explainability | 2026 | ICLR | Einndoven University, Tilburg University | B |
| 174 | Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models | 2026 | ICML | Carnegie Mellon University; EarthDynamics.ai; Max Planck Institute for Software | B |
| 175 | Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability | 2026 | ICML | Virginia Polytechnic Institute and State University | B |
| 176 | Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insights | 2026 | ICLR | Japan Advanced Institute of Science and Technology; Northwestern University; RIK | B |
| 177 | Targeted Neuron Modulation via Contrastive Pair Search | 2026 | Nous Research | C | |
| 178 | Target-Oriented Pretraining Data Selection via Neuron-Activated Graph | 2026 | ICML | ByteDance; ByteDance Inc.; UC Santa Cruz; University of California Santa Cruz; U | B |
| 179 | Taming Polysemanticity in LLMs: Theory-Grounded Feature Recovery via Sparse Autoencoders | 2026 | ICLR | Shanghai Jiaotong University; University of California, San Diego; Yale; Yale Un | B |
| 180 | Talent or Luck? Evaluating Attribution Bias in Large Language Models. | 2026 | ACL | George Mason University; Bronx High School of Science; University of Washington | B |
| 181 | Tackling the XAI Disagreement Problem with Adaptive Feature Grouping | 2026 | ICLR | CortAIx Lab, Thales; Thales Digital Solutions | B |
| 182 | TabReX: Tabular Referenceless eXplainable Evaluation. | 2026 | ACL | Adobe Systems (United States); Adobe Gastroenterology | B |
| 183 | TT-Sparse: Learning Sparse Rule Models with Differentiable Truth Tables | 2026 | ICML | Continental Automotive Singapore; Nanyang Technological University | B |
| 184 | TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG. | 2026 | ACL | University of Technology Sydney | B |
| 185 | TN-SHAP-G: Graph-Structured Tensor Network Surrogates for Shapley Values and Interactions | 2026 | ICML | Montreal Institute for Learning Algorithms, University of Montreal, University o | B |
| 186 | TIMESLIVER : SYMBOLIC-LINEAR DECOMPOSITION FOR EXPLAINABLE TIME SERIES CLASSIFICATION | 2026 | ICLR | Northwestern University; Northwestern University, Northwestern University | B |
| 187 | TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts. | 2026 | ACL | Tianjin University of Science and Technology; Tianjin University of Technology | B |
| 188 | TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning | 2026 | ICML | Shandong University | B |
| 189 | Synthesising Counterfactual Explanations via Label-Conditional Gaussian Mixture Variational Autoencoders | 2026 | ICLR | Imperial College London; JPMorganChase; King's College London | B |
| 190 | Syntax vs. Semantics: How Transformers Learn Deep Dependencies | 2026 | ICML | Beijing University of Posts and Telecommunications | B |
| 191 | Synergizing Stylometrics with Semantics: Dual-Path Framework for LLM Detection and Attribution. | 2026 | ACL | Chongqing Key Laboratory of Image Cognition; School of Artificial Intelligence a | B |
| 192 | Symmetry Reveals the In-Context Classifier: Transformers Implement Mean-Shift Dynamics | 2026 | ICML | Boston University; Boston University, Google Research | B |
| 193 | Symmetries in language statistics shape the geometry of model representations | 2026 | ICML | EPFL - EPF Lausanne; Johns Hopkins University; UC Berkeley | B |
| 194 | SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models | 2026 | ICML | Department of Computer Science, University of Washington; University of Washingt | B |
| 195 | SuperMAN: Interpretable and Expressive Networks over Temporally Sparse Heterogeneous Data | 2026 | ICLR | Aalborg University; Imperial College London; Meta; National Center of Excellence | B |
| 196 | Summaries as Centroids for Interpretable and Scalable Text Clustering | 2026 | ICLR | York University | B |
| 197 | Structural Inference: Interpreting Small Language Models with Susceptibilities | 2026 | ICLR | Timaeus; University of Melbourne | B |
| 198 | Stretching Beyond the Obvious: A Gradient-Free Framework to Unveil the Hidden Landscape of Visual Invariance | 2026 | ICLR | EPFL; Harvard Medical School; Harvard Medical School and MIT; International High | B |
| 199 | Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs | 2026 | ICLR | ELLIS Institute & MPI Intelligent Systems, Tübingen AI Center; EPFL; ETH Zurich; | B |
| 200 | Stop Hardening Everything: A Training-Free Neuron-Level Defense for Neural Ranking Models. | 2026 | ACL | State Key Laboratory of AI Safety; Chinese Academy of Sciences; University of Ch | B |
| 201 | Step-Resolved Data Attribution for Looped Transformers | 2026 | ICML | Google; Harvard; Hasso Plattner Institute & Alphabet; Technical University of Mu | B |
| 202 | Step-Level Sparse Autoencoder for Reasoning Process Interpretation | 2026 | ICML | City University of Hong Kong; Li Auto Inc.; University of Science and Technology | B |
| 203 | Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning | 2026 | ICML | Anthropic; Google DeepMind; Harvard University, Harvard University; Independent; | B |
| 204 | Steering LLM Thinking with Budget Guidance. | 2026 | ACL | University of Massachusetts Amherst; Zhejiang University; MIT-IBM Watson AI Lab | B |
| 205 | Steering Evaluation-Aware Language Models To Act Like They Are Deployed | 2026 | ICLR | Astra Fellow; DeepMind; Jump Trading; Northeastern University | B |
| 206 | Steering Away from Refusal: A Black-box Jailbreak Method Based on First-Token Distribution. | 2026 | ACL | State Key Laboratory of AI Safety, Institute of Computing Technology; Chinese Ac | B |
| 207 | Steering Autoregressive Music Generation with Recursive Feature Machines | 2026 | ICLR | UC San Diego; University of California, San Diego | B |
| 208 | Steer Like the LLM: Activation Steering that Mimics Prompting | 2026 | ICML | Nokia Bell Labs | BC |
| 209 | State-Dependent Safety Failures in Multi-Turn Language Model Interaction | 2026 | ICML | A*STAR; Beijing Electronic Science and Technology Institute; Nanyang Technologic | B |
| 210 | Stable and Explainable Personality Trait Evaluation in Large Language Models with Internal Activations. | 2026 | ACL | Georgia Institute of Technology; South China University of Technology | B |
| 211 | Stabilizing Equation Learning via Zero-Point Constraints | 2026 | ICML | Central China Normal University; Wuhan University of Technology | B |
| 212 | Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs | 2026 | ICML | Alibaba Group; East China Normal University; Linköping University; The Universit | B |
| 213 | Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities | 2026 | ICML | Fundação Getúlio Vargas (FGV); Independent; MATS; MRM Investments; Saarland Univ | B |
| 214 | Splat Regression Models | 2026 | ICLR | Massachusetts Institute of Technology | B |
| 215 | Spilled Energy in Large Language Models | 2026 | ICLR | Sapienza University of Rome; University of Roma "La Sapienza" | B |
| 216 | SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation. | 2026 | ACL | Nankai University; Microsoft (Finland) | B |
| 217 | Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation. | 2026 | ACL | Jilin University; The University of Tokyo; Key Laboratory of Ancient Chinese Scr | B |
| 218 | Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models | 2026 | ICLR | Department of Computer Science, ETHZ - ETH Zurich; ETH Zurich, Stanford; ETH Zür | B |
| 219 | Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders | 2026 | ICML | Independent; University College London, University of London | B |
| 220 | Sparse and Faithful Local Explanations with Piecewise Linear Surrogates | 2026 | ICML | Sichuan University | B |
| 221 | Sparse Relaxed-Lasso Steering: Automatic Sparse Autoencoder Feature Selection for Precise Image Editing | 2026 | ICML | Chinese Academy of Sciences, Chinese Academy of Sciences; Institute of Software, | B |
| 222 | Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning | 2026 | ICLR | Facebook; Meta; Stanford; University of Oxford | B |
| 223 | Sparse Bayesian Deep Functional Learning with Structured Region Selection | 2026 | ICML | Shanghai University of Finance and Economics; Yale University | B |
| 224 | Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech | 2026 | ICML | College of William & Mary; College of William and Mary; William & Mary | B |
| 225 | Sparse Autoencoders are Topic Models | 2026 | ICML | TUM; Technical University of Munich | B |
| 226 | Sparse Autoencoders Trained on the Same Data Learn Different Features | 2026 | ICLR | EleutherAI | B |
| 227 | Sparling: End-to-End Spatial Concept Learning via Extremely Sparse Activations | 2026 | ICLR | Massachusetts Institute of Technology; Massachussets Institute of Technology; Un | B |
| 228 | Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability | 2026 | ICLR | Charles University; FAR.AI; Imperial College London; MATS Research; independent/ | B |
| 229 | Singular Vectors of Attention Heads Align with Features | 2026 | ICML | Boston University; Boston University, Boston University | B |
| 230 | Simul-COMET: A Quality Metric for Simultaneous Interpretation in Distant Language Pair Considering Word Order Difference. | 2026 | ACL | Seikei University; Nara Institute of Science and Technology | B |
| 231 | Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence | 2026 | ICLR | Berkeley; Independent; University of Chicago | B |
| 232 | Shortcut-Resistant CAM Distillation for Long-Tailed Recognition | 2026 | ICML | Huazhong University of Science and Technology; Nanjing University of Aeronautics | B |
| 233 | Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms | 2026 | ICML | Yonsei University | B |
| 234 | Shapley Neuron Values for Continual Learning: Which Neurons Matter Most? | 2026 | ICML | Aarhus University | B |
| 235 | Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images | 2026 | ICLR | Beijing Jiaotong University; Microsoft Research Asia; Shenzhen University | B |
| 236 | Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language | 2026 | ICLR | Apple; Apple, Carnegie Mellon University; Computer Science and Artificial Intell | B |
| 237 | SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping. | 2026 | ACL | Association for Computational Linguistics | B |
| 238 | SelfReflect: Can LLMs Communicate Their Internal Answer Distribution? | 2026 | ICLR | Apple; Eberhard-Karls-Universität Tübingen; University of Tübingen | B |
| 239 | Self-Explaining Hate Speech Detection with Moral Rationales. | 2026 | ACL | Portland State University | B |
| 240 | Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS. | 2026 | ACL | Queen Mary University of London; Chongqing Key Laboratory of Big Data Intelligen | B |
| 241 | Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNs | 2026 | ICLR | Iowa State University; University of Electronic Science and Technology of China | B |
| 242 | Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection. | 2026 | ACL | VNU University of Science, Vietnam; EDF LAB Singapore Pte Ltd (Singapore) | B |
| 243 | Selective Concept Bottleneck Models Without Predefined Concepts | 2026 | ICML | University Freiburg; University of Freiburg; University of Freiburg, Anthropic A | B |
| 244 | Segment-Level Attribution for Selective Learning of Long Reasoning Traces | 2026 | ICLR | University of Southern California | B |
| 245 | Seeing to Generalize: How Visual Data Corrects Binding Shortcuts | 2026 | ICML | Centro Nacional de Inteligencia Artificial; Pontifical Catholic University of Ch | B |
| 246 | Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems. | 2026 | ACL | State Key Laboratory of Complex System Modeling and Simulation Technology; Chine | B |
| 247 | Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs | 2026 | ICLR | Amazon; Arizona State University; ByteDance Inc.; Microsoft; Pennsylvania State | B |
| 248 | See the Emotion: A Facial Emoji Proxy Modeling for EEG Emotion Recognition | 2026 | ICML | Anhui University; Hefei University of Technology; MBZUAI; University of Electron | B |
| 249 | Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories | 2026 | ICML | ByteDance; Peking University | B |
| 250 | SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity. | 2026 | ACL | Vrije Universiteit Amsterdam; Wageningen University & Research; Association for | B |
| 251 | Scalable and Interpretable Representation Alignment with Ordinal Similarity | 2026 | ICML | Helmholtz AI, Technical University of Munich; Helmholtz Munich / TUM; Helmholtz | B |
| 252 | SafeSeek: Universal Attribution of Safety Circuits in Language Models | 2026 | ICML | Hong Kong Polytechnic University; Intelligent Science & Technology Academy of CA | B |
| 253 | SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering. | 2026 | ACL | Kathmandu University | B |
| 254 | Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework. | 2026 | ACL | Alibaba Group (China); The State Key Laboratory of Blockchain and Data Security, | B |
| 255 | SVD as a Fast Interpretability Method for Transformers | 2026 | ICML | Heidelberg University | B |
| 256 | STEM: Scaling Transformers with Embedding Modules (Authors) | 2026 | Meta AI,Carnegie Mellon University | C | |
| 257 | ST-TGExplainer: Disentangling Stability and Transition Patterns for Temporal GNN Interpretability | 2026 | ICML | Griffith University; Hangzhou Dianzi University; Royal Melbourne Institute of Te | B |
| 258 | SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models. | 2026 | ACL | Northeastern University, Shenyang, China; CAS Key Laboratory of Behavioral Scien | B |
| 259 | SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks. | 2026 | ACL | Oracle | B |
| 260 | SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language Models. | 2026 | ACL | Zhejiang Lab; Zhejiang University of Technology | B |
| 261 | SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models. | 2026 | ACL | Xidian University; University of Science and Technology of China; Xi'an Jiaotong | B |
| 262 | SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models. | 2026 | ACL | University of Maryland, College Park | B |
| 263 | SLASH the Sink: Sharpening Structural Attention Inside LLMs | 2026 | ICML | IGSNRR, Chinese Academy of Sciences, Beijing, China; Shanghai Jiao Tong Universi | B |
| 264 | SIM-CoT: Supervised Implicit Chain-of-Thought | 2026 | ICLR | Fudan University; Microsoft; Nanyang Technological University; Shanghai AI Labor | B |
| 265 | SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning | 2026 | ICML | Edith Cowan University; Huazhong University of Science and Technology; Zhejiang | B |
| 266 | SCOUT: Selective Coupling via Optimal Unbalanced Transport for Interpretable Text Classification. | 2026 | ACL | Accessible Space; Zhejiang University; Hangzhou Pujian Medical Technology Co., L | B |
| 267 | SAVOIR: Learning Social Savoir-Faire via Shapley-based Reward Attribution. | 2026 | ACL | University of Hong Kong; Harbin Institute of Technology; Association for Computa | B |
| 268 | SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs | 2026 | ICLR | Alibaba Group; National University of Singapore; University of Science and Techn | B |
| 269 | SAEs-BrainMap: Unveiling the Emergence of Specialized Concepts in Deep Models via Brain Alignment | 2026 | ICML | BIT; Beijing Institute of Technology; Westlake University | B |
| 270 | SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders | 2026 | ICML | CENTAI; Intesa Sanpaolo AI Research; Polytechnic Institute of Turin; University | B |
| 271 | SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection. | 2026 | ACL | Georgia Institute of Technology; New Jersey Institute of Technology; Chinese Uni | B |
| 272 | SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training | 2026 | ICLR | MIT; Meituan; Peking University | BC |
| 273 | RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing | 2026 | ICML | Brown University; TU Wien | B |
| 274 | Role-Sensitive Neurons: A Neuron-Level Gain Control Mechanism for Confidence Steering. | 2026 | ACL | National Taiwan University | B |
| 275 | Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models | 2026 | ICML | Beijing University of Posts and Telecommunications; Shandong University of Scien | B |
| 276 | Robust Equation Structure Learning with Adaptive Refinement | 2026 | ICLR | Department of Computer Science and Engineering, The Chinese University of Hong K | B |
| 277 | Rhetorical Questions in LLM Representations: A Linear Probing Study. | 2026 | ACL | Independent Age; University of Cincinnati; Amazon | B |
| 278 | Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models. | 2026 | ACL | Ministry of Education | B |
| 279 | Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics | 2026 | ICML | Centrale Supélec; CentraleSupélec; IRT Saint Exupery & Mila; IRT Saint Exupéry & | B |
| 280 | Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis | 2026 | ICML | Wuhan University | B |
| 281 | Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity | 2026 | ICLR | CENIA; Centro Nacional de Inteligencia Artificial; Pontificia Universidad Catoli | B |
| 282 | Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry | 2026 | ICLR | PAII Inc.; Ping An Technology; Pingan Group; Pingan Technology; University of Co | BC |
| 283 | Rethinking LLM Reasoning: From Explicit Trajectories to Latent Representations | 2026 | Harbin Institute of Technology | C | |
| 284 | Responsible Text-to-Image Diffusion: Interpretable and Linearly Controllable Semantics for Fair and Safe Generation | 2026 | ICML | Kyungpook National University; Queen's University | B |
| 285 | Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection | 2026 | ICLR | University of Hong Kong | B |
| 286 | Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal Transport | 2026 | ICLR | University of California San Diego; University of California, San Diego | B |
| 287 | RepoShapley: Shapley-Enhanced Context Filtering for Repository-Level Code Completion. | 2026 | ACL | Chinese University of Hong Kong, Shenzhen; ♯Shenzhen Future Network of Intellige | B |
| 288 | Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics | 2026 | ICML | Johns Hopkins University; University of Amsterdam | B |
| 289 | Reinforcement Learning Towards Broadly and Persistently Beneficial Models | 2026 | OpenAI | C | |
| 290 | Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs | 2026 | ICLR | Electronic Engineering, Tsinghua University, Tsinghua University; Tsinghua Unive | B |
| 291 | Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies. | 2026 | ACL | Chinese University of Hong Kong; Nara Institute of Science | B |
| 292 | Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models | 2026 | ICLR | State University of New York at Stony Brook; Stony Brook; Stony Brook University | B |
| 293 | Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought | 2026 | ICML | Cornell University; Goodfire; Goodfire AI; Harvard University; Harvard Universit | B |
| 294 | Reasoning Fails Where Step Flow Breaks | 2026 | Shanghai Jiao Tong University / Fuzhou University / Jilin University | C | |
| 295 | Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing. | 2026 | ACL | Shanghai Jiao Tong University; The University of Sydney; Shenzhen Campus of Sun | B |
| 296 | Real-Time Visual Attribution Streaming in Thinking Model | 2026 | ICML | Amazon; Yonsei University; Yonsei University, Stanford University | B |
| 297 | Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting | 2026 | ICML | Fraunhofer FIT; Ludwig-Maximilians-Universität München; University of Maryland, | B |
| 298 | Rashomon Sets of Falling Trees | 2026 | ICML | Department of Computer Science, Duke University; Duke; Duke University; Universi | B |
| 299 | RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs? | 2026 | ICLR | Allen Institute for AI; Berkeley; Independent; University of California, Berkele | B |
| 300 | RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering. | 2026 | ACL | Tongji University; IGI Global Scientific Publishing (United States) | B |
| 301 | RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models | 2026 | ICLR | Seoul National University | B |
| 302 | REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control. | 2026 | ACL | National University of Singapore | B |
| 303 | RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited Data | 2026 | ICML | Forschungszentrum Juelich GmbH; Forschungszentrum Jülich, LMU Munich, MCML | B |
| 304 | RAP-ID: Mechanistic Prompt Injection Detection via Impostor Behavior Analysis. | 2026 | ACL | SK Innovation (South Korea) | B |
| 305 | RAG-KT: Cross-platform Explainable Knowledge Tracing with Multi-view Fusion Retrieval Generation. | 2026 | ACL | Inner Mongolia University | B |
| 306 | R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling. | 2026 | ACL | Shanghai Key Laboratory of Trustworthy Computing; East China Normal University; | B |
| 307 | Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects | 2026 | ICML | Hanyang University; Hanyang Universty | B |
| 308 | Query Circuits: Explaining How Language Models Answer User Prompts | 2026 | ICML | University of Oxford; University of Oxford / Martian | B |
| 309 | Quasi-Monte Carlo Methods Enable Extremely Low-Dimensional Deep Generative Models | 2026 | ICLR | Duke University; New York University | B |
| 310 | Quantifying LLM Attention-Head Stability: Implications for Circuit Universality | 2026 | ICML | INRIA; MILA, Quebec, Canada; McGill University; McGill University, McGill Univer | B |
| 311 | Quantifying Cross-Attention Interaction in Transformers for Interpreting TCR-pMHC Binding | 2026 | ICLR | Tulane University; Tulane University School of Medicine | B |
| 312 | Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness. | 2026 | ACL | University of Southern California; University of California, Los Angeles | B |
| 313 | Provably Explaining Neural Additive Models | 2026 | ICLR | Hebrew University of Jerusalem; Masaryk University; Technical University of Muni | B |
| 314 | Prototype-Grounded Concept Models for Verifiable Concept Alignment | 2026 | ICML | IBM Research; KU Leuven | B |
| 315 | Prototype Transformer: Towards Language Model Architectures Interpretable by Design | 2026 | ICML | Relational Intelligence; TU Wien; TU Wien and University of Oxford; Technische U | B |
| 316 | ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting | 2026 | ICLR | Alibaba Group; Renmin University of China | B |
| 317 | Protein Circuit Tracing via Cross-layer Transcoders | 2026 | ICML | Georgia Institute of Technology; Georgia Tech | B |
| 318 | Propaganda AI: An Analysis of Semantic Divergence in Large Language Models | 2026 | ICLR | Singapore Management University | B |
| 319 | Prompt Injection as Role Confusion | 2026 | ICML | Massachusetts Institute of Technology; NBC News | B |
| 320 | Profiling the Irrational Agent: Cognitive Modeling of LLM Behaviors in Sequential Jailbreaks | 2026 | ICML | Chinese Academy of Science; Chinese Academy of Sciences; Institute of Informatio | B |
| 321 | Probing the Safety Robustness of LLMs in Latent Space. | 2026 | ACL | Tsinghua University; Shanghai Artificial Intelligence Laboratory; Fudan Universi | B |
| 322 | Probing the Plasticity and Correlation of LLM Value Systems: LLM Value Rankings are Not Stable. | 2026 | ACL | Hong Kong University of Science and Technology (Guangzhou) | B |
| 323 | Probing the Inductive Bias of Neural Networks through Learning Random Cellular Automata | 2026 | ICML | Johannes-Gutenberg Universität Mainz; University of Mainz | B |
| 324 | Probing the Geometry of Diffusion Models with the String Method | 2026 | ICML | CFM; Capital Fund Management; New York University | B |
| 325 | Probing for Reading Times. | 2026 | ACL | ETH Zurich; Toyota Technological Institute; University College London | B |
| 326 | Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups. | 2026 | ACL | Politecnico di Milano; University of Science and Technology of China | B |
| 327 | Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing. | 2026 | ACL | University of Macau | B |
| 328 | Probing Rotary Position Embeddings through Frequency Entropy | 2026 | ICLR | Ehime University; NTT; NTT Human Informatics Laboratories; NTT corporation | B |
| 329 | Probing RLVR Training Instability through the Lens of Objective-Level Hacking | 2026 | ICML | Alibaba; Alibaba Group; Peking University; Renmin University of China; Tsinghua | B |
| 330 | Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation. | 2026 | ACL | Johns Hopkins University; Chinese University of Hong Kong; University of Chicago | B |
| 331 | Probing Cross-modal Information Hubs in Audio-Visual LLMs | 2026 | ICML | Chung-Ang University; KAIST; KAIST, Korea Advanced Institute of Science & Techno | B |
| 332 | Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio. | 2026 | ACL | Chinese University of Hong Kong; Chinese University of Hong Kong, Shenzhen; Stan | B |
| 333 | ProbeLLM: Automating Principled Diagnosis of LLM Failures | 2026 | ICML | IBM Research; LMU Munich; LMU Munich, MCML; MIT; Massachusetts Institute of Tech | B |
| 334 | Probabilistically-routed Bayesian Additive Spanning Trees for Learning on Constrained Domains | 2026 | ICML | Eli Lilly and Company; Florida State University; Medical University of South Car | B |
| 335 | ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs. | 2026 | ACL | Peking University; Key Laboratory of High Confidence Software Technologies, Mini | B |
| 336 | ProConMV: Provenance-Enabled Conceptual Framework for Interpretable Multi-View Diabetic Retinopathy Diagnosis | 2026 | ICML | Harbin Institute of Technology; Shenzhen University; The University of Nottingha | B |
| 337 | Priors in time: Missing inductive biases for language model interpretability | 2026 | ICLR | Bau Lab; Boston University; Harvard University; Kempner Institute, Harvard Unive | B |
| 338 | Priority-Aware Shapley Value | 2026 | ICML | CMU, Carnegie Mellon University; Carnegie Mellon University; Ohio State Universi | B |
| 339 | Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization. | 2026 | ACL | McGill University; Mila - Quebec AI Institute; Mohamed bin Zayed University of A | B |
| 340 | Precise and Interpretable Editing of Code Knowledge in Large Language Models | 2026 | ICLR | BIFOLD & TU Berlin; Heidelberg College; Heidelberg University; Ruprecht-Karls-Un | B |
| 341 | Pre-training Limited Memory Language Models with Internal and External Knowledge | 2026 | ICLR | Cornell University; Department of Computer Science, Cornell University | BC |
| 342 | Position: Your VLM May Not Be Thinking with Interleaved Images | 2026 | ICML | Fudan University | B |
| 343 | Position: When AI Decides Who Gets an Organ: Multi-Agentic AI Systems in Transplant Medicine Risk Amplifying Disparities Without Targeted Explainability and Deployment Strategies | 2026 | ICML | University Health Network; University of California, Irvine; University of Toron | B |
| 344 | Position: Use Sparse Autoencoders to Discover Unknowns | 2026 | ICML | Cornell; Cornell Tech; Cornell University; UC Berkeley | B |
| 345 | Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces! | 2026 | ICML | Arizona State University; Yale | B |
| 346 | Position: Multi-Agent Explainability Needs Contracts Before Methods | 2026 | ICML | Dartmouth; Dartmouth College | B |
| 347 | Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance | 2026 | ICML | Technical University of Munich; University of Exeter; University of Florida; Uni | B |
| 348 | Position: Interpretability in Deep Time Series Models Demands Semantic Alignment | 2026 | ICML | IBM Research; KU Leuven; SUPSI - University of Applied Sciences Southern Switzer | B |
| 349 | Position: Interpretability Can Be Actionable | 2026 | ICML | Google DeepMind; Harvard University; Kempner Institute, Harvard University; Mila | B |
| 350 | Position: In Defense of Information Leakage in Concept-based Models | 2026 | ICML | University of Oxford | B |
| 351 | Position: Genomic Model Research Must Move Beyond Anecdotal Evaluation of Interpretability Methods | 2026 | ICML | Living Systems Institute, University of Exeter; University of Electronic Science | B |
| 352 | Position: Explanation Stability Is a Property of the Model–Method Pair, Not the Model | 2026 | ICML | Duke-NUS medical School; Singapore Health Services (SingHealth) | B |
| 353 | Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods | 2026 | ICML | Bosch Research North America; Duke; Google Research; Harvard; Harvard University | B |
| 354 | Position: Don't Just "Fix it in Post'': A Science of AI Must Study Learning Dynamics | 2026 | ICML | EleutherAI; Humans&/CMU; Kempner Institute, Harvard University; Max Planck Insti | B |
| 355 | Position: Causality is Key for Interpretability Claims to Generalise | 2026 | ICML | Boston University; Cold Spring Harbor Laboratory; Max Planck Institute for Intel | B |
| 356 | Position: Behavioral Systems Require Behavioral Tests | 2026 | ICML | Dartmouth College; MIT Media Lab; Massachusetts Institute of Technology (MIT) | B |
| 357 | Position: Accountable Deployment of Agentic AI Demands Layered, System-Level Interpretability | 2026 | ICML | Dalhousie University; Ontario Tech University; Vector Institute; Vector Institut | B |
| 358 | PolySHAP: Extending KernelSHAP with Interaction-Informed Polynomial Regression | 2026 | ICLR | Bielefeld University; Claremont McKenna College; New York University | B |
| 359 | PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding | 2026 | ICML | The Cyprus Institute; University of Athens & Archimedes AI, Athena RC; Universit | B |
| 360 | Persona Features Control Emergent Misalignment | 2026 | ICLR | Independent Researcher; Massachusetts Institute of Technology; OpenAI; Stanford | B |
| 361 | PerfCoder: Large Language Models for Interpretable Code Performance Optimization. | 2026 | ACL | University of Alberta; University of Victoria; Huawei Technologies Ltd., Toronto | B |
| 362 | Patterning: The Dual of Interpretability | 2026 | ICML | The University of Melbourne; Timaeus | B |
| 363 | Patronus: Interpretable Diffusion Models with Prototypes | 2026 | ICLR | Technical University of Denmark | B |
| 364 | PathwayLLM: Explainable Clinical Trajectory Modeling with Structured Pathways for Sepsis Prediction | 2026 | ICML | The Second Affiliated Hospital of Zhejiang Chinese Medical University; Xiamen Un | B |
| 365 | Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN | 2026 | ICLR | FAR AI; FAR.AI | B |
| 366 | Partial Soft-Matching Distance For Neural Representational Comparison With Partial Unit Correspondence | 2026 | ICLR | New York University; University of California, San Diego | B |
| 367 | Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation. | 2026 | ACL | Technische Universität Berlin; German Research Centre for Artificial Intelligenc | B |
| 368 | Paradigm Shift of GNN Explainer from Label Space to Prototypical Representation Space | 2026 | ICLR | Central South University; Griffith University; Hong Kong Polytechnic University; | B |
| 369 | PV-SQL: Synergizing Database Probing and Rule-based Verification for Text-to-SQL Agents. | 2026 | ACL | Purdue University West Lafayette | B |
| 370 | PROBE: PROcess-Based BEnchmark for Hallucination Detection. | 2026 | ACL | NVIDIA; Chinese University of Hong Kong | B |
| 371 | PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues. | 2026 | ACL | Indian Institute of Technology Patna; Indian Institute of Technology Jodhpur | B |
| 372 | PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations. | 2026 | ACL | Hong Kong University of Science and Technology (Guangzhou); NYU Shanghai; Dongbe | B |
| 373 | PR-XAI: PageRank-Based Feature Attribution for Transformers. | 2026 | ACL | Simon Fraser University; University of California, Berkeley | B |
| 374 | PINNfluence: Interpreting PINNs through Influence Functions | 2026 | ICML | Fraunhofer HHI; Fraunhofer HHI, Berlin; Fraunhofer HHI, Einsteinufer 37, 10587 B | B |
| 375 | PGRF-Net: A Prototype-Guided Relational Fusion Network for Diagnostic Multivariate Time-Series Anomaly Detection | 2026 | ICLR | Yonsei University | B |
| 376 | PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes. | 2026 | ACL | Beijing Institute of Technology; Beihang University; Baidu (China) | B |
| 377 | PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans | 2026 | ICLR | Auburn University; Carnegie Mellon University; None | B |
| 378 | PAR: Training-Free Positional Perturbation and Attention Recycling for Faithful OCR. | 2026 | ACL | Shanghai Jiao Tong University; Higher Education Commission; University of Hong K | B |
| 379 | Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets | 2026 | ICML | Amazon; MATS; Redwood Research | B |
| 380 | Origo: Interpretable Multi-physics PDE Foundation Model through Neural Operator Splitting | 2026 | ICML | Beijing University of Post and Telecommunications; NCEPU; UIC; University of Ken | B |
| 381 | Optimal Transport Group Counterfactual Explanations | 2026 | ICML | LMU; Technical University of Madrid; Universidad Politécnica de Madrid; Universi | B |
| 382 | Open Sourcing Monitorability Evaluations | 2026 | Lab post (OpenAI) | OpenAI | A |
| 383 | One Probe Won’t Catch Them All: Towards Targeted Deception Detection | 2026 | ICML | Decode Research; Equivariant labs; LASR Labs; Lambda; UK AISI | B |
| 384 | One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models | 2026 | ICML | Stanford University | B |
| 385 | One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework. | 2026 | ACL | Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University; Jili | B |
| 386 | Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models. | 2026 | ACL | Qatar Computing Research Institute, HBKU, Doha, Qatar; Association for Computati | B |
| 387 | On the Variability of Concept Activation Vectors | 2026 | ICML | Julius-Maximilians-Universität Würzburg - CAIDAS; University of Würzburg | B |
| 388 | On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders | 2026 | ICML | Columbia University; Stanford; Stanford University | B |
| 389 | On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy | 2026 | ICLR | MIT; Peking University | B |
| 390 | On the Accuracy of Newton Step and Influence Function Data Attributions | 2026 | ICML | Massachusetts Institute of Technology | B |
| 391 | On The Geometry and Topology of Representations: the Manifolds of Modular Addition | 2026 | ICLR | Leiden University, Dept. of Mathematics, Leiden University; McGill University; M | B |
| 392 | On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs | 2026 | ICML | Apple; Apple Inc.; Harvard University, Apple | B |
| 393 | On Predictability of Reinforcement Learning Dynamics for Large Language Models | 2026 | ICLR | The Hong Kong University of Science and Technology; Tsinghua University, Tsinghu | B |
| 394 | On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM Agents | 2026 | CUHK (James Cheng's Group) | C | |
| 395 | OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation. | 2026 | ACL | Northeastern University; University of Southern California; Stony Brook Universi | B |
| 396 | ODASim: Ordered, Distinctive and Absolute Semantic Similarity for Code Explanation Evaluation. | 2026 | ACL | IBM Research - India | B |
| 397 | OCP: Outlier-Centric Probing for Dynamic Structured Pruning of LLMs. | 2026 | ACL | The Hong Kong University of Science and Technology (Guangzhou); National Univers | B |
| 398 | Nonparametric Data Attribution for Diffusion Models | 2026 | ICML | National University of Singapore; SMU, Singapore; Sea AI Lab; Tencent | B |
| 399 | Neuronal Insights into LLM Attacks: Targeted Neuron Tuning for Precise and Robust Vulnerability Patching. | 2026 | ACL | Tianjin University of Science and Technology; Tianjin University of Technology | B |
| 400 | Neuron-Level Analysis of Cultural Understanding in Large Language Models | 2026 | ICLR | The University of Tokyo; The University of Tokyo / Riken; University of Liverpoo | B |
| 401 | Neuron-Aware Active Few-Shot Learning for LLMs. | 2026 | ACL | University of Pittsburgh | B |
| 402 | Neuro-Fuzzy Concept Learning for Interpretable Large Multimodal Models | 2026 | ICML | Indian Institute of Technology Indore; Indian Institute of Technology, Indore | B |
| 403 | Neural–Evolutionary Symbolic Regression with Global Constraints: Constraint-Aware Decoding and Reward Shaping | 2026 | ICML | Beihang University; Beijing University of Aeronautics and Astronautics; Columbia | B |
| 404 | Neural+Symbolic Approaches for Interpretable Actor-Critic Reinforcement Learning | 2026 | ICLR | Maincode; Monash University | B |
| 405 | Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings | 2026 | ICML | German Research Center for AI; Max Planck Institute for Informatics; Technische | B |
| 406 | NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-Neurons. | 2026 | ACL | State Key Laboratory of General Artificial Intelligence; Peking University | B |
| 407 | NeuRAG: End-to-End Neural Knowledge Augmentation via Hyper-Neurons. | 2026 | ACL | Nankai University | B |
| 408 | Negative Pre-activations Differentiate Syntax | 2026 | ICLR | Computer Science and Artificial Intelligence Laboratory, Electrical Engineering | B |
| 409 | Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations (Authors) | 2026 | Anthropic | C | |
| 410 | Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations | 2026 | Lab post (Anthropic) | Anthropic | A |
| 411 | Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences | 2026 | ICLR | DeepMind; ENS Paris-Saclay; EPFL; Harvard University; Massachusetts Institute of | B |
| 412 | NSF-CoT: Neuro-Symbolic Formal Verification of Chain-of-Thought Faithfulness in Contextual Question Answering. | 2026 | ACL | The University of Texas at San Antonio | B |
| 413 | NIMO: a Nonlinear Interpretable MOdel | 2026 | ICLR | University of Basel | B |
| 414 | NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels | 2026 | ICML | National University of Singapore; Tsinghua University; University of Science and | B |
| 415 | NEAT: Neuron-Based Early Exit for Large Reasoning Models. | 2026 | ACL | Northeastern University | B |
| 416 | Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies | 2026 | ICLR | City University of Hong Kong; Huawei Technologies Ltd. | B |
| 417 | Multilingual Routing in Mixture-of-Experts | 2026 | ICLR | Fudan University; Google & University of California, Los Angeles; University of | B |
| 418 | Multi-scale Explainer for Graph Neural Networks | 2026 | ICML | Shanxi University; Southeast University; Suzhou University | B |
| 419 | Multi-ReduNet: Interpretable Class-Wise Decomposition of ReduNet | 2026 | ICLR | National University of Singapore | B |
| 420 | Motion Attribution for Video Generation | 2026 | ICML | MIT; NVIDIA; NVIDIA; UMich; Princeton University; Princeton University NVIDIA; U | B |
| 421 | More Than What Was Chosen: LLM-based Explainable Recommendation Beyond Noisy User Preferences | 2026 | ICLR | KAIST; SK Telecom; SK Telecom, KAIST | B |
| 422 | More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing | 2026 | ICML | City University of Hong Kong; University of Science and Technology of China | B |
| 423 | Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning | 2026 | ICML | Harvard; Harvard University | B |
| 424 | Modeling Hierarchical Thinking in Large Reasoning Models | 2026 | ICML | Sharif University of Technology; University of California, Riverside | B |
| 425 | Mixture of Concept Bottleneck Experts | 2026 | ICML | Department of Information Engineering and Mathematical Sciences, University of S | B |
| 426 | Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization | 2026 | ICLR | EPFL; EPFL - EPF Lausanne; Massachusetts Institute of Technology | B |
| 427 | Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling | 2026 | ICML | KAIST; KAIST AI; KRAFTON; Korea Advanced Institute of Science & Technology | B |
| 428 | Missingness Bias Calibration in Feature Attribution Explanations | 2026 | ICLR | University of Pennsylvania; University of Texas at Austin | B |
| 429 | MidSteer: Optimal Affine Framework for Steering Generative Models | 2026 | ICML | Huawei London; MVP Lab; Quantum Light; Queen Mary University of London; Universi | B |
| 430 | MicroC-KT: Modeling Community Effect via Learning Micro-Environment for Evidence-Grounded Explainable Knowledge Tracing. | 2026 | ACL | Inner Mongolia University; Jilin University | B |
| 431 | MetaOthello: A Controlled Study of Multiple World Models in Transformers | 2026 | ICML | University of Michigan - Ann Arbor; University of Vermont | B |
| 432 | MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers | 2026 | ICML | Apple; University of Texas at Austin | B |
| 433 | Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMs. | 2026 | ACL | Kyoto University; University of Tokyo; National Institute of Informatics; RIKEN | B |
| 434 | Medical Interpretability and Knowledge Maps of Large Language Models | 2026 | ICLR | Lumos AI | B |
| 435 | MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning. | 2026 | ACL | National University of Singapore; Chinese University of Hong Kong; Hunan Univers | B |
| 436 | MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis. | 2026 | ACL | Stanford University; Shenzhen University; Renmin University of China; Xi'an Jiao | B |
| 437 | Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation | 2026 | ICML | Wilfrid Laurier University | B |
| 438 | Mechanistic Interpretability of Text-to-Image Diffusion Models via Cross-Attention Interventions. | 2026 | ACL | School of Computer Science; University of Oklahoma | B |
| 439 | Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy. | 2026 | ACL | Sharif University of Technology | B |
| 440 | Mechanistic Interpretability as Statistical Estimation: A Variance Analysis | 2026 | ICML | CNRS; LIG / Université Grenoble Alpes; Université Grenoble Alpes | B |
| 441 | Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders. | 2026 | ACL | Stanford University | B |
| 442 | Mechanistic Insights into Deferred Semantic Drift in LLMs. | 2026 | ACL | Dalian University of Technology; Key Laboratory of Social Computing and Cognitiv | B |
| 443 | Mechanistic Detection and Mitigation of Hallucination in Large Reasoning Models | 2026 | ICLR | Renmin University of China; Renmin University of China, Gaoling School of Artifi | B |
| 444 | Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units;:ICML 2026 | 2026 | Peking University | C | |
| 445 | Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units | 2026 | ICML | Peking University | B |
| 446 | Mechanistic Anomaly Detection via Functional Attribution | 2026 | ICML | University of Melbourne | B |
| 447 | Mechanisms of Introspective Awareness | 2026 | ICML | Anthropic; Constellation Research Center; Massachusetts Institute of Technology | AB |
| 448 | Measuring and Mitigating Post-Hoc Rationalization in Reverse Chain-of-Thought Generation | 2026 | ICML | BOSS Zhipin; BOSS Zhipin Career Science Lab; Peking University; University of El | B |
| 449 | Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos. | 2026 | ACL | Harbin Institute of Technology | B |
| 450 | MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows. | 2026 | ACL | MoE Key Lab of Artificial Intelligence, X-LANCE Lab; Shanghai Jiao Tong Universi | B |
| 451 | Maximum Likelihood Reinforcement Learning | 2026 | Carnegie Mellon University | C | |
| 452 | Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs | 2026 | ICLR | Australian Institute for Machine Learning (AIML); Nanjing University of Science | B |
| 453 | Mask-to-Correct⁺: Leveraging Retriever Diversity for Masking-guided Faithful Fact Correction. | 2026 | ACL | Indian Association for the Cultivation of Science | B |
| 454 | Markovian Transformers for Informative Language Modeling | 2026 | ICLR | Stanford University | B |
| 455 | Mapping Semantic & Syntactic Relationships with Geometric Rotation | 2026 | ICLR | Fuel iX; TELUS Digital | B |
| 456 | Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs | 2026 | ICLR | NAVER AI Lab; Seoul National University | B |
| 457 | Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution | 2026 | ICML | INEEJI; KAIST; KAIST/INEEJI; Korea Advanced Institute of Science and Technology | B |
| 458 | Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing. | 2026 | ACL | University of Delaware | B |
| 459 | MOD-SR: Unifying Multimodal Learning and Direct Optimization with Gradient-Guided Diffusion Model for Symbolic Regression | 2026 | ICML | Shanghai Jiao Tong University; Shanghai Jiaotong University | B |
| 460 | MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models. | 2026 | ACL | University of Science and Technology of China; Beijing Institute for General Art | B |
| 461 | MICLIP: Learning to Interpret Representation in Vision Models | 2026 | ICLR | ShanghaiTech University | B |
| 462 | MCLE-Mol: Empowering LLM with Molecular Comprehension and Low-Cost Continual Evolution for Interpretable Property Prediction. | 2026 | ACL | East China University of Science and Technology | B |
| 463 | MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation | 2026 | ICML | Peking University | B |
| 464 | MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference. | 2026 | ACL | Chung-Ang University; Adobe Research, USA; Adobe Research | B |
| 465 | M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding | 2026 | ICLR | East China Normal University; Hangzhou Medical College; National University of S | B |
| 466 | LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrieval. | 2026 | ACL | University of Colorado System; University of Colorado Boulder; Loyola University | B |
| 467 | Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration | 2026 | ICML | Shanghai Jiao Tong University; Shanghai Jiaotong University; Shanghai artificial | B |
| 468 | Logit Distance Bounds Representational Similarity | 2026 | ICML | Helmholtz AI, Technical University of Munich; IT University of Copenhagen; Max P | B |
| 469 | LogicXGNN: Grounded Logical Rules for Explaining Graph Neural Networks | 2026 | ICLR | McGill University; McGill University, McGill University; University of Toronto; | B |
| 470 | Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models. | 2026 | ACL | University of Hong Kong; Fudan University; LMU Munich; Tsinghua University; Tech | B |
| 471 | Locate then Correct: Debiasing Attention Heads in CLIP | 2026 | ICML | A*STAR; Deakin University; Nanyang Technological University; School of Computer | B |
| 472 | Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis | 2026 | ICLR | Japan Advanced Institute of Science and Technology; RIKEN / Tohoku Univ.; Univer | B |
| 473 | Localizing Memorized Regions in Diffusion Models via Coordinate-Wise Curvature Differences | 2026 | ICML | Hanyang University | B |
| 474 | Local Mechanisms of Compositional Generalization | 2026 | ICML | Apple | B |
| 475 | Linguistic Properties and Model Scale in Brain Encoding: From Small to Compressed Language Models | 2026 | ICML | Amazon; GE HealthCare; Indian Institute of Technology, Delhi; International Inst | B |
| 476 | LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them? | 2026 | ICLR | Amazon; Georgia Institute of Technology | B |
| 477 | Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models | 2026 | ICLR | California Institute of Technology; Caltech; Facebook | B |
| 478 | Lightweight and Interpretable Transformer via Unrolling of Mixed Graph Algorithms for Traffic Forecast | 2026 | ICML | Tsinghua University; Tsinghua University, Tsinghua University; York University | B |
| 479 | Lightweight and Faithful Visual Condition Checking in Behavior Trees via Expert-Regularized Reinforcement Learning. | 2026 | ACL | University of Toronto | B |
| 480 | Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling | 2026 | ICLR | Autodesk; New York University; New York University Tandon School of Engineering; | B |
| 481 | Lending Eyesight to Language Models: Modeling and Probing Human scanpath through Transformer Decoder. | 2026 | ACL | Hong Kong Polytechnic University; Mohamed bin Zayed University of Artificial Int | B |
| 482 | Left–Right Symmetry Breaking in CLIP-Style Vision-Language Models Trained on Synthetic Spatial-Relation Data | 2026 | ICML | Toyota Motor Corporation | B |
| 483 | Learning to Weight Parameters for Training Data Attribution | 2026 | ICLR | EPFL; EPFL - EPF Lausanne; State University of New York at Stony Brook | B |
| 484 | Learning to Interpret Weight Differences in Language Models | 2026 | ICLR | MIT; Massachusetts Institute of Technology | B |
| 485 | Learning multimodal dictionary decompositions with group-sparse autoencoders | 2026 | ICLR | Dolby; Dolby Labs; Georgia Institute of Technology | B |
| 486 | Learning for Highly Faithful Explainability | 2026 | ICLR | Beijing Institute of Technology; Microsoft | B |
| 487 | Learning a Generative Meta-Model of LLM Activations | 2026 | ICML | Anthropic; Electrical Engineering & Computer Science Department; UC Berkeley; Un | B |
| 488 | Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling. | 2026 | ACL | Zhejiang University; State Key Laboratory of Transvascular Implantation Devices | B |
| 489 | Learning Through Dialogue: Engagement and Efficacy Matter More Than Explanations. | 2026 | ACL | Centre for Tactile Internet with Human-in-the-Loop; National University of Singa | B |
| 490 | Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs | 2026 | ICML | AE Studio; Agency Enterprise Studio; Independent; Princeton University | B |
| 491 | Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability | 2026 | ICLR | University of Maryland, College Park | B |
| 492 | Learning Protein Structure-Function Relationships through Knowledge-guided Representation Decomposition | 2026 | ICML | Fudan University; ICT; Shanghai Smart Logic Technology Co., Ltd.; Tsinghua Unive | B |
| 493 | Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies | 2026 | ICLR | ELLIS Institute; MPI for Intelligent Systems; Max Planck Institute for Intellige | B |
| 494 | Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding. | 2026 | ACL | Nanyang Technological University | B |
| 495 | Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement Learning | 2026 | ICML | Huawei Technologies; National University of Singapore; Zhejiang University | B |
| 496 | Learning Explicit Single-Cell Dynamics Using ODE Representations | 2026 | ICLR | Chan Zuckerberg Initiative; IST Austria CZI; ISTA; Institute of Science and Tech | B |
| 497 | Learning Efficient and Interpretable Multi-Agent Communication | 2026 | ICLR | BIGAI; Shandong University | B |
| 498 | Learning Concept Bottleneck Models from Mechanistic Explanations | 2026 | ICLR | Computer Science and Artificial Intelligence Laboratory, Electrical Engineering | B |
| 499 | Learning Coherent Representations: A Topological Approach to Interpretability | 2026 | ICML | NTNU; Norwegian University of Science and Technology | B |
| 500 | Learning AND–OR Templates for Compositional Representation in Art and Design | 2026 | ICLR | Beijing Electronic Science and Technology Institute; Beijing Institute for Gener | B |
| 501 | Learn from A Rationalist: Distilling Intermediate Interpretable Rationales | 2026 | ICML | University of Alberta | B |
| 502 | LatentQA: Teaching LLMs to Decode Activations Into Natural Language | 2026 | ICLR | Transluce, UC Berkeley; UC Berkeley; University of California, Berkeley | B |
| 503 | LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs | 2026 | ICML | McGill University; Meta; Mila (Montreal)/McGill University; Mila - Quebec AI Ins | B |
| 504 | Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts | 2026 | ICLR | Ohio State University, Columbus; Xi'an Jiaotong University | B |
| 505 | Latent Planning Emerges with Scale | 2026 | ICLR | Anthropic; University of Amsterdam | B |
| 506 | Latent Concept Disentanglement in Transformer-based Language Models | 2026 | ICLR | Google; Google Research; USC; University of Oxford; University of Southern Calif | B |
| 507 | LassoFlexNet: a Flexible Neural Architecture for Tabular Data | 2026 | ICML | Borealis AI; NACE.AI; Qube Research and Technologies; RBC Borealis AI | B |
| 508 | Large Vision-Language Models Get Lost in Attention | 2026 | ICML | Alibaba Group; Beijing Academy of Artificial Intelligence(BAAl) ; Beijing Univer | B |
| 509 | Large Language Model Agents Are Not Always Faithful Self-Evolvers | 2026 | ICML | Harbin Institute of Technology; Singapore Management University | B |
| 510 | Language Models are Injective and Hence Invertible | 2026 | ICLR | EPFL; EPFL - EPF Lausanne; National and Kapodistrian University of Athens; Sapie | BC |
| 511 | Language Models Represent and Transform Concepts with Shared Geometry | 2026 | Georgia Tech | C | |
| 512 | Language Model Circuits Are Sparse in the Neuron Basis | 2026 | ICML | Google; Stanford University; Transluce / MIT; University of California, Berkeley | B |
| 513 | Label-Free Mitigation of Spurious Correlations in VLMs using Sparse Autoencoders | 2026 | ICLR | State University of New York at Buffalo; State University of New York, Buffalo; | B |
| 514 | Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference. | 2026 | ACL | UCLouvain; F.R.S.-FNRS | B |
| 515 | LaVCa: LLM-assisted Visual Cortex Captioning | 2026 | ICLR | Fukui Computer Holdings,Inc; Nagoya Institute of Technology; University of Osaka | B |
| 516 | LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking | 2026 | ICLR | Baidu; Baidu Inc; Institute of automation, Chinese academy of science, Chinese A | B |
| 517 | LLMs Process Lists With General Filter Heads | 2026 | ICLR | Northeastern University | B |
| 518 | LLMs Lean on Priors, Not Programming Language Semantics | 2026 | ICML | Cisco; Google; University of Texas at Austin | B |
| 519 | LLM-induced Rationales for More Compact Explainable Style Classification Models. | 2026 | ACL | Bentley University; Metropolitan University | B |
| 520 | LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines. | 2026 | ACL | Stanford University; Independent Age; University of California, Irvine; Chinese | B |
| 521 | LLM Self-Recognition: Steering and Retrieving Activation Signatures | 2026 | ICML | FU Berlin; Freie Universitat Berlin; Freie Universität Berlin | B |
| 522 | LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals. | 2026 | ACL | Technion – Israel Institute of Technology | B |
| 523 | LEAF: Towards Lightweight Explainable Hateful Video Detection via Self-Grounding CoT Guided Stage-Wise Distillation. | 2026 | ACL | University of Electronic Science and Technology of China; Sichuan Youjianzhihui | B |
| 524 | LDEDE: LRP-Driven Efficient Detection and Editing Framework for LLM Privacy Neurons. | 2026 | ACL | Information Engineering University, Zhengzhou, 450001, Henan, China; State Key L | B |
| 525 | LAFaCT: Attribution-based Localization and Focused Sequential Analysis of Fact-Critical Tokens for Hallucination Detection. | 2026 | ACL | University of Science and Technology of China | B |
| 526 | Knowing the Unknown: Interpretable Open-World Object Detection via Concept Decomposition Model | 2026 | ICML | Northwest Polytechnical University; Northwest Polytechnical University Xi'an; No | B |
| 527 | Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement | 2026 | ICML | George Mason University | B |
| 528 | KODA: Contrastive Representation Comparison and Alignment for Vision-Language Foundation Models | 2026 | ICML | Chinese University of Hong Kong (CUHK); The Chinese University of Hong Kong | B |
| 529 | KANO: Kolmogorov-Arnold Neural Operator | 2026 | ICLR | Caltech; MIT; Massachusetts Institute of Technology; UC Santa Barbara; Universit | B |
| 530 | Joint Distribution–Informed Shapley Values for Sparse Counterfactual Explanations | 2026 | ICLR | King/Microsoft; Technical University of Denmark; University of Copenhagen | B |
| 531 | JX4MEI: Multimodal Semantically-Enhanced LLM for Joint Multimodal Emotion-Intent Explanation and Classification. | 2026 | ACL | Northeastern University | B |
| 532 | Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility | 2026 | ICLR | Brown University; DeepMind; Johns Hopkins University | B |
| 533 | Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models | 2026 | ICML | TU Dortmund University / Lamarr Institute; University of Tübingen, TU Dortmund | B |
| 534 | Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training | 2026 | University of Minnesota / Peking University / Amazon | C | |
| 535 | Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers. | 2026 | ACL | Department of Computer Science University of Texas at Dallas | B |
| 536 | Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization. | 2026 | ACL | University of North Carolina at Chapel Hill; University of North Carolina Health | B |
| 537 | Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective. | 2026 | ACL | National Key Laboratory for Multimedia Information Processing; Peking University | B |
| 538 | Investigating Counterfactual Unfairness in LLMs towards Identities through Humor. | 2026 | ACL | Yonsei University; Korea Advanced Institute of Science and Technology; Seoul Nat | B |
| 539 | Introspection Adapters: Training LLMs to Report Their Learned Behaviors | 2026 | ICML | Anthropic; Anthropic Fellows | B |
| 540 | Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry | 2026 | ICLR | Brown University; Harvard University; Stanford; University of Michigan; York Uni | B |
| 541 | Interpreting and Steering State-Space Models via Activation Subspace Bottlenecks | 2026 | ICML | AppViewX; Indian Institute of Technology Hyderabad; Microsoft Research | B |
| 542 | Interpreting Physics in Video World Models | 2026 | ICML | AMI Labs; Advanced Machine Intelligence; Brown University; FAR.AI; Goodfire AI; | B |
| 543 | Interpreting Genomic Language Models using Sparse Autoencoders | 2026 | ICML | University of Pennsylvania; University of Pennsylvania, University of Pennsylvan | B |
| 544 | Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation. | 2026 | ACL | Arizona State University | B |
| 545 | Interpretable Semantic Gradients in SSD: A PCA Sweep Approach and a Case Study on AI Discourse. | 2026 | ACL | Warsaw University of Technology | B |
| 546 | Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation | 2026 | ICML | Technical University of Munich; Technische Universität München | B |
| 547 | Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation. | 2026 | ACL | Beijing University of Posts and Telecommunications | B |
| 548 | Interpretable Neural ODEs for Gene Regulatory Network Discovery under Perturbations | 2026 | ICML | Columbia University; Columbia University & New York Genome Center; Genentech / C | B |
| 549 | Interpretable Functional Koopman Learning with Non-Markovian Closure for Spatiotemporal Systems | 2026 | ICML | Fudan University | B |
| 550 | Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit | 2026 | ICML | Google; Google DeepMind; Massachusetts Institute of Technology; Stanford Univers | BC |
| 551 | Interpretable Coreference Resolution Evaluation Using Explicit Semantics. | 2026 | ACL | Sapienza University of Rome; Babelscape | B |
| 552 | Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning | 2026 | ICLR | MPI Informatics; Max Planck Institute for Informatics; Saarland Informatics Camp | B |
| 553 | Interpretability from the Ground Up: Stakeholder-Centric Design of Automated Scoring in Educational Assessments. | 2026 | ACL | Stanford University | B |
| 554 | Interpretability and Generalization Bounds for Learning Spatial Physics | 2026 | ICML | Google; OpenAI; Sandia National Laboratories | B |
| 555 | Interpretability Transfer from Language to Vision via Sparse Autoencoders | 2026 | ICML | IIT Kanpur; Lambda; Samsung; University of Bath | B |
| 556 | Interpretability Driven Evolutionary Approach for the Design of Biological Sequences | 2026 | ICML | Northwestern University | B |
| 557 | Internal Planning in Language Models: Characterizing Horizon and Branch Awareness | 2026 | ICLR | Carnegie Mellon University | B |
| 558 | InteracSPARQL : An Interactive System for SPARQL Query Refinement Using Natural Language Explanations. | 2026 | ACL | University of Waterloo | B |
| 559 | Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers | 2026 | ICML | University of Delaware; University of Virginia | B |
| 560 | Information Flow Reveals When to Trust Language Models | 2026 | ICML | HKUST(GZ); Jilin University; The Hong Kong University of Science and Technology | B |
| 561 | InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition :ICML 2026 | 2026 | ByteDance | C | |
| 562 | Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback | 2026 | ICML | AstraZeneca; Google DeepMind / University of Cambridge; University of Cambridge; | B |
| 563 | Influence Dynamics and Stagewise Data Attribution | 2026 | ICLR | Independent; Timaeus; University College London, University of London; Universit | B |
| 564 | Inferring the Invisible: Neuro-Symbolic Rule Discovery for Missing Value Imputation | 2026 | ICLR | Texas A&M University; The Chinese University of Hong Kong; The Chinese Universit | B |
| 565 | Induction Heads Interpolate N-Grams | 2026 | ICML | EPFL; Indian Institute of Technology, Madras | B |
| 566 | In-context learning of representations can be explained by induction circuits | 2026 | ICLR | Northeastern University | B |
| 567 | In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations | 2026 | ICLR | MPI-SWS; Max Planck Institute for Software Systems; Max Planck Institute for Sof | B |
| 568 | Improving Adversarial Robustness of Attribution via Implicit Regularization | 2026 | ICML | Brown University; KTH; KTH Royal Institute of Technology | B |
| 569 | Imagination Helps Visual Reasoning, But Not Yet in Latent Space | 2026 | ICML | Beijing Jiaotong University; Institute of automation, Chinese academy of science | B |
| 570 | ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks | 2026 | ICLR | Academia Sinica; Beever AI; CCHUML; Center for Intelligent Multidimensional Data | B |
| 571 | INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs. | 2026 | ACL | State Key Laboratory of AI Safety; Chinese Academy of Sciences; School of Advanc | B |
| 572 | IDEA: An Interpretable and Editable Decision-Making Framework for LLMs via Verbal-to-Numeric Calibration. | 2026 | ACL | The Hong Kong University of Science and Technology (Guangzhou); Huawei Technolog | B |
| 573 | ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical Coding. | 2026 | ACL | Pennsylvania State University; Stony Brook University | B |
| 574 | I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data? | 2026 | ICLR | Adelaide University; The University of Adelaide; The University of Melbourne; Un | B |
| 575 | How do LLMs Compute Verbal Confidence? | 2026 | ICML | Google; Google DeepMind; Google DeepMind / University of Cambridge | B |
| 576 | How can embedding models bind concepts? | 2026 | ICML | KAIST & Cortiq; University of Tübingen | B |
| 577 | How Transformers Represent Hierarchies: A Local-to-Global Mechanism | 2026 | ICML | Yale University | B |
| 578 | How Transformers Learn Causal Structures In-Context: Explainable Mechanism Meets Theoretical Guarantee | 2026 | ICLR | Yale; Yale University | B |
| 579 | How To Open the Black Box: Modern Models for Mechanistic Interpretability | 2026 | ICLR | Patsnap; University of British Columbia | B |
| 580 | How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess | 2026 | ICML | University of California, San Diego | B |
| 581 | How Language Models Process Negation | 2026 | ICML | USC; USC Information Sciences Institute; University of Southern California | B |
| 582 | How Hard Is Science? | 2026 | ICML | University of Cambridge | B |
| 583 | How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning | 2026 | ICML | Saarland University; Universität des Saarlandes | B |
| 584 | How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning | 2026 | ICML | IBM Research; WeChat AI, Tencent; WeChat AI, Tencent Inc. | B |
| 585 | How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability | 2026 | ICLR | Department of Computer Sciences, University of Wisconsin - Madison; University o | B |
| 586 | How Do Language Models Speak Languages? A Case Study on Unintended Code-Switching | 2026 | ICML | Alibaba Cloud; Alibaba Group; Beijing Automobile Works; UniTTEC Co. Ltd.; Univer | B |
| 587 | How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study. | 2026 | ACL | Beijing Institute of Technology; Key Laboratory of Computing Power Network and I | B |
| 588 | Hinge Regression Tree: A Newton Method for Oblique Regression Tree Splitting | 2026 | ICLR | Harbin Institute of Technology; Harbin Institute of Technology, Shenzhen | B |
| 589 | Hierarchical Concept-based Interpretable Models | 2026 | ICLR | University of Cambridge | B |
| 590 | Hierarchical Causal Abduction: A Foundation Framework for Explainable Model Predictive Control | 2026 | ICML | Technische Universität Chemnitz | B |
| 591 | Hide&Seek: Learning to Explain in an End-to-End Differentiable Network | 2026 | ICML | University of Technology Sydney | B |
| 592 | Hidden in Plain Sight -- Class Competition Focuses Attribution Maps | 2026 | ICML | CISPA Helmholtz; CISPA Helmholtz Center; Max Planck Institute for Informatics | B |
| 593 | Hidden Breakthroughs in Language Model Training | 2026 | ICLR | Google Research; Harvard University; Kempner Institute, Harvard University | B |
| 594 | HiPPO Zoo: Explicit Memory Mechanisms for Interpretable State Space Models | 2026 | ICML | Duke University | B |
| 595 | Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs | 2026 | ICLR | United States Military Academy; University of Florida; University of Oklahoma | B |
| 596 | Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video Detection | 2026 | ICML | HKUST(GZ), CUHK; Shanghai Jiao Tong University; The Chinese University of Hong K | B |
| 597 | Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs | 2026 | ICLR | UMass Amherst; University of Massachusetts at Amherst; University of Massachuset | B |
| 598 | Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation Shaping | 2026 | ICML | Nanyang Technological University; Sichuan University; Zhejiang University | B |
| 599 | Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization | 2026 | ICLR | Sapienza University of Rome; University of Genoa; University of Genoa sAIfer L | B |
| 600 | Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing | 2026 | ICML | Columbia University; Northwestern University | B |
| 601 | Hallucination Reduction with CASAL: Contrastive Activation Steering for Amortized Learning | 2026 | ICLR | Facebook; Independent; Meta; Meta FAIR; New York University; University of Cambr | B |
| 602 | Hallucination Begins Where Saliency Drops | 2026 | ICLR | Alibaba Group; Dalian Martime University; Institute of automation, Chinese acade | B |
| 603 | Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders | 2026 | Tsinghua University | C | |
| 604 | Guaranteed Optimal Compositional Explanations for Neurons | 2026 | ICML | MIT; University of California, Santa Cruz | B |
| 605 | Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation | 2026 | ICLR | DFKI; German Research Center for AI; Saarland University, Universität des Saarla | B |
| 606 | Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test | 2026 | ICLR | MBZUAI; University of Maryland, College Park | B |
| 607 | Graph Signal Processing Meets Mamba2: Adaptive Filter Bank via Delta Modulation | 2026 | ICLR | KAIST; Korea Advanced Institute of Science & Technology; Korea Advanced Institut | B |
| 608 | Graph Explorer: Training Faithful KG Agents with Visibility-Grounded Supervision. | 2026 | ACL | Boston University; University of Washington | B |
| 609 | GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs. | 2026 | ACL | University of North Carolina at Chapel Hill; The University of | B |
| 610 | Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency | 2026 | KAUST | C | |
| 611 | Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning | 2026 | ICML | Devoteam | B |
| 612 | Geometric Collapse: When Vision Models Fail to Verify Physical Causality | 2026 | ICML | CUHK; Department of Computer Science and Engineering, The Chinese University of | B |
| 613 | Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models | 2026 | ICML | Northwestern; Northwestern University; Northwestern University, Northwestern Uni | B |
| 614 | Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline. | 2026 | ACL | Xi'an Jiaotong University; Hefei University of Technology; CSIRO; Northwestern P | B |
| 615 | GenSR: Symbolic regression based on equation generative space | 2026 | ICLR | Eastern Institute of Technology, Ningbo; Imperial College London; Shanghai Jiaot | B |
| 616 | Gauge-invariant representation holonomy | 2026 | ICLR | Athena RC; Athena Research Center | B |
| 617 | GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning | 2026 | ICML | Sony AI; Sony Group Corporation; Stanford University | B |
| 618 | GRASP: Awakening Latent Spatial Reasoning in LVLMs via Training-free Geometric Rectification | 2026 | ICML | Henan Univeristy; Soochow University; Suzhou University | B |
| 619 | GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning | 2026 | ICLR | Apple; Apple ML Research; Apple MLR and Mila; Meta | B |
| 620 | GNN Explanations that do not Explain and How to find Them | 2026 | ICLR | Fondazione Bruno Kessler; TU Wien; University of Trento | B |
| 621 | GAVEL: Towards Rule-Based Safety through Activation Monitoring | 2026 | ICLR | Amrita Vishwa Vidyapeetham (Deemed University); Ben Gurion University of the Neg | B |
| 622 | GARLIC: Graph Attention-based Relational Learning of Multivariate Time Series in Intensive Care | 2026 | ICLR | Department of Informatics, University of Zurich, University of Zurich; ETH Zuric | B |
| 623 | GALAX: Graph-Augmented Language Model for Explainable Reinforcement-Guided Subgraph Reasoning in Precision Medicine | 2026 | ICLR | Washington University; Washington University in Saint Louis; Washington Universi | B |
| 624 | Functional Decomposition and Shapley Interactions for Interpreting Survival Models | 2026 | ICML | Bielefeld University; Leibniz Institute for Prevention Research and Epidemiology | B |
| 625 | Functional Attention: From Pairwise Affinities to Functional Correspondences : ICML 2026 | 2026 | Technical University of Munich + University of Oxford + UT Austin | C | |
| 626 | Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition | 2026 | ICLR | Salesforce AI Research; University of Southern California | B |
| 627 | From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution. | 2026 | ACL | Technical University of Applied Sciences Würzburg-Schweinfurt; JetBrains Researc | B |
| 628 | From Weights to Activations: Is Steering the Next Frontier of Adaptation? | 2026 | ACL | Saarland University; German Research Centre for Artificial Intelligence; Centre | B |
| 629 | From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment. | 2026 | ACL | Technical University of Munich; Munich Center for Machine Learning; Lund Univers | B |
| 630 | From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets | 2026 | ICML | Department of Computer Science, Duke University; Duke; Duke University; Universi | B |
| 631 | From RAG to Agentic RAG for Faithful Islamic Question Answering. | 2026 | ACL | Qatar Computing Research Institute, HBKU, Qatar; Hamad Bin Khalifa University; A | B |
| 632 | From Nodes to Narratives: Explaining Graph Neural Networks with LLMs and Graph Context. | 2026 | ACL | University of Illinois Chicago; Indian Institute of Technology Kharagpur | B |
| 633 | From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts? | 2026 | ACL | Boston University; Harvard University; Mila – Quebec AI Institute; University of | B |
| 634 | From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models. | 2026 | ACL | Tokyo University of Science | B |
| 635 | From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models. | 2026 | ACL | Tianjin University; Alibaba Group (China) | B |
| 636 | From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models. | 2026 | ACL | Tongji University; University of Wisconsin–Madison | B |
| 637 | From Growing to Looping: A Unified View of Iterative Computation in LLMs | 2026 | ICML | Google; Helmholtz AI | Technical University of Munich; Helmholtz Munich; KTH Sto | B |
| 638 | From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations. | 2026 | ACL | Tata Consultancy Services Research; Indian Institute of Technology Patna; Univer | B |
| 639 | From Data Statistics to Feature Geometry: How Correlations Shape Superposition | 2026 | ICLR | Imperial College; Imperial College London | B |
| 640 | From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers | 2026 | ICLR | Meta AI; New York University | B |
| 641 | From Basis to Basis: Gaussian Particle Representation for Interpretable PDE Operators | 2026 | ICML | The Hong Kong University of Science and Technology; The Hong Kong University of | B |
| 642 | From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards. | 2026 | ACL | William & Mary; Anytime AI | B |
| 643 | Fresh in memory: Training-order recency is linearly encoded in language model activations | 2026 | ICLR | University of Cambridge | B |
| 644 | Formalizing the Binding Problem | 2026 | ICML | Carnegie Mellon University; Mila - Quebec AI Institute, University of Pennsylvan | B |
| 645 | Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees | 2026 | ICLR | Hebrew University of Jerusalem | B |
| 646 | Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based Learning | 2026 | ICML | Amazon; IIT Hyderabad; Indian Institute of Technology, Hyderabad | B |
| 647 | Forget-It-All: Multi-Concept Machine Unlearning via Concept-Aware Neuron Masking | 2026 | ICML | Clemson University; University of Arizona; University of Georgia; University of | B |
| 648 | Forest Before Trees: Latent Superposition for Efficient Visual Reasoning. | 2026 | ACL | Mohamed bin Zayed University of Artificial Intelligence; Fudan University; Renmi | B |
| 649 | ForensicConcept: Transferable Forensic Concepts for AIGI Detection | 2026 | ICML | Tencent Youtu Lab; Westlake University; Xiamen University | B |
| 650 | Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs. | 2026 | ACL | AI Singapore | B |
| 651 | Focus and Dilution: The Multi-stage Learning Process of Attention | 2026 | ICML | Shanghai Jiao Tong University; Shanghai Jiaotong University | B |
| 652 | FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models | 2026 | ICLR | Delineate Inc.; Microsoft; Stevens Institute of Technology; University of Centra | B |
| 653 | Flow-Disentangled Feature Importance | 2026 | ICLR | SUN YAT-SEN UNIVERSITY; St. Jude Children's Research Hospital; University of Hon | B |
| 654 | Fix the Mind, Not the Move: Interpretable AI Assistance via Knowledge-Gap Localization | 2026 | ICML | University of Southern California | B |
| 655 | First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation | 2026 | ICLR | University of South Florida | B |
| 656 | FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models. | 2026 | ACL | University of California, Los Angeles | B |
| 657 | Fine-grained Analysis of Brain-LLM Alignment through Input Attribution | 2026 | ICML | Carnegie Mellon University; Goethe University Frankfurt; Sony AI | B |
| 658 | Fiction Flows: A Replication and Reinterpretation of Narrative Sequentiality. | 2026 | ACL | Cornell University; McGill University | B |
| 659 | Features Emerge as Discrete States: The First Application of SAEs to 3D Representations | 2026 | ICLR | Stony Brook University; University of Cambridge; University of Cambridge & Googl | B |
| 660 | Feature segregation by signed weights in artificial vision systems and biological models | 2026 | ICLR | Harvard Medical School; Harvard Medical School, Harvard University | B |
| 661 | Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers | 2026 | ICML | The Chinese University of Hong Kong | B |
| 662 | Fast Retrieval and Slow Reasoning for Explainable Multimodal Sentiment Analysis. | 2026 | ACL | Anhui Province Key Laboratory of Affective Computing and Advanced Intelligent Ma | B |
| 663 | Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process | 2026 | ICML | Google; Google DeepMind; Google Deepmind; Google Inc.; UT Austin; University of | B |
| 664 | False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize. | 2026 | ACL | National University of Singapore; Peking University; University of California, D | B |
| 665 | FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning | 2026 | ICLR | Ant Group; AntGroup; Shanghai Jiao Tong University; Shanghai Jiaotong University | B |
| 666 | Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation. | 2026 | ACL | ETH Zurich; Mohamed bin Zayed University of Artificial Intelligence | B |
| 667 | Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence. | 2026 | ACL | The University of Texas at Austin; Universidad del Noreste; Scripps MD Anderson | B |
| 668 | Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning | 2026 | Lab post (Anthropic) | Anthropic | A |
| 669 | Faithfulness Under the Distribution: A New Look at Attribution Evaluation | 2026 | ICLR | MI2.AI - WarsawTech; Suzhou Yierqi; University of Technology Sydney; University | B |
| 670 | Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs. | 2026 | ACL | Shanghai Jiao Tong University; Hong Kong University of Science and Technology; A | B |
| 671 | Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance. | 2026 | ACL | Tel Aviv University | B |
| 672 | Faithful Persona Steering under Incongruity via Dual-Stream Refinement. | 2026 | ACL | National Yang Ming Chiao Tung University | B |
| 673 | Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions | 2026 | ICLR | Ant Group; National Certification Technology (Hangzhou) Co., Ltd; Zhejiang Norma | B |
| 674 | FaithLens: Detecting and Explaining Faithfulness Hallucination. | 2026 | ACL | Fudan University; Tsinghua University; University of Illinois Urbana-Champaign; | B |
| 675 | FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning | 2026 | ICLR | Arizona State University; Department of Computer Science, University of North Ca | B |
| 676 | Factorized Scheduling Principle: Learning Interpretable and Transferable Policies via Structured Additive Functions | 2026 | ICML | Sejong University | B |
| 677 | FLAIR: Steering LLM Mathematical Problem Solving based on A Fuzzy-Logic-AssIsted Reasoner. | 2026 | ACL | Central China Normal University; University of Wollongong; University of Wiscons | B |
| 678 | FIPN: Forward Self-Organizing Interpretable Polynomial Networks for Time Series Forecasting | 2026 | ICML | College of Information Communication Technology University of Suwon; Linyi Unive | B |
| 679 | FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation. | 2026 | ACL | University of Oregon | B |
| 680 | FAME: Formal Abstract Minimal Explanation for Neural Networks | 2026 | ICLR | Airbus; Airbus SAS, IRT Saint Exupéry; Hebrew University of Jerusalem | B |
| 681 | Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks | 2026 | ICML | Emory University; Griffith University (Australia); Michigan State University; Pe | B |
| 682 | Exploring and Distilling Multi-Dimensional Clues for Interpretable Social Bot Detection. | 2026 | ACL | Singapore Management University; Beijing Language and Culture University | B |
| 683 | Exploring Layer Activation Dynamic of CoT via Knowledge Probe. | 2026 | ACL | Southeast University | B |
| 684 | Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts | 2026 | ICLR | Microsoft Research; Microsoft Research Asia; Shanghai Artificial Intelligence La | B |
| 685 | Exploring Accurate and Transparent Domain Adaptation in Predictive Healthcare via Concept-Grounded Orthogonal Inference | 2026 | ICML | Stevens Institute of Technology; University of Massachusetts Chan Medical School | B |
| 686 | Explanations are a Means to an End: Decision Theoretic Explanation Evaluation | 2026 | ICML | Northwestern University; UCSD | B |
| 687 | Explanation Quality Assessment as Ranking with Listwise Rewards. | 2026 | ACL | Université d'Artois; AIL Research (United States) | B |
| 688 | Explaining Sources of Uncertainty in Automated Fact-Checking. | 2026 | ACL | University of Copenhagen | B |
| 689 | Explaining Grokking and Information Bottleneck through Neural Collapse Emergence | 2026 | ICLR | The University of Tokyo; Tokyo University, Tokyo Institute of Technology | B |
| 690 | Explaining Concept Shift with Interpretable Feature Attribution | 2026 | ICML | CMU, Carnegie Mellon University; Carnegie Mellon University; School of Computer | B |
| 691 | Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection. | 2026 | ACL | Griffith University | B |
| 692 | Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets | 2026 | ICLR | Alibaba Group; Nanyang Technological University; Zhejiang University; Zhejiang U | B |
| 693 | Explainable Quantum Program Repair with Verifiable Proof Traces. | 2026 | ACL | Singapore Management University | B |
| 694 | Explainable Mixture Models through Differentiable Rule Learning | 2026 | ICLR | CISPA Helmholtz; CISPA Helmholtz Center for Information Security; Universität de | B |
| 695 | Explainable Forensics of Manipulated Segments in Untrimmed Long Videos | 2026 | ICML | Independent Researcher; Nanjing University; Nanjing University of Aeronautics an | B |
| 696 | Explainable Federated Learning via Global–Local Attribution Alignment | 2026 | ICML | DEVCOM Army Research Laboratory; Virginia Polytechnic Institute and State Univer | B |
| 697 | Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI. | 2026 | ACL | University of Oregon; Adobe Research, CA, USA | B |
| 698 | Explainable $ K $-means Neural Networks for Multi-view Clustering | 2026 | ICLR | Fudan University; Shanghai University | B |
| 699 | Explain the Synth: Interpretable Evaluation of LLM Data Synthesis. | 2026 | ACL | Australian Regenerative Medicine Institute; Monash University; Centre de Recherc | B |
| 700 | ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns | 2026 | ICML | Amazon; ByteDance; Bytedance Seed & PSU; Fudan University; Soochow University; T | B |
| 701 | Expert Heads: Robust Evidence Identification for Large Language Models | 2026 | ICLR | Jilin University; Renmin University of China; Soochow University | B |
| 702 | Expand Neurons, Not Parameters | 2026 | ICML | Computer Science and Artificial Intelligence Laboratory, Electrical Engineering | B |
| 703 | Exactly Computing do-Shapley Values | 2026 | ICML | Barcelona Supercomputing Center; Bielefeld University; Claremont McKenna College | B |
| 704 | Exact Functional ANOVA Decomposition for Categorical Inputs Models | 2026 | ICML | EDF & Sorbonne Université; Université Toulouse Paul Sabatier Institut de Mathéma | B |
| 705 | Exact Functional ANOVA Decomposition for Categorical Inputs | 2026 | ICML | EDF & Sorbonne Université; Université Toulouse Paul Sabatier Institut de Mathéma | B |
| 706 | ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability. | 2026 | ACL | Mohamed bin Zayed University of Artificial Intelligence; Tokyo Institute of Tech | B |
| 707 | ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection | 2026 | ICLR | Alibaba Group; Tencent; University of Bath; University of Cambridge; University | B |
| 708 | ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior | 2026 | ICML | LMU Munich; Ludwig-Maximilians-Universität München; Saarland University | B |
| 709 | Evolution of Concepts in Language Model Pre-Training | 2026 | ICLR | Fudan University; Shanghai Artificial Intelligence Laboratory | B |
| 710 | Evidential Reasoning Advances Interpretable Real-World Disease Screening | 2026 | ICML | Hong Kong Polytechnic University; The Hong Kong Polytechnic University; Tsinghua | B |
| 711 | Evidential Copula Concept Embedding Models | 2026 | ICML | Macquarie University; Shanghai University; Tongji University | B |
| 712 | Evian: Towards Explainable Visual Instruction-tuning Data Auditing. | 2026 | ACL | The Hong Kong University of Science and Technology (Guangzhou); ByteDance | B |
| 713 | Evaluating and Steering Modality Preferences in Multi-modal LLMs | 2026 | ICML | Harbin Institute of Technology; Harbin Institute of Technology (Shenzhen); Harbi | B |
| 714 | Evaluating SAE interpretability without generating explanations | 2026 | ICLR | EleutherAI | B |
| 715 | Evaluating Data Influence in Meta Learning | 2026 | ICLR | KAUST; King Abdullah University of Science and Technology; MBZUAI; Shanghai Arti | B |
| 716 | Escaping Low-Rank Traps: Interpretable Visual Concept Learning via Implicit Vector Quantization | 2026 | ICLR | Fudan University; Fudan University; Shanghai AI Lab; Fudan university; Hong Kong | B |
| 717 | Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning | 2026 | ICLR | MPI-SP; Max Planck Institute; Max Planck Institute for Security and Privacy; Seo | B |
| 718 | Ensembling Sparse Autoencoders | 2026 | ICML | University of Washington | B |
| 719 | EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace Method | 2026 | ICLR | Penn State | B |
| 720 | Endogenous Resistance to Activation Steering in Language Models | 2026 | ICML | AE Studio; Agency Enterprise Studio; Independent; Princeton University | B |
| 721 | Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models | 2026 | ICLR | Georgia Institute of Technology | B |
| 722 | EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning | 2026 | ICLR | Chinese University of Hong Kong, The Chinese University of Hong Kong; Microsoft; | B |
| 723 | Emotion Concepts and their Function in a Large Language Model | 2026 | Lab post (Anthropic) | Anthropic | A |
| 724 | EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness. | 2026 | ACL | South China University of Technology | B |
| 725 | Emergent Analogical Reasoning in Transformers | 2026 | ICML | The University of Tokyo; The University of Tokyo, The University of Tokyo; Unive | B |
| 726 | Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought | 2026 | ICLR | EECS, UC Berkeley; Meta AI Research; UC San Diego; University of California - Be | B |
| 727 | Emergence and Localisation of Semantic Role Circuits in LLMs. | 2026 | ACL | University of Manchester | B |
| 728 | Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models. | 2026 | ACL | Yonsei University | B |
| 729 | Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models | 2026 | ICML | University of Birmingham; University of Leicester | B |
| 730 | Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models. | 2026 | ACL | ETH Zurich; Mohamed bin Zayed University of Artificial Intelligence; University | B |
| 731 | Efficient Hallucination Detection for LLMs Using Uncertainty-Aware Attention Heads | 2026 | ICML | FusionBrain; FusionBrain Lab; Independent researcher; MBZUAI; Mohamed bin Zayed | B |
| 732 | Efficient Estimation of Kernel Surrogate Models for Task Attribution | 2026 | ICLR | Northeastern University | B |
| 733 | Effective Reasoning Chains Reduce Intrinsic Dimensionality | 2026 | ICML | Google; Google DeepMind; University of North Carolina at Chapel Hill; University | B |
| 734 | EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI. | 2026 | ACL | LMU Klinikum; Ludwig-Maximilians-Universität München | B |
| 735 | EGG-SR: Embedding Symbolic Equivalence into Symbolic Regression via Equality Graph | 2026 | ICLR | Purdue University; University of Texas at El Paso | B |
| 736 | EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions. | 2026 | ACL | Georgia Institute of Technology; Virginia Tech | B |
| 737 | ECSEL: Explainable Classification via Signomial Equation Learning | 2026 | ICML | University of Amsterdam | B |
| 738 | Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP | 2026 | ICLR | Fraunhofer HHI; Oxford & Fraunhofer HHI; Technical University of Berlin; Univers | B |
| 739 | Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure | 2026 | ICML | Alibaba Group; Harbin Institute of Technology; Harbin Institute of Technology (S | B |
| 740 | Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers | 2026 | ICLR | University of California, Berkeley; University of Chicago | B |
| 741 | Dynamic Reflections: Probing Video Representations with Text Alignment | 2026 | ICLR | DeepMind; Google; Google DeepMind; Princeton University | Google DeepMind; Stanf | B |
| 742 | Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models | 2026 | ICLR | East China Normal University | B |
| 743 | Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models | 2026 | ICML | Seoul National University | B |
| 744 | Domain Restriction via SAE Multi-Layer Transitions | 2026 | ICML | Technion - Israel Institute of Technology; Technion - Israel Institute of Techno | B |
| 745 | Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking | 2026 | ICML | Corning Inc.; The University of Sydney; University of Central Florida; Universit | B |
| 746 | Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders | 2026 | ICLR | Shenzhen Research Institute of Big Data; The Chinese University of Hong Kong; Un | B |
| 747 | Do Transformers Grok Succinct Algorithms? Mechanistic Evidence for Counting Circuits. | 2026 | ACL | Nanjing University | B |
| 748 | Do Sparse Autoencoders Identify Reasoning Features in Language Models? | 2026 | ICML | UC Berkeley; University of California, Berkeley | B |
| 749 | Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models. | 2026 | ACL | Laboratory for Social and Neural Systems Research; National Center for Theoretic | B |
| 750 | Do Neural Operators Forget Geometry? The Forgetting Hypothesis in Deep Operator Learning | 2026 | ICML | Tsinghua University | B |
| 751 | Do Language Models Track Entities Across State Changes? | 2026 | ICML | Boston University; Monash University; University of Vienna | B |
| 752 | Do LLMs “Feel”? Emotion Circuits Discovery and Control | 2026 | ICML | MBZUAI; Mohamed bin Zayed University of Artificial Intelligence; NYU Shanghai & | BC |
| 753 | Do LLMs Signal When They’re Right? Evidence from Neuron Agreement | 2026 | ICML | Fudan University; Harbin Institute of Technology; University of Hong Kong | B |
| 754 | Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox | 2026 | ICML | University of Southern California | B |
| 755 | Do Activation Verbalization Methods Convey Privileged Information? | 2026 | ICML | Kempner Institute, Harvard University; Northeastern; Northeastern University | B |
| 756 | Distributionally Robust Causal Abstractions | 2026 | ICML | University of Warwick; University of Warwick & The Alan Turing Institute | B |
| 757 | Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs | 2026 | ICML | Peking University | B |
| 758 | Dissecting the Safety Circuit: Neuronal Intervention for Transferable Adversarial Attacks on VLMs | 2026 | ICML | Chongqing University; Huawei Technologies Ltd.; Nanyang Technological University | B |
| 759 | Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers | 2026 | ICML | Beijing University of Posts and Telecommunications; Heidelberg University, Unive | B |
| 760 | Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding | 2026 | ICML | University of Auckland; University of Electronic Science and Technology of China | B |
| 761 | Disentangling Latent Risk Pathways via Bayesian Hypergraph Inference | 2026 | ICML | Yale University | B |
| 762 | Disentangling Geometry, Performance, and Training in Language Models | 2026 | ICML | Carnegie Mellon University; USC; University of California, Los Angeles | B |
| 763 | Disentangled Representation Learning for Parametric Partial Differential Equations | 2026 | ICLR | Amazon; International Business Machines; Lehigh University | B |
| 764 | Discovering and Steering Interpretable Concepts in Large Generative Music Models | 2026 | ICLR | Dartmouth College; Massachusetts Institute of Technology | B |
| 765 | Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models. | 2026 | ACL | Johns Hopkins University; University of Edinburgh | B |
| 766 | Discovering a Shared Logical Subspace: Steering LLM Logical Reasoning via Alignment of Natural-Language and Symbolic Views. | 2026 | ACL | University of Illinois Urbana-Champaign | B |
| 767 | Discovering Interpretable Algorithms by Decompiling Transformers to RASP | 2026 | ICML | NYU / AI2; Saarland University; University of Oxford; Universität des Saarlandes | B |
| 768 | Discovering Implicit Large Language Model Alignment Objectives | 2026 | ICML | Stanford University; Stanford University & Apple; Stanford University // Virtue | B |
| 769 | Directly Optimizing Natural Language Explanations for Behavioral Faithfulness: Simulatability and Recoverability | 2026 | ICML | International Institute of Information Technology - Hyderabad; University of Nor | B |
| 770 | Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning | 2026 | ICML | Fudan University; Shanghai Innovation Institute | B |
| 771 | Diffusion-CAM: Faithful Visual Explanations for dMLLMs. | 2026 | ACL | Shanghai Jiao Tong University; Sun Yat-sen University; Northwestern University; | B |
| 772 | Dialectical Structured Reasoning for Explainable Multimodal Fake News Detection. | 2026 | ACL | University of Science and Technology Beijing; National University of Singapore; | B |
| 773 | Dialectic-Med: Mitigating Diagnostic Hallucinations via Counterfactual Adversarial Multi-Agent Debate. | 2026 | ACL | Xi’an Jiaotong-Liverpool University | B |
| 774 | Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers | 2026 | ICML | Korea University; Seoul National University | B |
| 775 | Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution | 2026 | ICML | Arizona State University; Google DeepMind; Johns Hopkins University; University | B |
| 776 | Diagnosing Generalization Failures from Representational Geometry Markers | 2026 | ICLR | Center of Computational Neuroscience, Flatiron Institute; Google DeepMind; Harva | B |
| 777 | Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation | 2026 | ICML | University of Illinois Urbana-Champaign; University of Southern California; Yale | B |
| 778 | Detecting What Queries Seek: Steering LLM Safety with FFN Output Activation Monitoring. | 2026 | ACL | Northeastern University | B |
| 779 | Detecting Invariant Manifolds in ReLU-Based RNNs | 2026 | ICLR | Central Institute of Mental Health; Dept. Theoretical Neuroscience, Central Inst | B |
| 780 | Density-Guided Robust Counterfactual Explanations on Tabular Data under Model Multiplicity | 2026 | ICML | Central South University | B |
| 781 | Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning | 2026 | ICML | Harvard University; Yale University | B |
| 782 | Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success | 2026 | ICML | Sapienza University of Rome; University of California, San Diego | B |
| 783 | Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring | 2026 | ICLR | AITRICS; Korea Advanced Institute of Science & Technology; Sung Kyun Kwan Univer | B |
| 784 | Deliberate Evolution: Agentic Reasoning for Sample-Efficient Symbolic Regression with LLMs | 2026 | ICML | HKBU / RIKEN; HKBU / Stanford; Hong Kong Baptist University; Lenovo Group Limite | B |
| 785 | Deep neural networks divide and conquer dihedral multiplication | 2026 | ICML | Leiden University, Dept. of Mathematics, Leiden University; McGill University; M | B |
| 786 | Deep networks learn to parse uniform-depth context-free languages from local statistics | 2026 | ICML | EPFL; International Higher School for Advanced Studies Trieste | B |
| 787 | Deep Single-Index Fréchet Regression | 2026 | ICML | University of California, Davis; Waymo | B |
| 788 | Deconstructing Positional Information: From Attention Logits to Training Biases | 2026 | ICLR | Chinese Academy of Sciences; Institute of Information Engineering; Shanghai Jiao | B |
| 789 | Deconstructing Guidance: A Semantic Hierarchy for Precise Diffusion Model Editing | 2026 | ICLR | Korea University | B |
| 790 | Decomposition of Concept-Level Rules in Visual Scenes | 2026 | ICLR | Fudan University | B |
| 791 | Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning | 2026 | ICLR | Saarland University; Universität des Saarlandes | B |
| 792 | Decomposing Query-Key Feature Interactions Using Contrastive Covariances | 2026 | ICML | Google; Harvard University; Technion, Technion | B |
| 793 | Decomposing LLM Computation with Jets | 2026 | ICLR | AWS; Massachusetts Institute of Technology; University College London; Universit | B |
| 794 | DecodeShare: Tracing the Shared Pathways of LLM Decode-Time Decisions | 2026 | ICML | Department of Computer Science, Duke University; Duke University; Duke Universit | B |
| 795 | Deciphering Cultural Representations in Large Language Models via Sparse Autoencoders. | 2026 | ACL | Mohamed bin Zayed University of Artificial Intelligence; University of Toronto | B |
| 796 | Debugging Concept Bottleneck Models through Removal and Retraining | 2026 | ICLR | Cornell University | B |
| 797 | De-Anonymization at Scale via Tournament-Style Attribution. | 2026 | ACL | Peking University; Beihang University | B |
| 798 | Data-Aware and Scalable Sensitivity Analysis for Decision Tree Ensembles | 2026 | ICLR | IIT Bombay; Indian Institute of Technology Bombay, Mumbai, India; Indian Institu | B |
| 799 | DVI-DTM: Dual-View Representation Learning for Interpretable Short Text Dynamic Topic Modeling. | 2026 | ACL | University of Chinese Academy of Sciences; Beijing University of Posts and Telec | B |
| 800 | DRIV-EX: Counterfactual Explanations for Driving LLMs. | 2026 | ACL | Institut des langues et cultures d'Europe, Amérique, Afrique, Asie et Australie; | B |
| 801 | DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole Slide Image Survival Prediction | 2026 | ICML | IIAI; Imperial colleges London; National University of Singapore; Peking Union M | B |
| 802 | DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models. | 2026 | ACL | Southeast University; Shanghai Jiao Tong University; East China Normal Universit | B |
| 803 | DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders | 2026 | ICML | Alibaba Group; The University of Hong Kong; University of Hong Kong; the Univers | B |
| 804 | DISSOLVR: An Interpretable and Fast Framework for Aqueous and Organic Solubility Prediction | 2026 | ICML | IIT Delhi; Indian Institute of Technology, Delhi | B |
| 805 | DEFT: Demystifying VLN Failures via a Unified Dual-View Explainability Framework for LLM-based Agents. | 2026 | ACL | Institute of Software; Chinese Academy of Sciences | B |
| 806 | DB-KSVD: Scalable Alternating Optimization for Disentangling High-Dimensional Embedding Spaces | 2026 | ICML | Google DeepMind; Stanford University; Stanford University/ ddx.inc | B |
| 807 | DAVE: Distribution-aware Attribution via ViT Gradient Decomposition | 2026 | ICML | Jagiellonian University; Jagiellonian University in Krakow; Jagiellonian Univers | B |
| 808 | Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings | 2026 | ICLR | CNRS; ENS Paris Saclay; Harvard University; IRT Saint-Exupery | B |
| 809 | CrisPrune: Combining Contextual Relevance and Intrinsic Saliency for Efficient Visual Token Pruning in MLLMs. | 2026 | ACL | Tongji University; Ant Group | B |
| 810 | Credal Concept Bottleneck Models for Epistemic-Aleatoric Uncertainty Decomposition. | 2026 | ACL | Université d'Artois; AIL Research (United States) | B |
| 811 | Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs. | 2026 | ACL | University of Notre Dame | B |
| 812 | Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System. | 2026 | ACL | Bangalore , India | B |
| 813 | Counterfactual Explanations on Robust Perceptual Geodesics | 2026 | ICLR | QIMR Berghofer Medical Research Institute; Queensland University of Technology; | B |
| 814 | Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs. | 2026 | ACL | Yunnan University; National University of Singapore | B |
| 815 | CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features | 2026 | ICML | Department of Computer Science, University College London, University of London; | B |
| 816 | Controlled LLM Training on Spectral Sphere | 2026 | MSRA (Baining Guo et al.) | C | |
| 817 | Controllable and Explainable Personality Sliders for LLMs at Inference Time | 2026 | ICML | University College Dublin (UCD); University of Cambridge; University of Cambridg | B |
| 818 | Controllable Molecule Generation via Sparse Representation Editing: An Interpretability-Driven Perspective | 2026 | ICML | Department of Computing, The Hong Kong Polytechnic University; Hong Kong Polytec | B |
| 819 | Controllable LLM Reasoning via Sparse Autoencoder-Based Steering. | 2026 | ACL | University of Science and Technology of China; Alibaba Group (United States) | BC |
| 820 | Contribution Weights: A Geometrical Analysis of Self-Attention Transformers | 2026 | ICML | Imperial College London; University College London, University of London | B |
| 821 | Contrastive Symbolic Regression: Aligned Representations, Adaptive Prediction, and Diverse Ensembles | 2026 | ICML | Michigan State University; Victoria University of Wellington | B |
| 822 | Continuous Interpretive Steering for Scalar Diversity. | 2026 | ACL | Sungkyunkwan University | B |
| 823 | ContextCheck: Sentence-Level Faithfulness Verification with Context-Aware Disambiguation. | 2026 | ACL | Microsoft (United States); University of Michigan; The University of Texas at Au | B |
| 824 | Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding. | 2026 | ACL | Tencent; McGill University; Wuhan University; Université de Montréal; Tsinghua U | B |
| 825 | Context Attribution with Multi-Armed Bandit Optimization. | 2026 | ACL | University of Notre Dame | B |
| 826 | Constructing Interpretable Features from Compositional Neuron Groups. | 2026 | ACL | Tel Aviv University; ESI Group (France) | B |
| 827 | Concepts' Information Bottleneck Models | 2026 | ICLR | University of Amsterdam; University of Hull; University of Oslo; University of t | B |
| 828 | Concept-TRAK: Understanding how diffusion models learn concepts through concept attribution | 2026 | ICLR | Sony; Sony AI; Sony Group Corporation; University of Pennsylvania | B |
| 829 | Concept Concentration for Faithful Representation Intervention | 2026 | ICML | CUHK; Carnegie Mellon University & MBZUAI; HKBU / RIKEN; Johns Hopkins Universit | B |
| 830 | ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution | 2026 | ICML | Tel Aviv University; The Open University of Israel | B |
| 831 | Compressed Sensing for Capability Localization in Large Language Models | 2026 | ICML | Carnegie Mellon University; Thinking Machines Lab | B |
| 832 | Compositional Steering of Large Language Models with Steering Tokens. | 2026 | ACL | University of Edinburgh; iMinds | B |
| 833 | Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models | 2026 | ICML | Helmholtz AI, Technical University of Munich; KAIST & Cortiq; University of Tübi | B |
| 834 | Composable Sparse Subnetworks via Maximum-Entropy Principle | 2026 | ICLR | Polytechnic of Turin; Sapienza University of Rome; University of Roma "La Sapien | B |
| 835 | Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors. | 2026 | ACL | Zhejiang University; Palo Alto Networks; OPPO Research Institute; Qilu Universit | B |
| 836 | Comparing the learning dynamics of in-context learning and fine-tuning in language models | 2026 | ICLR | Harvard University; University College London; University College London, Univer | B |
| 837 | Comparing Human and Large Language Model Interpretation of Implicit Information. | 2026 | ACL | Politecnico di Milano | B |
| 838 | Compact Example-Based Explanations for Language Models. | 2026 | ACL | University of Vienna | B |
| 839 | CombinationTS: A Modular Framework for Understanding Time-Series Forecasting Models | 2026 | ICML | , Chinese Academy of Sciences; Beijing Normal-Hong Kong Baptist University; Comp | B |
| 840 | Colorful Talks with Graphs: Human-Interpretable Graph Encodings for Large Language Models. | 2026 | ACL | University of Illinois Chicago | B |
| 841 | Cognitive models can reveal interpretable value trade-offs in language models | 2026 | ICLR | Google DeepMind; Harvard University; Johns Hopkins University; Kempner Institute | B |
| 842 | Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement | 2026 | ICML | Artificial Intelligence Institute of South Carolina; Indian Institute of Technol | B |
| 843 | CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation | 2026 | ICML | Institute of Automation, Chinese Academy of Sciences; Institute of Information E | B |
| 844 | CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs | 2026 | ICLR | Monash University; Southeast University; southeast university | B |
| 845 | CoSToM: Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models. | 2026 | ACL | National University of Singapore | B |
| 846 | CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment. | 2026 | ACL | Queen's University; Cornell University | B |
| 847 | Clustered Influence Functions | 2026 | ICML | Eötvös Lorand University; Eötvös Loránd University | B |
| 848 | CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation. | 2026 | ACL | University of Waterloo; Hong Kong University of Science and Technology; William | B |
| 849 | Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models | 2026 | ICLR | Duke University; Duke University / Apple; Google; Simon Fraser University | B |
| 850 | CircuitSynth: Reliable Synthetic Data Generation. | 2026 | ACL | University of Oxford | B |
| 851 | CircuitPrint: Mechanistic Circuit Fingerprints for Large Language Models | 2026 | ICML | Hunan University | B |
| 852 | Circuit Insights: Towards Interpretability Beyond Activations | 2026 | ICLR | Fraunhofer HHI; Fraunhofer HHI, Fraunhofer IAIS; Fraunhofer Heinrich Hertz Insti | B |
| 853 | CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization. | 2026 | ACL | The Hong Kong University of Science and Technology (Guangzhou); Chinese Universi | B |
| 854 | Characterizing, Evaluating, and Optimizing Complex Reasoning | 2026 | Shanghai AI Lab + SJTU + CUHK | C | |
| 855 | Characterizing the Expressivity of Local Attention in Transformers :arXiv 2026.05(ACL 2026 Best Paper) | 2026 | ETH Zürich | C | |
| 856 | Characterizing Pattern Matching and Its Limits on Compositional Task Structures | 2026 | ICLR | Amazon; KAIST; KAIST AI; Korea Advanced Institute of Science & Technology; LG AI | B |
| 857 | Characteristic Root Analysis and Regularization for Linear Time Series Forecasting | 2026 | ICLR | Bosch; Bosch (China) Investment Co., Ltd.; Bosch Cooperate Research; Robert Bosc | B |
| 858 | Chain-of-Thought Reasoning In The Wild Is Not Always Faithful | 2026 | ICML | Google DeepMind; MATS(Neel/Nanda); ML Alignment & Theory Scholars; Poseidon Rese | B |
| 859 | Chain-of-Relations: Faithful and Efficient LLM Reasoning over Knowledge Graphs via Relation-Centric Exploration. | 2026 | ACL | Sun Yat-sen University | B |
| 860 | Certified Evaluation of Model-Level Explanations for Graph Neural Networks | 2026 | ICLR | Indian Statistical Institute | B |
| 861 | Certified Circuits: Stability Guarantees for Mechanistic Circuits | 2026 | ICML | CISPA Helmholtz Center; CISPA Helmholtz Center for Infomation Security; MPI for | B |
| 862 | CausalXRL: Explainable Reinforcement Learning through Causal Graph Reasoning | 2026 | ICML | State University of New York at Stony Brook; Stony Brook University | B |
| 863 | CausalX: A Unified and Causally-Interpretable Plug-and-Play Model for Multi-modal Spatio-Temporal Forecasting | 2026 | ICML | Shandong University; Zhejiang University of Technology | B |
| 864 | CausalGaze: Unveiling Hallucinations via Counterfactual Graph Intervention in Large Language Models. | 2026 | ACL | National University of Defense Technology; Anhui Province Key Laboratory of Cybe | B |
| 865 | CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution | 2026 | ICML | Google; Google DeepMind; SK hynix Research; Seoul National University | B |
| 866 | Causal-Steer: Disentangled Continuous Style Control without Parallel Corpora | 2026 | ICLR | Zhejiang University | B |
| 867 | Causal Interpretation of Neural Network Computations with Contribution Decomposition | 2026 | ICLR | Stanford University | B |
| 868 | Capturing Visual Environment Structure Correlates with Control Performance | 2026 | ICLR | Department of Computer Science, University of Illinois at Urbana-Champaign; Toyo | B |
| 869 | Capacity without Access: Reinterpreting the Mid-Depth Spectral Plateau in LLMs | 2026 | ICML | Chung-Ang University; Chung-Ang university; TTA | B |
| 870 | Can SAEs reveal and mitigate racial biases of LLMs in healthcare? | 2026 | ICLR | Northeastern University | B |
| 871 | Can LLMs Reason Soundly in Law? Auditing Inference Patterns for Legal Judgment | 2026 | ICLR | Beijing Institute for General Artificial Intelligence; Shanghai Artificial Intel | B |
| 872 | Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs | 2026 | Yale University | C | |
| 873 | CSI: An Investigative Multi-Agent Framework for Explainable Short Video Fake News Detection. | 2026 | ACL | Inner Mongolia University | B |
| 874 | CRISP: Persistent Concept Unlearning via Sparse Autoencoders. | 2026 | ACL | Technion – Israel Institute of Technology; University of Zagreb | B |
| 875 | CRISP: Compressing Redundancy in Chain-of-Thought via Intrinsic Saliency Pruning. | 2026 | ACL | College of Artificial Intelligence; Nanjing University of Aeronautics and Astron | B |
| 876 | COCOGEC: Counterfactual Generation for Robust Grammatical Error Correction. | 2026 | ACL | East China Normal University | B |
| 877 | CNN Interpretability with Multivector Tucker Saliency Maps for Self-Supervised Models | 2026 | ICLR | ENS Ulm; Ecole normale supérieure | B |
| 878 | CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark. | 2026 | ACL | Saarland University; German Research Centre for Artificial Intelligence; Centre | B |
| 879 | CLUE: Conflict-guided Localization for LLM Unlearning Framework | 2026 | ICLR | Department of Computer Science and Engineering, The Chinese University of Hong K | B |
| 880 | CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally | 2026 | ICLR | University of Tübingen; University of Tübingen, Max Planck Institute for Intelli | B |
| 881 | CLARITree: Cholesky and Lookahead Accelerations for Regression with Interpretable Piecewise Linear Trees | 2026 | ICML | Duke; Duke University; University of British Columbia | B |
| 882 | CB-SLICE: Concept-Based Interpretable Error Slice Discovery | 2026 | ICML | University of Cambridge; University of Oxford | B |
| 883 | CAMP: Coherent Alignment of Multimodal Prototypes for Explainable Complementary Learning | 2026 | ICML | J.P. Morgan Chase; JPMorgan Chase; Lancaster University; University College Dubl | B |
| 884 | C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders | 2026 | ICML | College of Computer Science, Chongqing University; Renmin University of China; U | B |
| 885 | Budget Alignment: Making Models Reason in the User's Language | 2026 | ICLR | Department of Computer Science; ETH Zürich University of Gro | B |
| 886 | Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions | 2026 | ICML | Seoul National University | B |
| 887 | Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers | 2026 | Microsoft Research / ETH Zürich / KRAFTON | C | |
| 888 | Bridging Radiology and Pathology Foundation Models via Concept-Based Multimodal Co-Adaptation | 2026 | ICLR | Hong Kong Polytechnic University; Stanford University; University of Hong Kong; | B |
| 889 | Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM Generation. | 2026 | ACL | Beijing Normal University | B |
| 890 | Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection? | 2026 | ICLR | Saarland University; Saarland University, Saarland University; Saarland Universi | B |
| 891 | Bridging Explainability and Embeddings: BEE Aware of Spuriousness | 2026 | ICLR | Bitdefender; Bitdefender, Bucharest, Romania; Mila; University of Bucharest | B |
| 892 | Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression | 2026 | ICML | Heidelberg University; IWR Heidelberg | B |
| 893 | Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models | 2026 | ICML | National University of Singapore; Shanghai Jiaotong University; Shenzhen Institu | B |
| 894 | Block Recurrent Dynamics in Vision Transformers | 2026 | ICLR | Harvard University; Universität Osnabrück | B |
| 895 | Biases in the Blind Spot: Detecting What LLMs Fail to Mention | 2026 | ICML | Independent; Poseidon Research; UCL; University College London, University of Lo | B |
| 896 | Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying Prompts | 2026 | ICLR | Ocean University of China; Xiamen University; vivo | B |
| 897 | Beyond Static Personas: Situational Personality Steering for Large Language Models. | 2026 | ACL | University of Science and Technology of China; Singapore Management University | B |
| 898 | Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme Understanding. | 2026 | ACL | National University of Defense Technology | B |
| 899 | Beyond Self-Report: Bridging the Intention-Behavior Gap in Critical Thinking Assessment via Interpretable Multi-Agent System. | 2026 | ACL | Tsinghua University | B |
| 900 | Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering. | 2026 | ACL | Xi'an Jiaotong University; National University of Singapore; Sichuan University; | B |
| 901 | Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization. | 2026 | ACL | Zhejiang Normal University; Huawei Technologies; GS1 Hong Kong; Sun Yat-sen Univ | B |
| 902 | Beyond Linear Probes: Dynamic Safety Monitoring for Language Models | 2026 | ICLR | Queen Mary University of London; University of Oxford; University of Oxford / Ma | B |
| 903 | Beyond Fixed Biases: Decoding the Role of Reasoning Uncertainty in MLLM Modality Conflicts | 2026 | ICML | KAUST; MBZUAI; Peking University; South China University of Technology; Universi | B |
| 904 | Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring | 2026 | ICML | MBZUAI; Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong Universi | B |
| 905 | Beyond Evidence: Belief-Chain Conditioning for Persuasive Misinformation Debunking Explanation. | 2026 | ACL | Academia Sinica; National Tsing Hua University; Pennsylvania State University | B |
| 906 | Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs | 2026 | ICLR | Westlake University; Zhejiang University | B |
| 907 | Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks. | 2026 | ACL | Université de Reims Champagne-Ardenne; Chochoy Conseil | B |
| 908 | Beyond Black-Box Interventions: Latent Probing for Faithful Retrieval-Augmented Generation. | 2026 | ACL | Xiamen University; Alibaba Group (China); Jilin University; University of Chines | B |
| 909 | Beyond Additive Decompositions: Interpretability Through Separability | 2026 | ICML | University of Copenhagen | B |
| 910 | Beyond Accuracy and Complexity: The Effective Information Criterion for Structurally Stable Symbolic Regression | 2026 | ICML | Tsinghua University; Tsinghua University, Tsinghua University | B |
| 911 | Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations. | 2026 | ACL | University of Southern California; California Southern University | B |
| 912 | Behavior Learning (BL) | 2026 | ICLR | Eberhard-Karls-Universität Tübingen; Xi'an Jiaotong University; Xiamen Universit | B |
| 913 | Bayesian Neural Networks for Functional ANOVA Model | 2026 | ICLR | Seoul National University; University of Twente | B |
| 914 | Bayesian Influence Functions for Hessian-Free Data Attribution | 2026 | ICLR | Timaeus; Timaeus Research; University of Colorado at Boulder; University of Melb | B |
| 915 | Bayesian Gated Non-Negative Contrastive Learning | 2026 | ICML | MBZUAI; Mohamed bin Zayed University of Artificial Intelligence | B |
| 916 | Base Models Know How to Reason, Thinking Models Learn When | 2026 | ICML | Google DeepMind; Oxford; Poseidon Research; University of Oxford | BC |
| 917 | BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation. | 2026 | ACL | Independent University, Bangladesh | B |
| 918 | BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking | 2026 | ICML | Carnegie Mellon University | B |
| 919 | Awakening Dormant Experts: Counterfactual Routing to Mitigate MoE Hallucinations. | 2026 | ACL | Xi'an Jiaotong University; China Telecom; Beijing Foreign Studies University | B |
| 920 | Automatically Finding Reward Model Biases | 2026 | ICML | Google DeepMind; Massachusetts Institute of Technology; Poseidon Research | B |
| 921 | Automatic Image-Level Morphological Trait Annotation for Organismal Images | 2026 | ICLR | Ohio State University; The Ohio State University; The Ohio State University, Col | B |
| 922 | Automatic Construction of Clinical Scoring Systems with LLM Agents | 2026 | ICML | University of Cambridge; University of Cambridge and UCLA | B |
| 923 | Automated Knowledge Component Generation and Interpretable Knowledge Tracing in Coding Problems. | 2026 | ACL | University of Pittsburgh; University of Massachusetts Amherst | B |
| 924 | Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers | 2026 | ICLR | University of Bristol | B |
| 925 | AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning. | 2026 | ACL | University of Notre Dame; Uniphore | B |
| 926 | Authorship Attribution in Multilingual Machine-Generated Texts. | 2026 | ACL | University of Calabria; Kempelen Institute of Intelligent Technologies | B |
| 927 | Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions | 2026 | ICML | Centre for Credible AI, University of Warsaw; Children Clinical Hospital, Medica | B |
| 928 | AudioStealer: Extracting Audio Prompts via Shapley Value-Guided Query Search. | 2026 | ACL | Hong Kong Polytechnic University; New York University | B |
| 929 | Attribution-Guided Decoding | 2026 | ICLR | Fraunhofer HHI; Fraunhofer Heinrich Hertz Institut; Fraunhofer Heinrich Hertz In | B |
| 930 | Attribution-Based Analysis and Optimization of Modular Agentic Workflows. | 2026 | ACL | Shanghai Jiao Tong University | B |
| 931 | Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models. | 2026 | ACL | Center for Scalable Data Analytics and Artificial Intelligence; Technische Unive | B |
| 932 | Attributing Response to Context: A Jensen–Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation | 2026 | ICLR | Department of Computer Science, University College London; Nanyang Technological | B |
| 933 | Attentive Multi-Layer Fusion for Vision Transformers | 2026 | ICML | Aignostics; Google DeepMind, TU Munich; TU Berlin; Technical University of Munic | B |
| 934 | Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models | 2026 | ICML | Wroclaw University of Science and Technology; Wroclaw University of Science and | B |
| 935 | At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization | 2026 | ICML | INRIA; McGill University; McGill University, McGill University; Meta; Sapient In | B |
| 936 | Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility. | 2026 | ACL | University of Maryland, College Park | B |
| 937 | Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination | 2026 | ICML | Tiktok; University of California, San Diego; University of Illinois at Urbana-Ch | B |
| 938 | Are Reasoning LLMs Robust to Interventions on their Chain-of-Thought? | 2026 | ICLR | Helmholtz AI; Technische Universität München | B |
| 939 | Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering. | 2026 | ACL | Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry | B |
| 940 | Arboreal Neural Network | 2026 | ICML | Beijing University of Technology; Du Xiaoman Technology(BeiJing); DuXiaoman Tech | B |
| 941 | Analysing the Safety Pitfalls of Steering Vectors. | 2026 | ACL | Technical University of Munich | B |
| 942 | An Odd Estimator for Shapley Values | 2026 | ICML | Bielefeld University; Claremont McKenna College; UC Berkeley; UC Berkeley, Snork | B |
| 943 | An Information-Theoretic Parameter-Free Bayesian Framework for Probing Labeled Dependency Trees from Attention Score | 2026 | ICLR | Beijing University of Post and Telecommunication; Beijing University of Posts an | B |
| 944 | All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection. | 2026 | ACL | University of Manchester; Stevens Institute of Technology; Columbia University; | B |
| 945 | All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs | 2026 | ICML | Northwestern University; Peking University; Rutgers; Rutgers University; TU Darm | B |
| 946 | Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation | 2026 | Stanford University | C | |
| 947 | Aligning What LLMs Do and Say: Towards Self-Consistent Explanations. | 2026 | ACL | Technion – Israel Institute of Technology; University of Edinburgh | B |
| 948 | Aligned, Orthogonal or In-conflict: When Can We Safely Optimize Chain-of-Thought? | 2026 | Lab post (Google DeepMind) | Google DeepMind | A |
| 949 | Algorithmic Recourse of In-Context Learning for Tabular Data | 2026 | ICML | HKUST(GZ); KAUST; King Abdullah University of Science and Technology (KAUST); MB | B |
| 950 | Aitchison Embeddings for Learning Compositional Graph Representations | 2026 | ICML | Natera; University of Peloponnese; Yale University; École Normale Supérieure (EN | B |
| 951 | Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations. | 2026 | ACL | Ludwig-Maximilians-Universität München; UCLouvain; University of Vienna; Georget | B |
| 952 | Aggregate Models, Not Explanations: Improving Feature Importance Estimation | 2026 | ICML | Inria; Roche Pharma Research & Early Development (pRED); Institut de Mathémathiq | B |
| 953 | AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction | 2026 | ICML | Fudan University; Shanghai Jiao Tong University; Shanghai Jiaotong University; T | B |
| 954 | AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems? | 2026 | ICLR | Beijing University of Post and Telecommunications; Guangdong OPPO Mobile Telecom | B |
| 955 | Adversarial Vulnerability from Interference Between Features in Superposition | 2026 | ICML | Bogazici University; Imperial College London | B |
| 956 | Addressing divergent representations from causal interventions on neural networks | 2026 | ICLR | Stanford University | B |
| 957 | AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations. | 2026 | ACL | Zhejiang University; University of Illinois Chicago; Chinese University of Hong | B |
| 958 | Adaptive Node Feature Selection for Graph Neural Networks | 2026 | ICML | Rice University | B |
| 959 | Adaptive Concept Discovery for Interpretable Few-Shot Text Classification | 2026 | ICLR | HKUST; Hong Kong University of Science and Technology | B |
| 960 | Adalina: Adaptive Linear Approximation for the Shapley Value and Beyond | 2026 | ICML | National University of Singapore; University of Waterloo | B |
| 961 | ActivationReasoning: Logical Reasoning in Latent Activation Spaces | 2026 | ICLR | Adobe, hessian.AI; CS Department, TU Darmstadt; German Research Center for AI; M | B |
| 962 | Activation Steering with a Feedback Controller | 2026 | ICLR | Hanoi University of Science and Technology; National University of Singapore | B |
| 963 | Activation Steering for Chain-of-Thought Compression. | 2026 | ACL | University of Southern California; California Southern University | B |
| 964 | Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers | 2026 | ICML | Anthropic; ENS Paris-Saclay / EPFL; EPFL - EPF Lausanne; Independent; MIT Tegmar | BC |
| 965 | Activation Decomposition and Steering for LLM Backdoor Remediation. | 2026 | ACL | Macquarie University | B |
| 966 | Accumulating Context Changes the Beliefs of Language Models | 2026 | Carnegie Mellon University & Princeton University & Standfor | C | |
| 967 | AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features | 2026 | ICLR | Ohio State University, Columbus | B |
| 968 | AURA: Visually Interpretable Affective Understanding via Robust Archetypes | 2026 | ICML | Queen Mary University of London; Queen Mary University of London & Xi'an Jiaoton | B |
| 969 | ATEX-CF: Attack-Informed Counterfactual Explanations for Graph Neural Networks | 2026 | ICLR | AI Institute - University of Central Florida; Aalborg University; Bowling Green | B |
| 970 | ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack | 2026 | ICLR | Korea University | B |
| 971 | APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation. | 2026 | ACL | Tianjin University; Zhejiang University; North University of China; Institute of | B |
| 972 | AIR: Post-training Data Selection for Reasoning via Attention Head Influence | 2026 | ICML | Alibaba Group; Beihang University; Beihang University; Beijing University of Ae | B |
| 973 | AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives. | 2026 | ACL | Arizona State University; Clemson University; Washington University in St. Louis | B |
| 974 | ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall | 2026 | ICLR | Beijing University of Aeronautics and Astronautics; HKUST (GZ); The Hong Kong Un | B |
| 975 | A Syntactic and Semantic Probe into Language Evolution based on Large Language Models. | 2026 | ACL | Dalian University of Technology | B |
| 976 | A Probabilistic Hard Concept Bottleneck for Steerable Generative Models | 2026 | ICLR | Saarland University, Saarland University; Saarland University, Universität des S | B |
| 977 | A Positive Case for Faithfulness: Explanations Help Predict Model Behavior | 2026 | ICML | Arcadia Impact; Google DeepMind; UC Berkeley; University of Oxford | B |
| 978 | A Narrowing Geometry in Contaminated Reasoning | 2026 | ICML | Institute of automation, Chinese academy of science; Institute of automation, Ch | B |
| 979 | A Mechanistic Perspective and Difficulty Metric for Unlearning. | 2026 | ACL | University of Massachusetts Lowell | B |
| 980 | A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies | 2026 | ICML | Amazon; The University of Texas at Austin; University of Texas at Austin; Univer | B |
| 981 | A Lightweight Explainable Guardrail for Prompt Safety. | 2026 | ACL | University of Arizona | B |
| 982 | A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments | 2026 | ICLR | Dartmouth College; Massachusetts Institute of Technology; University of Californ | B |
| 983 | A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insights | 2026 | ICML | International Higher School for Advanced Studies Trieste; SISSA (IT); SISSA (Int | B |
| 984 | A Factorized Low-Rank RNN Framework for Uncovering Independent Neural Latent Dynamics and Connectivity | 2026 | ICML | Emory University; Georgia Institute of Technology; Georgia Institute of Technolo | B |
| 985 | A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle | 2026 | ICML | Fudan University; Xi'an Jiaotong University | B |
| 986 | A Counterfactual Explanation Framework for Retrieval Models. | 2026 | ACL | University of California San Diego; University of Liverpool | B |
| 987 | A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models | 2026 | ICLR | Microsoft Research; The University of Tokyo | B |
| 988 | A Capacity-Based Rationale for Multi-Head Attention | 2026 | ICML | Computer Science and Artificial Intelligence Laboratory, Electrical Engineering | B |
| 989 | A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents | 2026 | ICML | CapitalOne; Fraunhofer HHI; Indiana University at Bloomington; Northeastern Univ | B |
| 990 | A 'Diff' Tool for AI: Finding Behavioral Differences in New Models | 2026 | Lab post (Anthropic) | Anthropic | A |
| 991 | $\texttt{ShaplEIG}$: Bayesian Experimental Design for Shapley Value Estimation | 2026 | ICML | Bielefeld University; LMU; LMU Munich; LMU Munich, MCML; Lamarr Institute, TU Do | B |
| 992 | Can Training Only a Single Transformer Block Match or Even Surpass Full-Parameter RL? | 2025 | — | C | |
| 993 | [Survey] Interpreting Language Models Through Concept Descriptions: A Survey | 2025 | — | C | |
| 994 | [MASK]ED - Language Modeling for Explainable Classification and Disentangling of Socially Unacceptable Discourse. | 2025 | EMNLP | CY Cergy Paris Université | B |
| 995 | Zero-shot protein stability prediction by inverse folding models: a free energy interpretation. | 2025 | NeurIPS | Technical University of Denmark; University of | B |
| 996 | Zero-Shot Natural Language Explanations. | 2025 | ICLR | Vrije Universiteit Brussel, Brussels, Belgium | B |
| 997 | XDAC: XAI-Driven Detection and Attribution of LLM-Generated News Comments in Korean. | 2025 | ACL | National Security Research Institute; Korea Advanced Institute of Science and Te | B |
| 998 | XAIguiFormer: explainable artificial intelligence guided transformer for brain disorder identification. | 2025 | ICLR | B | |
| 999 | X2-DFD: A framework for explainable and extendable Deepfake Detection. | 2025 | NeurIPS | School of Data Science, The Chinese University of Hong Kong, Shenzhen, Guangdong | B |
| 1000 | X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning. | 2025 | EMNLP | Rochester Institute of Technology; DEVCOM Army Research Laboratory; United State | B |
| 1001 | Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA. | 2025 | ACL | University of California, Santa Cruz; Carnegie Mellon University | B |
| 1002 | Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers. | 2025 | ICLR | FZI Research Center for Information Technology; Karlsruhe Institute of Technolog | B |
| 1003 | Who's the Author? How Explanations Impact User Reliance in AI-Assisted Authorship Attribution. | 2025 | EMNLP | University of Maryland | B |
| 1004 | Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. | 2025 | ICML | Pennsylvania State University; Duke University; University of Washington; Nanyan | B |
| 1005 | Where Did That Come From? Sentence-Level Error-Tolerant Attribution. | 2025 | EMNLP | Mila - Quebec Artificial Intelligence Institute; McGill University; Bar-Ilan Uni | B |
| 1006 | When Models Manipulate Manifolds: The Geometry of a Counting Task | 2025 | Lab post (Anthropic) | Anthropic | A |
| 1007 | When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations. | 2025 | ACL | Columbia University; Nanyang Technological University; Meta; Rutgers, The State | B |
| 1008 | What’s the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns | 2025 | Center for Information and Language Processing, LMU Munich | C | |
| 1009 | What should a neuron aim for? Designing local objective functions based on information theory. | 2025 | ICLR | Faculty of Physics, Institute for the Dynamics of Complex Systems, University of | B |
| 1010 | What makes an Ensemble (Un) Interpretable? | 2025 | ICML | Hebrew University of Jerusalem | B |
| 1011 | What Makes a Reward Model a Good Teacher? An Optimization Perspective | 2025 | Princeton University | C | |
| 1012 | What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models. | 2025 | ICML | Dana Foundation; Harvard University | B |
| 1013 | Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model | 2025 | Shanghai Jiao Tong University | C | |
| 1014 | Wasserstein Distances, Neuronal Entanglement, and Sparsity. | 2025 | ICLR | Institute of Science and Technology Austria; Neural Magic; Red Hat | B |
| 1015 | Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations. | 2025 | ICLR | Microsoft Research | B |
| 1016 | WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification. | 2025 | EMNLP | Monash University | B |
| 1017 | WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data. | 2025 | ACL | National University of Singapore; Institute for Infocomm Research; AI Singapore; | B |
| 1018 | Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting. | 2025 | NeurIPS | Toyota Technological Institute at Chicago; University of California, Berkeley | B |
| 1019 | Verbosity-Aware Rationale Reduction: Sentence-Level Rationale Reduction for Efficient and Effective Reasoning. | 2025 | ACL | University of Illinois Urbana-Champaign; Kyungpook National University | B |
| 1020 | Variational Counterfactual Intervention Planning to Achieve Target Outcomes. | 2025 | ICML | University of Science and Technology of China; Nanyang Technological University | B |
| 1021 | Validating Mechanistic Interpretations: An Axiomatic Approach. | 2025 | ICML | University of Wisconsin–Madison; Colorado State University; Scale AI; Carnegie M | B |
| 1022 | VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models. | 2025 | NeurIPS | Nanchang University | B |
| 1023 | VISA: Retrieval Augmented Generation with Visual Source Attribution. | 2025 | ACL | University of Waterloo; CSIRO | B |
| 1024 | VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware Tree. | 2025 | NeurIPS | School of Software, Xi’an Jiaotong University; College of Computer Science and T | B |
| 1025 | V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models. | 2025 | EMNLP | Tongji University; University of Wisconsin–Madison | B |
| 1026 | Using Shapley interactions to understand how models use structure. | 2025 | ACL | Harvard University Press | B |
| 1027 | Unveiling the Magic of Code Reasoning through Hypothesis Decomposition and Amendment | 2025 | USTC (ICLR 2025) | C | |
| 1028 | Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiency. | 2025 | EMNLP | Tianjin University; GMT Technology (Shenzhen) Co., Ltd. (China) | B |
| 1029 | Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders. | 2025 | ACL | University of Science and Technology of China; Sichuan University | B |
| 1030 | Unstructured Evidence Attribution for Long Context Query Focused Summarization. | 2025 | EMNLP | University of Michigan | B |
| 1031 | Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection. | 2025 | ICML | Jinan University; University of Macau; Nanyang Technological University | B |
| 1032 | Unlearning-based Neural Interpretations. | 2025 | ICLR | Massachusetts Institute of Technology; University of Oxford; Pioneer Centre for | B |
| 1033 | Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment. | 2025 | ICML | York University; State Research Center of Virology and Biotechnology VECTOR; Har | B |
| 1034 | Understanding and Mitigating Hallucination in Large Vision-Language Models via Modular Attribution and Intervention. | 2025 | ICLR | B | |
| 1035 | Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs. | 2025 | EMNLP | Beijing Institute for General Artificial Intelligence; Wuhan University | BC |
| 1036 | Understanding and Improving Adversarial Robustness of Neural Probabilistic Circuits. | 2025 | NeurIPS | University of Illinois Urbana-Champaign | B |
| 1037 | Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron. | 2025 | ICLR | Mila - Quebec AI Institute, Montreal, QC, Canada; University of Montreal, QC, Ca | B |
| 1038 | Understanding Refusal in Language Models with Sparse Autoencoders. | 2025 | EMNLP | Nanyang Technological University; Singapore University of Technology and Design; | B |
| 1039 | Understanding Neural Networks Through Sparse Circuits | 2025 | Lab post (OpenAI) | OpenAI | A |
| 1040 | Understanding How Value Neurons Shape the Generation of Specified Values in LLMs. | 2025 | EMNLP | Provable Responsible AI and Data Analytics (PRADA) Lab; King Abdullah University | B |
| 1041 | Uncovering Gaps in How Humans and LLMs Interpret Subjective Language. | 2025 | ICLR | UC Berkeley | B |
| 1042 | UTILITY: Utilizing Explainable Reinforcement Learning to Improve Reinforcement Learning. | 2025 | ICLR | B | |
| 1043 | UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective | 2025 | HKU Team | C | |
| 1044 | Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training. | 2025 | NeurIPS | Zhejiang University; Shandong University | B |
| 1045 | Turning Logic Against Itself: Probing Model Defenses Through Contrastive Questions. | 2025 | EMNLP | Ubiquitous Knowledge Processing Lab (UKP Lab); hessian.AI; Technische Universitä | B |
| 1046 | Tree-of-Quote Prompting Improves Factuality and Attribution in Multi-Hop and Medical Reasoning. | 2025 | EMNLP | University of Oxford; Saarland University; All India Institute of Medical Scienc | B |
| 1047 | Treble Counterfactual VLMs: A Causal Approach to Hallucination. | 2025 | EMNLP | University of Southern California; National University of Singapore; University | B |
| 1048 | Transformer-Based Spatial-Temporal Counterfactual Outcomes Estimation. | 2025 | ICML | National University of Defense Technology; PLA Academy of Military Science | B |
| 1049 | Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders. | 2025 | NeurIPS | Tohoku University; RIKEN | BC |
| 1050 | Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language Models. | 2025 | EMNLP | State Key Lab of AI Safety, Institute of Computing Technology, CAS; Chinese Acad | B |
| 1051 | Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy. | 2025 | EMNLP | T-Tech; HSE University | B |
| 1052 | Tracing the Thoughts of a Large Language Model | 2025 | Lab post (Anthropic) | Anthropic | AC |
| 1053 | Tracing the Roots: Leveraging Temporal Dynamics in Diffusion Trajectories for Origin Attribution. | 2025 | NeurIPS | Imperial College London; Apple | B |
| 1054 | Towards the Pedagogical Steering of Large Language Models for Tutoring: A Case Study with Modeling Productive Failure. | 2025 | ACL | ETH Zurich; École Polytechnique; ETH AI Center; Association for Computational Li | B |
| 1055 | Towards counterfactual fairness through auxiliary variables. | 2025 | ICLR | University of Maryland, College Park, ^2Clemson University | B |
| 1056 | Towards an Explainable Comparison and Alignment of Feature Embeddings. | 2025 | ICML | Chinese University of Hong Kong | B |
| 1057 | Towards a Mechanistic Explanation of Diffusion Model Generalization. | 2025 | ICML | University of British Columbia; Alberta Machine Intelligence Institute | B |
| 1058 | Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures. | 2025 | ICLR | OpenMOSS Team, School of Computer Science, Fudan University | B |
| 1059 | Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization. | 2025 | ICLR | B | |
| 1060 | Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis. | 2025 | ICML | University of Hong Kong; Chinese University of Hong Kong, Shenzhen | B |
| 1061 | Towards Synergistic Path-based Explanations for Knowledge Graph Completion: Exploration and Evaluation. | 2025 | ICLR | Massachusetts Institute of Technology, Cambridge, MA, USA | B |
| 1062 | Towards Robustness and Explainability of Automatic Algorithm Selection. | 2025 | ICML | Hong Kong Polytechnic University; Chongqing University | B |
| 1063 | Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration. | 2025 | ICML | Shanghai University; Tencent Youtu Lab | B |
| 1064 | Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control. | 2025 | ICLR | Independent Researchers; Google DeepMind | B |
| 1065 | Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few. | 2025 | NeurIPS | School of Artificial Intelligence, Beijing University of Posts and Telecommunica | BC |
| 1066 | Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders. | 2025 | NeurIPS | University of Wisconsin–Madison; Queen Mary University of London; University of | B |
| 1067 | Towards Global-level Mechanistic Interpretability: A Perspective of Modular Circuits of Large Language Models. | 2025 | ICML | University of Virginia; Florida State University | B |
| 1068 | Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models. | 2025 | EMNLP | Nanyang Technological University; Institute of High Performance Computing; Agenc | B |
| 1069 | Towards Explaining the Power of Constant-depth Graph Neural Networks for Structured Linear Programming. | 2025 | ICLR | B | |
| 1070 | Towards Explainable Temporal Reasoning in Large Language Models: A Structure-Aware Generative Framework. | 2025 | ACL | Wuhan University; The Hong Kong University of Science and Technology (Guangzhou) | B |
| 1071 | Towards Explainable Hate Speech Detection. | 2025 | ACL | Universitätsmedizin Greifswald; Universität Greifswald | B |
| 1072 | Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning. | 2025 | ICML | Nanyang Technological University | B |
| 1073 | Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales. | 2025 | EMNLP | Harbin Institute of Technology; Pengcheng Laboratory, Shenzhen, China; Shaoguan | B |
| 1074 | Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness. | 2025 | ACL | Chinese Academy of Sciences | B |
| 1075 | Towards Automated Knowledge Integration From Human-Interpretable Representations. | 2025 | ICLR | University of Cambridge | B |
| 1076 | Towards Attributions of Input Variables in a Coalition. | 2025 | ICML | Sun Yat-sen University | B |
| 1077 | Towards Achieving Concept Completeness for Textual Concept Bottleneck Models. | 2025 | EMNLP | Sorbonne Université; Ekimetrics | B |
| 1078 | Toward Real-world Text Image Forgery Localization: Structured and Interpretable Data Synthesis. | 2025 | NeurIPS | Sun Yat-sen University; Beihang University; Guangzhou University; Peng Cheng Lab | B |
| 1079 | Toward Inclusive Language Models: Sparsity-Driven Calibration for Systematic and Interpretable Mitigation of Social Biases in LLMs. | 2025 | EMNLP | George Mason University | B |
| 1080 | Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language Models. | 2025 | EMNLP | University of Illinois Urbana-Champaign, IL, USA | B |
| 1081 | TopInG: Topologically Interpretable Graph Learning via Persistent Rationale Filtration. | 2025 | ICML | Rutgers, The State University of New Jersey | B |
| 1082 | TokenShapley: Token Level Context Attribution with Shapley Value. | 2025 | ACL | TikTok; Princeton University | B |
| 1083 | To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts. | 2025 | ICLR | Duke University | B |
| 1084 | To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models. | 2025 | ICML | Work done during an . Morgan AI Research; J.P. Morgan | B |
| 1085 | To See a World in a Spark of Neuron: Disentangling Multi-Task Interference for Training-Free Model Merging. | 2025 | EMNLP | Xiamen University Malaysia | B |
| 1086 | TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research. | 2025 | EMNLP | Martian; Apart Research; NVIDIA; Thoughtworks; Hubei Polytechnic University | B |
| 1087 | TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop. | 2025 | NeurIPS | School of Computing, University of Connecticut; System Security Department, NEC | B |
| 1088 | ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations. | 2025 | EMNLP | The University of Sydney | B |
| 1089 | Thought Anchors: Which LLM Reasoning Steps Matter? | 2025 | DeepSeek | C | |
| 1090 | ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models. | 2025 | EMNLP | Universidad Católica Santo Domingo | B |
| 1091 | Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation. | 2025 | ICLR | Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, | B |
| 1092 | TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding. | 2025 | ACL | University of Waterloo; State Research Center of Virology and Biotechnology VECT | B |
| 1093 | The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It. | 2025 | EMNLP | University of Trento; EU Business School, Munich; Munich Center for Machine Lear | B |
| 1094 | The Two Paradigms of LLM Detection: Authorship Attribution vs Authorship Verification. | 2025 | ACL | Leipzig University | B |
| 1095 | The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs. | 2025 | EMNLP | Japan Advanced Institute of Science and Technology; RIKEN | B |
| 1096 | The Superposition of Diffusion Models Using the Itô Density Estimator. | 2025 | ICLR | University of Toronto; Vector Institute; University of Oxford; Mila - Quebec AI | B |
| 1097 | The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas. | 2025 | EMNLP | Chinese Academy of Sciences; University of Chinese Academy of Sciences; Zhonggua | B |
| 1098 | The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability? | 2025 | NeurIPS | ETH Zurich; École Polytechnique Fédérale de Lausanne | B |
| 1099 | The Logical Implication Steering Method for Conditional Interventions on Transformer Generation. | 2025 | ICML | Salesforce | B |
| 1100 | The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination | 2025 | UIUC / Columbia University / Northwestern University / Stanford University | C | |
| 1101 | The Knowledge Microscope: Features as Better Analytical Lenses than Neurons. | 2025 | ACL | Institute for Complex Systems; Chinese Academy of Sciences; University of Chines | B |
| 1102 | The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information Steering. | 2025 | ICML | Rutgers, The State University of New Jersey; Stanford University; Google DeepMin | B |
| 1103 | The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness. | 2025 | NeurIPS | Microsoft; ELLIS Institute Tübingen; MPI for Intelligent Systems; Tübingen AI Ce | B |
| 1104 | The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization. | 2025 | NeurIPS | Michigan State University | B |
| 1105 | The Emergence of Abstract Thought in Large Language Models Beyond Any Language | 2025 | National University of Singapore / Peking University | C | |
| 1106 | The Coverage Principle: How Pre-Training Enables Post-Training | 2025 | Microsoft Research/Princeton | C | |
| 1107 | The Computational Complexity of Circuit Discovery for Inner Interpretability. | 2025 | ICLR | University of Bristol; Goethe University Frankfurt; Memorial University of Newfo | B |
| 1108 | The Anatomy of Evidence: An Investigation Into Explainable ICD Coding. | 2025 | ACL | Fraunhofer Institute for Applied and Integrated Security; Lamarr Institute for M | B |
| 1109 | Test-Time Steering for Lossless Text Compression via Weighted Product of Experts. | 2025 | EMNLP | University of British Columbia; State Research Center of Virology and Biotechnol | B |
| 1110 | Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models. | 2025 | NeurIPS | Rutgers University | B |
| 1111 | Temporal Misalignment in ANN-SNN Conversion and its Mitigation via Probabilistic Spiking Neurons. | 2025 | ICML | Department of ML, MBZUAI, Abu Dhabi, UAE; City University of Macau; Jilin Univer | B |
| 1112 | Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations. | 2025 | NeurIPS | α X-LANCE Lab, Department of Computer Science and Engineering; MoE Key Lab of Ar | BC |
| 1113 | Taming Hyperparameter Sensitivity in Data Attribution: Practical Selection Without Costly Retraining. | 2025 | NeurIPS | University of Michigan Ann Arbor; University of Illinois Urbana-Champaign | B |
| 1114 | Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers. | 2025 | EMNLP | National Institute of Informatics; The University of Tokyo; National Taiwan Univ | B |
| 1115 | TS-LIF: A Temporal Segment Spiking Neuron Network for Time Series Forecasting. | 2025 | ICLR | Nanyang Technological University; University of Chinese Academy of Sciences; Ten | B |
| 1116 | TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection. | 2025 | EMNLP | National University of Singapore | B |
| 1117 | TIMING: Temporality-Aware Integrated Gradients for Time Series Explanation. | 2025 | ICML | AITRICS; Korea Advanced Institute of Science and Technology | B |
| 1118 | TAO: Using test-time compute to train efficient LLMs without labeled data | 2025 | databricks | C | |
| 1119 | SynC-LLM: Generation of Large-Scale Synthetic Circuit Code with Hierarchical Language Models. | 2025 | EMNLP | Hong Kong University of Science and Technology | B |
| 1120 | Supervised and Unsupervised Probing of Shortcut Learning: Case Study on the Emergence and Evolution of Syntactic Heuristics in BERT. | 2025 | ACL | KU Leuven | B |
| 1121 | Superposition Yields Robust Neural Scaling. | 2025 | NeurIPS | Massachusetts Institute of Technology | B |
| 1122 | SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines | 2025 | ByteDance | C | |
| 1123 | Steering off Course: Reliability Challenges in Steering Language Models. | 2025 | ACL | The Ohio State University; University of Washington; Fastino AI; Allen Institute | B |
| 1124 | Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models. | 2025 | ACL | Georgia Institute of Technology; Apple (United States) | B |
| 1125 | Steering Protein Language Models. | 2025 | ICML | Tencent AI Lab | B |
| 1126 | Steering Protein Family Design through Profile Bayesian Flow. | 2025 | ICLR | Institute of AI Industry Research (AIR), Tsinghua University; School of Pharmace | B |
| 1127 | Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction. | 2025 | ICLR | Université de Montréal, ^ 2 Mila, ^ 3 Dreamfold, ^ 4 Duke University, ^ 5 McGill | B |
| 1128 | Steering Large Language Models between Code Execution and Textual Reasoning. | 2025 | ICLR | Microsoft | B |
| 1129 | Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect. | 2025 | EMNLP | University of Mannheim; Ho Technical University | B |
| 1130 | Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation. | 2025 | EMNLP | Southeast University; Key Laboratory of New Generation Artificial Intelligence T | B |
| 1131 | Steering LLM Reasoning Through Bias-Only Adaptation. | 2025 | EMNLP | T-Tech; Central University | BC |
| 1132 | Steering Information Utility in Key-Value Memory for Language Model Post-Training. | 2025 | NeurIPS | Rice University | BC |
| 1133 | Steering Generative Models with Experimental Data for Protein Fitness Optimization. | 2025 | NeurIPS | California Institute of Technology; Mathematical Sciences; Microsoft Corporation | B |
| 1134 | SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models. | 2025 | EMNLP | Virginia Tech | B |
| 1135 | Start Smart: Leveraging Gradients For Enhancing Mask-based XAI Methods. | 2025 | ICLR | B | |
| 1136 | Spurious Forgetting in Continual Learning of Language Models | 2025 | South China University of Technology | C | |
| 1137 | Splitting & Integrating: Out-of-Distribution Detection via Adversarial Gradient Attribution. | 2025 | ICML | Suzhou University of Technology; University of Malaya; University of Technology | B |
| 1138 | SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking. | 2025 | ICLR | School of Artificial Intelligence, University of Chinese Academy of Sciences; In | B |
| 1139 | Speech Act Patterns for Improving Generalizability of Explainable Politeness Detection Models. | 2025 | ACL | Bentley University | B |
| 1140 | SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder | 2025 | USTC + Huawei | C | |
| 1141 | SparseMVC: Probing Cross-view Sparsity Variations for Multi-view Clustering. | 2025 | NeurIPS | China University of Geosciences; The Hong Kong University of Science and Technol | B |
| 1142 | Sparse autoencoders reveal selective remapping of visual concepts during adaptation. | 2025 | ICLR | Institute of Computational Biology, Computational Health Center, Helmholtz Munic | B |
| 1143 | Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs. | 2025 | EMNLP | Brown University; Drexel University | B |
| 1144 | Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. | 2025 | ICLR | Northeastern University | BC |
| 1145 | Sparse Autoencoders, Again? | 2025 | ICML | Fudan University; Amazon Web Services | B |
| 1146 | Sparse Autoencoders for Hypothesis Generation. | 2025 | ICML | Cornell University | B |
| 1147 | Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models. | 2025 | ICLR | Institute for Human Centered Design; Max Planck Institute for Biological Cyberne | B |
| 1148 | Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models. | 2025 | NeurIPS | Technical University of Munich; Munich Center for Machine Learning; Munich Data | B |
| 1149 | Sparse Autoencoders Do Not Find Canonical Units of Analysis. | 2025 | ICLR | Durham University; Decode Research; Apollo Research | B |
| 1150 | Sparse Autoencoder Features for Classifications and Transferability. | 2025 | EMNLP | Harvard University; Mass General Brigham; Boston Children's Hospital; Johns Hopk | B |
| 1151 | Sound Logical Explanations for Mean Aggregation Graph Neural Networks. | 2025 | NeurIPS | University of Oxford | B |
| 1152 | Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment. | 2025 | EMNLP | Indian Institute of Technology Kharagpur; Eindhoven University of Technology | B |
| 1153 | Solving Satisfiability Modulo Counting Exactly with Probabilistic Circuits. | 2025 | ICML | Purdue University | B |
| 1154 | Small Changes, Big Impact: How Manipulating a Few Neurons Can Drastically Alter LLM Aggression. | 2025 | ACL | Konkuk University; Electronics and Telecommunications Research Institute | B |
| 1155 | Since Faithfulness Fails: The Performance Limits of Neural Causal Discovery. | 2025 | ICML | University of Warsaw; Poznan University of Technology, Poznan, Poland; Polish Ac | B |
| 1156 | Should I Trust You? Detecting Deception in Negotiations using Counterfactual RL. | 2025 | ACL | University of Maryland, College Park; Northwestern University; The University of | B |
| 1157 | Shedding Light on Time Series Classification using Interpretability Gated Networks. | 2025 | ICLR | B | |
| 1158 | Shapley-Guided Utility Learning for Effective Graph Inference Data Valuation. | 2025 | ICLR | Rensselaer Polytechnic Institute, Troy, NY, United States | B |
| 1159 | Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents. | 2025 | NeurIPS | Antai College of Economics and Management, Shanghai Jiao Tong University, Shangh | B |
| 1160 | ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference | 2025 | ByteDance Seed | C | |
| 1161 | Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers. | 2025 | ACL | ENS Paris-Saclay; Université Paris-Saclay; Northeastern University; Princeton Un | B |
| 1162 | Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors. | 2025 | ICLR | School of Informatics, University of Edinburgh; Huawei Technologies Co., Ltd; Sc | B |
| 1163 | SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models. | 2025 | ICML | Massachusetts Institute of Technology; FAIR | B |
| 1164 | Self-Supervised Discovery of Neural Circuits in Spatially Patterned Neural Responses with Graph Neural Networks. | 2025 | NeurIPS | Department of Artificial Intelligence; Hanyang University | B |
| 1165 | Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models. | 2025 | ACL | Chinese Academy of Sciences | B |
| 1166 | Self-Critique and Refinement for Faithful Natural Language Explanations. | 2025 | EMNLP | University of Copenhagen | B |
| 1167 | Scaling up Test-Time Compute with Latent Reasoning | 2025 | ELLIS Institute Tübingen | C | |
| 1168 | Scaling and evaluating sparse autoencoders. | 2025 | ICLR | OpenAI | B |
| 1169 | Scaling Sparse Feature Circuits For Studying In-Context Learning. | 2025 | ICML | ETH Zurich; Georgia Institute of Technology | B |
| 1170 | Scaling Probabilistic Circuits via Monarch Matrices. | 2025 | ICML | University of California, Los Angeles | B |
| 1171 | Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow Matching. | 2025 | NeurIPS | ♣ The Leiden Institute of Advanced Computer Science (LIACS), Leiden University; | B |
| 1172 | Scalable Mechanistic Neural Networks. | 2025 | ICLR | Institute of Science and Technology Austria | B |
| 1173 | Salvage: Shapley-distribution Approximation Learning Via Attribution Guided Exploration for Explainable Image Classification. | 2025 | ICLR | B | |
| 1174 | Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning. | 2025 | EMNLP | Hanyang University | B |
| 1175 | SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning. | 2025 | EMNLP | Shanghai Jiao Tong University; Alibaba Group (China) | B |
| 1176 | SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior. | 2025 | ICML | Allen Institute for Artificial Intelligence | B |
| 1177 | Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety. | 2025 | ACL | Carnegie Mellon University | B |
| 1178 | SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations. | 2025 | ICLR | University of Chicago, ^ 2 Virtue AI, ^ 3 University of Illinois, Urbana-Champai | B |
| 1179 | SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals. | 2025 | EMNLP | University of Illinois Urbana-Champaign | B |
| 1180 | STSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex. | 2025 | NeurIPS | Neuroscience Interdepartmental Program, Stanford University, Stanford, CA; Depar | B |
| 1181 | STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment. | 2025 | EMNLP | Nanyang Technological University; Harbin Institute of Technology | B |
| 1182 | SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning. | 2025 | NeurIPS | Zhejiang University | B |
| 1183 | SPEX: Scaling Feature Interaction Explanations for LLMs. | 2025 | ICML | University of California, Berkeley | B |
| 1184 | SHARP: Steering Hallucination in LVLMs via Representation Engineering. | 2025 | EMNLP | State Key Laboratory of Pattern Recognition; State Key Laboratory of Multimodal | B |
| 1185 | SHAP zero Explains Biological Sequence Models with Near-zero Marginal Cost for Future Queries. | 2025 | NeurIPS | Georgia Institute of Technology | B |
| 1186 | SEK: Self-Explained Keywords Empower Large Language Models for Code Generation. | 2025 | ACL | Zhejiang University | B |
| 1187 | SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling. | 2025 | EMNLP | EPFL | B |
| 1188 | SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs. | 2025 | NeurIPS | University of Electronic Science and Technology of China (UESTC); Shenzhen Insti | B |
| 1189 | SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation | 2025 | Southeast University | C | |
| 1190 | SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text Generation. | 2025 | ICLR | Sorbonne Université, CNRS, ISIR, F-75005 Paris, France; Miles Team, LAMSADE, Uni | B |
| 1191 | SCDTour: Embedding Axis Ordering and Merging for Interpretable Semantic Change Detection. | 2025 | EMNLP | University of Liverpool | B |
| 1192 | SCBench: A KV Cache-Centric Analysis of Long-Context Methods | 2025 | Microsoft | C | |
| 1193 | SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders. | 2025 | ICML | Warsaw University of Technology; IDEAS Research Institute | B |
| 1194 | SAKE: Steering Activations for Knowledge Editing. | 2025 | ACL | ÅF (Switzerland); AXA; Sorbonne Université; Polish Academy of Sciences | B |
| 1195 | SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs. | 2025 | EMNLP | Texas A&M University; University of Milano-Bicocca; Hamad bin Khalifa University | B |
| 1196 | SAEs Are Good for Steering - If You Select the Right Features. | 2025 | EMNLP | Technion – Israel Institute of Technology; Boston University | B |
| 1197 | SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders. | 2025 | NeurIPS | Peking University | BC |
| 1198 | SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. | 2025 | ICML | Decode Research; University College London; MATS Research; Anthropic (United Sta | BC |
| 1199 | SAE-V: Interpreting Multimodal Models for Enhanced Alignment. | 2025 | ICML | Peking University | B |
| 1200 | SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models. | 2025 | EMNLP | New Jersey Institute of Technology; Rutgers, The State University of New Jersey; | B |
| 1201 | Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset. | 2025 | ACL | University of Cambridge; SB Intuitions; Tohoku University; RIKEN; NetMind.AI; Un | B |
| 1202 | Route Sparse Autoencoder to Interpret Large Language Models. | 2025 | EMNLP | University of Science and Technology of China; Douyin Co., Ltd | B |
| 1203 | Robustly identifying concepts introduced during chat fine-tuning using crosscoders | 2025 | EPFL、ETHZ、Ecole Normale Supérieure Paris-Saclay、Université P | C | |
| 1204 | Rhetorical Device-Aware Sarcasm Detection with Counterfactual Data Augmentation. | 2025 | ACL | Tongji University; Ministry of Education, Shanghai 201804, China; DC Arts & Huma | B |
| 1205 | Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability. | 2025 | NeurIPS | Tel-Aviv University | B |
| 1206 | Revisiting LLM Value Probing Strategies: Are They Robust and Expressive? | 2025 | EMNLP | University of Michigan; LG AI Research | B |
| 1207 | Revisiting In-context Learning Inference Circuit in Large Language Models. | 2025 | ICLR | Japan Advanced Institute of Science and Technology; RIKEN | B |
| 1208 | Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing. | 2025 | ACL | University of Chinese Academy of Sciences; Institute for Complex Systems; Chines | B |
| 1209 | Retrieve to Explain: Evidence-driven Predictions for Explainable Drug Target Identification. | 2025 | ACL | BenevolentAI (United Kingdom) | B |
| 1210 | RetrievalAttention: ACCELERATING LONG-CONTEXT LLM INFERENCE VIA VECTOR RETRIEVAL | 2025 | Microsoft | C | |
| 1211 | Retrieval Head Mechanistically Explains Long-Context Factuality. | 2025 | ICLR | Peking University; University of Washington; University of Edinburgh | BC |
| 1212 | Rethinking and Improving Autoformalization: Towards a Faithful Metric and a Dependency Retrieval-based Approach. | 2025 | ICLR | B | |
| 1213 | Rethinking Visual Counterfactual Explanations Through Region Constraint. | 2025 | ICLR | University of Warsaw; Warsaw University of Technology; University of Warsaw, War | B |
| 1214 | Rethinking Shapley Value for Negative Interactions in Non-convex Games. | 2025 | ICLR | B | |
| 1215 | Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words. | 2025 | ICLR | The University of Tokyo | B |
| 1216 | Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective | 2025 | Zhejiang University / Tencent | C | |
| 1217 | Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates. | 2025 | NeurIPS | School of Computer Science and Technology; Xi’an Jiaotong University; School of | B |
| 1218 | Restyling Unsupervised Concept Based Interpretable Networks with Generative Models. | 2025 | ICLR | Sorbonne Université; Institut Systèmes Intelligents et de Robotique; Helmholtz Z | B |
| 1219 | Residualized Similarity for Faithfully Explainable Authorship Verification. | 2025 | EMNLP | Department of Computer Science; Department of Linguistics; Institute for Advance | B |
| 1220 | Rescorla-Wagner Steering of LLMs for Undesired Behaviors over Disproportionate Inappropriate Context. | 2025 | EMNLP | University of Illinois Urbana-Champaign | B |
| 1221 | Rescaled Influence Functions: Accurate Data Attribution in High Dimension. | 2025 | NeurIPS | EECS, MIT, Cambridge, MA | B |
| 1222 | Representational Similarity via Interpretable Visual Concepts. | 2025 | ICLR | University of Edinburgh | B |
| 1223 | Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition. | 2025 | NeurIPS | Nanjing University of Aeronautics and Astronautics, Nanjing, China | B |
| 1224 | Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering | 2025 | Xiaomi | C | |
| 1225 | Reinforced Learning Explicit Circuit Representations for Quantum State Characterization from Local Measurements. | 2025 | ICML | University of Hong Kong | B |
| 1226 | Regional Explanations: Bridging Local and Global Variable Importance. | 2025 | NeurIPS | J.P. Morgan AI Research; LaMME, ENSIIE, University Paris Saclay | B |
| 1227 | Refining Attention for Explainable and Noise-Robust Fact-Checking with Transformers. | 2025 | EMNLP | EURECOM | B |
| 1228 | Redundancy Undermines the Trustworthiness of Self-Interpretable GNNs. | 2025 | ICML | University of Electronic Science and Technology of China; Iowa State University | B |
| 1229 | Reducing Hallucinations in Large Vision-Language Models via Latent Space Steering. | 2025 | ICLR | Stanford University | B |
| 1230 | Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation. | 2025 | NeurIPS | Indian Institute of Information Technology Dharwad; Indian Institute of Technolo | B |
| 1231 | Reconsidering Faithfulness in Regular, Self-Explainable and Domain Invariant GNNs. | 2025 | ICLR | University of Trento | B |
| 1232 | Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation. | 2025 | ACL | Westlake University; Hangzhou Dianzi University | B |
| 1233 | Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought. | 2025 | NeurIPS | Meta AI | B |
| 1234 | Reasoning Models Don't Always Say What They Think | 2025 | Lab post (Anthropic) | Anthropic | AC |
| 1235 | Reasoning Elicitation in Language Models via Counterfactual Feedback. | 2025 | ICLR | Harvard University; Microsoft Research Cambridge; Cornell Tech | BC |
| 1236 | Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference. | 2025 | ACL | Idiap Research Institute | B |
| 1237 | ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning | 2025 | Baichuan AI / Tongji University / University of Edinburgh / Zhejiang University | C | |
| 1238 | ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability. | 2025 | ICLR | Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, | B |
| 1239 | ReCoVeR the Target Language: Language Steering without Sacrificing Task Performance. | 2025 | EMNLP | University of Cambridge; University of Würzburg | B |
| 1240 | Rationalize and Align: Enhancing Writing Assistance with Rationale via Self-Training for Improved Alignment. | 2025 | ACL | National University of Singapore | B |
| 1241 | Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability. | 2025 | ACL | University of Science and Technology of China; Beijing University of Posts and T | B |
| 1242 | RankSHAP: Shapley Value Based Feature Attributions for Learning to Rank. | 2025 | ICLR | Manning College of Information and Computer Sciences; University of Massachusett | B |
| 1243 | RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models | 2025 | Houmo AI | C | |
| 1244 | RISE: Radius of Influence based Subgraph Extraction for 3D Molecular Graph Explanation. | 2025 | ICML | Stony Brook University; New Jersey Institute of Technology; Arizona State Univer | B |
| 1245 | RD-MCSA: A Multi-Class Sentiment Analysis Approach Integrating In-Context Classification Rationales and Demonstrations. | 2025 | EMNLP | Beijing Institute of Mathematical Sciences and Applications; Renmin University o | B |
| 1246 | RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals. | 2025 | ICML | University of Chicago | B |
| 1247 | R.I.P.: Better Models by Survival of the Fittest Prompts | 2025 | Meta + New York University + UC Berkeley | C | |
| 1248 | Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models. | 2025 | Tsinghua University | C | |
| 1249 | Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization | 2025 | The University of Tokyo | C | |
| 1250 | Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering. | 2025 | EMNLP | Iowa State University | B |
| 1251 | Quantifying Misattribution Unfairness in Authorship Attribution. | 2025 | ACL | Stony Brook University; University of Pennsylvania | B |
| 1252 | QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models | 2025 | New York University | C | |
| 1253 | QPM: Discrete Optimization for Globally Interpretable Image Classification. | 2025 | ICLR | Institute for Information Processing (tnt); Leibniz University Hannover; Intel L | B |
| 1254 | QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models. | 2025 | EMNLP | Inspired Spine; University of Science and Technology of China; ByteDance; Fudan | B |
| 1255 | Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment | 2025 | Seoul National University | C | |
| 1256 | PychoAgent: Psychology-driven LLM Agents for Explainable Panic Prediction on Social Media during Sudden Disaster Events. | 2025 | EMNLP | National University of Defense Technology; Tsinghua University; Department of Ps | B |
| 1257 | Provably Robust Explainable Graph Neural Networks against Graph Perturbation Attacks. | 2025 | ICLR | Cranberry-Lemon University; University of the Witwatersrand; Illinois Institute | B |
| 1258 | Provably Accurate Shapley Value Estimation via Leverage Score Sampling. | 2025 | ICLR | New York University | B |
| 1259 | ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering. | 2025 | EMNLP | Shandong University; Harvard University | B |
| 1260 | ProtoPairNet: Interpretable Regression through Prototypical Pair Reasoning. | 2025 | NeurIPS | University of Maine; University of New Hampshire at Manchester; Dartmouth Colleg | B |
| 1261 | ProtoLens: Advancing Prototype Learning for Fine-Grained Interpretability in Text Classification. | 2025 | ACL | University University | B |
| 1262 | PropXplain: Can LLMs Enable Explainable Propaganda Detection? | 2025 | EMNLP | University of Toronto | B |
| 1263 | Programming Refusal with Conditional Activation Steering. | 2025 | ICLR | University of Pennsylvania; IBM Research | B |
| 1264 | Probing the Latent Hierarchical Structure of Data via Diffusion Models. | 2025 | ICLR | Institute of Physics, EPFL; EPFL | B |
| 1265 | Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks. | 2025 | ACL | Zhejiang University; Zhejiang Normal University | B |
| 1266 | Probing for Arithmetic Errors in Language Models. | 2025 | EMNLP | ETH Zrich | B |
| 1267 | Probing and Boosting Large Language Models Capabilities via Attention Heads. | 2025 | EMNLP | Harbin Institute of Technology; Peng Cheng Laboratory | B |
| 1268 | Probing Visual Language Priors in VLMs. | 2025 | ICML | University of Michigan | B |
| 1269 | Probing Subphonemes in Morphology Models. | 2025 | ACL | Ben-Gurion University of the Negev | B |
| 1270 | Probing Semantic Routing in Large Mixture-of-Expert Models. | 2025 | EMNLP | Intel Labs; Intel (United Arab Emirates); Oracle | B |
| 1271 | Probing Relative Interaction and Dynamic Calibration in Multi-modal Entity Alignment. | 2025 | ACL | Northeastern University | B |
| 1272 | Probing Political Ideology in Large Language Models: How Latent Political Representations Generalize Across Tasks. | 2025 | EMNLP | University of Chicago | B |
| 1273 | Probing Neural Combinatorial Optimization Models. | 2025 | NeurIPS | Singapore Management University; Massachusetts Institute of; Technology | B |
| 1274 | Probing Narrative Morals: A New Character-Focused MFT Framework for Use with Large Language Models. | 2025 | EMNLP | McGill University | B |
| 1275 | Probing Logical Reasoning of MLLMs in Scientific Diagrams. | 2025 | EMNLP | University of Pittsburgh | B |
| 1276 | Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set. | 2025 | ACL | EU Business School, Munich; Munich Center for Machine Learning | B |
| 1277 | Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding. | 2025 | EMNLP | University of Wisconsin–Madison | B |
| 1278 | Probing Equivariance and Symmetry Breaking in Convolutional Networks. | 2025 | NeurIPS | AMLab, University of Amsterdam; QurAI, University of Amsterdam; Independent Rese | B |
| 1279 | Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models. | 2025 | ICLR | College of Cyber Science, Nankai University; Independent Researcher; Alibaba Gro | B |
| 1280 | Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing. | 2025 | ICLR | University of Minnesota; University of North Carolina at Charlotte | B |
| 1281 | PrimeX: A Dataset of Worldview, Opinion, and Explanation. | 2025 | EMNLP | Apple (United States); University of Southern California | B |
| 1282 | Prediction via Shapley Value Regression. | 2025 | ICML | KTH Royal Institute of Technology | B |
| 1283 | Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMs. | 2025 | ICLR | University of Science and Technology of China; Department of Computer Science an | B |
| 1284 | Practical do-Shapley Explanations with Estimand-Agnostic Causal Inference. | 2025 | NeurIPS | Barcelona Supercomputing Center | B |
| 1285 | Practical Guide for Model Selection for Real‑World Use Cases (Blog) | 2025 | Openai | C | |
| 1286 | Position-aware Automatic Circuit Discovery. | 2025 | ACL | Technion – Israel Institute of Technology; Northeastern University | BC |
| 1287 | Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs. | 2025 | EMNLP | Computer Science Lab, SRI; Johns Hopkins University | B |
| 1288 | PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration | 2025 | Renmin University of China | C | |
| 1289 | Pixel-level Certified Explanations via Randomized Smoothing. | 2025 | ICML | Helmholtz Center for Information Security | B |
| 1290 | Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis. | 2025 | EMNLP | The Hong Kong University of Science and Technology (Guangzhou); Hong Kong Univer | B |
| 1291 | Personalized Text Generation with Contrastive Activation Steering. | 2025 | ACL | Chinese Academy of Sciences; University of Chinese Academy of Sciences; Northeas | B |
| 1292 | Personality Alignment of Large Language Models | 2025 | Zhejiang University | C | |
| 1293 | Persona Vectors: Monitoring and Controlling Character Traits in Language Models | 2025 | Lab post (Anthropic) | Anthropic | A |
| 1294 | Performative Validity of Recourse Explanations. | 2025 | NeurIPS | Tübingen AI Center, ^2University of Tübingen; Korteweg-de Vries Institute for Ma | B |
| 1295 | ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization | 2025 | Meta AI | C | |
| 1296 | ParamΔ for Direct Weight Mixing: Post-Train Large Language Model at Zero Cost | 2025 | Meta | C | |
| 1297 | ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation. | 2025 | NeurIPS | School of Computer Science and Engineering, Northeastern University, China; Depa | B |
| 1298 | PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-Encoder. | 2025 | ACL | National University of Singapore; Harbin Institute of Technology | B |
| 1299 | POISONING ATTACKS ON LLMS REQUIRE A NEAR-CONSTANT NUMBER OF POISON SAMPLES | 2025 | Anthropic | C | |
| 1300 | Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning. | 2025 | NeurIPS | École Normale Supérieure Paris-Saclay; École Normale Supérieure | B |
| 1301 | Optimal Information Retention for Time-Series Explanations. | 2025 | ICML | Beijing Jiaotong University; Beijing Emergency Medical Center | B |
| 1302 | OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution. | 2025 | EMNLP | University of Calabria | B |
| 1303 | Open-World Authorship Attribution. | 2025 | ACL | National University of Singapore | B |
| 1304 | Open-Sourcing Circuit-Tracing Tools | 2025 | Lab post (Anthropic) | Anthropic | A |
| 1305 | Open Problems in Mechanistic Interpretability | 2025 | Lab post (Google DeepMind) | Google DeepMind | A |
| 1306 | One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models. | 2025 | NeurIPS | École Polytechnique Fédérale de Lausanne; Northeastern University | B |
| 1307 | One Wave To Explain Them All: A Unifying Perspective On Feature Attribution. | 2025 | ICML | Électricité de France (France); Centre Observation, Impacts, Énergie; Hôpital Sa | B |
| 1308 | One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs. | 2025 | NeurIPS | Beijing University of Posts and Telecommunications; Shihezi University | B |
| 1309 | On the Versatility of Sparse Autoencoders for In-Context Learning. | 2025 | EMNLP | University of Southern California | B |
| 1310 | On the Biology of a Large Language Model | 2025 | Lab post (Anthropic) | Anthropic | A |
| 1311 | On Synthesizing Data for Context Attribution in Question Answering. | 2025 | ACL | NEC Laboratories Europe; NEC Laboratories America; NUC Corporation (Japan); Univ | B |
| 1312 | On Relation-Specific Neurons in Large Language Models. | 2025 | EMNLP | Ludwig-Maximilians-Universität München; Technical University of Munich; Google D | B |
| 1313 | On Minimizing Adversarial Counterfactual Error in Adversarial Reinforcement Learning. | 2025 | ICLR | Singapore Management University; Rutgers, The State University of New Jersey | B |
| 1314 | On Logic-based Self-Explainable Graph Neural Networks. | 2025 | NeurIPS | INSA Lyon; EPITA | B |
| 1315 | On Explaining Equivariant Graph Networks via Improved Relevance Propagation. | 2025 | ICML | Texas A&M University; University of Houston | B |
| 1316 | OmniKV: Dynamic Context Selection for Efficient Long-Context LLMs | 2025 | Ant Group | C | |
| 1317 | OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature. | 2025 | EMNLP | University of Massachusetts Amherst; University of Maryland, College Park | B |
| 1318 | ON THE ROLE OF ATTENTION HEADS IN LARGE LANGUAGE MODEL SAFETY | 2025 | Tongyi Lab (ICLR Oral) | C | |
| 1319 | Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning. | 2025 | ICLR | The University of Texas at Austin; University of California San Diego; Universit | B |
| 1320 | Not Lost After All: How Cross-Encoder Attribution Challenges Position Bias Assumptions in LLM Summarization. | 2025 | EMNLP | Dalhousie University; State Research Center of Virology and Biotechnology VECTOR | B |
| 1321 | Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human Diversity. | 2025 | EMNLP | University of Illinois Urbana-Champaign | B |
| 1322 | Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions Explainability. | 2025 | ACL | Corti; University of Copenhagen; IT University of Copenhagen; LUT University | B |
| 1323 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context. | 2025 | EMNLP | Imperial College London | B |
| 1324 | No Black Boxes: Interpretable and Interactable Predictive Healthcare with Knowledge-Enhanced Agentic Causal Discovery. | 2025 | EMNLP | Stevens Institute of Technology | B |
| 1325 | Neurons as Detectors of Coherent Sets in Sensory Dynamics. | 2025 | NeurIPS | Center for Computational Neuroscience, Flatiron Institute, Simons Foundation, Ne | B |
| 1326 | NeuronTune: Towards Self-Guided Spurious Bias Mitigation. | 2025 | ICML | University of Virginia | B |
| 1327 | NeuronMerge: Merging Models via Functional Neuron Groups. | 2025 | ACL | Zhejiang University; Alibaba Group (China); Zhejiang Gongshang University | B |
| 1328 | Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models. | 2025 | ICLR | University of Southern California, CA, USA; University of California, Riverside, | B |
| 1329 | Neuron-Level Sequential Editing for Large Language Models. | 2025 | ACL | University of Science and Technology of China; National University of Singapore; | B |
| 1330 | Neuron-Level Differentiation of Memorization and Generalization in Large Language Models. | 2025 | EMNLP | National Taiwan University; National Tsing Hua University | B |
| 1331 | Neuron based Personality Trait Induction in Large Language Models. | 2025 | ICLR | Gaoling School of Artificial Intelligence, Renmin University of China; Tongyi La | B |
| 1332 | Neuron Platonic Intrinsic Representation From Dynamics Using Contrastive Learning. | 2025 | ICLR | Peking University; University of Georgia; Chinese Academy of Sciences; Institute | B |
| 1333 | Neuron Empirical Gradient: Discovering and Quantifying Neurons' Global Linear Controllability. | 2025 | ACL | Advanced Institute of Industrial Technology; The University of Tokyo | B |
| 1334 | Neuron Activation Modulation for Text Style Transfer: Guiding Large Language Models. | 2025 | ACL | Beijing University of Posts and Telecommunications | B |
| 1335 | NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning. | 2025 | EMNLP | Institute for Bulgarian Language | B |
| 1336 | Neural Interpretable PDEs: Harmonizing Fourier Insights with Attention for Scalable and Interpretable Physics Discovery. | 2025 | ICML | Lehigh University; Global Embedded Technologies (United States) | B |
| 1337 | Neural Causal Graph for Interpretable and Intervenable Classification. | 2025 | ICLR | B | |
| 1338 | NeurFlow: Interpreting Neural Networks through Neuron Groups and Functional Interactions. | 2025 | ICLR | Institute for Learning Innovation; Hanoi University of Science and Technology; U | B |
| 1339 | NetFormer: An interpretable model for recovering dynamical connectivity in neuronal population dynamics. | 2025 | ICLR | B | |
| 1340 | Negative Results for Sparse Autoencoders on Downstream Tasks and Deprioritising SAE Research | 2025 | Lab post (Google DeepMind) | Google DeepMind | A |
| 1341 | NavBench: Probing Multimodal Large Language Models for Embodied Navigation. | 2025 | NeurIPS | The University of Adelaide; The University of Queensland; Mohamed bin Zayed Univ | B |
| 1342 | Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability. | 2025 | ICLR | University of Technology Sydney; University of Sydney | B |
| 1343 | NarratEX Dataset: Explaining the Dominant Narratives in News Texts. | 2025 | EMNLP | Universidade do Porto; Athens University of Economics and Business | B |
| 1344 | NarGINA: Towards Accurate and Interpretable Children's Narrative Ability Assessment via Narrative Graphs. | 2025 | ACL | Nanjing Normal University | B |
| 1345 | NOBLE - Neural Operator with Biologically-informed Latent Embeddings to Capture Experimental Variability in Biological Neuron Models. | 2025 | NeurIPS | ETH Zürich; California Institute of Technology; University of Alberta; Alberta M | B |
| 1346 | Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems. | 2025 | EMNLP | Technische Universität Berlin; German Research Centre for Artificial Intelligenc | B |
| 1347 | Multi-Level Explanations for Generative Language Models. | 2025 | ACL | Harvard University; IBM Research; Merck Research Labs | B |
| 1348 | Multi-Domain Explainability of Preferences. | 2025 | EMNLP | Decision Sciences International Corporation (United States); IIBM Research | B |
| 1349 | Multi-Attribute Steering of Language Models via Targeted Intervention. | 2025 | ACL | University of North Carolina at Chapel Hill; University of North Carolina Health | B |
| 1350 | Monet: Mixture of Monosemantic Experts for Transformers. | 2025 | ICLR | Korea University, ^2KAIST, ^3AIGEN Sciences | B |
| 1351 | MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correction. | 2025 | EMNLP | Carnegie Mellon University; Hong Kong University of Science and Technology | B |
| 1352 | Models of Heavy-Tailed Mechanistic Universality. | 2025 | ICML | The University of Melbourne; University of California, Berkeley; International C | B |
| 1353 | Model Unlearning via Sparse Autoencoder Subspace Guided Projections. | 2025 | EMNLP | University of Hong Kong; Chinese University of Hong Kong, Shenzhen | B |
| 1354 | Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws. | 2025 | ICML | Texas A&M University; Oracle; Indiana University; Google; University of Florida | B |
| 1355 | Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models. | 2025 | ACL | University of Notre Dame; University of Pennsylvania; Georgia Institute of Techn | B |
| 1356 | MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines. | 2025 | ACL | Charles University; Sorbonne Université | B |
| 1357 | Mixture of Experts Made Intrinsically Interpretable. | 2025 | ICML | University of Oxford; National University of Singapore | B |
| 1358 | Mitigating Spurious Correlations via Counterfactual Contrastive Learning. | 2025 | EMNLP | University of Amsterdam; Peking University | B |
| 1359 | Mitigating Overthinking in Large Reasoning Models via Manifold Steering. | 2025 | NeurIPS | Institute of Artificial Intelligence, Beihang University, Beijing 100191, China; | B |
| 1360 | Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values? | 2025 | University of Washington | C | |
| 1361 | MicroEdit: Neuron-level Knowledge Disentanglement and Localization in Lifelong Model Editing. | 2025 | EMNLP | Jilin University; Engineering Research Center of Knowledge-Driven Human-Machine | B |
| 1362 | MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning | 2025 | University of Illinois Urbana-Champaign | C | |
| 1363 | Metric-Driven Attributions for Vision Transformers. | 2025 | ICLR | B | |
| 1364 | MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs. | 2025 | EMNLP | Yale University; Google (United States); University of Toronto | B |
| 1365 | MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social Media. | 2025 | EMNLP | Beijing University of Technology; Wuhan University; Pitié-Salpêtrière Hospital | B |
| 1366 | MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models. | 2025 | EMNLP | Chinese University of Hong Kong; University of International Relations; Tencent | B |
| 1367 | MemeIntel: Explainable Detection of Propagandistic and Hateful Memes. | 2025 | EMNLP | Qatar Computing Research Institute, Qatar; Blackbird.AI; APAVI.AI | B |
| 1368 | MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models. | 2025 | ICLR | Department of Electrical and Computer Engineering, University of Southern Califo | B |
| 1369 | Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization. | 2025 | ICML | University of Maryland; Georgia Institute of Technology; University of Bristol; | B |
| 1370 | Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models. | 2025 | EMNLP | EU Business School, Munich; Munich Center for Machine Learning | B |
| 1371 | Mechanistic Permutability: Match Features Across Layers. | 2025 | ICLR | T-Tech, ^2 Moscow Institute of Physics and Technologies, ^3 HSE University | B |
| 1372 | Mechanistic PDE Networks for Discovery of Governing Equations. | 2025 | ICML | Institute of Science and Technology, Klosterneuburg, Austria; University of Amst | B |
| 1373 | Mechanistic Interpretability of RNNs emulating Hidden Markov Models. | 2025 | NeurIPS | Institute of Neuroinformatics, University of Zurich; ETH Zurich | B |
| 1374 | Mechanistic Interpretability of Emotion Inference in Large Language Models. | 2025 | ACL | University of Southern California; University of California, Los Angeles | B |
| 1375 | Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations. | 2025 | EMNLP | Stanford University | B |
| 1376 | Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models. | 2025 | NeurIPS | Zidi Xiong, Harvard University; Harvard University | B |
| 1377 | Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks. | 2025 | EMNLP | University of Zurich | B |
| 1378 | Measuring and Guiding Monosemanticity. | 2025 | NeurIPS | Technische Universität Darmstadt; German Research Center for AI (DFKI) | B |
| 1379 | Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse. | 2025 | ICLR | Singapore University of Technology and Design; DSO National Laboratories | B |
| 1380 | Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps. | 2025 | EMNLP | Technion – Israel Institute of Technology; University of Utah | BC |
| 1381 | Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning. | 2025 | NeurIPS | Hong Kong University of Science and Technology; Mila - Québec AI Institute; Univ | B |
| 1382 | MarathiEmoExplain: A Dataset for Sentiment, Emotion, and Explanation in Low-Resource Marathi. | 2025 | EMNLP | Indian Institute of Technology Jammu | B |
| 1383 | Make Information Diffusion Explainable: LLM-based Causal Framework for Diffusion Prediction. | 2025 | NeurIPS | Tianjin University | B |
| 1384 | MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Model. | 2025 | EMNLP | Seoul National University; AIGENDRUG Co., Ltd. (South Korea) | B |
| 1385 | MPRF: Interpretable Stance Detection through Multi-Path Reasoning Framework. | 2025 | EMNLP | University of Chinese Academy of Sciences; State Key Laboratory of AI Safety; Ch | B |
| 1386 | MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking | 2025 | Lab post (Google DeepMind) | Google DeepMind | A |
| 1387 | MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models. | 2025 | NeurIPS | University of Rochester; Purdue University; NVIDIA | B |
| 1388 | MMDEND: Dendrite-Inspired Multi-Branch Multi-Compartment Parallel Spiking Neuron for Sequence Modeling. | 2025 | ACL | Institute for Complex Systems; Chinese Academy of Sciences; University of Chines | B |
| 1389 | MLIP Arena: Advancing Fairness and Transparency in Machine Learning Interatomic Potentials via an Open, Accessible Benchmark Platform. | 2025 | NeurIPS | University of California, Berkeley; Lawrence Berkeley National Laboratory; Imper | B |
| 1390 | MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models. | 2025 | ICLR | School of Artificial Intelligence, University of Chinese Academy of Sciences; Th | B |
| 1391 | MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion Prediction. | 2025 | NeurIPS | Renmin University of China; University of Southern California; Amazon Web Servic | B |
| 1392 | MIB: A Mechanistic Interpretability Benchmark. | 2025 | ICML | University of Maryland; Brown University; Technion – Israel Institute of Technol | B |
| 1393 | MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation. | 2025 | EMNLP | Southern States University | B |
| 1394 | MELODI: Exploring Memory Compression for Long Contexts | 2025 | DeepMind | C | |
| 1395 | MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis. | 2025 | ACL | University of California, Santa Barbara; NEC Laboratories Europe, Heidelberg, Ge | B |
| 1396 | MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data. | 2025 | ACL | Maastricht University | B |
| 1397 | MAPLE: Enhancing Review Generation with Multi-Aspect Prompt LEarning in Explainable Recommendation. | 2025 | ACL | Department of Computer Science and Information Engineering; National Cheng Kung | B |
| 1398 | MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset. | 2025 | ACL | University of Colorado Boulder; University of Chicago; Johannes Gutenberg Univer | B |
| 1399 | MAGE: Model-Level Graph Neural Networks Explanations via Motif-based Graph Generation. | 2025 | ICLR | Iowa State University | B |
| 1400 | LucidPPN: Unambiguous Prototypical Parts Network for User-centric Interpretable Computer Vision. | 2025 | ICLR | Jagiellonian University; Institute of Applied Psychology | B |
| 1401 | Low-Rank Adapting Models for Sparse Autoencoders. | 2025 | ICML | Massachusetts Institute of Technology | B |
| 1402 | Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models | 2025 | — | C | |
| 1403 | LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data. | 2025 | ACL | Digital Development Communications International (United States); The Hong Kong | B |
| 1404 | Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs. | 2025 | EMNLP | National Centre for Atmospheric Science | B |
| 1405 | Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs. | 2025 | ACL | University of Toronto; Technische Universität Darmstadt; Microsoft Research | B |
| 1406 | LittleBit: Ultra Low-Bit Quantization via Latent Factorization | 2025 | Samsung Research | C | |
| 1407 | Linguistic Neuron Overlap Patterns to Facilitate Cross-lingual Transfer on Low-resource Languages. | 2025 | EMNLP | Beijing Foreign Studies University; King's College London | B |
| 1408 | LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss. | 2025 | NeurIPS | Apple, ^2Sapienza, ^ | B |
| 1409 | LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference. | 2025 | EMNLP | EU Business School, Munich; Munich Center for Machine Learning; Faculty of Compu | B |
| 1410 | Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language Models. | 2025 | ACL | Centre for Argument Technology; University of Dundee | B |
| 1411 | Leveraging Variation Theory in Counterfactual Data Augmentation for Optimized Active Learning. | 2025 | ACL | University of Notre Dame | B |
| 1412 | Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility. | 2025 | ACL | University of Massachusetts Amherst; Northwestern University | B |
| 1413 | Less is More: Explainable and Efficient ICD Code Prediction with Clinical Entities. | 2025 | ACL | The University of Sydney; CSIRO Data61; Beamtree | B |
| 1414 | Learning to cluster neuronal function. | 2025 | NeurIPS | Institute of Computer Science and Campus Institute Data Science, University Gött | B |
| 1415 | Learning to Steer: Input-dependent Steering for Multimodal LLMs. | 2025 | NeurIPS | ISIR, Sorbonne Université, Paris, France | B |
| 1416 | Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention. | 2025 | ACL | Hong Kong Polytechnic University | B |
| 1417 | Learning and aligning single-neuron invariance manifolds in visual cortex. | 2025 | ICLR | University of Göttingen, Germany; University of Tübingen, Institute for Neurobio | B |
| 1418 | Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning. | 2025 | ACL | Adobe | B |
| 1419 | Learning Multi-Level Features with Matryoshka Sparse Autoencoders. | 2025 | ICML | Independent Researchers (MATS); Google DeepMind | BC |
| 1420 | Learning Interpretable Hierarchical Dynamical Systems Models from Time Series Data. | 2025 | ICLR | Central Institute of Mental Health (CIMH); Interdisciplinary Center for Scientif | B |
| 1421 | Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression | 2025 | Nanyang Technological University | C | |
| 1422 | Learning Counterfactual Outcomes Under Rank Preservation. | 2025 | NeurIPS | Beijing Technology and Business University; Peking University; Zhejiang Universi | B |
| 1423 | LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching. | 2025 | NeurIPS | Munich Center for Machine Learning (MCML), LMU Munich, Germany; Aarhus Universit | B |
| 1424 | LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language Models. | 2025 | NeurIPS | Chinese Academy of Sciences; University of Chinese Academy of Sciences | B |
| 1425 | Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations. | 2025 | EMNLP | University of the District of Columbia; University of Missouri; Columbia Univers | B |
| 1426 | Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion. | 2025 | NeurIPS | Seoul National University | B |
| 1427 | Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability. | 2025 | EMNLP | Northwestern University; New Jersey Institute of Technology; University of Brist | B |
| 1428 | Large Language Models are Interpretable Learners. | 2025 | ICLR | UCLA; Google Research | B |
| 1429 | Large Language Models Are Cross-Lingual Knowledge-Free Reasoners | 2025 | Nanjing University | C | |
| 1430 | LaMAGIC2: Advanced Circuit Formulations for Language Model-Based Analog Topology Generation. | 2025 | ICML | Duke University; University of California, Los Angeles; IBM T. J. Watson Researc | B |
| 1431 | LORE: Continual Logit Rewriting Fosters Faithful Generation. | 2025 | EMNLP | William & Mary | B |
| 1432 | LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering. | 2025 | ACL | Ludwig-Maximilians-Universität München; Munich Research Center, Huawei Technolog | B |
| 1433 | LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing. | 2025 | ACL | Polytechnic University of Bari; Sapienza University of Rome | B |
| 1434 | LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations. | 2025 | EMNLP | University of Oxford; Trinity College Dublin | B |
| 1435 | LLM Interpretability with Identifiable Temporal-Instantaneous Representation. | 2025 | NeurIPS | Carnegie Mellon University; Mohamed bin Zayed University of Artificial Intellige | B |
| 1436 | LIMEFLDL: A Local Interpretable Model-Agnostic Explanations Approach for Label Distribution Learning. | 2025 | ICML | Nanjing University of Science and Technology; Hong Kong Polytechnic University; | B |
| 1437 | LIME: Less Is More for MLLM Evaluation. | 2025 | ACL | University of Manchester | B |
| 1438 | LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement Learning. | 2025 | ICLR | Carnegie Mellon University | B |
| 1439 | LDIR: Low-Dimensional Dense and Interpretable Text Embeddings with Relative Representations. | 2025 | ACL | Shenzhen University | B |
| 1440 | LAQuer: Localized Attribution Queries in Content-grounded Generation. | 2025 | ACL | Bar-Ilan University; University of North Carolina at Chapel Hill | B |
| 1441 | Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models. | 2025 | ACL | National University of Singapore | B |
| 1442 | Know Thyself by Knowing Others: Learning Neuron Identity from Population Context. | 2025 | NeurIPS | University of Pennsylvania, ^ 2 Columbia University, ^ 3 McGill University, ^ 4 | B |
| 1443 | Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents | 2025 | Tsinghua University / Moonshot AI | C | |
| 1444 | Kernel Density Steering: Inference-Time Scaling via Mode Seeking for Image Restoration. | 2025 | NeurIPS | Google, ^2Washington University in St. Louis | B |
| 1445 | KLay: Accelerating Arithmetic Circuits for Neurosymbolic AI. | 2025 | ICLR | Centre for Applied Autonomous Sensor Systems; Örebro University | B |
| 1446 | Just a Scratch: Enhancing LLM Capabilities for Self-harm Detection through Intent Differentiation and Emoji Interpretation. | 2025 | ACL | Fondazione Bruno Kessler; Indian Institute of Technology Patna; Indian Institute | B |
| 1447 | JoPA: Explaining Large Language Model's Generation via Joint Prompt Attribution. | 2025 | ACL | Pennsylvania State University | B |
| 1448 | Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations. | 2025 | ICML | University of Bristol | B |
| 1449 | J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning | 2025 | Meta | C | |
| 1450 | Iterative Vectors: In-Context Gradient Steering without Backpropagation. | 2025 | ICML | Peking University, School of Intelligence Science and Technology, State Key Labo | B |
| 1451 | It Helps to Take a Second Opinion: Teaching Smaller LLMs To Deliberate Mutually via Selective Rationale Optimisation. | 2025 | ICLR | Media and Data Science Research Lab, Adobe | B |
| 1452 | Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness. | 2025 | ICLR | Institute of Computing Technology; University of Chinese Academy of Sciences; Un | B |
| 1453 | Investigating the Impact of Conceptual Metaphors on LLM-based NLI through Shapley Interactions. | 2025 | EMNLP | Leibniz University Hannover; EU Business School, Munich; Bielefeld University; S | B |
| 1454 | Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution. | 2025 | ACL | Indian Institute of Technology Delhi; Independent Researcher; University of Tart | B |
| 1455 | Investigating Pattern Neurons in Urban Time Series Forecasting. | 2025 | ICLR | National University of Singapore | B |
| 1456 | Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors. | 2025 | EMNLP | Nara Institute of Science and Technology; Mohamed bin Zayed University of Artifi | B |
| 1457 | Investigating Context Faithfulness in Large Language Models: The Roles of Memory Strength and Evidence Style. | 2025 | ACL | Department of Computer Science; Iowa State University | B |
| 1458 | Introducing Nested Learning: A new ML paradigm for continual learning | 2025 | Google Research / University of Southern California (USC) | C | |
| 1459 | Intrinsic User-Centric Interpretability through Global Mixture of Experts. | 2025 | ICLR | EPFL | B |
| 1460 | Interpreting the Second-Order Effects of Neurons in CLIP. | 2025 | ICLR | University of California, Berkeley | B |
| 1461 | Interpreting Language Reward Models via Contrastive Explanations. | 2025 | ICLR | Imperial College London; J.P. Morgan AI Research | B |
| 1462 | Interpreting CLIP with Hierarchical Sparse Autoencoders. | 2025 | ICML | University of Warsaw | B |
| 1463 | Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety. | 2025 | EMNLP | Georgia Institute of Technology | B |
| 1464 | Interpretable and Parameter Efficient Graph Neural Additive Models with Random Fourier Features. | 2025 | NeurIPS | Fujitsu Research of India | B |
| 1465 | Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational Pathology. | 2025 | ICLR | University of Electronic Science and Technology of China | B |
| 1466 | Interpretable Unsupervised Joint Denoising and Enhancement for Real-World low-light Scenarios. | 2025 | ICLR | Tsinghua University | B |
| 1467 | Interpretable Text Embeddings and Text Similarity Explanation: A Survey. | 2025 | EMNLP | University of Zurich; University of Stuttgart | B |
| 1468 | Interpretable Next-token Prediction via the Generalized Induction Head. | 2025 | NeurIPS | Microsoft Research; Seoul National University; Stanford University | B |
| 1469 | Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximization. | 2025 | EMNLP | University of Massachusetts Amherst | B |
| 1470 | Interpretable Causal Representation Learning for Biological Data in the Pathway Space. | 2025 | ICLR | CIMA University of Navarra, CCUN, IdiSNA, Pamplona, Spain; TECNUN, University of | B |
| 1471 | Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks. | 2025 | ICLR | The Hong Kong University of Science and Technology; Sun Yat-Sen Memorial Hospita | B |
| 1472 | Interpretability Analysis of Arithmetic In-Context Learning in Large Language Models. | 2025 | EMNLP | University of Tübingen | B |
| 1473 | Interpret and Improve In-Context Learning via the Lens of Input-Label Mappings. | 2025 | ACL | Inspired Spine; University of Science and Technology of China; Independent Age; | B |
| 1474 | Interpolating Neural Network-Tensor Decomposition (INN-TD): a scalable and interpretable approach for large-scale physics-based problems. | 2025 | ICML | Northwestern University; Institute of Computational Modeling | B |
| 1475 | Internal Value Alignment in LLM through Controlled Value Vector Activation | 2025 | — | C | |
| 1476 | InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized Rationales. | 2025 | ICLR | University of Virginia | B |
| 1477 | InstaSHAP: Interpretable Additive Models Explain Shapley Values Instantly. | 2025 | ICLR | University of Southern California | B |
| 1478 | InfoCons: Identifying Interpretable Critical Concepts in Point Clouds via Information Theory. | 2025 | ICML | School of Computer Science, Fudan University, China | B |
| 1479 | Influence Functions for Scalable Data Attribution in Diffusion Models. | 2025 | ICLR | University of Cambridge; University of Toronto; Max Planck Institute for Intelli | B |
| 1480 | Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models | 2025 | Tsinghua University + CMU (ICLR 2025) | C | |
| 1481 | Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning. | 2025 | ICML | Imperial College London | B |
| 1482 | Inducing Argument Facets for Faithful Opinion Summarization. | 2025 | EMNLP | Shandong University | B |
| 1483 | InCoDe: Interpretable Compressed Descriptions For Image Generation. | 2025 | ICLR | B | |
| 1484 | In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention. | 2025 | ICML | Yale University; Nanjing University | B |
| 1485 | Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanations. | 2025 | EMNLP | Peking University; Fuzhou University | B |
| 1486 | Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templates. | 2025 | EMNLP | University of Notre Dame; Amazon | B |
| 1487 | Improving LLM Reasoning through Interpretable Role-Playing Steering. | 2025 | EMNLP | EU Business School, Munich; Northwestern University; Munich Center for Machine L | B |
| 1488 | Improving Instruction-Following in Language Models through Activation Steering. | 2025 | ICLR | ETH Zurich; Microsoft Research (India) | B |
| 1489 | Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced Optimization. | 2025 | ACL | Harbin Institute of Technology; Peng Cheng Laboratory; Northeastern University, | B |
| 1490 | Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE. | 2025 | ACL | NASK National Research Institute; Vrije Universiteit Amsterdam | B |
| 1491 | Improved Representation Steering for Language Models. | 2025 | NeurIPS | Stanford University | B |
| 1492 | Identifying interactions across brain areas while accounting for individual-neuron dynamics with a Transformer-based variational autoencoder. | 2025 | NeurIPS | Carnegie Mellon University | B |
| 1493 | Identifying and Answering Questions with False Assumptions: An Interpretable Approach. | 2025 | EMNLP | Department of Computer Science; University of Arizona | B |
| 1494 | Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework. | 2025 | EMNLP | Hong Kong University of Science and Technology | B |
| 1495 | Identification of Multiple Logical Interpretations in Counter-Arguments. | 2025 | EMNLP | Tohoku University; RIKEN; Beyond Reason; Ricoh; Japan Advanced Institute of Scie | B |
| 1496 | IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response Theory. | 2025 | ACL | University of Science and Technology of China; Institute of Artificial Intellige | B |
| 1497 | IRIS: Interpretable Retrieval-Augmented Classification for Long Interspersed Document Sequences. | 2025 | ACL | Duke University | B |
| 1498 | IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector. | 2025 | NeurIPS | Computer Science and Engineering, Hong Kong University of Science and Technology | B |
| 1499 | ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs. | 2025 | ACL | Peking University | B |
| 1500 | ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests. | 2025 | NeurIPS | Gaoling School of Artificial Intelligence, Renmin University of China | B |
| 1501 | IBCircuit: Towards Holistic Circuit Discovery with Information Bottleneck. | 2025 | ICML | Chinese University of Hong Kong; Alibaba Group (China); The Hong Kong University | B |
| 1502 | I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts. | 2025 | ICML | University of Pennsylvania; University of North Texas; University of Science and | B |
| 1503 | I2AM: Interpreting Image-to-Image Latent Diffusion Models via Bi-Attribution Maps. | 2025 | ICLR | Dongguk University | B |
| 1504 | I-GUARD: Interpretability-Guided Parameter Optimization for Adversarial Defense. | 2025 | EMNLP | King's College London | B |
| 1505 | HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks. | 2025 | ICLR | Stanford University; ♠ Confirm Labs; ♣ Ghent University | B |
| 1506 | Hybrid Re-matching for Continual Learning with Parameter-Efficient Tuning | 2025 | Nankai University & Tsinghua University | C | |
| 1507 | How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations. | 2025 | ICLR | Max Planck Institute for Informatics, Saarland Informatics Campus, Germany, ^2In | B |
| 1508 | How to Generalize the Detection of AI-Generated Text: Confounding Neurons. | 2025 | EMNLP | Italian institute for Genomic Medicine | B |
| 1509 | How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders. | 2025 | EMNLP | RIKEN; Tohoku University | B |
| 1510 | How Programming Concepts and Neurons Are Shared in Code Language Models. | 2025 | ACL | Munich Center for Machine Learning; Sorbonne Université | B |
| 1511 | How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation. | 2025 | EMNLP | Fudan University; University of Southern California | B |
| 1512 | How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis. | 2025 | EMNLP | University of Oxford; Jagiellonian University; Harvard University | B |
| 1513 | How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training. | 2025 | ACL | Zhejiang University; National University of Singapore; ♢Zhejiang Key Laboratory | BC |
| 1514 | High-order Interactions Modeling for Interpretable Multi-Agent Q-Learning. | 2025 | NeurIPS | School of Management and Engineering; Nanjing University; School of Information | B |
| 1515 | High-dimensional neuronal activity from low-dimensional latent dynamics: a solvable model. | 2025 | NeurIPS | University College London; École Polytechnique Fédérale de Lausanne; Shanghai Ji | B |
| 1516 | High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity. | 2025 | ICLR | Dalian University of Technology; vivo Mobile Communication Co., Ltd | B |
| 1517 | Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion Trajectory. | 2025 | NeurIPS | Fudan University; The University of Hong Kong | B |
| 1518 | HiBug2: Efficient and Interpretable Error Slice Discovery for Comprehensive Model Debugging. | 2025 | ICLR | The Chinese University of Hong Kong | B |
| 1519 | HeadMap: Locating and Enhancing Knowledge Circuits in LLMs. | 2025 | ICLR | B | |
| 1520 | Head Pursuit: Probing Attention Specialization in Multimodal Transformers. | 2025 | NeurIPS | Sapienza University of Rome, Italy; Institute of Science and Technology, Austria | B |
| 1521 | Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models. | 2025 | EMNLP | University of Science, VNU-HCM; Indiana University | B |
| 1522 | HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models. | 2025 | ICLR | Georgia Institute of Technology; University of Oregon; University of Illinois Ur | B |
| 1523 | H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs | 2025 | Tsinghua University | C | |
| 1524 | Gumbel Counterfactual Generation From Language Models. | 2025 | ICLR | Bar-Ilan University; Allen Institute for Artificial Intelligence | B |
| 1525 | Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups. | 2025 | EMNLP | London School of Economics and Political Science | B |
| 1526 | GraphNarrator: Generating Textual Explanations for Graph Neural Networks. | 2025 | ACL | Emory University | B |
| 1527 | Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models. | 2025 | ICML | Monash University; Nanjing University of Science and Technology; Shanghai Jiao T | B |
| 1528 | Graph-Guided Textual Explanation Generation Framework. | 2025 | EMNLP | University of Copenhagen; University of Mannheim; University of Technology Nurem | B |
| 1529 | Graph Inverse Style Transfer for Counterfactual Explainability. | 2025 | ICML | Sapienza University of Rome | B |
| 1530 | Gradient-based Explanations for Deep Learning Survival Models. | 2025 | ICML | Leibniz Institute for Prevention Research and Epidemiology - BIPS; University of | B |
| 1531 | GradPS: Resolving Futile Neurons in Parameter Sharing Network for Multi-Agent Reinforcement Learning. | 2025 | ICML | Xiamen University; Key Laboratory of Multimedia Trusted Perception and Efficient | B |
| 1532 | Going Beyond Static: Understanding Shifts with Time-Series Attribution. | 2025 | ICLR | Tsinghua University, Department of Computer Science and Technology, Beijing, Chi | B |
| 1533 | Gnothi Seauton: Empowering Faithful Self-Interpretability in Black-Box Transformers. | 2025 | ICLR | School of Artificial Intelligence, Shanghai Jiao Tong University; Efficient and | B |
| 1534 | Generating Likely Counterfactuals Using Sum-Product Networks. | 2025 | ICLR | Faculty of Electrical Engineering, Czech Technical University | B |
| 1535 | Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution. | 2025 | ACL | Chinese University of Hong Kong; Massachusetts Institute of Technology; Associat | B |
| 1536 | Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow. | 2025 | ACL | Simon Fraser University; University of California, Berkeley | B |
| 1537 | Gemma Scope 2 | 2025 | Lab post (Google DeepMind) | Google DeepMind | A |
| 1538 | Gaussian Mixture Counterfactual Generator. | 2025 | ICLR | B | |
| 1539 | GSE: Group-wise Sparse and Explainable Adversarial Attacks. | 2025 | ICLR | Department for AI in Society, Science, and Technology, Zuse Institute Berlin, Ge | B |
| 1540 | GEMS: Generation-Based Event Argument Extraction via Multi-perspective Prompts and Ontology Steering. | 2025 | ACL | University of Electronic Science and Technology of China | B |
| 1541 | GEFA: A General Feature Attribution Framework Using Proxy Gradient Estimation. | 2025 | ICML | Freie Universität Berlin | B |
| 1542 | From Synapses to Dynamics: Obtaining Function from Structure in a Connectome Constrained Model of the Head Direction Circuit. | 2025 | NeurIPS | Institute of Cognitive and Brain Sciences; Massachusetts Institute of Technology | B |
| 1543 | From Shortcuts to Balance: Attribution Analysis of Speech-Text Feature Utilization in Distinguishing Original from Machine-Translated Texts. | 2025 | EMNLP | Center for Language Studies; University of Groningen; Universitat d’Alacant; Uni | B |
| 1544 | From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers | 2025 | The University of Tokyo & RIKEN AIP | C | |
| 1545 | From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models. | 2025 | EMNLP | Microsoft | B |
| 1546 | From Probability to Counterfactuals: the Increasing Complexity of Satisfiability in Pearl's Causal Hierarchy. | 2025 | ICLR | Saarland University, Germany; University of Lübeck, Germany | B |
| 1547 | From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature Selection. | 2025 | ICML | ETH Zurich | B |
| 1548 | From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment. | 2025 | ACL | Xiamen University | B |
| 1549 | From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language Models. | 2025 | ICML | Columbia University; Ginkgo Bioworks, Inc. (United States) | B |
| 1550 | From Imitation to Introspection: Probing Self-Consciousness in Language Models. | 2025 | ACL | Shanghai Artificial Intelligence Laboratory; Tongji University; Fudan University | B |
| 1551 | From GNNs to Trees: Multi-Granular Interpretability for Graph Neural Networks. | 2025 | ICLR | Zhejiang University; Nanyang Technological University | B |
| 1552 | From Black-box to Causal-box: Towards Building More Interpretable Models. | 2025 | NeurIPS | Causal Artificial Intelligence Lab; Columbia University | B |
| 1553 | Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages. | 2025 | ICML | Moody Foundation; University of Notre Dame; MIT-IBM Watson AI Lab, IBM Research | B |
| 1554 | Forward Knows Efficient Backward Path: Saliency-Guided Memory-Efficient Fine-tuning of Large Language Models. | 2025 | ACL | Korea University | B |
| 1555 | Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents. | 2025 | EMNLP | University of Maryland; Adobe Research | B |
| 1556 | Focus On This, Not That! Steering LLMs with Adaptive Feature Specification. | 2025 | ICML | University of Oxford; University of Illinois Urbana-Champaign; University of Chi | B |
| 1557 | FlowSearch: Advancing deep research with dynamic structured knowledge flow | 2025 | Shanghai AI Lab | C | |
| 1558 | FlowMixer: A Depth-Agnostic Neural Architecture for Interpretable Spatiotemporal Forecasting. | 2025 | NeurIPS | New York University in Abu Dhabi; New York University | B |
| 1559 | FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models. | 2025 | EMNLP | School of Intelligent Systems Engineering, Sun Yat-Sen University; DP Technology | B |
| 1560 | Fix False Transparency by Noise Guided Splatting. | 2025 | NeurIPS | Case Western Reserve University | B |
| 1561 | FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation. | 2025 | ACL | Technische Universität Berlin; German Research Centre for Artificial Intelligenc | BC |
| 1562 | First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training | 2025 | Shanghai Jiao Tong University & Zhongguancun Academy | C | |
| 1563 | Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State Tracking. | 2025 | ACL | Key Laboratory of High Confidence Software Technology (PKU), MOE, China; Peking | B |
| 1564 | Final-Model-Only Data Attribution with a Unifying View of Gradient-Based Methods. | 2025 | NeurIPS | IBM Research; Merck Research Labs | B |
| 1565 | FinGrAct: A Framework for FINe-GRrained Evaluation of ACTionability in Explainable Automatic Fact-Checking. | 2025 | EMNLP | Université de Sherbrooke | B |
| 1566 | FiDeLiS: Faithful Reasoning in Large Language Models for Knowledge Graph Question Answering. | 2025 | ACL | National University of Singapore; University of Science and Technology of China | B |
| 1567 | Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations. | 2025 | NeurIPS | University of Maryland, College Park | B |
| 1568 | Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders. | 2025 | ACL | Skolkovo Institute of Science and Technology; AI Foundation and Algorithm Lab; A | B |
| 1569 | Feature Responsiveness Scores: Model-Agnostic Explanations for Recourse. | 2025 | ICLR | Haverford College | B |
| 1570 | Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models. | 2025 | EMNLP | New Jersey Institute of Technology; University of California, Santa Barbara; Geo | B |
| 1571 | FastCAV: Efficient Computation of Concept Activation Vectors for Explaining Deep Neural Networks. | 2025 | ICML | Institute of Data Science, German Aerospace Center, Jena, Germany; Friedrich Sch | B |
| 1572 | Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models. | 2025 | NeurIPS | University of Oxford; Polytechnique Montréal | B |
| 1573 | FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models. | 2025 | ICLR | School of Electronic and Computer Engineering, Peking University; Peking Univers | B |
| 1574 | FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation. | 2025 | ACL | Xiamen University; Hong Kong Polytechnic University; Migu Meland Co., Ltd; Sooch | B |
| 1575 | Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations. | 2025 | ACL | University of Manchester; Idiap Research Institute; University of Sheffield | B |
| 1576 | Faithful Group Shapley Value. | 2025 | NeurIPS | The Ohio State University; Carnegie Mellon University | B |
| 1577 | Faithful Dynamic Imitation Learning from Human Intervention with Dynamic Regret Minimization. | 2025 | NeurIPS | Southeast University; Nanjing University of Science and Technology | B |
| 1578 | FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge. | 2025 | EMNLP | Seoul National University; Adobe Research; LG AI Research | B |
| 1579 | FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows". | 2025 | ICLR | Salesforce AI Research; University of Texas at Austin | B |
| 1580 | Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs | 2025 | — | C | |
| 1581 | Fairness on Principal Stratum: A New Perspective on Counterfactual Fairness. | 2025 | ICML | Peking University; Carnegie Mellon University; Sun Yatsen University; Renmin Uni | B |
| 1582 | FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering. | 2025 | ACL | Zhejiang University; Zhejiang Key Laboratory of Medical Imaging Artificial Intel | B |
| 1583 | Factor Graph-based Interpretable Neural Networks. | 2025 | ICLR | Dalian University of Technology, ^ 2 Jilin University, ^ 3 RMIT University | B |
| 1584 | Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning. | 2025 | NeurIPS | MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, USTC; | B |
| 1585 | Fact Recall, Heuristics or Pure Guesswork? Precise Interpretations of Language Models for Fact Completion. | 2025 | ACL | Chalmers University of Technology; University of Gothenburg; Linköping Universit | B |
| 1586 | FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models. | 2025 | EMNLP | Zhongguancun Laboratory; Renmin University of China; The Hong Kong University of | B |
| 1587 | FaVe: Factored and Verified Search Rationale for Long-form Answer. | 2025 | ACL | Seoul National University | B |
| 1588 | FaCT: Faithful Concept Traces for Explaining Neural Network Decisions. | 2025 | NeurIPS | Max Planck Institute for Informatics, Saarland Informatics Campus, Saarbrücken, | B |
| 1589 | FLARE: Faithful Logic-Aided Reasoning and Exploration. | 2025 | EMNLP | University of Copenhagen; University of Edinburgh; Miniml.AI; Cohere (Canada); N | B |
| 1590 | F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. | 2025 | ACL | MoE Key Lab of Artificial Intelligence, X-LANCE Lab, School of Computer Science; | B |
| 1591 | F-Fidelity: A Robust Framework for Faithfulness Evaluation of Explainable AI. | 2025 | ICLR | NEC Laboratories America, Princeton, United States; University of California, Sa | B |
| 1592 | Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward Pass. | 2025 | EMNLP | TU Dresden; ScaDS.AI Dresden/Leipzig, Germany | B |
| 1593 | Exploring the Translation Mechanism of Large Language Models | 2025 | Harbin Institute of Technology (Shenzhen) / Pengcheng Laboratory | C | |
| 1594 | Exploring Explanations Improves the Robustness of In-Context Learning. | 2025 | ACL | CyberAgent (Japan) | B |
| 1595 | Explanations of GNN on Evolving Graphs via Axiomatic Layer edges. | 2025 | ICLR | B | |
| 1596 | Explaining, Fast and Slow: Abstraction and Refinement of Provable Explanations. | 2025 | ICML | Hebrew University of Jerusalem | B |
| 1597 | Explaining the role of Intrinsic Dimensionality in Adversarial Training. | 2025 | ICML | Qatar Computing Research Institute, HBKU, Doha, Qatar; Dalhousie University | B |
| 1598 | Explaining novel senses using definition generation with open language models. | 2025 | EMNLP | University of Oslo; KU Leuven | B |
| 1599 | Explaining Puzzle Solutions in Natural Language: An Exploratory Study on 6x6 Sudoku. | 2025 | ACL | University of Colorado Boulder; University of Colorado System | B |
| 1600 | Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation. | 2025 | ICLR | The Blavatnik School of Computer Science, Tel Aviv University | B |
| 1601 | Explaining Matters: Leveraging Definitions and Semantic Expansion for Sexism Detection. | 2025 | ACL | University of Warwick; University of Leeds | B |
| 1602 | Explaining Length Bias in LLM-Based Preference Evaluations. | 2025 | EMNLP | University of Southern California | B |
| 1603 | Explaining Differences Between Model Pairs in Natural Language through Sample Learning. | 2025 | EMNLP | University of North Carolina at Chapel Hill | B |
| 1604 | Explainably Safe Reinforcement Learning. | 2025 | NeurIPS | Masaryk University; Graz University of Technology; Technical University of Munic | B |
| 1605 | Explainable Text Classification with LLMs: Enhancing Performance through Dialectical Prompting and Explanation-Guided Training. | 2025 | EMNLP | School of Computing and Artificial Intelligence; Southwestern University of Fina | B |
| 1606 | Explainable Reinforcement Learning from Human Feedback to Improve Alignment. | 2025 | NeurIPS | Department of Electrical Engineering, Pennsylvania State University; Department | B |
| 1607 | Explainable Hallucination through Natural Language Inference Mapping. | 2025 | ACL | University of Bonn; University of Sheffield; Lamarr Institute for Machine Learni | B |
| 1608 | Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation. | 2025 | ACL | King's College London; Southeast University; The Alan Turing Institute | B |
| 1609 | Explainable Concept Generation through Vision-Language Preference Learning for Understanding Neural Networks' Internal Representations. | 2025 | ICML | School of Computing and Augmented Intelligence, Arizona; State University, Tempe | B |
| 1610 | Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics. | 2025 | EMNLP | University of California San Diego; Adobe Research | B |
| 1611 | Explainability and Interpretability of Multilingual Large Language Models: A Survey. | 2025 | EMNLP | University of Cambridge | B |
| 1612 | Explain then Rank: Scale Calibration of Neural Rankers Using Natural Language Explanations from LLMs. | 2025 | ACL | Snowflake Inc. (United States); Dataminr (United States) | B |
| 1613 | Explain Yourself, Briefly! Self-Explaining Neural Networks with Concise Sufficient Reasons. | 2025 | ICLR | IBM Research; The Hebrew University of Jerusalem; Bar-Ilan University | B |
| 1614 | ExpProof : Operationalizing Explanations for Confidential Models with ZKPs. | 2025 | ICML | UC San Diego; Stanford University | B |
| 1615 | Exogenous Isomorphism for Counterfactual Identifiability. | 2025 | ICML | East China Normal University | B |
| 1616 | Exact Computation of Any-Order Shapley Interactions for Graph Neural Networks. | 2025 | ICLR | Universitat de València; University of Padua | B |
| 1617 | ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation. | 2025 | ACL | University of Massachusetts Amherst | B |
| 1618 | ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning. | 2025 | NeurIPS | University of Texas at Austin | B |
| 1619 | Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language Models. | 2025 | ICML | Shenzhen Campus of Sun Yat-Sen University, School of Cyber Science and Technolog | B |
| 1620 | Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable? | 2025 | ICLR | Laboratoire d'Informatique de Grenoble; GIPSA-Lab | B |
| 1621 | Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition. | 2025 | ICML | University of Michigan | B |
| 1622 | Evaluation of Attribution Bias in Generator-Aware Retrieval-Augmented Large Language Models. | 2025 | ACL | Leiden University; University of Strathclyde; eBay; University of Amsterdam | B |
| 1623 | Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration. | 2025 | ACL | Korea Advanced Institute of Science and Technology; KT Corporation; Sogang Unive | B |
| 1624 | Evaluating Neuron Explanations: A Unified Framework with Sanity Checks. | 2025 | ICML | CSE, UC San Diego, CA, USA; HDSI, UC San Diego | B |
| 1625 | Evaluating Chain-of-Thought Monitorability | 2025 | Lab post (OpenAI) | OpenAI | A |
| 1626 | Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis. | 2025 | ACL | University of Chinese Academy of Sciences; Chinese Academy of Sciences; Tsinghua | B |
| 1627 | Enhancing the Comprehensibility of Text Explanations via Unsupervised Concept Discovery. | 2025 | ACL | Chinese Academy of Sciences; University of Chinese Academy of Sciences | B |
| 1628 | Enhancing Uncertainty Estimation and Interpretability with Bayesian Non-negative Decision Layer. | 2025 | ICLR | Xidian University; Xi'an Jiaotong University; Asia School of Business | B |
| 1629 | Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering Perspective. | 2025 | ICML | The University of Queensland; The University of Melbourne; Mohamed bin Zayed Uni | B |
| 1630 | Enhancing Training Data Attribution with Representational Optimization. | 2025 | NeurIPS | Carnegie Mellon University; University of Toronto; Vector Institute | B |
| 1631 | Enhancing Recommendation Explanations through User-Centric Refinement. | 2025 | EMNLP | Renmin University of China; Huawei | B |
| 1632 | Enhancing Pre-trained Representation Classifiability can Boost its Interpretability. | 2025 | ICLR | Key Lab of Intell. Info. Process., Inst. of Comput. Tech., CAS; University of Ch | B |
| 1633 | Enhancing Performance of Explainable AI Models with Constrained Concept Refinement. | 2025 | ICML | University of Michigan; Princeton University | B |
| 1634 | Enhancing Interpretable Image Classification Through LLM Agents and Conditional Concept Bottleneck Models. | 2025 | ACL | Monash University | B |
| 1635 | Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategy. | 2025 | EMNLP | Nara Institute of Science and Technology; University of the Philippines Diliman | B |
| 1636 | Enhancing Graph Of Thought: Enhancing Prompts with LLM Rationales and Dynamic Temperature Control. | 2025 | ICLR | B | |
| 1637 | Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data. | 2025 | ICLR | School of Computing, University of Georgia; Department of Radiology, Massachuset | B |
| 1638 | Enhancing Chain-of-Thought Reasoning via Neuron Activation Differential Analysis. | 2025 | EMNLP | Renmin University of China; iFLYTEK AI Research | B |
| 1639 | Enhancing Automated Interpretability with Output-Centric Feature Descriptions. | 2025 | ACL | Tel Aviv University; R Group | BC |
| 1640 | Enhanced Noun-Noun Compound Interpretation through Textual Enrichment. | 2025 | EMNLP | Brandeis University | B |
| 1641 | EnSToM: Enhancing Dialogue Systems with Entropy-Scaled Steering Vectors for Topic Maintenance. | 2025 | ACL | Pohang University of Science and Technology | B |
| 1642 | Emergent Introspective Awareness in Large Language Models | 2025 | Lab post (Anthropic) | Anthropic | AC |
| 1643 | Emergence and Evolution of Interpretable Concepts in Diffusion Models. | 2025 | NeurIPS | University of Southern California | B |
| 1644 | Eliminating Position Bias of Language Models: A Mechanistic Approach. | 2025 | ICLR | Texas A&M University | B |
| 1645 | Efficiently Verifiable Proofs of Data Attribution. | 2025 | NeurIPS | Morgan Stanley Machine Learning Research and Harvard Business School; Google Res | B |
| 1646 | Efficient and Generalizable Mixed-Precision Quantization via Topological Entropy | 2025 | Shanxi University & Northeastern University | C | |
| 1647 | Efficient and Accurate Explanation Estimation with Distribution Compression. | 2025 | ICLR | University of Warsaw; Munich Center for Machine Learning; Warsaw University of T | B |
| 1648 | Efficient Neuron Segmentation in Electron Microscopy by Affinity-Guided Queries. | 2025 | ICLR | B | |
| 1649 | Efficient Dictionary Learning with Switch Sparse Autoencoders. | 2025 | ICLR | Massachusetts Institute of Technology; University of Oxford | BC |
| 1650 | Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition. | 2025 | ICLR | CSAIL, MIT; Center for Computational Biology | B |
| 1651 | Effective and Efficient Time-Varying Counterfactual Prediction with State-Space Models. | 2025 | ICLR | B | |
| 1652 | Editable Concept Bottleneck Models. | 2025 | ICML | King Abdullah University of Science and Technology; Shanghai Jiao Tong Universit | B |
| 1653 | Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs. | 2025 | NeurIPS | Key Lab of Intell. Info. Process., Inst. of Comput. Tech., CAS; University of Ch | B |
| 1654 | EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations. | 2025 | ACL | Seoul National University; Coxwave | B |
| 1655 | ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations. | 2025 | ACL | University of Southern California; Stanford University | B |
| 1656 | EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification. | 2025 | NeurIPS | King Abdullah University of Science and Technology; Harbin Institute of Technolo | B |
| 1657 | E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel Time. | 2025 | ICML | Dartmouth College | B |
| 1658 | Dynamic Steering With Episodic Memory For Large Language Models. | 2025 | ACL | Deakin University; Meta | B |
| 1659 | Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models | 2025 | East China Normal University | C | |
| 1660 | Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models. | 2025 | ACL | Ping An (China); University of Electronic Science and Technology of China | B |
| 1661 | Dually Self-Improved Counterfactual Data Augmentation Using Large Language Model. | 2025 | ACL | Beijing Institute of Technology; Harbin Institute of Technology | B |
| 1662 | Dual-Path Counterfactual Integration for Multimodal Aspect-Based Sentiment Classification. | 2025 | EMNLP | Chinese Academy of Sciences; China Mobile (China); ByteDance; GMT Technology (Sh | B |
| 1663 | Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs. | 2025 | EMNLP | Zhejiang University; Shanghai Innovation Institute; Southeast University; Shangh | B |
| 1664 | Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic Inference. | 2025 | ACL | King's College London | B |
| 1665 | Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains. | 2025 | ACL | Peking University | B |
| 1666 | Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation. | 2025 | ACL | Korea Advanced Institute of Science and Technology | B |
| 1667 | Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models. | 2025 | EMNLP | Hanyang University | B |
| 1668 | DocVXQA: Context-Aware Visual Explanations for Document Question Answering. | 2025 | ICML | Universitat Autònoma de Barcelona; UiT The Arctic University of Norway; Inria | B |
| 1669 | Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing. | 2025 | EMNLP | Munich Center for Machine Learning | B |
| 1670 | Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations? | 2025 | ICLR | Heidelberg University | B |
| 1671 | Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs | 2025 | — | C | |
| 1672 | Do Large Language Models Have "Emotion Neurons"? Investigating the Existence and Role. | 2025 | ACL | Konkuk University | B |
| 1673 | Do LLMs Behave as Claimed? Investigating How LLMs Follow Their Own Claims using Counterfactual Questions. | 2025 | EMNLP | Harbin Institute of Technology | B |
| 1674 | Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs. | 2025 | EMNLP | Kyoto University | B |
| 1675 | Dissecting Persona-Driven Reasoning in Language Models via Activation Patching. | 2025 | EMNLP | Independent Researchers | B |
| 1676 | Disentangling Superpositions: Interpretable Brain Encoding Model with Sparse Concept Atoms. | 2025 | NeurIPS | University of California, Berkeley | B |
| 1677 | Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition. | 2025 | NeurIPS | Kyung Hee University; Korea University | B |
| 1678 | Discursive Circuits: How Do Language Models Understand Discourse Relations? | 2025 | EMNLP | National University of Singapore | B |
| 1679 | Discovering Influential Neuron Path in Vision Transformers. | 2025 | ICLR | ShanghaiTech University, ^2Tencent PCG | B |
| 1680 | Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing. | 2025 | ACL | Language Science (South Korea); University of Edinburgh | B |
| 1681 | Diffusion Counterfactual Generation with Semantic Abduction. | 2025 | ICML | Imperial College London | B |
| 1682 | Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion Models. | 2025 | ICLR | The University of Sydney, ^ 2 City University of Hong Kong | B |
| 1683 | Differentially Private Steering for Large Language Model Alignment. | 2025 | ICLR | Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science an | B |
| 1684 | Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework. | 2025 | ACL | Tencent; Tsinghua University; Chinese Academy of Sciences; Huazhong University o | B |
| 1685 | Detecting Misbehavior in Frontier Reasoning Models | 2025 | Lab post (OpenAI) | OpenAI | A |
| 1686 | Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation. | 2025 | EMNLP | Massachusetts Institute of Technology; Nanyang Technological University | B |
| 1687 | Dense SAE Latents Are Features, Not Bugs. | 2025 | NeurIPS | ETH Zürich; University of Sheffield | B |
| 1688 | Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling. | 2025 | NeurIPS | University of Electronic Science and Technology of China; The Chinese University | B |
| 1689 | Dementia Through Different Eyes: Explainable Modeling of Human and LLM Perceptions for Early Awareness. | 2025 | EMNLP | Decision Sciences International Corporation (United States); University of Washi | B |
| 1690 | DeepSeek-V3.2 and DeepSeekMath-V2 | 2025 | DeepSeek | C | |
| 1691 | DeepLayout: Learning Neural Representations of Circuit Placement Layout. | 2025 | ICML | Peking University; Business Innovation Centre (Czechia); Wuhan University; Chine | B |
| 1692 | DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale. | 2025 | ICLR | The Chinese University of Hong Kong; Shanghai Jiao Tong University | B |
| 1693 | Deep Linear Probe Generators for Weight Space Learning. | 2025 | ICLR | School of Computer Science and Engineering; The Hebrew University of Jerusalem, | B |
| 1694 | Deep Learning Alternatives Of The Kolmogorov Superposition Theorem. | 2025 | ICLR | University of Pennsylvania | B |
| 1695 | Deep Bayesian Filter for Bayes-Faithful Data Assimilation. | 2025 | ICML | Preferred Networks (Japan) | B |
| 1696 | Decoupling Memories, Muting Neurons: Towards Practical Machine Unlearning for Large Language Models. | 2025 | ACL | Huazhong University of Science and Technology | B |
| 1697 | Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains. | 2025 | ACL | Chinese University of Hong Kong; Massachusetts Institute of Technology | B |
| 1698 | Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis. | 2025 | ACL | The Hong Kong University of Science and Technology (Guangzhou); Hong Kong Univer | B |
| 1699 | Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval. | 2025 | EMNLP | Sungkyunkwan University | B |
| 1700 | Decision Information Meets Large Language Models: The Future of Explainable Operations Research. | 2025 | ICLR | Department of Computer Science, City University of Hong Kong; Huawei Noah’s Ark | B |
| 1701 | DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction. | 2025 | ACL | Graduate School of Artificial Intelligence, POSTECH, Republic of Korea; Republic | B |
| 1702 | Data-centric Prediction Explanation via Kernelized Stein Discrepancy. | 2025 | ICLR | Dalhousie University | B |
| 1703 | Data Shapley in One Training Run. | 2025 | ICLR | Princeton University | B |
| 1704 | DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models. | 2025 | EMNLP | Fudan University; Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Ton | B |
| 1705 | DO I KNOW THIS ENTITY? KNOWLEDGE AWARENESS AND HALLUCINATIONS IN LANGUAGE MODELS | 2025 | Universitat Politècnica de Catalunya | C | |
| 1706 | DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models. | 2025 | NeurIPS | University of Catania; University of Central Florida | B |
| 1707 | DCBM: Data-Efficient Visual Concept Bottleneck Models. | 2025 | ICML | University of Mannheim; Clausthal University of Technology; Max Planck Institute | B |
| 1708 | DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models. | 2025 | NeurIPS | Carnegie Mellon University; University of Michigan | B |
| 1709 | DAPO: An Open-Source LLM Reinforcement Learning System at Scale | 2025 | Tsinghua University / ByteDance Seed | C | |
| 1710 | DAPI: Domain Adaptive Toxicity Probe Vector Intervention, for Fine-Grained Detoxification. | 2025 | ACL | Sungkyunkwan University | B |
| 1711 | D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models | 2025 | The Ohio State University / USTC | C | |
| 1712 | Cyclic Counterfactuals under Shift-Scale Interventions. | 2025 | NeurIPS | Indian Statistical Institute | B |
| 1713 | Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models. | 2025 | ACL | Mohamed bin Zayed University of Artificial Intelligence; University of Notre Dam | B |
| 1714 | Cross-Lingual Generalization and Compression: From Language-Specific to Shared Neurons. | 2025 | ACL | Heidelberg University; Association for Computational Linguistics | B |
| 1715 | Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction. | 2025 | EMNLP | Key Laboratory of Aerospace Information Security and Trusted Computing; Ministry | B |
| 1716 | Cracking Factual Knowledge: A Comprehensive Analysis of Degenerate Knowledge Neurons in Large Language Models. | 2025 | ACL | Chinese Academy of Sciences | B |
| 1717 | Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models. | 2025 | ACL | Artificial Intelligence in Medicine (Canada) | B |
| 1718 | Counterfactual reasoning: an analysis of in-context emergence. | 2025 | NeurIPS | Max Planck Institute for Intelligent Systems; ETH Zurich; University of Cambridg | B |
| 1719 | Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation. | 2025 | ICML | University of Illinois Chicago; University of Michigan; LG AI Research (South Ko | B |
| 1720 | Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models. | 2025 | NeurIPS | Renmin University of China; Microsoft Research Asia | B |
| 1721 | Counterfactual Realizability. | 2025 | ICLR | Causal Artificial Intelligence Lab; Columbia University | B |
| 1722 | Counterfactual Identifiability via Dynamic Optimal Transport. | 2025 | NeurIPS | Imperial College London, UK | B |
| 1723 | Counterfactual Graphical Models: Constraints and Inference. | 2025 | ICML | Universidad Autonoma de Manizales | B |
| 1724 | Counterfactual Generative Modeling with Variational Causal Inference. | 2025 | ICLR | University of California, Berkeley | B |
| 1725 | Counterfactual Evolution of Multimodal Datasets via Visual Programming. | 2025 | NeurIPS | Zhejiang University; National University of Singapore; Nanyang Technological Uni | B |
| 1726 | Counterfactual Effect Decomposition in Multi-Agent Sequential Decision Making. | 2025 | ICML | Max Planck Institute for Software Systems, Germany | B |
| 1727 | Counterfactual Contrastive Learning with Normalizing Flows for Robust Treatment Effect Estimation. | 2025 | ICML | Shanxi University; Agency for Science, Technology and Research | B |
| 1728 | Counterfactual Concept Bottleneck Models. | 2025 | ICLR | Università della Svizzera italiana; IBM Research; Work conducted while employed | B |
| 1729 | Correcting on Graph: Faithful Semantic Parsing over Knowledge Graphs with Large Language Models. | 2025 | ACL | Huazhong University of Science and Technology | B |
| 1730 | CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models. | 2025 | ICML | Duke University | B |
| 1731 | Contrastive Prompting Enhances Sentence Embeddings in LLMs through Inference-Time Steering. | 2025 | ACL | Nanjing University | B |
| 1732 | Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models. | 2025 | EMNLP | Xi'an Jiaotong University | B |
| 1733 | Continued Pretraining and Interpretability-Based Evaluation for Low-Resource Languages: A Galician Case Study. | 2025 | ACL | Citrix (Switzerland) | B |
| 1734 | Context-DPO: Aligning Language Models for Context-Faithfulness. | 2025 | ACL | University of Chinese Academy of Sciences; Microsoft (Finland); University of Ca | B |
| 1735 | Context Steering: Controllable Personalization at Inference Time. | 2025 | ICLR | UC Berkeley | BC |
| 1736 | Context Copying Modulation: The Role of Entropy Neurons in Managing Parametric and Contextual Knowledge Conflicts. | 2025 | EMNLP | Sorbonne Université; Institut Systèmes Intelligents et de Robotique; BNP Paribas | B |
| 1737 | Constrain Alignment with Sparse Autoencoders. | 2025 | ICML | King's College London | B |
| 1738 | Constitutional AI: Harmlessness from AI Feedback | 2025 | Anthropic | C | |
| 1739 | Connectome Mapping: Shape-Memory Network via Interpretation of Contextual Semantic Information. | 2025 | ICLR | B | |
| 1740 | ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning. | 2025 | ICLR | University of Edinburgh, ^2Samsung AI Research Centre, Cambridge | B |
| 1741 | ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features. | 2025 | ICML | Virginia Tech; Georgia Institute of Technology | B |
| 1742 | Concept-Centric Token Interpretation for Vector-Quantized Generative Models. | 2025 | ICML | University of Georgia; New Jersey Institute of Technology; New York University | B |
| 1743 | Concept Bottleneck Large Language Models. | 2025 | ICLR | University of California San Diego | B |
| 1744 | Concept Bottleneck Language Models For Protein Design. | 2025 | ICLR | University of California San Diego; Guide Labs; Department of Computer Science, | B |
| 1745 | ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability. | 2025 | ACL | Université Toulouse III - Paul Sabatier; Université Toulouse-I-Capitole; IRT M2P | B |
| 1746 | Computing Circuits Optimization via Model-Based Circuit Genetic Evolution. | 2025 | ICLR | B | |
| 1747 | Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders. | 2025 | ICML | Australian National University; Nazarbayev University; Cold Spring Harbor Labora | B |
| 1748 | Compositional Generalisation for Explainable Hate Speech Detection. | 2025 | EMNLP | University of Edinburgh; Cohere (Canada) | B |
| 1749 | Colloquial Singaporean English Style Transfer with Fine-Grained Explainable Control. | 2025 | ACL | Singapore Management University; DSO National Laboratories; Australian National | B |
| 1750 | Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs | 2025 | — | C | |
| 1751 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models. | 2025 | ACL | The Hong Kong University of Science and Technology (Guangzhou); Tencent; Beijing | B |
| 1752 | CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models. | 2025 | ACL | The University of Sydney | B |
| 1753 | CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark. | 2025 | ICLR | Tokyo Institute of Technology; School of Engineering Westlake Univeristy | B |
| 1754 | CoDy: Counterfactual Explainers for Dynamic Graphs. | 2025 | ICML | Karlsruhe Institute of Technology; University of Amsterdam | B |
| 1755 | CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis. | 2025 | ACL | Shenzhen Research Institute of Big Data; Chinese University of Hong Kong, Shenzh | B |
| 1756 | CoCoA: A Minimum Bayes Risk Framework Bridging Confidence and Consistency for Uncertainty Quantification in LLMs | 2025 | MBZUAI | C | |
| 1757 | CiteEval: Principle-Driven Citation Evaluation for Source Attribution. | 2025 | ACL | AWS AI Labs; Orby AI; Google | B |
| 1758 | Circuits Updates (monthly series, Jan-Dec 2025) | 2025 | Lab post (Anthropic) | Anthropic | A |
| 1759 | CircuitFusion: Multimodal Circuit Representation Learning for Agile Chip Design. | 2025 | ICLR | The Hong Kong University of Science and Technology | B |
| 1760 | Circuit Transformer: A Transformer That Preserves Logical Equivalence. | 2025 | ICLR | University College London | B |
| 1761 | Circuit Tracing: Revealing Computational Graphs in Language Models | 2025 | Lab post (Anthropic) | Anthropic | AC |
| 1762 | Circuit Stability Characterizes Language Model Generalization. | 2025 | ACL | Carnegie Mellon University | B |
| 1763 | Circuit Representation Learning with Masked Gate Modeling and Verilog-AIG Alignment. | 2025 | ICLR | The Chinese University of Hong Kong; Shanghai Artificial Intelligence Laboratory | B |
| 1764 | Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models. | 2025 | ACL | EU Business School, Munich; Munich Center for Machine Learning; University of Os | B |
| 1765 | Circuit Complexity Bounds for RoPE-based Transformer Architecture. | 2025 | EMNLP | Middle Tennessee State University; Stevens Institute of Technology; The Universi | B |
| 1766 | CheXalign: Preference fine-tuning in chest X-ray interpretation models without human feedback. | 2025 | ACL | University of Oxford; Stanford University | B |
| 1767 | ChatGPT and the art of post-training | 2025 | Former OpenAI researchers | C | |
| 1768 | ChartLens: Fine-grained Visual Attribution in Charts. | 2025 | ACL | Adobe Systems (United States); University of Maryland, College Park | B |
| 1769 | Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models. | 2025 | ICLR | Department of Computer Science, Northwestern University, Evanston, IL 60208, USA | B |
| 1770 | Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety | 2025 | Lab post (Joint (OpenAI + DeepMind + Anthropic + others)) | Joint (OpenAI + DeepMind + Anthropic + others) | AC |
| 1771 | Certifying Counterfactual Bias in LLMs. | 2025 | ICLR | UIUC, ^2 Amazon, ^3 Oracle Health | B |
| 1772 | Causally Reliable Concept Bottleneck Models. | 2025 | NeurIPS | Università della Svizzera Italiana; University of Liechtenstein; IBM Research; S | B |
| 1773 | Causality Meets the Table: Debiasing LLMs for Faithful TableQA via Front-Door Intervention. | 2025 | NeurIPS | Anhui University | B |
| 1774 | Causal Logistic Bandits with Counterfactual Fairness Constraints. | 2025 | ICML | Iowa State University; Mohamed bin Zayed University of Artificial Intelligence | B |
| 1775 | Causal Explanation-Guided Learning for Organ Allocation. | 2025 | NeurIPS | Vrije Universiteit Brussel; King's College London | B |
| 1776 | Causal Attribution Analysis for Continuous Outcomes. | 2025 | ICML | Beijing Technology and Business University; Alibaba Group (China) | B |
| 1777 | Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework. | 2025 | NeurIPS | Technische Universität Berlin, Germany; UMI Lab, ATB Potsdam, Germany; Fraunhofe | B |
| 1778 | CapeX: Category-Agnostic Pose Estimation from Textual Point Explanation. | 2025 | ICLR | Tel Aviv University | B |
| 1779 | Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs. | 2025 | EMNLP | Osnabrück University | B |
| 1780 | Can LLMs Explain Themselves Counterfactually? | 2025 | EMNLP | Ruhr University Bochum; University Alliance Ruhr Research Center for Trustworthy | B |
| 1781 | Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs. | 2025 | ACL | Southeast University | B |
| 1782 | Can Input Attributions Explain Inductive Reasoning in In-Context Learning? | 2025 | ACL | Tohoku University; Mohamed bin Zayed University of Artificial Intelligence; RIKE | B |
| 1783 | Calibrating LLM Confidence by Probing Perturbed Representation Stability. | 2025 | EMNLP | Michigan State University; Independent Researcher; JPMorgan AI Research; Henry F | B |
| 1784 | CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners. | 2025 | EMNLP | National University of Singapore; University of California, Los Angeles | B |
| 1785 | CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic. | 2025 | NeurIPS | College of Computer Science and Technology, Zhejiang University, China; Research | B |
| 1786 | COSDA: Counterfactual-based Susceptibility Risk Framework for Open-Set Domain Adaptation. | 2025 | ICML | Ocean University of China; National University of Defense Technology; University | B |
| 1787 | CONDA: Adaptive Concept Bottleneck for Foundation Models Under Distribution Shifts. | 2025 | ICLR | University of Wisconsin-Madison, Madison, USA | B |
| 1788 | COMRECGC: Global Graph Counterfactual Explainer through Common Recourse. | 2025 | ICML | University of Illinois Chicago | B |
| 1789 | CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection. | 2025 | ACL | Yunnan University; National University of Singapore | B |
| 1790 | CLIX: Cross-Lingual Explanations of Idiomatic Expressions. | 2025 | ACL | University of Colorado Boulder | B |
| 1791 | CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction. | 2025 | ACL | Tsinghua University; Huazhong University of Science and Technology; ByteDance; P | B |
| 1792 | CEAES: Bidirectional Reinforcement Learning Optimization for Consistent and Explainable Essay Assessment. | 2025 | ACL | Guangdong University of Foreign Studies | B |
| 1793 | CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environments. | 2025 | EMNLP | École Polytechnique Fédérale de Lausanne | B |
| 1794 | CAIR: Counterfactual-based Agent Influence Ranker for Agentic AI Workflows. | 2025 | EMNLP | Fujitsu (United Kingdom) | B |
| 1795 | Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency. | 2025 | EMNLP | Abu Dhabi University; Abu Dhabi Health Services | B |
| 1796 | Bridging Relevance and Reasoning: Rationale Distillation in Retrieval-Augmented Generation. | 2025 | ACL | City University of Hong Kong; University of Science and Technology of China; Hua | B |
| 1797 | Bridging Brains and Concepts: Interpretable Visual Decoding from fMRI with Semantic Bottlenecks. | 2025 | NeurIPS | Department of Biomedicine and Prevention; University of Rome Tor Vergata; Viale | B |
| 1798 | Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization. | 2025 | ICLR | School of Computer Science and Technology, HUST; Faculty of Artificial Intellige | B |
| 1799 | Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders. | 2025 | EMNLP | Siebel School of Computing and Data Science; University of Illinois Urbana-Champ | B |
| 1800 | Bounds on the computational complexity of neurons due to dendritic morphology. | 2025 | NeurIPS | University of Washington; Allen Institute | B |
| 1801 | BottleHumor: Self-Informed Humor Explanation using the Information Bottleneck Principle. | 2025 | ACL | University of British Columbia; State Research Center of Virology and Biotechnol | B |
| 1802 | Boosting the visual interpretability of CLIP via adversarial fine-tuning. | 2025 | ICLR | B | |
| 1803 | Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation. | 2025 | ACL | University of Chinese Academy of Sciences; Chinese Academy of Sciences; Departme | B |
| 1804 | Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing. | 2025 | NeurIPS | Department of Electronic Engineering, Tsinghua University; School of Computer Sc | B |
| 1805 | BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments | 2025 | Fudan University / Shanghai AI Lab | C | |
| 1806 | Bilinear MLPs enable weight-based mechanistic interpretability. | 2025 | ICLR | University of Antwerp; University of Antwerp, sqIRL/IDLab; Apollo Research | B |
| 1807 | Beyond single neurons: population response geometry in digital twins of mouse visual cortex. | 2025 | ICLR | B | |
| 1808 | Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels. | 2025 | EMNLP | University of Washington; Université Paris Cité | B |
| 1809 | Beyond Topological Self-Explainable GNNs: A Formal Explainability Perspective. | 2025 | ICML | University of Trento | B |
| 1810 | Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing. | 2025 | EMNLP | Peking University | B |
| 1811 | Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms. | 2025 | ACL | Zhejiang University; Tencent; National University of Singapore | B |
| 1812 | Beyond Linear Steering: Unified Multi-Attribute Control for Language Models. | 2025 | EMNLP | University of Oxford | B |
| 1813 | Beyond Last-Click: An Optimal Mechanism for Ad Attribution. | 2025 | NeurIPS | Gaoling School of Artificial Intelligence; Renmin University of China; School of | B |
| 1814 | Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness. | 2025 | ICLR | Peking University; MIT CSAIL; New York University; MIT EECS, CSAIL | B |
| 1815 | Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders. | 2025 | EMNLP | Northwestern University; University of Georgia; New Jersey Institute of Technolo | B |
| 1816 | Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence. | 2025 | ICML | The University of Tokyo | B |
| 1817 | Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits. | 2025 | NeurIPS | Indian Institute of Technology Kanpur (IIT Kanpur) | B |
| 1818 | Beyond Circuit Connections: A Non-Message Passing Graph Transformer Approach for Quantum Error Mitigation. | 2025 | ICLR | B | |
| 1819 | Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases. | 2025 | ACL | Center for Data Science; Department of Linguistics; New York University | B |
| 1820 | Better Training Data Attribution via Better Inverse Hessian-Vector Products. | 2025 | NeurIPS | University of Toronto; Vector Institute for Artificial Intelligence; Tübingen AI | B |
| 1821 | Beneath the Facade: Probing Safety Vulnerabilities in LLMs via Auto-Generated Jailbreak Prompts. | 2025 | EMNLP | School of Computing, KAIST; Graduate School of Data Science, KAIST | B |
| 1822 | Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks. | 2025 | EMNLP | Korea University; Soongsil University | B |
| 1823 | Bayesian Concept Bottleneck Models with LLM Priors. | 2025 | NeurIPS | University of California, San Francisco; Microsoft Research; National University | B |
| 1824 | BEDAA: Bayesian Enhanced DeBERTa for Uncertainty-Aware Authorship Attribution. | 2025 | ACL | University of Manchester; Imperial College London | B |
| 1825 | BANMIME : Misogyny Detection with Metaphor Explanation on Bangla Memes. | 2025 | EMNLP | Chittagong University of Engineering & Technology; Dhaka International Universit | B |
| 1826 | AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders. | 2025 | ICML | Stanford University | B |
| 1827 | Automating Steering for Safe Multimodal Large Language Models. | 2025 | EMNLP | Zhejiang University; National University of Singapore | B |
| 1828 | Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation. | 2025 | ACL | King University; Peking University | B |
| 1829 | Automatically Identifying Local and Global Circuits with Linear Computation Graphs | 2025 | Anthropic | C | |
| 1830 | AutoCT: Automating Interpretable Clinical Trial Prediction with LLM Agents. | 2025 | EMNLP | University of Pennsylvania; Massachusetts Institute of Technology; Oracle; Santa | B |
| 1831 | Auditing Language Models for Hidden Objectives | 2025 | Lab post (Anthropic) | Anthropic | A |
| 1832 | Attribution and Application of Multiple Neurons in Multimodal Large Language Models. | 2025 | EMNLP | Beijing Language and Culture University | B |
| 1833 | AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution. | 2025 | ICLR | University of Toronto, Vector Institute | B |
| 1834 | Attention with Dependency Parsing Augmentation for Fine-Grained Attribution. | 2025 | ACL | Chinese Academy of Sciences; Institute of Computing Technology, CAS, Beijing 100 | B |
| 1835 | Attention Consistency for LLMs Explanation. | 2025 | EMNLP | HB Studio; INALCO; Sorbonne Université; University of Washington; VitaSight | B |
| 1836 | Associative memory and dead neurons. | 2025 | ICLR | AIRI; Skolkovo Institute of Science and Technology | BC |
| 1837 | Artificial Kuramoto Oscillatory Neurons. | 2025 | ICLR | Bernstein Center for Computational Neuroscience Tübingen; Tübingen AI Center | B |
| 1838 | Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) | 2025 | OpenAI | C | |
| 1839 | Artifcial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) | 2025 | University of Washington | C | |
| 1840 | Around the World in 24 Hours: Probing LLM Knowledge of Time and Place. | 2025 | ACL | Universität Hamburg; Bocconi University | B |
| 1841 | Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. | 2025 | ICML | Massachusetts Institute of Technology | B |
| 1842 | Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation. | 2025 | EMNLP | Shanghai Maritime University; École Polytechnique Fédérale de Lausanne; Xi’an Ji | B |
| 1843 | Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice. | 2025 | ACL | Università della Svizzera italiana; National Institute of Informatics; Brown Uni | B |
| 1844 | Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models. | 2025 | ICML | Harvard University; Kempner Institute | B |
| 1845 | ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation [Technical Report] | 2025 | CUHK (Shenzhen) | C | |
| 1846 | Angular Steering: Behavior Control via Rotation in Activation Space. | 2025 | NeurIPS | National University of Singapore | B |
| 1847 | Analyze Feature Flow to Enhance Interpretation and Steering in Language Models. | 2025 | ICML | T-Tech; FXM Research | B |
| 1848 | Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation? | 2025 | EMNLP | University of Sheffield | B |
| 1849 | AnalogGenie: A Generative Engine for Automatic Discovery of Analog Circuit Topologies. | 2025 | ICLR | Northeastern University, ^ 2; The George Washington University | B |
| 1850 | AnalogGenie-Lite: Enhancing Scalability and Precision in Circuit Topology Discovery through Lightweight Graph Modeling. | 2025 | ICML | Department of Electrical and Computer Engineering, Northeastern University, Bost | B |
| 1851 | An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks. | 2025 | ICML | Max Planck Institute for Intelligent Systems; ELLIS Institute Tübingen | B |
| 1852 | An Approach to Technical AGI Safety and Security | 2025 | Lab post (Google DeepMind) | Google DeepMind | A |
| 1853 | All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning | 2025 | CMU / Cornell University | C | |
| 1854 | Aligned at the Start: Conceptual Groupings in LLM Embeddings | 2025 | Virginia Tech | C | |
| 1855 | Africa Health Check: Probing Cultural Bias in Medical LLMs. | 2025 | EMNLP | Georgia Institute of Technology; Google (United States) | B |
| 1856 | Adversarial Attacks on Data Attribution. | 2025 | ICLR | University of Michigan; University of Illinois Urbana-Champaign | B |
| 1857 | Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization. | 2025 | ICML | Universit´e de Montreal 2 Mila - Qu´ebec AI; Institute 3 HEC Montr´eal 4 Univers | B |
| 1858 | Adaptive Transformer Programs: Bridging the Gap Between Performance and Interpretability in Transformers. | 2025 | ICLR | B | |
| 1859 | Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimates. | 2025 | EMNLP | Northeastern University | B |
| 1860 | Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search. | 2025 | NeurIPS | Mohamed bin Zayed University of Artificial Intelligence (MBZUAI); University of | B |
| 1861 | AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness. | 2025 | ACL | Beijing University of Posts and Telecommunications; Hong Kong Baptist University | B |
| 1862 | Active feature acquisition via explainability-driven ranking. | 2025 | ICML | Boston University; Southern Illinois University School of Medicine; Worcester Po | B |
| 1863 | Activation Steering Decoding: Mitigating Hallucination in Large Vision-Language Models through Bidirectional Hidden State Intervention. | 2025 | ACL | Hong Kong Polytechnic University; State Key Laboratory of Pattern Recognition; S | B |
| 1864 | Accelerating Training with Neuron Interaction and Nowcasting Networks. | 2025 | ICLR | Samsung – SAIT AI Lab, Montreal; Concordia University; Université de Montréal; M | B |
| 1865 | Abstract Counterfactuals for Language Model Agents. | 2025 | NeurIPS | King’s College in London | B |
| 1866 | AUTOCIRCUIT-RL: Reinforcement Learning-Driven LLM for Automated Circuit Topology Generation. | 2025 | ICML | IBM Almaden Research Center; IBM Research - Thomas J. Watson Research Center | B |
| 1867 | AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text. | 2025 | ICLR | ♡University of Washington; ♠Allen Institute for Artificial Intelligence | B |
| 1868 | ADIFF: Explaining audio difference using natural language. | 2025 | ICLR | Carnegie Mellon University | B |
| 1869 | A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders. | 2025 | NeurIPS | LASR Labs; University College London; Tübingen AI Center, University of Tübingen | B |
| 1870 | A data and task-constrained mechanistic model of the mouse outer retina shows robustness to contrast variations. | 2025 | NeurIPS | University of Tübingen; University of Washington; Institute for Ophthalmic Resea | B |
| 1871 | A Versatile Influence Function for Data Attribution with Non-Decomposable Loss. | 2025 | ICML | University of Illinois Urbana-Champaign; Carnegie Mellon University | B |
| 1872 | A Unified Framework for Provably Efficient Algorithms to Estimate Shapley Values. | 2025 | NeurIPS | Global Technology Applied Research, JPMorganChase, New York, NY 10001, USA | B |
| 1873 | A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models. | 2025 | EMNLP | Northwestern University; University of Georgia; New Jersey Institute of Technolo | B |
| 1874 | A Survey on Large Language Models with Multilingualism | 2025 | Beijing Jiaotong University + Université de Montréal | C | |
| 1875 | A Survey of Context Engineering for Large Language Models | 2025 | Chinese Academy of Sciences | C | |
| 1876 | A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMs. | 2025 | EMNLP | Nvidia (United Kingdom); University of Groningen | B |
| 1877 | A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression | 2025 | Renmin University of China | C | |
| 1878 | A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI. | 2025 | ACL | EU Business School, Munich; Munich Center for Machine Learning; University of Ca | B |
| 1879 | A Quantum Circuit-Based Compression Perspective for Parameter-Efficient Learning. | 2025 | ICLR | Graduate Institute of Applied Physics, National Taiwan University, Taipei, Taiwa | B |
| 1880 | A New Approach to Backtracking Counterfactual Explanations: A Unified Causal Framework for Efficient Model Interpretability. | 2025 | ICML | Technical University of Munich; Munich Center for Machine Learning; École Polyte | B |
| 1881 | A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations. | 2025 | EMNLP | University of Maryland | B |
| 1882 | A Lens into Interpretable Transformer Mistakes via Semantic Dependency. | 2025 | ICML | The University of Sydney; Hong Kong Baptist University | B |
| 1883 | A Hierarchy of Graphical Models for Counterfactual Inferences. | 2025 | NeurIPS | Causal Artificial Intelligence Lab; Columbia University | B |
| 1884 | A General Framework for Producing Interpretable Semantic Text Embeddings. | 2025 | ICLR | National University of Singapore; Harbin Institute of Technology (Shenzhen) | B |
| 1885 | A General Framework for Inference-time Scaling and Steering of Diffusion Models. | 2025 | ICML | Department of Computer Science, New York University; Columbia University; Center | B |
| 1886 | A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability. | 2025 | ACL | King University | B |
| 1887 | A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models. | 2025 | EMNLP | Department of Computer Science; University of Maryland, College Park | B |
| 1888 | A Causal Lens for Evaluating Faithfulness Metrics. | 2025 | EMNLP | University of North Carolina at Chapel Hill; University of North Carolina Health | B |
| 1889 | "I've Decided to Leak": Probing Internals Behind Prompt Leakage Intents. | 2025 | EMNLP | Tsinghua University; Ant Group (China) | B |
| 1890 | A Survey of Multilingual & Factual Recall Research | 2024 | — | C | |
| 1891 | xTower: A Multilingual LLM for Explaining and Correcting Translation Errors. | 2024 | EMNLP | Instituto de Telecomunicações; Instituto Superior Técnico, Universidade de Lisbo | B |
| 1892 | xMIL: Insightful Explanations for Multiple Instance Learning in Histopathology. | 2024 | NeurIPS | Berlin Institute for the Foundations of Learning and Data; Technische Universitä | B |
| 1893 | shapiq: Shapley Interactions for Machine Learning. | 2024 | NeurIPS | LMU Munich; University of Warsaw; Munich Center for Machine Learning; Warsaw Uni | B |
| 1894 | dattri: A Library for Efficient Data Attribution. | 2024 | NeurIPS | University of Illinois Urbana-Champaign | B |
| 1895 | Yuan 2.0-M32: Mixture of Experts with Attention Router | 2024 | — | C | |
| 1896 | XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs. | 2024 | EMNLP | University of California, Santa Barbara; Nanyang Technological University | B |
| 1897 | Xmodel-1.5: an 1B-Scale Multilingual LLM. Authors: Wang Qun, Liu Yang, Lin Qingquan, Jiang Ling, from XiaoduoAI. Summary: Xmodel-1.5 is a new 1B-parameter multilingual large model developed by the AI Lab of Xiaoduo Technology, pretrained on roughly 2 trillion tokens. The model shows strong performance across many languages, standing out especially on Thai, Arabic and French, while also performing well on Chinese and English. The authors additionally release a Thai evaluation dataset containing hundreds of questions annotated by students of the Integrated Innovation program at Chulalongkorn University, providing a valuable resource for future Thai NLP research. Although the results are encouraging, the authors acknowledge substantial room for improvement. They hope this work advances multilingual AI research and promotes better cross-lingual understanding across a range of NLP tasks. The model and code are publicly available on GitHub. | 2024 | — | C | |
| 1898 | XRec: Large Language Models for Explainable Recommendation. | 2024 | EMNLP | University of Hong Kong | B |
| 1899 | XDetox: Text Detoxification with Token-Level Toxicity Explanations. | 2024 | EMNLP | Hanyang University | B |
| 1900 | X-ACE: Explainable and Multi-factor Audio Captioning Evaluation. | 2024 | ACL | University of Science and Technology of China; Inspired Spine; University of Cal | B |
| 1901 | What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation. | 2024 | ICML | University College London; Google DeepMind (United Kingdom) | B |
| 1902 | What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models. | 2024 | EMNLP | Integrated Vision and Language Lab , KAIST | B |
| 1903 | What does the Knowledge Neuron Thesis Have to do with Knowledge? | 2024 | ICLR | University of Toronto, ^2University of Waterloo, ^3Stevens Institute of Technolo | B |
| 1904 | What Would Gauss Say About Representations? Probing Pretrained Image Models using Synthetic Gaussian Benchmarks. | 2024 | ICML | Massachusetts Institute of Technology; IBM Research | B |
| 1905 | What Matters in Memorizing and Recalling Facts? Multifaceted Benchmarks for Knowledge Probing in Language Models. | 2024 | EMNLP | Advanced Institute of Industrial Technology; The University of Tokyo | B |
| 1906 | What Makes and Breaks Safety Fine-tuning? A Mechanistic Study. | 2024 | NeurIPS | Five AI; University of Michigan; Harvard University; University of Oxford; Max P | B |
| 1907 | What Does Parameter-free Probing Really Uncover? | 2024 | ACL | University of Helsinki | B |
| 1908 | What Do Language Models Hear? Probing for Auditory Representations in Language Models. | 2024 | ACL | Massachusetts Institute of Technology | B |
| 1909 | WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models. | 2024 | NeurIPS | Michigan State University; IBM Research | B |
| 1910 | Visual Pinwheel Centers Act as Geometric Saliency Detectors. | 2024 | NeurIPS | Fudan University; Frontiers Center for Brain Science of the Ministry of Educatio | B |
| 1911 | Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving. | 2024 | EMNLP | University of Manchester; Idiap Research Institute | B |
| 1912 | VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance. | 2024 | NeurIPS | UC San Diego | B |
| 1913 | VISIT: Visualizing and Interpreting the Semantic Information Flow of Transformers | 2024 | — | C | |
| 1914 | VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation. | 2024 | ACL | University of Waterloo; IN.AI Research♡; oiled science; content machine | B |
| 1915 | VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models. | 2024 | ACL | University of California, Los Angeles | B |
| 1916 | Utilizing Human Behavior Modeling to Manipulate Explanations in AI-Assisted Decision Making: The Good, the Bad, and the Scary. | 2024 | NeurIPS | Department of Computer Science; Purdue University | B |
| 1917 | Using Natural Language Explanations to Improve Robustness of In-context Learning. | 2024 | ACL | Weco AI; University College London; University of Edinburgh | B |
| 1918 | Unveiling Factual Recall Behaviors of Large Language Models through Knowledge Neurons. | 2024 | EMNLP | State Key Laboratory of Multimodal Artificial Intelligence Systems; Chinese Acad | B |
| 1919 | Unsupervised Distractor Generation via Large Language Model Distilling and Counterfactual Contrastive Decoding. | 2024 | ACL | King University; Peking University | B |
| 1920 | Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models. | 2024 | EMNLP | Institute for Complex Systems; Chinese Academy of Sciences; University of Chines | B |
| 1921 | Unified Lexical Representation for Interpretable Visual-Language Alignment. | 2024 | NeurIPS | Amazon Web Services; Fudan University | B |
| 1922 | Unelicitable Backdoors via Cryptographic Transformer Circuits. | 2024 | NeurIPS | Contramont Research; Institute of Mathematics and Computer Science, University o | B |
| 1923 | Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective. | 2024 | NeurIPS | The University of Tokyo | B |
| 1924 | Understanding Faithfulness and Reasoning of Large Language Models on Plain Biomedical Summaries. | 2024 | EMNLP | Data61 | B |
| 1925 | Understanding Emergent Abilities of Language Models from the Loss Perspective | 2024 | Zhipu AI | C | |
| 1926 | Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense. | 2024 | NeurIPS | Hong Kong University of Science and Technology; Pennsylvania State University | B |
| 1927 | Unconditional stability of a recurrent neural circuit implementing divisive normalization. | 2024 | NeurIPS | Courant Institute of Mathematical Sciences; Centre for Nano and Soft Matter Scie | B |
| 1928 | UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models. | 2024 | ICLR | Department of Artificial Intelligence; Sungkyunkwan University; Republic of Kore | B |
| 1929 | Trust Regions for Explanations via Black-Box Probabilistic Certification. | 2024 | ICML | IBM Research, Yorktown Heights, NY, USA; University of Tübingen | B |
| 1930 | Tree-of-Counterfactual Prompting for Zero-Shot Stance Detection. | 2024 | ACL | The University of Texas at Dallas; Human Language Technology Research Institute | B |
| 1931 | Transferable and Efficient Non-Factual Content Detection via Probe Training with Offline Consistency Checking. | 2024 | ACL | Renmin University of China; Tsinghua University | B |
| 1932 | Transcoders find interpretable LLM feature circuits. | 2024 | NeurIPS | Yale University; Columbia University | B |
| 1933 | Training for Stable Explanation for Free. | 2024 | NeurIPS | Harbin Institute of Technology; Key Laboratory of Trustworthy Distributed Comput | B |
| 1934 | Training Data Attribution via Approximate Unrolling. | 2024 | NeurIPS | University of Toronto; Vector Institute; NVIDIA | B |
| 1935 | Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech. | 2024 | ACL | Indian Institute of Technology Delhi; Association for Computational Linguistics | B |
| 1936 | Towards a Greek Proverb Atlas: Computational Spatial Exploration and Attribution of Greek Proverbs. | 2024 | EMNLP | Athens University of Economics and Business; Athena Research and Innovation Cent | B |
| 1937 | Towards Verifiable Generation: A Benchmark for Knowledge-aware Language Model Attribution. | 2024 | ACL | Nanyang Technological University; Fudan University; University of California, Sa | B |
| 1938 | Towards Robust Fidelity for Evaluating Explainability of Graph Neural Networks. | 2024 | ICLR | College Information Sciences and Technology, The Pennsylvania State University, | B |
| 1939 | Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights. | 2024 | EMNLP | Monash University | B |
| 1940 | Towards Next-Generation Logic Synthesis: A Scalable Neural Circuit Generation Framework. | 2024 | NeurIPS | Inspired Spine; University of Science and Technology of China; Noah’s Ark Lab, H | B |
| 1941 | Towards Neuron Attributions in Multi-Modal Large Language Models. | 2024 | NeurIPS | University of Science and Technology of China; National University of Singapore; | B |
| 1942 | Towards Multi-dimensional Explanation Alignment for Medical Classification. | 2024 | NeurIPS | Provable Responsible AI and Data Analytics (PRADA) Lab; King Abdullah University | B |
| 1943 | Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models. | 2024 | EMNLP | Apart Research; University of Oxford | B |
| 1944 | Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation. | 2024 | ICML | King Abdullah University of Science and Technology; Harbin Institute of Technolo | B |
| 1945 | Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering. | 2024 | ACL | University of Zurich; ETH Zurich; University of Regensburg; Swiss Finance Instit | B |
| 1946 | Towards Faithful XAI Evaluation via Generalization-Limited Backdoor Watermark. | 2024 | ICLR | Tsinghua University, Department of Mechanical Engineering, Beijing, China | B |
| 1947 | Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering. | 2024 | EMNLP | Harbin Institute of Technology; Queen Mary University of London; XtalPi (China) | B |
| 1948 | Towards Faithful Explanations: Boosting Rationalization with Shortcuts Discovery. | 2024 | ICLR | : State Key Laboratory of Cognitive Intelligence, University of Science and Tech | B |
| 1949 | Towards Explainable Computerized Adaptive Testing with Large Language Model. | 2024 | EMNLP | University of Science and Technology of China; Institute of Artificial Intellige | B |
| 1950 | Towards Explainable Chinese Native Learner Essay Fluency Assessment: Dataset, Tasks, and Method. | 2024 | EMNLP | East China Normal University; Microsoft Research (India) | B |
| 1951 | Towards Characterizing Domain Counterfactuals for Invertible Latent Causal Models. | 2024 | ICLR | Elmore Family School of Electrical and Computer Engineering; Purdue University | B |
| 1952 | Towards Best Practices of Activation Patching in Language Models: Metrics and Methods. | 2024 | ICLR | Google DeepMind | B |
| 1953 | Towards Artwork Explanation in Large-scale Vision Language Models. | 2024 | ACL | Nara Institute of Science and Technology; Google; EleutherAI; University of Cali | B |
| 1954 | Towards 3D Molecule-Text Interpretation in Language Models. | 2024 | ICLR | niversity of Science and Technology of China; ational University of Singapore; o | B |
| 1955 | TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving Scenes. | 2024 | NeurIPS | Institute of Computing Technology; Chinese Academy of Sciences; University of Ch | B |
| 1956 | TimeX++: Learning Time-Series Explanations with Information Bottleneck. | 2024 | ICML | Nanjing University; Microsoft Research (India); Pennsylvania State University; F | B |
| 1957 | Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability. | 2024 | EMNLP | Peking University | B |
| 1958 | The motion planning neural circuit in goal-directed navigation as Lie group operator search. | 2024 | NeurIPS | Center for Life Sciences; McGovern Institute for Brain Research; Peking Universi | B |
| 1959 | The mechanistic basis of data dependence and abrupt learning in an in-context classification task. | 2024 | ICLR | Informatics Labs, NTT Research Inc; Center for Brain Science, Harvard University | B |
| 1960 | The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models. | 2024 | ACL | University College London | B |
| 1961 | The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and Analysis. | 2024 | EMNLP | King's College London | B |
| 1962 | The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI. | 2024 | EMNLP | Technical University of Munich; University of Michigan | B |
| 1963 | The Illusion of Competence: Evaluating the Effect of Explanations on Users' Mental Models of Visual Question Answering Systems. | 2024 | EMNLP | Bielefeld University | B |
| 1964 | The Geometry of Concepts: Sparse Autoencoder Feature Structure | 2024 | MIT | C | |
| 1965 | The Expressive Leaky Memory Neuron: an Efficient and Expressive Phenomenological Neuron Model Can Solve Long-Horizon Tasks. | 2024 | ICLR | University of Tübingen, Germany; Max Planck Institute for Intelligent Systems, T | B |
| 1966 | The Effect of Weight Precision on the Neuron Count in Deep ReLU Networks. | 2024 | ICML | Rutgers, The State University of New Jersey | B |
| 1967 | The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value Factorization. | 2024 | NeurIPS | National University of Defense Technology; Istituto Nazionale di Fisica Nucleare | B |
| 1968 | The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models. | 2024 | ICLR | ◆ Chinese University of Hong Kong; Peking University; ▶ National University of S | B |
| 1969 | The Bayesian sampling in a canonical recurrent circuit with a diversity of inhibitory interneurons. | 2024 | NeurIPS | Southwestern Medical Center | B |
| 1970 | TextGenSHAP: Scalable Post-Hoc Explanations in Text Generation with Long Documents. | 2024 | ACL | University of Southern California; Google Cloud AI Research, Sunnyvale, CA | B |
| 1971 | Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game. | 2024 | ICLR | University of California, Berkeley; Carnegie Mellon University | B |
| 1972 | Temporally Consistent Factuality Probing for Large Language Models. | 2024 | EMNLP | Indian Institute of Technology Delhi; Wipro | B |
| 1973 | Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs. | 2024 | ICLR | Georgia Institute of Technology; University of California, Berkeley; ⋄Microsoft | B |
| 1974 | Technical Report: Enhancing LLM Reasoning with Reward-guided Tree Search | 2024 | Jinhao Jiang, Zhipeng Chen, Yingqian Min, Jie Chen, Xiaoxue | C | |
| 1975 | Technical Report on Slow Thinking with LLMs: II - Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems | 2024 | Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng | C | |
| 1976 | Teaching Small Language Models Reasoning through Counterfactual Distillation. | 2024 | EMNLP | Zhejiang University; Ant Group (China) | B |
| 1977 | TaPERA: Enhancing Faithfulness and Interpretability in Long-Form Table QA by Content Planning and Execution-based Reasoning. | 2024 | ACL | Yale University; Zhejiang University; Allen Institute for Artificial Intelligenc | B |
| 1978 | TVE: Learning Meta-attribution for Transferable Vision Explainer. | 2024 | ICML | Rice University; Wake Forest University; New Jersey Institute of Technology; Tex | B |
| 1979 | TELLER: A Trustworthy Framework for Explainable, Generalizable and Controllable Fake News Detection. | 2024 | ACL | City University of Hong Kong; Nanyang Technological University; University of El | B |
| 1980 | SyntaxShap: Syntax-aware Explainability Method for Text Generation. | 2024 | ACL | ETH Zurich | B |
| 1981 | Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation. | 2024 | EMNLP | University of California, Los Angeles | B |
| 1982 | Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation. | 2024 | ICML | Apple, Cupertino, CA, USA; Meta, Menlo xPark, CA, USA (*Work done ) | B |
| 1983 | SummaCoz: A Dataset for Improving the Interpretability of Factual Consistency Detection for Summarization. | 2024 | EMNLP | Iowa State University | B |
| 1984 | Successor Heads: Recurring, Interpretable Attention Heads In The Wild. | 2024 | ICLR | University of Cambridge | B |
| 1985 | StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements. | 2024 | EMNLP | University of Washington; Allen Institute for Artificial Intelligence | B |
| 1986 | Style-Specific Neurons for Steering LLMs in Text Style Transfer. | 2024 | EMNLP | Technical University of Munich; Munich Center for Machine Learning; EU Business | B |
| 1987 | Structured Matrix Basis for Multivariate Time Series Forecasting with Interpretable Dynamics. | 2024 | NeurIPS | Harbin Institute of Technology | B |
| 1988 | Structure Your Data: Towards Semantic Graph Counterfactuals. | 2024 | ICML | National Technical University of Athens | B |
| 1989 | Stochastic Concept Bottleneck Models. | 2024 | NeurIPS | Department of Computer Science; ETH Zurich | B |
| 1990 | Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution. | 2024 | NeurIPS | Stanford University; University of Washington | B |
| 1991 | Steering Llama 2 via Contrastive Activation Addition. | 2024 | ACL | Center for Human Genetics | BC |
| 1992 | Standardized Interpretable Fairness Measures for Continuous Risk Scores. | 2024 | ICML | SCHUFA Holding AG, Germany | B |
| 1993 | Speaking Your Language: Spatial Relationships in Interpretable Emergent Communication. | 2024 | NeurIPS | University of Southampton | B |
| 1994 | SparseFit: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations. | 2024 | ACL | ETH Zurich; University of Edinburgh; University College London | B |
| 1995 | Sparse Autoencoders Find Highly Interpretable Features in Language Models. | 2024 | ICLR | EleutherAI; MATS; Bristol AI Safety Centre; Apollo Research | B |
| 1996 | Socratic Human Feedback (SoHF): Expert Steering Strategies for LLM Code Generation. | 2024 | EMNLP | Amazon Web Services; Amazon | B |
| 1997 | Social Bias Probing: Fairness Benchmarking for Language Models. | 2024 | EMNLP | ETH Zurich; University of Copenhagen | B |
| 1998 | Single-Model Attribution of Generative Models Through Final-Layer Inversion. | 2024 | ICML | Ruhr University Bochum; Universität Hamburg | B |
| 1999 | Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair. | 2024 | EMNLP | Nara Institute of Science and Technology | B |
| 2000 | Sign Gradient Descent-based Neuronal Dynamics: ANN-to-SNN Conversion Beyond ReLU Network. | 2024 | ICML | Department of Computer Science & Engineering, Seoul Na-; tional University, Seou | B |
| 2001 | ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors. | 2024 | EMNLP | Tsinghua University; Peking University; Beihang University | B |
| 2002 | Shaping the distribution of neural responses with interneurons in a recurrent circuit model. | 2024 | NeurIPS | Flatiron Institute; New York University | B |
| 2003 | Semantics or spelling? Probing contextual word embeddings with orthographic noise. | 2024 | ACL | Cornell University | B |
| 2004 | Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP. | 2024 | EMNLP | Seoul National University; LG AI Research (South Korea) | B |
| 2005 | SelfIE: Self-Interpretation of Large Language Model Embeddings. | 2024 | ICML | Columbia University | B |
| 2006 | Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models | 2024 | Alibaba Qwen Team | C | |
| 2007 | Self-Supervised Interpretable End-to-End Learning via Latent Functional Modularity. | 2024 | ICML | Korea Advanced Institute of Science and Technology | B |
| 2008 | Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations. | 2024 | EMNLP | LFI - Learning, Fuzzy and Intelligent systems (France); Ekimetrics (36 Rue La Fa | B |
| 2009 | Selective Explanations. | 2024 | NeurIPS | Harvard University; IBM Research | B |
| 2010 | Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability. | 2024 | EMNLP | Hong Kong University of Science and Technology; Tencent AI Lab | B |
| 2011 | SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation. | 2024 | NeurIPS | University of Cambridge; University of Hong Kong; Google DeepMind (United Kingdo | B |
| 2012 | Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation | 2024 | Jointly conducted by researchers from the MAPLE Lab at Westlake University, South China University of Technology, Peking University, and the Westlake Institute for Advanced Study | C | |
| 2013 | Scaling Tractable Probabilistic Circuits: A Systems Perspective. | 2024 | ICML | National University of Singapore; University of California, Los Angeles | B |
| 2014 | Scaling Synthetic Data Creation with 1,000,000,000 Personas | 2024 | — | C | |
| 2015 | Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension | 2024 | University of Maryland + Microsoft | C | |
| 2016 | Scaling Continuous Latent Variable Models as Probabilistic Integral Circuits. | 2024 | NeurIPS | Eindhoven University of Technology; University of Edinburgh | B |
| 2017 | SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales. | 2024 | EMNLP | Purdue University; University of Illinois Urbana-Champaign; University of Southe | B |
| 2018 | Saliency-driven Experience Replay for Continual Learning. | 2024 | NeurIPS | University of Catania; University of Modena and Reggio Emilia | B |
| 2019 | Saliency strikes back: How filtering out high frequencies improves white-box explanations. | 2024 | ICML | Harvard University; Brown University | B |
| 2020 | SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation. | 2024 | ICLR | Michigan State University, ^ ‡ University of Pennsylvania, ^§IBM Research | B |
| 2021 | Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations. | 2024 | EMNLP | Singapore University of Technology and Design; Indian Institute of Technology Kh | B |
| 2022 | STORYSUMM: Evaluating Faithfulness in Story Summarization. | 2024 | EMNLP | Columbia International University; University of Missouri; Columbia University | B |
| 2023 | START: A Generalized State Space Model with Saliency-Driven Token-Aware Transformation. | 2024 | NeurIPS | Southeast University | B |
| 2024 | SPIN: Sparsifying and Integrating Internal Neurons in Large Language Models for Text Classification. | 2024 | ACL | University of Toronto; Technical University of Munich | B |
| 2025 | SOInter: A Novel Deep Energy-Based Interpretation Method for Explaining Structured Output Models. | 2024 | ICLR | Sharif University of Technology | B |
| 2026 | SIN: Selective and Interpretable Normalization for Long-Term Time Series Forecasting. | 2024 | ICML | Nanjing University | B |
| 2027 | SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning. | 2024 | NeurIPS | University of Maryland; Clemson University; Rutgers, The State University of New | B |
| 2028 | SEER: Facilitating Structured Reasoning and Explanation via Reinforcement Learning. | 2024 | ACL | Shanghai Artificial Intelligence Laboratory; Chinese Academy of Sciences; Shangh | B |
| 2029 | Reverse Thinking Makes LLMs Stronger Reasoners | 2024 | C | ||
| 2030 | Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods. | 2024 | ACL | University of Copenhagen; Language Science (South Korea) | B |
| 2031 | Revealing Personality Traits: A New Benchmark Dataset for Explainable Personality Recognition on Dialogues. | 2024 | EMNLP | Renmin University of China; Independent Age | B |
| 2032 | Retrieval-Guided Reinforcement Learning for Boolean Circuit Minimization. | 2024 | ICLR | University of Calgary | B |
| 2033 | Rethinking the symmetry-preserving circuits for constrained variational quantum algorithms. | 2024 | ICLR | B | |
| 2034 | Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits. | 2024 | ICML | Princeton University; East China Normal University; Stanford University; Columbi | B |
| 2035 | Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition. | 2024 | ICLR | Seoul National University | B |
| 2036 | Representing Molecules as Random Walks Over Interpretable Grammars. | 2024 | ICML | MIT-IBM Watson AI Lab, IBM Research; Massachusetts Institute of Technology | B |
| 2037 | Representation Surgery: Theory and Practice of Affine Steering. | 2024 | ICML | Bar-Ilan University; Google | B |
| 2038 | Relational Concept Bottleneck Models. | 2024 | NeurIPS | Scuola Normale Superiore | B |
| 2039 | RegExplainer: Generating Explanations for Graph Neural Networks in Regression Tasks. | 2024 | NeurIPS | New Jersey Institute of Technology; Florida International University; Arizona St | B |
| 2040 | Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning. | 2024 | ICLR | Monash University; Griffith University | B |
| 2041 | Rationales for Answers to Simple Math Word Problems Confuse Large Language Models. | 2024 | ACL | Sichuan University; University of California, Berkeley; Anthropic; Robotics and | B |
| 2042 | Rationale-Aware Answer Verification by Pairwise Self-Evaluation. | 2024 | EMNLP | Asahi Shimbun Company (Japan) | B |
| 2043 | RORA: Robust Free-Text Rationale Evaluation. | 2024 | ACL | Johns Hopkins University | B |
| 2044 | RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation. | 2024 | ICML | Northwestern University | B |
| 2045 | RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation. | 2024 | EMNLP | Seoul National University | B |
| 2046 | RDRec: Rationale Distillation for LLM-based Recommendation. | 2024 | ACL | Graduate School of Engineering; Interdisciplinary Graduate School; University of | B |
| 2047 | RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations. | 2024 | ACL | Stanford University | B |
| 2048 | RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering. | 2024 | ICLR | B | |
| 2049 | RA2FD: Distilling Faithfulness into Efficient Dialogue Systems. | 2024 | EMNLP | Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory | B |
| 2050 | Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing. | 2024 | NeurIPS | Department of Computer Science; University of California San Diego | B |
| 2051 | QORA: Zero-Shot Transfer via Interpretable Object-Relational Model Learning. | 2024 | ICML | Texas A&M University | B |
| 2052 | Q-Probe: A Lightweight Approach to Reward Maximization for Language Models. | 2024 | ICML | New York University; Courant Institute of Mathematical Sciences | B |
| 2053 | Putting Gale & Shapley to Work: Guaranteeing Stability Through Learning. | 2024 | NeurIPS | Penn State University, USA; University of Leeds, UK | B |
| 2054 | Provably Better Explanations with Optimized Aggregation of Feature Attributions. | 2024 | ICML | University of Oxford | B |
| 2055 | Prospector Heads: Generalized Feature Attribution for Large Models & Data. | 2024 | ICML | Stanford University; Art Institute of Portland; University of Waterloo | B |
| 2056 | Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks. | 2024 | ACL | Rutgers, The State University of New Jersey; AWS AI Labs | B |
| 2057 | Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal Features. | 2024 | ICLR | University of California Berkeley, CA, USA | B |
| 2058 | Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions. | 2024 | ICML | JPMorganChase AI Research | B |
| 2059 | Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models | 2024 | Laura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kama | C | |
| 2060 | Probing the Uniquely Identifiable Linguistic Patterns of Conversational AI Agents. | 2024 | ACL | University of Manchester; Manchester University | B |
| 2061 | Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games. | 2024 | ACL | Apple (United States) | B |
| 2062 | Probing the Emergence of Cross-lingual Alignment during LLM Training. | 2024 | ACL | University of Edinburgh | B |
| 2063 | Probing the Decision Boundaries of In-context Learning in Large Language Models. | 2024 | NeurIPS | Department of Computer Science; University of California, Los Angeles | B |
| 2064 | Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction. | 2024 | EMNLP | Brandeis University | B |
| 2065 | Probing Social Bias in Labor Market Text Generation by ChatGPT: A Masked Language Model Approach. | 2024 | NeurIPS | University of Alberta; Lancaster University; Concordia University; Tsinghua Univ | B |
| 2066 | Probing Language Models for Pre-training Data Detection. | 2024 | ACL | Institute of Artificial Intelligence, School of Computer Science and Technology; | B |
| 2067 | Probabilistic Generating Circuits - Demystified. | 2024 | ICML | Saarland University | B |
| 2068 | Probabilistic Constrained Reinforcement Learning with Formal Interpretability. | 2024 | ICML | Imperial College London | B |
| 2069 | Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models. | 2024 | ICML | Rutgers, The State University of New Jersey | B |
| 2070 | Presentations are not always linear! GNN meets LLM for Text Document-to-Presentation Transformation with Attribution. | 2024 | EMNLP | Microsoft; Adobe Research | B |
| 2071 | Predictive, scalable and interpretable knowledge tracing on structured domains. | 2024 | ICLR | University of Tübingen, 2; Tübingen AI Center, 4 | B |
| 2072 | Pre-trained Language Models Return Distinguishable Probability Distributions to Unfaithfully Hallucinated Texts. | 2024 | EMNLP | Korea University | B |
| 2073 | Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience. | 2024 | ICML | Institute for Neuroscience; Goethe University Frankfurt; New York University; he | BC |
| 2074 | Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs | 2024 | Liyi Chen, Panrong Tong, Zhongming Jin, Ying Sun, Jieping Ye | C | |
| 2075 | Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models. | 2024 | EMNLP | KU Leuven; Sailplane AI | B |
| 2076 | Piecewise Linear Parametrization of Policies: Towards Interpretable Deep Reinforcement Learning. | 2024 | ICLR | McGill University, Montreal, Canada | B |
| 2077 | Persuasiveness of Generated Free-Text Rationales in Subjective Decisions: A Case Study on Pairwise Argument Ranking. | 2024 | EMNLP | University of Pittsburgh; Microsoft | B |
| 2078 | Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization. | 2024 | NeurIPS | Pennsylvania State University | B |
| 2079 | Peering into the Mind of Language Models: An Approach for Attribution in Contextual Question Answering. | 2024 | ACL | Adobe Systems (United States) | B |
| 2080 | Paying More Attention to Source Context: Mitigating Unfaithful Translations from Large Language Model. | 2024 | ACL | Harbin Institute of Technology; Peng Cheng Laboratory | B |
| 2081 | Path Choice Matters for Clear Attributions in Path Methods. | 2024 | ICLR | Tsinghua University | B |
| 2082 | Partial observation can induce mechanistic mismatches in data-constrained models of neural dynamics. | 2024 | NeurIPS | Harvard University | B |
| 2083 | Parameter Competition Balancing for Model Merging | 2024 | Guodong Du, Junlin Lee, Jing Li, Runhua Jiang, Yifei Guo, Sh | C | |
| 2084 | PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning. | 2024 | ACL | Shenzhen University; Nanyang Technological University | B |
| 2085 | PRIME: Prioritizing Interpretability in Failure Mode Extraction. | 2024 | ICLR | Department of Computer Science, University of Maryland | B |
| 2086 | PREALIGN | 2024 | NLP Group, Nanjing University | C | |
| 2087 | PEDANTS: Cheap but Effective and Interpretable Answer Equivalence. | 2024 | EMNLP | University of Maryland, College Park | B |
| 2088 | PE: A Poincare Explanation Method for Fast Text Hierarchy Generation. | 2024 | EMNLP | East China Normal University; NPPA Key Laboratory of Publishing Integration Deve | B |
| 2089 | P-MMEval | 2024 | Qwen Team | C | |
| 2090 | Optimization Algorithm Design via Electric Circuits. | 2024 | NeurIPS | Stanford University; Rice University | B |
| 2091 | Optimal ablation for interpretability. | 2024 | NeurIPS | Harvard University | B |
| 2092 | Ontologically Faithful Generation of Non-Player Character Dialogues. | 2024 | EMNLP | Johns Hopkins University; Microsoft | B |
| 2093 | Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment | 2024 | Keming Lu, Bowen Yu, Fei Huang, Yang Fan, Runji Lin, Chang Z | C | |
| 2094 | On the Tractability of SHAP Explanations under Markovian Distributions. | 2024 | ICML | Nantes Université | B |
| 2095 | On the Similarity of Circuits across Languages: a Case Study on the Subject-verb Agreement Task. | 2024 | EMNLP | Universitat Politècnica de Catalunya; Meta; Association for Computational Lingui | B |
| 2096 | On the Feasibility of Single-Pass Full-Capacity Learning in Linear Threshold Neurons with Binary Input Vectors. | 2024 | ICML | Syracuse University | B |
| 2097 | On the Expressive Power of Tree-Structured Probabilistic Circuits. | 2024 | NeurIPS | Department of Computer Science; University of Illinois Urbana-Champaign | B |
| 2098 | On Mechanistic Knowledge Localization in Text-to-Image Generative Models. | 2024 | ICML | University of Maryland; Adobe Research | B |
| 2099 | On Measuring Faithfulness or Self-consistency of Natural Language Explanations. | 2024 | ACL | Heidelberg University | B |
| 2100 | On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey | 2024 | Zhejiang University + Harbin Institute of Technology — Survey on Synthetic Data Generation with LLMs | C | |
| 2101 | On Gradient-like Explanation under a Black-box Setting: When Black-box Explanations Become as Good as White-box. | 2024 | ICML | Freie Universität Berlin | B |
| 2102 | On Evaluating Explanation Utility for Human-AI Decision Making in NLP. | 2024 | EMNLP | Kahlert School of Computing, University of Utah; University of Utah | B |
| 2103 | O1 Replication Journey: A Strategic Progress Report | 2024 | Shanghai Jiao Tong University | C | |
| 2104 | Not All Language Model Features Are Linear | 2024 | MIT | C | |
| 2105 | Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics Discovery. | 2024 | NeurIPS | Lehigh University; Global Engineering and Materials (United States); Johns Hopki | B |
| 2106 | Non-asymptotic Approximation Error Bounds of Parameterized Quantum Circuits. | 2024 | NeurIPS | Wuhan University; National University of Singapore; Hubei Key Laboratory of Comp | B |
| 2107 | Neurons in Large Language Models: Dead, N-gram, Positional. | 2024 | ACL | Meta; Universitat Politècnica de Catalunya | B |
| 2108 | Neuronal Competition Groups with Supervised STDP for Spike-Based Classification. | 2024 | NeurIPS | École Centrale de Lille; CRIStAL, University of Lille | B |
| 2109 | Neuron-Level Knowledge Attribution in Large Language Models. | 2024 | EMNLP | National Centre for Atmospheric Science | B |
| 2110 | Neuron-Enhanced AutoEncoder Matrix Completion and Collaborative Filtering: Theory and Practice. | 2024 | ICLR | The Chinese University of Hong Kong, Shenzhen, China; University of Texas at Arl | B |
| 2111 | Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine Translation. | 2024 | EMNLP | Language Technology Lab; University of Amsterdam | B |
| 2112 | Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization. | 2024 | ICLR | City University of Hong Kong; The University of Tokyo | B |
| 2113 | NeurRev: Train Better Sparse Neural Network Practically via Neuron Revitalization. | 2024 | ICLR | B | |
| 2114 | Nearest Neighbor Speculative Decoding for LLM Generation and Attribution. | 2024 | NeurIPS | Cohere (Canada); Meta; University of Chicago; Carnegie Mellon University; Univer | B |
| 2115 | Navigating the OverKill in Large Language Models | 2024 | — | C | |
| 2116 | Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics. | 2024 | NeurIPS | German Cancer Research Center; ETH Zurich; Heidelberg University; University of | B |
| 2117 | Natural Counterfactuals With Necessary Backtracking. | 2024 | NeurIPS | Chinese University of Hong Kong; University of California San Diego; Rutgers, Th | B |
| 2118 | NDOT: Neuronal Dynamics-based Online Training for Spiking Neural Networks. | 2024 | ICML | Department of Machine Learning, MBZUAI, Abu Dhabi, UAE; Technology Innovation In | B |
| 2119 | NALA: an Effective and Interpretable Entity Alignment Method. | 2024 | EMNLP | Northeastern University | B |
| 2120 | NAISR: A 3D Neural Additive Model for Interpretable Shape Representation. | 2024 | ICLR | University of North Carolina at Chapel Hill, ^2Wake Forest School of Medicine, ^ | B |
| 2121 | MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering. | 2024 | NeurIPS | Tsinghua University; Pharmolix Inc | B |
| 2122 | Multiply-Robust Causal Change Attribution. | 2024 | ICML | Massachusetts Institute of Technology; Amazon | B |
| 2123 | Multi-Aspect Controllable Text Generation with Disentangled Counterfactual Augmentation. | 2024 | ACL | Nanjing University | B |
| 2124 | MorphGrower: A Synchronized Layer-by-layer Growing Approach for Plausible Neuronal Morphology Generation. | 2024 | ICML | School of Artificial Intelligence & De-; partment of Computer Science and Engine | B |
| 2125 | Model Reconstruction Using Counterfactual Explanations: A Perspective From Polytope Theory. | 2024 | NeurIPS | University of Maryland | B |
| 2126 | Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation. | 2024 | EMNLP | University of Groningen; Northeastern University | B |
| 2127 | MoE-CT: A Novel Approach For Large Language Models Training With Resistance To Catastrophic Forgetting | 2024 | Alibaba | C | |
| 2128 | Mixture of a Million Experts | 2024 | DeepMind | C | |
| 2129 | Mitigating Privacy Seesaw in Large Language Models: Augmented Privacy Neuron Editing via Activation Patching. | 2024 | ACL | Tianjin University | B |
| 2130 | Mitigating Language Bias of LMMs in Social Intelligence Understanding with Virtual Counterfactual Calibration. | 2024 | EMNLP | Monash University; Beihang University | B |
| 2131 | Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination. | 2024 | ACL | Seoul National University; LG AI Research; University of Michigan | B |
| 2132 | MindMerger: Efficient Boosting LLM Reasoning in non-English Languages. Authors: Zixian Huang, Wenhao Zhu, Gong Cheng, Lei Li, Fei Yuan, respectively from the State Key Laboratory for Novel Software Technology at Nanjing University, Carnegie Mellon University, and Shanghai AI Laboratory. | 2024 | — | C | |
| 2133 | Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models. | 2024 | NeurIPS | Korea Advanced Institute of Science and Technology | B |
| 2134 | MetaGPT: Merging Large Language Models Using Model Exclusive Task Arithmetic | 2024 | Meta | C | |
| 2135 | MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing. | 2024 | ACL | Indraprastha Institute of Information Technology Delhi; Indian Institute of Tech | B |
| 2136 | Mediator Interpretation and Faster Learning Algorithms for Linear Correlated Equilibria in General Sequential Games. | 2024 | ICLR | Carnegie Mellon University, Pittsburgh, USA | B |
| 2137 | Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks. | 2024 | ICLR | University of Cambridge, UK; University College London, UK; EECS Department, Uni | B |
| 2138 | Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations. | 2024 | EMNLP | University of Toronto; McGill University; Mila – Québec AI Institute; University | B |
| 2139 | Mechanistic Neural Networks for Scientific Machine Learning. | 2024 | ICML | \addr Informatics Institute, University of Amsterdam; \addr Institute of Science | B |
| 2140 | Mechanistic Design and Scaling of Hybrid Architectures. | 2024 | ICML | Qure.ai | B |
| 2141 | Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals. | 2024 | EMNLP | Allen Institute for Artificial Intelligence; University of Washington; Universit | B |
| 2142 | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models. | 2024 | NeurIPS | University of Mannheim; Harvard University; Northeastern University | B |
| 2143 | Measuring Per-Unit Interpretability at Scale Without Humans. | 2024 | NeurIPS | University of Tübingen | B |
| 2144 | Mastering the game of Go without human knowledge | 2024 | C | ||
| 2145 | Many-Shot In-Context Learning | 2024 | Google DeepMind | C | |
| 2146 | Manifold Integrated Gradients: Riemannian Geometry for Feature Attribution. | 2024 | ICML | ARC Training Centre for Information Resilience (CIRES), Brisbane, Australia; The | B |
| 2147 | MambaLRP: Explaining Selective State Space Sequence Models. | 2024 | NeurIPS | Technische Universität Berlin; Link (Germany); Berlin Institute for the Foundati | B |
| 2148 | MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education. | 2024 | EMNLP | Rice University | B |
| 2149 | Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning. | 2024 | EMNLP | École Polytechnique Fédérale de Lausanne | B |
| 2150 | MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model. | 2024 | EMNLP | The Hong Kong University of Science and Technology (Guangzhou); Hong Kong Univer | B |
| 2151 | MIDGArD: Modular Interpretable Diffusion over Graphs for Articulated Designs. | 2024 | NeurIPS | The Advanced Reality Lab | B |
| 2152 | MG-Net: Learn to Customize QAOA with Circuit Depth Awareness. | 2024 | NeurIPS | The University of Sydney; Wuhan University; Nanyang Technological University | B |
| 2153 | MEQA: A Benchmark for Multi-hop Event-centric Question Answering with Explanations. | 2024 | NeurIPS | Xi’an Jiaotong-Liverpool University; University of Liverpool | B |
| 2154 | MARE: Multi-Aspect Rationale Extractor on Unsupervised Rationale Extraction. | 2024 | EMNLP | Hunan Provincial Key Lab on Bioinformatics, School of Computer Science and Engin | B |
| 2155 | Local vs. Global Interpretability: A Computational Complexity Perspective. | 2024 | ICML | Hebrew University of Jerusalem | B |
| 2156 | Local Feature Selection without Label or Feature Leakage for Interpretable Machine Learning Predictions. | 2024 | ICML | Leibniz University Hannover | B |
| 2157 | Linear Explanations for Individual Neurons. | 2024 | ICML | CSE, UC San Diego, CA, USA; HDSI, UC San Diego | B |
| 2158 | Lie Neurons: Adjoint-Equivariant Neural Networks for Semisimple Lie Algebras. | 2024 | ICML | University of Michigan, Ann Arbor, MI | B |
| 2159 | LiDAR: Sensing Linear Probing Performance in Joint Embedding SSL Architectures. | 2024 | ICLR | Apple | B |
| 2160 | Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding. | 2024 | NeurIPS | University of Illinois Urbana-Champaign; Carnegie Mellon University | B |
| 2161 | Leveraging Machine-Generated Rationales to Facilitate Social Meaning Detection in Conversations. | 2024 | ACL | Carnegie Mellon University | B |
| 2162 | Less is More: Fewer Interpretable Region via Submodular Subset Selection. | 2024 | ICLR | Institute of Information Engineering, Chinese Academy of Sciences, Beijing 10009 | B |
| 2163 | Legal Judgment Reimagined: PredEx and the Rise of Intelligent AI Interpretation in Indian Courts. | 2024 | ACL | Indian Institute of Technology Kanpur; Indian Institute of Science Education and | B |
| 2164 | Learning to Intervene on Concept Bottlenecks. | 2024 | ICML | Technische Universität Darmstadt; German Research Center for AI (DFKI) | B |
| 2165 | Learning the greatest common divisor: explaining transformer predictions. | 2024 | ICLR | Meta AI | B |
| 2166 | Learning interpretable control inputs and dynamics underlying animal locomotion. | 2024 | ICLR | B | |
| 2167 | Learning from Natural Language Explanations for Generalizable Entity Matching. | 2024 | EMNLP | ♢Northeastern University; Amazon | B |
| 2168 | Learning a Single Neuron Robustly to Distributional Shifts and Adversarial Label Noise. | 2024 | NeurIPS | University of Wisconsin–Madison | B |
| 2169 | Learning Interpretable Legal Case Retrieval via Knowledge-Guided Case Reformulation. | 2024 | EMNLP | Renmin University of China | B |
| 2170 | Learning Disentangled Semantic Spaces of Explanations via Invertible Neural Networks. | 2024 | ACL | University of Manchester; Idiap Research Institute | B |
| 2171 | Learnable Privacy Neurons Localization in Language Models. | 2024 | ACL | Zhejiang University | B |
| 2172 | Latent Logic Tree Extraction for Event Sequence Explanation from LLMs. | 2024 | ICML | Nanyang Technological University; Chinese University of Hong Kong, Shenzhen | B |
| 2173 | Latent Concept-based Explanation of NLP Models. | 2024 | EMNLP | Dalhousie University; Hamad bin Khalifa University; Independent Age | B |
| 2174 | Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment. | 2024 | ACL | Alibaba Inc | B |
| 2175 | Large Language Models Can Self-Improve in Long-context Reasoning | 2024 | Siheng Li, Cheng Yang, Zesen Cheng, Lemao Liu, Mo Yu, Yuyu Y | C | |
| 2176 | Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models. | 2024 | ACL | Renmin University of China; Microsoft Research Asia (China) | B |
| 2177 | Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication. | 2024 | NeurIPS | University of Pittsburgh; Honda Research Institute, USA; Carnegie Mellon Univers | B |
| 2178 | LaMAGIC: Language-Model-based Topology Generation for Analog Integrated Circuits. | 2024 | ICML | IBM T. J. Watson Research Center; Duke University; MIT-IBM Watson AI Lab; New Je | B |
| 2179 | LOFIT: Localized Fine-tuning on LLM Representations | 2024 | — | C | |
| 2180 | LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages. Authors: Yinquan Lu, Wenhao Zhu, Lei Li, Yu Qiao, and Fei Yuan. Summary: LLaMAX aims to extend the linguistic boundaries of large language models (LLMs) by enhancing translation capabilities to cover more than 100 languages. Through extensive multilingual continual pretraining on the LLaMA series, the work achieves translation support for over 100 languages. The team developed LLaMAX via a comprehensive analysis of training strategies such as vocabulary extension and data augmentation. Without sacrificing generalization ability, LLaMAX attains significantly higher translation performance than existing open-source LLMs, and performs on par with the dedicated translation model M2M-100-12B on the Flores-101 benchmark. | 2024 | — | C | |
| 2181 | LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study. | 2024 | EMNLP | University of Marburg; University of Mannheim | B |
| 2182 | LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression. | 2024 | ACL | Tsinghua University; Microsoft (Finland) | B |
| 2183 | LLMFactor: Extracting Profitable Factors through Prompts for Explainable Stock Movement Prediction. | 2024 | ACL | Tokyo University of Agriculture | B |
| 2184 | LLM Explainability via Attributive Masking Learning. | 2024 | EMNLP | Tel Aviv University | B |
| 2185 | LLM Circuit Analyses Are Consistent Across Training and Scale. | 2024 | NeurIPS | EleutherAI; University of Amsterdam; Brown University | B |
| 2186 | LILO: Learning Interpretable Libraries by Compressing and Documenting Code. | 2024 | ICLR | MIT CSAIL; MIT Brain and Cognitive Sciences; Harvey Mudd College | B |
| 2187 | LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation. | 2024 | ACL | Tianjin University; Tsinghua University; Baidu (China) | B |
| 2188 | Knowledge Mechanisms in Large Language Models: A Survey and Perspective. Authors: Mengru Wang, Yunzhi Yao, Ziwen Xu, Shuofei Qiao, Shumin Deng, Peng Wang, Xiang Chen, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang — respectively from Zhejiang University, the NUS-NLP Joint Lab at the National University of Singapore, the University of California, Los Angeles, and Alibaba Group. Summary: This survey examines knowledge mechanisms in large language models (LLMs), which are critical to advancing trustworthy AI. Starting from a novel taxonomy, it reviews the analysis of knowledge mechanisms along two axes: knowledge utilization and knowledge evolution. Knowledge utilization covers memorization, comprehension, application and creation mechanisms, while knowledge evolution focuses on the dynamic progression of knowledge within individual and group LLMs. The paper also discusses what knowledge LLMs have learned, the reasons for the fragility of parametric knowledge, and the potentially challenging "dark knowledge" hypothesis. The authors hope this work helps understand knowledge in LLMs and offers insights for future research. | 2024 | — | C | |
| 2189 | Knowledge Circuits in Pretrained Transformers. | 2024 | NeurIPS | Zhejiang University; National University of Singapore; Zhejiang Key Laboratory o | BC |
| 2190 | KernelSHAP-IQ: Weighted Least Square Optimization for Shapley Interactions. | 2024 | ICML | Bielefeld University | B |
| 2191 | KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache | 2024 | Rice University,Texas A&M University,Stevens Institute of Te | C | |
| 2192 | KAN:Kolmogorov-Arnold Networks | 2024 | MIT | C | |
| 2193 | Iterative Search Attribution for Deep Neural Networks. | 2024 | ICML | The University of Sydney; University of Malaya; CSIRO Data61; University of Woll | B |
| 2194 | Iteration Head: A Mechanistic Study of Chain-of-Thought. | 2024 | NeurIPS | Meta (United States); Laboratoire de Mathématiques d'Orsay; Centre Inria de Sacl | B |
| 2195 | Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-Rationalization. | 2024 | NeurIPS | School of Computer Science and Technology; Huazhong University of Science and Te | B |
| 2196 | Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching. | 2024 | ICLR | SERI MATS SERI MATS; OpenAI, 2023 | B |
| 2197 | Is Epistemic Uncertainty Faithfully Represented by Evidential Deep Learning Methods? | 2024 | ICML | Ghent University; Institute of Communications and Navigation, German Aerospace C | B |
| 2198 | Investigating the Impact of Model Instability on Explanations and Uncertainty. | 2024 | ACL | Machine Science | B |
| 2199 | Intriguing Properties of Data Attribution on Diffusion Models. | 2024 | ICLR | Singapore Management University; Sea AI Lab, Singapore | B |
| 2200 | Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis. | 2024 | EMNLP | National Centre for Atmospheric Science | B |
| 2201 | Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models. | 2024 | ACL | Microsoft (Finland); Purdue University | B |
| 2202 | Interpretable Sparse System Identification: Beyond Recent Deep Learning Techniques on Time-Series Prediction. | 2024 | ICLR | B | |
| 2203 | Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts. | 2024 | EMNLP | University of Illinois Urbana-Champaign; University of Wisconsin–Madison | B |
| 2204 | Interpretable Meta-Learning of Physical Systems. | 2024 | ICLR | École Normale Supérieure - PSL; Département d'Informatique | B |
| 2205 | Interpretable Mesomorphic Networks for Tabular Data. | 2024 | NeurIPS | Department of Representation Learning; University of Freiburg; Department of Mac | B |
| 2206 | Interpretable Lightweight Transformer via Unrolling of Learned Graph Smoothness Priors. | 2024 | NeurIPS | York University; Seattle University | B |
| 2207 | Interpretable Image Classification with Adaptive Prototype-based Vision Transformers. | 2024 | NeurIPS | Dartmouth College; Duke University; University of Maine | B |
| 2208 | Interpretable Generalized Additive Models for Datasets with Missing Values. | 2024 | NeurIPS | Department of Computer Science; Duke University; University of British Columbia | B |
| 2209 | Interpretable Diffusion via Information Decomposition. | 2024 | ICLR | University of California Riverside, ^2University of Southern California | B |
| 2210 | Interpretable Deep Clustering for Tabular Data. | 2024 | ICML | Technion – Israel Institute of Technology | B |
| 2211 | Interpretable Concept-Based Memory Reasoning. | 2024 | NeurIPS | Università della Svizzera italiana; University of Cambridge; Scuola Normale Supe | B |
| 2212 | Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents. | 2024 | NeurIPS | Technische Universität Darmstadt; Hessian Center for Artificial Intelligence; Ge | B |
| 2213 | Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional Understanding. | 2024 | EMNLP | MoE Key Laboratory of Brain-inspired; University of Science and Technology of Ch | B |
| 2214 | Interpretability-based Tailored Knowledge Editing in Transformers. | 2024 | EMNLP | The London College; University College London | B |
| 2215 | Interpretability of Language Models via Task Spaces. | 2024 | ACL | Universitat Pompeu Fabra | B |
| 2216 | Interpretability Illusions in the Generalization of Simplified Models. | 2024 | ICML | Princeton University | B |
| 2217 | Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification. | 2024 | NeurIPS | Xi’an Jiaotong-Liverpool University; University of Liverpool; Duke Kunshan Unive | B |
| 2218 | InterpreTabNet: Distilling Predictive Signals from Tabular Data by Salient Feature Interpretation. | 2024 | ICML | University of Toronto; State Research Center of Virology and Biotechnology VECTO | B |
| 2219 | InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques. | 2024 | NeurIPS | Universidad de Buenos Aires | B |
| 2220 | IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning. | 2024 | EMNLP | University of Maryland, College Park | B |
| 2221 | Inherently Interpretable Time Series Classification via Multiple Instance Learning. | 2024 | ICLR | Amazon Prime Video, UK | B |
| 2222 | Inference to the Best Explanation in Large Language Models. | 2024 | ACL | Idiap Research Institute | B |
| 2223 | InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance | 2024 | Fudan University | C | |
| 2224 | Incorporating Information into Shapley Values: Reweighting via a Maximum Entropy Approach. | 2024 | ICML | University of Minnesota | B |
| 2225 | In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering. | 2024 | ICML | New York University | B |
| 2226 | Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders. | 2024 | NeurIPS | Google DeepMind (United Kingdom) | B |
| 2227 | Improving Quotation Attribution with Fictional Character Embeddings. | 2024 | EMNLP | Laboratoire Lorrain de Recherche en Informatique et ses Applications | B |
| 2228 | Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining. | 2024 | ICML | Harvard University | B |
| 2229 | Improving LLM Attributions with Randomized Path-Integration. | 2024 | EMNLP | Tel Aviv University | B |
| 2230 | Improving Interpretation Faithfulness for Vision Transformers. | 2024 | ICML | King Abdullah University of Science and Technology; SDAIA-KAUST AI Center; Lehig | B |
| 2231 | Improving Alignment and Robustness with Circuit Breakers. | 2024 | NeurIPS | Gray Swan AI; Carnegie Mellon University; Center for AI Safety | B |
| 2232 | Improve Mathematical Reasoning in Language Models by Automated Process Supervision | 2024 | Google DeepMind | C | |
| 2233 | Image Inpainting via Tractable Steering of Diffusion Models. | 2024 | ICLR | Department of Computer Science, University of California, Los Angeles; Departmen | B |
| 2234 | IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware Neurons. | 2024 | NeurIPS | Tianjin University | B |
| 2235 | IPO: Interpretable Prompt Optimization for Vision-Language Models. | 2024 | NeurIPS | University of Amsterdam; University of Science and Technology of China | B |
| 2236 | INViTE: INterpret and Control Vision-Language Models with Text Explanations. | 2024 | ICLR | Columbia University, New York, NY, USA | B |
| 2237 | ICLEF: In-Context Learning with Expert Feedback for Explainable Style Transfer. | 2024 | ACL | Columbia University | B |
| 2238 | Hypothesis Testing the Circuit Hypothesis in LLMs. | 2024 | NeurIPS | Columbia University; University of Michigan; FAR AI | B |
| 2239 | How do Large Language Models Handle Multilingualism? | 2024 | Renmin University of China | C | |
| 2240 | How connectivity structure shapes rich and lazy learning in neural circuits. | 2024 | ICLR | University of Washington, Seattle, WA, USA; Allen Institute for Brain Science, S | B |
| 2241 | How Interpretable Are Interpretable Graph Neural Networks? | 2024 | ICML | Tencent AI Lab; Chinese University of Hong Kong; Hong Kong Baptist University | B |
| 2242 | How Do Large Language Models Acquire Factual Knowledge During Pretraining? | 2024 | KAIST, UCL, KT | C | |
| 2243 | How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition | 2024 | Alibaba | C | |
| 2244 | Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model Predictions. | 2024 | ICML | National University of Singapore; Institute for Infocomm Research | B |
| 2245 | HelmFluid: Learning Helmholtz Dynamics for Interpretable Fluid Prediction. | 2024 | ICML | School of Software, BNRist, Tsinghua; University | B |
| 2246 | HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models. | 2024 | EMNLP | University of Maryland; Microsoft Research | B |
| 2247 | Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning. | 2024 | ICLR | National University of Singapore, 2; Loyola Marymount University; Element, Inc., | B |
| 2248 | Harder Tasks Need More Experts: Dynamic Routing in MoE Models | 2024 | Peking University | C | |
| 2249 | Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL. | 2024 | ACL | Korea Advanced Institute of Science and Technology; Microsoft Research (India) | B |
| 2250 | HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language Model. | 2024 | NeurIPS | Carnegie Mellon University | B |
| 2251 | Grounding Language Plans in Demonstrations Through Counterfactual Perturbations. | 2024 | ICLR | Cranberry-Lemon University; University of the Witwatersrand | B |
| 2252 | Grokking of Implicit Reasoning in Transformers: A Mechanistic Journey to the Edge of Generalization. | 2024 | NeurIPS | The Ohio State University; Carnegie Mellon University | B |
| 2253 | GraphTrail: Translating GNN Predictions into Human-Interpretable Logical Rules. | 2024 | NeurIPS | Department of Computer Science & Engineering; Department of Computer Science; In | B |
| 2254 | Graph Neural Networks and Arithmetic Circuits. | 2024 | NeurIPS | Institute of Theoretical Computer Science; Leibniz University Hannover; School o | B |
| 2255 | Graph Neural Network Explanations are Fragile. | 2024 | ICML | Nanchang University; Illinois Institute of Technology; Milwaukee School of Engin | B |
| 2256 | Gradient-based Visual Explanation for Transformer-based CLIP. | 2024 | ICML | City University of Hong Kong; Sensetime (China) | B |
| 2257 | Going Beyond Neural Network Feature Similarity: The Network Feature Complexity and Its Interpretation Using Category Theory. | 2024 | ICLR | Department of Computer Science and Engineering and MoE Key Lab of Artificial Int | B |
| 2258 | Generating and Evaluating Plausible Explanations for Knowledge Graph Completion. | 2024 | ACL | NEC Laboratories Europe; University College London | B |
| 2259 | Generating In-Distribution Proxy Graphs for Explaining Graph Neural Networks. | 2024 | ICML | Florida International University; New Jersey Institute of Technology; University | B |
| 2260 | Gender Identity in Pretrained Language Models: An Inclusive Approach to Data Creation and Probing. | 2024 | EMNLP | University of Stuttgart | B |
| 2261 | Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 | 2024 | Google DeepMind | C | |
| 2262 | GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation. | 2024 | EMNLP | Georgia Institute of Technology | B |
| 2263 | GOAt: Explaining Graph Neural Networks via Graph Output Attribution. | 2024 | ICLR | Department of Electrical and Computer Engineering, University of Alberta; Soluti | B |
| 2264 | GNNBoundary: Towards Explaining Graph Neural Networks through the Lens of Decision Boundaries. | 2024 | ICLR | Ohio State University, Columbus, USA | B |
| 2265 | F²RL: Factuality and Faithfulness Reinforcement Learning Framework for Claim-Guided Evidence-Supported Counterspeech Generation. | 2024 | EMNLP | National University of Defense Technology; PLA Academy of Military Science | B |
| 2266 | From Neurons to Neutrons: A Case Study in Interpretability. | 2024 | ICML | The NSF AI Institute for Artificial Intelligence and Fundamental Interactions; M | B |
| 2267 | From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP. | 2024 | EMNLP | Mila Quebec AI Institute; McGill University; Saarland University; Pontificia Uni | B |
| 2268 | Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge | 2024 | Jiahuan Li, Yiqing Cao, Shujian Huang, Jiajun Chen, from the Dept. of Computer Science, Nanjing University | C | |
| 2269 | Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting. | 2024 | EMNLP | University of Oxford; Vienna University of Technology; Rayscape; University Coll | B |
| 2270 | Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion Perception. | 2024 | NeurIPS | Nanyang Technological University; Agency for Science, Technology and Research; N | B |
| 2271 | Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation. | 2024 | ACL | City University of Hong Kong; Chinese University of Hong Kong; Shenzhen Universi | B |
| 2272 | Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers. | 2024 | ACL | University of Science and Technology of China; Fudan University; Tsinghua Univer | B |
| 2273 | Finding Transformer Circuits With Edge Pruning. | 2024 | NeurIPS | Princeton University | B |
| 2274 | Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion Models. | 2024 | NeurIPS | German Research Centre for Artificial Intelligence; Technische Universität Darms | B |
| 2275 | Finding NEM-U: Explaining unsupervised representation learning through neural network generated explanation masks. | 2024 | ICML | University of Copenhagen; UiT The Arctic University of Norway; Norwegian Computi | B |
| 2276 | Finding Blind Spots in Evaluator LLMs with Interpretable Checklists. | 2024 | EMNLP | Indian Institute of Technology Madras | B |
| 2277 | FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents. | 2024 | EMNLP | Yale-NUS College | B |
| 2278 | Figuratively Speaking: Authorship Attribution via Multi-Task Figurative Language Modeling. | 2024 | ACL | Department of Computer Science, *; Department of Cognitive Science; Rensselaer P | B |
| 2279 | Federated Self-Explaining GNNs with Anti-shortcut Augmentations. | 2024 | ICML | University of Science and Technology of China; Institute of Artificial Intellige | B |
| 2280 | Federated Behavioural Planes: Explaining the Evolution of Client Behaviour in Federated Learning. | 2024 | NeurIPS | Università della Svizzera italiana; Syneos Health | B |
| 2281 | Feature Attribution with Necessity and Sufficiency via Dual-stage Perturbation Test for Causal Explanation. | 2024 | ICML | Guangdong University of Technology; Guangzhou Laboratory; University of Cambridg | B |
| 2282 | Faithfulness Measurable Masked Language Models. | 2024 | ICML | Mila - Quebec AI Institute | B |
| 2283 | Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach. | 2024 | EMNLP | Poznan University of Technology, Faculty of Computing and Telecommunications, Po | B |
| 2284 | Faithful and Efficient Explanations for Neural Networks via Neural Tangent Kernel Surrogate Models. | 2024 | ICLR | Pacific Northwest National Laboratory; University of California; Courant Institu | B |
| 2285 | Faithful Vision-Language Interpretation via Concept Bottleneck Models. | 2024 | ICLR | B | |
| 2286 | Faithful Rule Extraction for Differentiable Rule Learning Models. | 2024 | ICLR | University of Oxford, UK | B |
| 2287 | Faithful Persona-based Conversational Dataset Generation with Large Language Models. | 2024 | ACL | Google (United States) | B |
| 2288 | Faithful Logical Reasoning via Symbolic Chain-of-Thought. | 2024 | ACL | National University of Singapore; University of California, Santa Barbara; Unive | B |
| 2289 | Faithful Explanations of Black-box NLP Models Using LLM-generated Counterfactuals. | 2024 | ICLR | Technion – Israel Institute of Technology; Columbia University; Google | B |
| 2290 | Faithful Chart Summarization with ChaTS-Pi. | 2024 | ACL | Google (United States) | B |
| 2291 | FRVA: Fact-Retrieval and Verification Augmented Entailment Tree Generation for Explainable Question Answering. | 2024 | ACL | Shanxi University | B |
| 2292 | FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model. | 2024 | ACL | Seoul National University | B |
| 2293 | FFAM: Feature Factorization Activation Map for Explanation of 3D Detectors. | 2024 | NeurIPS | Sun Yat-sen University | B |
| 2294 | FANTAstic SEquences and Where to Find Them: Faithful and Efficient API Call Generation through State-tracked Constrained Decoding and Reranking. | 2024 | EMNLP | Texas A&M University; Amazon | B |
| 2295 | Exploring the trade-off between deep-learning and explainable models for brain-machine interfaces. | 2024 | NeurIPS | Robotics Research (United States); ETH Zurich; BioSurfaces (United States); Eind | B |
| 2296 | Explanations that reveal all through the definition of encoding. | 2024 | NeurIPS | New York University | B |
| 2297 | Explanation-aware Soft Ensemble Empowers Large Language Model In-context Learning. | 2024 | ACL | Google (United States); Cornell University | B |
| 2298 | Explaining and Improving Contrastive Decoding by Extrapolating the Probabilities of a Huge and Hypothetical LM. | 2024 | EMNLP | University of Massachusetts Amherst; Amazon; University of Southern California | B |
| 2299 | Explaining Time Series via Contrastive and Locally Sparse Perturbations. | 2024 | ICLR | Nanjing University, 2; Pennsylvania State University, 4; Tsinghua University, 5; | B |
| 2300 | Explaining Probabilistic Models with Distributional Values. | 2024 | ICML | Amazon | B |
| 2301 | Explaining Mixtures of Sources in News Articles. | 2024 | EMNLP | University of Southern California; University of California, Los Angeles; Stanfo | B |
| 2302 | Explaining Kernel Clustering via Decision Trees. | 2024 | ICLR | Technical University of Munich; Amazon Research | B |
| 2303 | Explaining Graph Neural Networks with Large Language Models: A Counterfactual Perspective on Molecule Graphs. | 2024 | EMNLP | University of Virginia; Florida State University | B |
| 2304 | Explaining Graph Neural Networks via Structure-aware Interaction Index. | 2024 | ICML | Yale University; VinAI Research | B |
| 2305 | Explaining Datasets in Words: Statistical Models with Natural Language Parameters. | 2024 | NeurIPS | All authors affiliated with UC Berkeley | B |
| 2306 | Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster. | 2024 | ACL | University of Edinburgh; Snap Inc. | B |
| 2307 | Exogenous Matching: Learning Good Proposals for Tractable Counterfactual Estimation. | 2024 | NeurIPS | East China Normal University | B |
| 2308 | Exact Soft Analytical Side-Channel Attacks using Tractable Circuits. | 2024 | ICML | Graz University of Technology; Institute of Information and Communication Techno | B |
| 2309 | Evaluating Readability and Faithfulness of Concept-based Explanations. | 2024 | EMNLP | Renmin University of China; University of Science and Technology of China; Kuais | B |
| 2310 | Evaluate then Cooperate: Shapley-based View Cooperation Enhancement for Multi-view Clustering. | 2024 | NeurIPS | University of Defence; Academy of Military Science | B |
| 2311 | Episodic Memory Retrieval from LLMs: A Neuromorphic Mechanism to Generate Commonsense Counterfactuals for Relation Extraction. | 2024 | ACL | Wuhan University | B |
| 2312 | Enhancing Training Data Attribution for Large Language Models with Fitting Error Consideration. | 2024 | EMNLP | Chinese Academy of Sciences; Institute of Computing Technology, CAS; University | B |
| 2313 | Enhancing Semantic Consistency of Large Language Models through Model Editing: An Interpretability-Oriented Approach. | 2024 | ACL | Tianjin University; The Sense Innovation and Research Center | B |
| 2314 | Enhancing Robustness of Graph Neural Networks on Social Media with Explainable Inverse Reinforcement Learning. | 2024 | NeurIPS | Key Laboratory of Trustworthy Distributed Computing and Service (BUPT); Beijing | B |
| 2315 | Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer Decomposition. | 2024 | EMNLP | Adobe Systems (United States) | B |
| 2316 | Enhancing Explainable Rating Prediction through Annotated Macro Concepts. | 2024 | ACL | Hong Kong Polytechnic University | B |
| 2317 | Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic Interpretations. | 2024 | ICLR | The Hong Kong University of Science and Technology, ^2University of Washington, | B |
| 2318 | End-to-End Neuro-Symbolic Reinforcement Learning with Textual Explanations. | 2024 | ICML | Beijing Institute for General Artificial Intelligence | B |
| 2319 | Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective. | 2024 | EMNLP | King's College London; Carnegie Mellon University; Mohamed bin Zayed University | B |
| 2320 | Embedded Named Entity Recognition using Probing Classifiers. | 2024 | EMNLP | Karlsruhe Institute of Technology; Technische Universität Dresden | B |
| 2321 | Eliciting Latent Predictions from Transformers with the Tuned Lens | 2024 | — | C | |
| 2322 | EiG-Search: Generating Edge-Induced Subgraphs for GNN Explanation in Linear Time. | 2024 | ICML | University of Alberta; Université de Montréal; Huawei | B |
| 2323 | Efficient and Interpretable Grammatical Error Correction with Mixture of Experts. | 2024 | EMNLP | National University of Singapore; Mohamed bin Zayed University of Artificial Int | B |
| 2324 | Efficient Sketches for Training Data Attribution and Studying the Loss Landscape. | 2024 | NeurIPS | Google DeepMind (United Kingdom); Imec the Netherlands | B |
| 2325 | Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization. | 2024 | ICLR | Center for Innovation in Data Engineering and Science, University of Pennsylvani | B |
| 2326 | EX-FEVER: A Dataset for Multi-hop Explainable Fact Verification. | 2024 | ACL | University of Chinese Academy of Sciences; State Key Laboratory of Pattern Recog | B |
| 2327 | ESCoT: Towards Interpretable Emotional Support Dialogue Systems. | 2024 | ACL | Renmin University of China; Independent Age | B |
| 2328 | EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free Probing. | 2024 | NeurIPS | State Key Lab of CAD&CG, Zhejiang University; China Mobile (Zhejiang) Research & | B |
| 2329 | ELAD: Explanation-Guided Large Language Models Active Distillation. | 2024 | ACL | Emory University; Emory and Henry College | B |
| 2330 | Dynamic Multi-granularity Attribution Network for Aspect-based Sentiment Analysis. | 2024 | EMNLP | State Key Laboratory of Cognitive Intelligence; University of Science and Techno | B |
| 2331 | Dynamic Evaluation of Large Language Models by Meta Probing Agents. | 2024 | ICML | Microsoft Research; University of Science and Technology of China | B |
| 2332 | Dynamic Discounted Counterfactual Regret Minimization. | 2024 | ICLR | B | |
| 2333 | Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection. | 2024 | EMNLP | Peking University | B |
| 2334 | Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations. | 2024 | EMNLP | Singapore Management University; National University of Singapore | B |
| 2335 | Does Large Language Model Contain Task-Specific Neurons? | 2024 | EMNLP | Faculty of Information Engineering and Automation; Kunming University of Science | B |
| 2336 | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations. | 2024 | ICML | Columbia University; University of California, Berkeley; New York University | B |
| 2337 | Do Llamas Work in English? On the Latent Language of Multilingual Transformers | 2024 | — | C | |
| 2338 | Do Large Language Models Latently Perform Multi-Hop Reasoning? | 2024 | C | ||
| 2339 | Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates. | 2024 | ICML | University of Wisconsin–Madison | B |
| 2340 | Do LLMs Build World Representations? Probing Through the Lens of State Abstraction. | 2024 | NeurIPS | McGill University | B |
| 2341 | Do Counterfactually Fair Image Classifiers Satisfy Group Fairness? - A Theoretical and Empirical Study. | 2024 | NeurIPS | Seoul National University | B |
| 2342 | Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional Semantics. | 2024 | ACL | Chinese University of Hong Kong; University of Cambridge | B |
| 2343 | Dissect Black Box: Interpreting for Rule-Based Explanations in Unsupervised Anomaly Detection. | 2024 | NeurIPS | Shanghai Artificial Intelligence Laboratory; Shenzhen University; Tsinghua Shenz | B |
| 2344 | Disentangling Interpretable Factors with Supervised Independent Subspace Principal Component Analysis. | 2024 | NeurIPS | BGI Genomics; Columbia University; New York Genome Center | B |
| 2345 | Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models' Understanding of Discourse Relations. | 2024 | ACL | National University of Singapore; Sichuan University; Institute for Infocomm Res | B |
| 2346 | Discovering plasticity rules that organize and maintain neural circuits. | 2024 | NeurIPS | Department of Physics; Department of Physiology and Biophysics; University of Wa | B |
| 2347 | Discovering Biases in Information Retrieval Models Using Relevance Thesaurus as Global Explanation. | 2024 | EMNLP | University of Massachusetts Amherst | B |
| 2348 | Digital Socrates: Evaluating LLMs through Explanation Critiques. | 2024 | ACL | Allen Institute for Artificial Intelligence | B |
| 2349 | Diffusion-TS: Interpretable Diffusion for General Time Series Generation. | 2024 | ICLR | Hefei University of Technology | B |
| 2350 | Dial BeInfo for Faithfulness: Improving Factuality of Information-Seeking Dialogue via Behavioural Fine-Tuning. | 2024 | EMNLP | LTL, University of Cambridge | B |
| 2351 | DetoxLLM: A Framework for Detoxification with Explanations. | 2024 | EMNLP | University of British Columbia | B |
| 2352 | Detecting, Explaining, and Mitigating Memorization in Diffusion Models. | 2024 | ICLR | University of Maryland; Zhejiang University; Sony Digital Audio Disc Corporation | B |
| 2353 | Designs for Enabling Collaboration in Human-Machine Teaming via Interactive and Explainable Systems. | 2024 | NeurIPS | MIT Lincoln Laboratory; The University of; Georgia Institute of Technology | B |
| 2354 | Designing Decision Support Systems using Counterfactual Prediction Sets. | 2024 | ICML | Max Planck Institute for Software Systems | B |
| 2355 | Denoising Diffusion Path: Attribution Noise Reduction with An Auxiliary Diffusion Model. | 2024 | NeurIPS | Shanghai Artificial Intelligence Laboratory; Fudan University; Institute of Scie | B |
| 2356 | DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search | 2024 | DeepSeek | C | |
| 2357 | Decomposing Co-occurrence Matrices into Interpretable Components as Formal Concepts. | 2024 | ACL | Japan Advanced Institute of Science and Technology; Tokyo Denki University; JSPS | B |
| 2358 | Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning | 2024 | Akshara Prabhakar, Thomas L. Griffiths, R. Thomas McCoy, respectively from | C | |
| 2359 | Data-faithful Feature Attribution: Mitigating Unobservable Confounders via Instrumental Variables. | 2024 | NeurIPS | University of Illinois Urbana-Champaign | B |
| 2360 | Data-Centric Explainable Debiasing for Improving Fairness in Pre-trained Language Models. | 2024 | ACL | Jilin University; New Jersey Institute of Technology; Key Laboratory of Symbolic | B |
| 2361 | Data Debugging with Shapley Importance over Machine Learning Pipelines. | 2024 | ICLR | Microsoft, USA; University of Amsterdam, The Netherlands | B |
| 2362 | Data Attribution for Text-to-Image Models by Unlearning Synthesized Images. | 2024 | NeurIPS | Carnegie Mellon University; Adobe Research; University of California, Berkeley | B |
| 2363 | Dancing in Chains: Reconciling Instruction Following and Faithfulness in Language Models. | 2024 | EMNLP | Stanford University; Samaya AI; Orby AI; AWS AI Labs; Google; NVIDIA; Denser.ai | B |
| 2364 | DU-Shapley: A Shapley Value Proxy for Efficient Dataset Valuation. | 2024 | NeurIPS | Inria; ENSAE Paris | B |
| 2365 | DISCRET: Synthesizing Faithful Explanations For Treatment Effect Estimation. | 2024 | ICML | Peking University; University of Pennsylvania; Harvard University | B |
| 2366 | DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning. | 2024 | NeurIPS | National University of Singapore; Institute for Infocomm Research; AI Singapore; | B |
| 2367 | DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection. | 2024 | ACL | School of Computer Science and Technology; Xi'an Jiaotong University; University | B |
| 2368 | Credit Attribution and Stable Compression. | 2024 | NeurIPS | Tel Aviv University; Technion and Google Research; Georgetown University and Goo | B |
| 2369 | Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions. | 2024 | NeurIPS | Berkeley College; University of California, Berkeley; Microsoft Research (United | B |
| 2370 | Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training. | 2024 | ICML | Nanjing University of Aeronautics and Astronautics; RIKEN Center for Advanced In | B |
| 2371 | Counterfactual Metarules for Local and Global Recourse. | 2024 | ICML | J.P. Morgan | B |
| 2372 | Counterfactual Image Editing. | 2024 | ICML | Department of Computer Science, Columbia Uni- | B |
| 2373 | Counterfactual Fairness by Combining Factual and Counterfactual Predictions. | 2024 | NeurIPS | Purdue University | B |
| 2374 | Counterfactual Density Estimation using Kernel Stein Discrepancies. | 2024 | ICLR | Carnegie Mellon University | B |
| 2375 | CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples. | 2024 | ACL | University of Wisconsin–Madison; Microsoft Research | B |
| 2376 | Cosmopedia: how to create large-scale synthetic data for pre-training | 2024 | AI2 | C | |
| 2377 | Controlling Risk of Retrieval-augmented Generation: A Counterfactual Prompting Framework. | 2024 | EMNLP | Institute of Computing Technology, Chinese Academy of Sciences; University of Ch | B |
| 2378 | Controlling Counterfactual Harm in Decision Support Systems Based on Prediction Sets. | 2024 | NeurIPS | Max Planck Institute for Software Systems | B |
| 2379 | Consistent Document-level Relation Extraction via Counterfactuals. | 2024 | EMNLP | EU Business School, Munich; Munich Center for Machine Learning | B |
| 2380 | Confidence Regulation Neurons in Language Models. | 2024 | NeurIPS | ETH Zurich; University of Sheffield | B |
| 2381 | Concept Bottleneck Generative Models. | 2024 | ICLR | Emory University, Atlanta, USA | B |
| 2382 | Concept Alignment | 2024 | Princeton | C | |
| 2383 | Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks. | 2024 | ICML | Computer and Information Science, University of Pennsylvania; Center for Brain S | B |
| 2384 | Complex priors and flexible inference in recurrent circuits with dendritic nonlinearities. | 2024 | ICLR | New York University | B |
| 2385 | Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals. | 2024 | ACL | University of Trieste; ETH Zurich; AREA Science Park; Max Planck Institute for I | B |
| 2386 | Compact Proofs of Model Performance via Mechanistic Interpretability. | 2024 | NeurIPS | Massachusetts Institute of Technology | B |
| 2387 | Codebook Features: Sparse and Discrete Interpretability for Neural Networks. | 2024 | ICML | Stanford University | B |
| 2388 | CodePlan: Unlocking Reasoning Potential in Large Language Models by Scaling Code-form Planning | 2024 | Tsinghua University + Ant Group | C | |
| 2389 | Coarse-to-Fine Concept Bottleneck Models. | 2024 | NeurIPS | Laboratoire d'Informatique, de Robotique et de Microélectronique de Montpellier; | B |
| 2390 | CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems. | 2024 | EMNLP | German Research Centre for Artificial Intelligence; Technische Universität Berli | B |
| 2391 | CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity. | 2024 | EMNLP | Intel (United States) | B |
| 2392 | CoSy: Evaluating Textual Explanations of Neurons. | 2024 | NeurIPS | Technische Universität Berlin; BIFOLD; UMI Lab, ATB Potsdam, Germany; Fraunhofer | B |
| 2393 | Cluster-Norm for Unsupervised Probing of Knowledge. | 2024 | EMNLP | Cadenza Labs; University of Cambridge; ENS Paris-Saclay; University of Californi | B |
| 2394 | ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs. | 2024 | EMNLP | University of Washington; University of California, Berkeley; Stanford Universit | B |
| 2395 | CircuitNet 2.0: An Advanced Dataset for Promoting Machine Learning Innovations in Realistic Chip Design Environment. | 2024 | ICLR | B | |
| 2396 | Circuit Component Reuse Across Tasks in Transformer Language Models. | 2024 | ICLR | Brown University; School of Medicine; University of Tübingen | B |
| 2397 | ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline | 2024 | Zhipu AI | C | |
| 2398 | ChartCheck: Explainable Fact-Checking over Real-World Chart Images. | 2024 | ACL | King's College London; University of Utah; University of Pennsylvania; TIB – Lei | B |
| 2399 | Causality-Inspired Spatial-Temporal Explanations for Dynamic Graph Neural Networks. | 2024 | ICLR | B | |
| 2400 | CausalGym: Benchmarking causal interpretability methods on linguistic tasks. | 2024 | ACL | Nielsen Engineering & Research (United States) | B |
| 2401 | Causal Discovery from Event Sequences by Local Cause-Effect Attribution. | 2024 | NeurIPS | CISPA Helmholtz Center; DER Security (United States); Institute for Structural P | B |
| 2402 | Causal Contrastive Learning for Counterfactual Regression Over Time. | 2024 | NeurIPS | Paris-Saclay University, CentraleSupélec, MICS Lab, Gif-sur-Yvette, France; Sain | B |
| 2403 | Causal Action Influence Aware Counterfactual Data Augmentation. | 2024 | ICML | ETH Zurich; Max Planck Institute for Intelligent Systems; University of Tübingen | B |
| 2404 | Can Large Language Models Mine Interpretable Financial Factors More Effectively? A Neural-Symbolic Factor Mining Agent Model. | 2024 | ACL | Renmin University of China; Faculty of Information Engineering and Automation; K | B |
| 2405 | Can Large Language Models Interpret Noun-Noun Compounds? A Linguistically-Motivated Study on Lexicalized and Novel Compounds. | 2024 | ACL | University of Bologna; Hong Kong Polytechnic University | B |
| 2406 | Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words? | 2024 | EMNLP | Google (United States) | B |
| 2407 | Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales? | 2024 | NeurIPS | Hong Kong Baptist University; Wuhan University | B |
| 2408 | Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments. | 2024 | ACL | Nanjing University; Microsoft | B |
| 2409 | Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring. | 2024 | EMNLP | King's College London; AQA; The Alan Turing Institute; University of Warwick | B |
| 2410 | COLEP: Certifiably Robust Learning-Reasoning Conformal Prediction via Probabilistic Circuits. | 2024 | ICLR | Delft University of Technology | B |
| 2411 | CLOMO: Counterfactual Logical Modification with Large Language Models. | 2024 | ACL | City University of Hong Kong; Tsinghua University; Seattle University; The Hong | B |
| 2412 | CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural Networks. | 2024 | ICML | The Hong Kong University of Science and Technology (Guangzhou); Beihang Universi | B |
| 2413 | CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems. | 2024 | ACL | Huawei Noah’s Ark Lab; Université de Montréal | B |
| 2414 | CF-OPT: Counterfactual Explanations for Structured Prediction. | 2024 | ICML | Polytechnique Montréal | B |
| 2415 | CAUSE: Counterfactual Assessment of User Satisfaction Estimation in Task-Oriented Dialogue Systems. | 2024 | ACL | Leiden University; University of Amsterdam | B |
| 2416 | Bridging Word-Pair and Token-Level Metaphor Detection with Explainable Domain Mining. | 2024 | ACL | Chinese Academy of Sciences | B |
| 2417 | Born Differently Makes a Difference: Counterfactual Study of Bias in Biography Generation from a Data-to-Text Perspective. | 2024 | ACL | Data61 | B |
| 2418 | Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image Classification. | 2024 | ICLR | University of Pennsylvania, PA, USA; Johns Hopkins University, Department of Bio | B |
| 2419 | Binding in hippocampal-entorhinal circuits enables compositionality in cognitive maps. | 2024 | NeurIPS | Neurosciences Institute; Université Paris-Saclay; University of California, Davi | B |
| 2420 | BiasWipe: Mitigating Unintended Bias in Text Classifiers through Model Interpretability. | 2024 | EMNLP | Indian Institute of Technology Patna | B |
| 2421 | Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations. | 2024 | EMNLP | Sichuan University; Singapore Management University; ♡Engineering Research Cente | B |
| 2422 | Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning. | 2024 | EMNLP | University of Illinois Urbana-Champaign; Vanderbilt University | B |
| 2423 | Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS | 2024 | Jinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che, Zengqi We | C | |
| 2424 | Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics. | 2024 | EMNLP | Unitelma Sapienza University; Sapienza University of Rome | B |
| 2425 | Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable? | 2024 | NeurIPS | ETH Zurich | B |
| 2426 | Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals. | 2024 | EMNLP | Beijing Normal University | B |
| 2427 | Beyond Accuracy: Ensuring Correct Predictions With Correct Rationales. | 2024 | NeurIPS | University of Delaware | B |
| 2428 | Benchmarking the Attribution Quality of Vision Models. | 2024 | NeurIPS | Technische Universität Darmstadt | B |
| 2429 | Benchmarking Deletion Metrics with the Principled Explanations. | 2024 | ICML | Elmore Family School of Electrical and Computer Engineering, Purdue University, | B |
| 2430 | Benchmarking Counterfactual Image Generation. | 2024 | NeurIPS | National and Kapodistrian University of Athens; Athena Research and Innovation C | B |
| 2431 | Beam Enumeration: Probabilistic Explainability For Sample Efficient Self-conditioned Molecular Design. | 2024 | ICLR | Laboratory of Artificial Chemical Intelligence (LIAC), Institut des Sciences et | B |
| 2432 | Bayesian Power Steering: An Effective Approach for Domain Adaptation of Diffusion Models. | 2024 | ICML | Department of Applied Mathematics, The Hong Kong; Polytechnic University, Hong K | B |
| 2433 | Balancing Speciality and Versatility: a Coarse to Fine Framework for Supervised Fine-tuning Large Language Model | 2024 | EleutherAI | C | |
| 2434 | Balanced Resonate-and-Fire Neurons. | 2024 | ICML | Institute of Robotics | B |
| 2435 | Backward Lens: Projecting Language Model Gradients into the Vocabulary Space | 2024 | Shahar Katz, Yonatan Belinkov:Technion – Israel Institute of | C | |
| 2436 | BAN: Detecting Backdoors Activated by Adversarial Neuron Noise. | 2024 | NeurIPS | Radboud University; Delft University of Technology; Vrije Universiteit Amsterdam | B |
| 2437 | BAM!Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts | 2024 | Cohere | C | |
| 2438 | B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable. | 2024 | NeurIPS | Max Planck Institute for Informatics; RTG Neuroexplicit Models of Language, Visi | B |
| 2439 | AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments. | 2024 | EMNLP | Princeton University | B |
| 2440 | Autaptic Synaptic Circuit Enhances Spatio-temporal Predictive Learning of Spiking Neural Networks. | 2024 | ICML | Peking University; Shenzhen Maternity and Child Healthcare Hospital | B |
| 2441 | Auditing Local Explanations is Hard. | 2024 | NeurIPS | University of Tübingen | B |
| 2442 | AttributionBench: How Hard is Automatic Attribution Evaluation? | 2024 | ACL | The Ohio State University | B |
| 2443 | Attribute Based Interpretable Evaluation Metrics for Generative Models. | 2024 | ICML | Yonsei University | B |
| 2444 | Attention Meets Post-hoc Interpretability: A Mathematical Perspective. | 2024 | ICML | Université Côte d'Azur; Institut de Biologie Valrose; Laboratoire Jean-Alexandre | B |
| 2445 | AttEXplore: Attribution for Explanation with model parameters eXploration. | 2024 | ICLR | B | |
| 2446 | Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability. | 2024 | ACL | Soongsil University; Adobe Research, USA | B |
| 2447 | Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics | 2024 | Yaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan Belinkov | C | |
| 2448 | Are self-explanations from Large Language Models faithful? | 2024 | ACL | Mila – Quebec AI Institute; Polytechnique Montréal; McGill University; CIFAR; Fa | B |
| 2449 | Arcee’s MergeKit: A Toolkit for Merging Large Language Models | 2024 | Toolkit developed by the Arcee team (Florida, USA), including Charles Goddard, Shamane Sir | C | |
| 2450 | Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution. | 2024 | ACL | Bocconi University | B |
| 2451 | Analysing the Generalisation and Reliability of Steering Vectors. | 2024 | NeurIPS | University College London; FAR AI; Athena Research and Innovation Center In Info | B |
| 2452 | An interpretable error correction method for enhancing code-to-code translation. | 2024 | ICLR | B | |
| 2453 | An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare Records. | 2024 | EMNLP | University of Copenhagen; Cork University Hospital | B |
| 2454 | An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs. | 2024 | ACL | Department of Computer Science; George Mason University | BC |
| 2455 | An Interpretable Evaluation of Entropy-based Novelty of Generative Models. | 2024 | ICML | Department of Computer Science and Engineering, The Chinese University of Hong K | B |
| 2456 | An Empirical Examination of Balancing Strategy for Counterfactual Estimation on Time Series. | 2024 | ICML | Jilin University; University of Southern California; University of California Sa | B |
| 2457 | Almost-Linear RNNs Yield Highly Interpretable Symbolic Codes in Dynamical Systems Reconstruction. | 2024 | NeurIPS | Central Institute of Mental Health; Heidelberg University | B |
| 2458 | Advancing Large Language Model Attribution through Self-Improving. | 2024 | EMNLP | Harbin Institute of Technology; Peng Cheng Laboratory; Northeastern University, | B |
| 2459 | Adaptive Quantization Error Reconstruction for LLMs with Mixed Precision | 2024 | Alibaba Cloud | C | |
| 2460 | Activation Scaling for Steering and Interpreting Language Models. | 2024 | EMNLP | Pioneer (United States); Pioneer Hi-Bred | B |
| 2461 | Accelerating Nash Equilibrium Convergence in Monte Carlo Settings Through Counterfactual Value Based Fictitious Play. | 2024 | NeurIPS | Huazhong University of Science and Technology; National Key Laboratory of Scienc | B |
| 2462 | Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe Sampling. | 2024 | NeurIPS | Université de Montréal | B |
| 2463 | Abstracted Shapes as Tokens - A Generalizable and Interpretable Model for Time-series Classification. | 2024 | NeurIPS | Rensselaer Polytechnic; Stony Brook University; Touro University California; IBM | B |
| 2464 | AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation. | 2024 | ACL | Department of Computer Science and Engineering, HKUST; Tencent AI Lab; Amazon; N | B |
| 2465 | AXCEL: Automated eXplainable Consistency Evaluation using LLMs. | 2024 | EMNLP | Amazon | B |
| 2466 | AR-Pro: Counterfactual Explanations for Anomaly Repair with Formal Properties. | 2024 | NeurIPS | Department of Computer and Information Science; University of Pennsylvania | B |
| 2467 | AGR: Reinforced Causal Agent-Guided Self-explaining Rationalization. | 2024 | ACL | Shanxi University | B |
| 2468 | A theoretical design of concept sets: improving the predictability of concept bottleneck models. | 2024 | NeurIPS | University of Cambridge | B |
| 2469 | A hierarchical decomposition for explaining ML performance discrepancies. | 2024 | NeurIPS | University of California, San Francisco; Center for Devices and Radiological Hea | B |
| 2470 | A Survey on Natural Language Counterfactual Generation. | 2024 | EMNLP | Nanyang Technological University; Shenzhen University | B |
| 2471 | A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis. | 2024 | ICLR | The Ohio State University; Amazon Alexa; Princeton University; Rensselaer Polyte | B |
| 2472 | A Robust Dual-debiasing VQA Model based on Counterfactual Causal Effect. | 2024 | EMNLP | Research & Development Institute; Northwestern Polytechnical University | B |
| 2473 | A Neural Network Approach for Efficiently Answering Most Probable Explanation Queries in Probabilistic Models. | 2024 | NeurIPS | The University of Texas at Dallas | B |
| 2474 | A Multimodal Automated Interpretability Agent. | 2024 | ICML | Massachusetts Institute of Technology | B |
| 2475 | A Metalearned Neural Circuit for Nonparametric Bayesian Inference. | 2024 | NeurIPS | Department of Computer Science; Princeton University; Department of Psychology | B |
| 2476 | A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. | 2024 | ICML | The University of Sydney | B |
| 2477 | A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task. | 2024 | ACL | University of Mannheim; Georgia Institute of Technology; Heinrich Heine Universi | B |
| 2478 | A Linear Algebraic Framework for Counterfactual Generation. | 2024 | ICLR | B | |
| 2479 | A Hierarchical Adaptive Multi-Task Reinforcement Learning Framework for Multiplier Circuit Design. | 2024 | ICML | University of Science and Technology of China; The Hong Kong University of Scien | B |
| 2480 | A Geometric Explanation of the Likelihood OOD Detection Paradox. | 2024 | ICML | University of Toronto | B |
| 2481 | A General Protocol to Probe Large Vision Models for 3D Physical Understanding. | 2024 | NeurIPS | University of Oxford; Shanghai Jiao Tong University | B |
| 2482 | A Dual-module Framework for Counterfactual Estimation over Time. | 2024 | ICML | University of Science and Technology of China; Nanyang Technological University; | B |
| 2483 | A Concept-Based Explainability Framework for Large Multimodal Models. | 2024 | NeurIPS | Sorbonne Université; Institut Polytechnique de Paris | B |
| 2484 | A Compositional Atlas for Algebraic Circuits. | 2024 | NeurIPS | University of California, Los Angeles; Universidade de São Paulo; Arizona State | B |
| 2485 | A Circuit Domain Generalization Framework for Efficient Logic Synthesis in Chip Design. | 2024 | ICML | Key Laboratory of Technology in GIP AS, University of Science and; Technology of | B |
| 2486 | A Causal Approach for Counterfactual Reasoning in Narratives. | 2024 | ACL | Hong Kong Polytechnic University | B |
| 2487 | A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution. | 2024 | EMNLP | University of Maryland, College Park | B |
| 2488 | "What Data Benefits My Classifier?" Enhancing Model Performance and Interpretability through Influence-Based Data Selection. | 2024 | ICLR | University of South Florida, Bellini College of Artificial Intelligence, Cyberse | B |
| 2489 | "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations? | 2024 | EMNLP | EU Business School, Munich; Munich Center for Machine Learning; University of Ca | B |
| 2490 | Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning | 2023 | — | C | |
| 2491 | Language Representation Projection: Can We Transfer Factual Knowledge across Languages in Multilingual Language Models? | 2023 | Tianjin University | C | |
| 2492 | Hyperpolyglot LLMs: Cross-Lingual Interpretability in Token Embeddings | 2023 | — | C | |
| 2493 | Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models | 2023 | UNC Chapel Hill / Google Research | C | |
| 2494 | Context-DPO | 2023 | Chinese Academy of Sciences / Microsoft | C | |
| 2495 | The Geometry of Multilingual Language Model Representations | 2022 | — | C | |
| 2496 | Discovering Low-rank Subspaces for Language-agnostic Multilingual Representations | 2021 | Shanghai Jiao Tong University | C | |
| 2497 | Paper: Enhancing Automated Interpretability with Output-Centric Feature Descriptions | Yoav Gur-Arieh, Roy Mayan, Chen Agassy (Tel Aviv University) | C | ||
| 2498 | Paper 2: FADE: Why Bad Descriptions Happen to Good Features | Tianjin University | C | ||
| 2499 | Scaling Still Matters Most in Model Training — An Interview with Baosong Yang, Head of Multilingual at Alibaba Tongyi Qwen | Tongyi Lab | C | ||
| 2500 | Multi-Turn Planning Techniques for LLM Agent RL Training: An Explainer Thread Covering Several Papers | — | C | ||
| 2501 | An Overnight Reversal: The Brazilian LLM That "Broke Into the Top Tier" Turned Out to Be a Re-Skinned Chinese Model | IT Company of the Rio de Janeiro City Government, Brazil | C | ||
| 2502 | ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning | Shanghai Jiao Tong University | C | ||
| 2503 | mHC: Manifold-Constrained Hyper-Connections | DeepSeek | C | ||
| 2504 | ZHEN: Check which papers this paper cites, and its "" experiments | — | C | ||
| 2505 | WorldPM | Qwen Team | C | ||
| 2506 | Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation | Zhejiang University | C | ||
| 2507 | When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers | Rensselaer Polytechnic Institute, USA | C | ||
| 2508 | When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs | Harvard + Amazon | C | ||
| 2509 | When AI builds itself : Anthropic Blog | Anthropic | C | ||
| 2510 | What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs | Signals and Interactive Systems Lab, University of Trento, Italy | C | ||
| 2511 | Weight Patching: Toward Source-Level Mechanistic Localization in LLMs | University of Science and Technology of China | C | ||
| 2512 | Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning | USTC | C | ||
| 2513 | Unlearning Isn’t Deletion: Investigating Reversibility of Machine Unlearning in LLMs | The Hong Kong Polytechnic University | C | ||
| 2514 | Understanding the Dark Side of LLMs’ Intrinsic Self-Correction | Oxford / Google DeepMind / Mila | C | ||
| 2515 | Understanding and Enforcing Weight Disentanglement in Task Arithmetic | Nanjing University | C | ||
| 2516 | Trust-Region Adaptive Policy Optimization | Tsinghua University / Ant Group | C | ||
| 2517 | Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models | Alibaba | C | ||
| 2518 | TriAttention: Efficient Long Reasoning with Trigonometric KV Compression | MIT / NVIDIA / Zhejiang University | C | ||
| 2519 | Trans-Zero | ByteDance | C | ||
| 2520 | Training-Free Looped Transformers :arXiv | University of Texas at Austin | C | ||
| 2521 | Training Transformers for KV Cache Compressibility | University of Oxford、Technion、AITHYRA、NVIDIA | C | ||
| 2522 | Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning | Harbin Institute of Technology | C | ||
| 2523 | Topology of Reasoning | The University of Tokyo / Google DeepMind | C | ||
| 2524 | Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning | Renmin University of China / BAAI / Kuaishou | C | ||
| 2525 | Token-Importance Guided Direct Preference Optimization | Chinese Academy of Sciences / ByteDance | C | ||
| 2526 | Titans: Learning to Memorize at Test Time | Google Research | C | ||
| 2527 | TileRT: Tile-Based Runtime for Ultra-Low-Latency LLM Inference | TileRT | C | ||
| 2528 | The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks :arXiv | LeCun's Team | C | ||
| 2529 | The Latent Space | NUS Team | C | ||
| 2530 | The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models | Oxford & Anthropic | C | ||
| 2531 | TTRL: Test-Time Reinforcement Learning | Tsinghua University / Shanghai AI Laboratory | C | ||
| 2532 | TOKEN ALIGNMENT HEADS: UNVEILING ATTENTION’S ROLE IN LLM MULTILINGUAL TRANSLATION | ByteDance | C | ||
| 2533 | Superposition, Memorization, and Double Descent | anthropic | C | ||
| 2534 | Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data | Anthropic,Berkerly | C | ||
| 2535 | Subliminal Learning Is Steering Vector Distillation | Stanford University | C | ||
| 2536 | Steering Language Models with Weight Arithmetic | Anthropic | C | ||
| 2537 | Spurious rewards: rethinking training signals in RLVR | University of Washington / Allen Institute for AI / UC Berkeley | C | ||
| 2538 | Sparse Crosscoders for Cross-Layer Features and Model Diffing | Anthropic | C | ||
| 2539 | Sparse Attention Post-Training for Mechanistic Interpretability | Max Planck Institute for Intelligent Systems (MPI-IS) + Oxford + | C | ||
| 2540 | Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance | Meta | C | ||
| 2541 | SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis | Renmin University of China / BAAI / DataCanvas | C | ||
| 2542 | Sigma-Moe-Tiny Technical Report | Microsoft | C | ||
| 2543 | Self-Distillation Enables Continual Learning | MIT、ETH | C | ||
| 2544 | Self-DC: When to Reason and When to Act Self Divide-and-Conquer for Compositional Questions | The Chinese University of Hong Kong | C | ||
| 2545 | Seed-Thinking-v1.5: Advancing Superb Reasoning Models with Reinforcement Learning | ByteDance | C | ||
| 2546 | Scaling sparse feature circuit finding for in-context learning | ETH Zurich | C | ||
| 2547 | Scaling and context steer LLMs along the same computational path as the human brain | Meta | C | ||
| 2548 | Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining | University of Chicago | C | ||
| 2549 | STRESS-TESTING MODEL SPECS REVEALS CHARACTER DIFFERENCES AMONG LANGUAGE MODELS | Anthropic | C | ||
| 2550 | SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning | Chinese Academy of Sciences / Meituan | C | ||
| 2551 | SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model Training | bilibili | C | ||
| 2552 | SPARSE FEATURE CIRCUITS | Northwestern University | C | ||
| 2553 | SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs | Zhejiang University / Ant Group | C | ||
| 2554 | SE-BENCH: Benchmarking Self-Evolution with Knowledge Internalization | Tsinghua University | C | ||
| 2555 | Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging | UC Berkeley | C | ||
| 2556 | Rethinking Thinking Tokens: LLMs as Improvement Operators | Meta | C | ||
| 2557 | Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution | Shanghai Jiao Tong University / Alibaba Tongyi Lab | C | ||
| 2558 | Reinforcement Learning with Rubric Anchors | Ant Group + Zhejiang University | C | ||
| 2559 | Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection | Tsinghua Shenzhen International Graduate School / Alibaba | C | ||
| 2560 | Recursive Multi-Agent Systems :arXiv | UIUC、Stanford University、NVIDIA、MIT | C | ||
| 2561 | Reasoning Models Struggle to Control their Chains of Thought:arXiv | OpenAI et al. | C | ||
| 2562 | Reasoning Models Generate Societies of Thought | Google / University of Chicago | C | ||
| 2563 | ReFT: Reasoning with Reinforced Fine-Tuning | ByteDance | C | ||
| 2564 | R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning | Renmin University of China / Beijing Institute of Technology / DataCanvas | C | ||
| 2565 | Quantize What Counts: More for Keys, Less for Values | Case Western Reserve University; Rice University; Meta | C | ||
| 2566 | Pretraining with hierarchical memories: separating long-tail and common knowledge | Apple | C | ||
| 2567 | Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter | Kimi (Moonshot AI) / Tsinghua University | C | ||
| 2568 | PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models | Fudan University / Soochow University / Shanghai AI Laboratory / Stony Brook University / CUHK | C | ||
| 2569 | OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization | Together AI、University of Sydney、UIUC | C | ||
| 2570 | Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights | MIT CSAIL | C | ||
| 2571 | Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models | ByteDance Seed / Peking University | C | ||
| 2572 | Navigating the Accuracy-Size Trade-Off with Flexible Model Merging | EPFL | C | ||
| 2573 | Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training | Xi'an Jiaotong University / CUHK / China Mobile / Nanyang Technological University | C | ||
| 2574 | Monitoring Monitorability | OpenAI | C | ||
| 2575 | Model Merging in Pre-training of Large Language Models | SEED | C | ||
| 2576 | Mixture-of-Depths Attention | Seed | C | ||
| 2577 | Memory in the Age of Al Agents | National University of Singapore, Renmin University of China | C | ||
| 2578 | Memory in the Age of AI Agents | National University of Singapore et al. | C | ||
| 2579 | Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning | Ruoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpe | C | ||
| 2580 | Memento-Skills: Let Agents Design Agents | Memento | C | ||
| 2581 | Mapping Post-Training Forgetting in Language Models at Scale | University of Tübingen | C | ||
| 2582 | Mamba-3: Improved Sequence Modeling using State Space Principles:ICLR | Carnegie Mellon University / Princeton University | C | ||
| 2583 | MODEL MERGING WITH FUNCTIONAL DUAL ANCHORS | CUHK / Westlake University | C | ||
| 2584 | MLP Memory: A Retriever-Pretrained Memory for Large Language Models | Shanghai Jiao Tong University | C | ||
| 2585 | MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection | Beihang University / IQuest Research / Shanghai Jiao Tong University / UBC / Langboat | C | ||
| 2586 | MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging | University of Electronic Science and Technology of China / Dalian University of Technology | C | ||
| 2587 | MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs | EPFL, Switzerland | C | ||
| 2588 | MAmmoTH2: Scaling Instructions from the Web | Carnegie Mellon University | C | ||
| 2589 | Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression | Northwestern University | C | ||
| 2590 | Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex | Tencent Hunyuan | C | ||
| 2591 | Linking forward-pass dynamics in Transformers and real-time human processing | Harvard University | C | ||
| 2592 | Linear representations in language models can change dramatically over a conversation | Google Deepmind | C | ||
| 2593 | Let’s Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model | University of Macau / Tiger Research | C | ||
| 2594 | Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs | Meta | C | ||
| 2595 | Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts. | Beijing Jiaotong University | C | ||
| 2596 | Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs :arXiv | University of Georgia | C | ||
| 2597 | Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs | University of Georgia / University of California | C | ||
| 2598 | Learning to Reason in 13 Parameters | Meta | C | ||
| 2599 | Learning is Forgetting: LLM Training As Lossy Compression | Princeton University | C | ||
| 2600 | Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning. Authors: Yun Qu, Yuhang Jiang, Boyuan Wang, Yixiu Mao et al., Prof. Xiangyang Ji's team, Tsinghua University | — | C | ||
| 2601 | Language as a Latent Variable for Reasoning Optimization | Tongyi, ZJU | C | ||
| 2602 | Language Models' Factuality Depends on the Language of Inquiry. | Harvard University | C | ||
| 2603 | LLaMAX2: Your Translation-Enhanced Model also Performs Well in Reasoning | Nanjing University / Shanghai AI Lab | C | ||
| 2604 | KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning | CMU | C | ||
| 2605 | Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation | Peking University | C | ||
| 2606 | Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference | Microsoft | C | ||
| 2607 | KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction | Seoul National University | C | ||
| 2608 | KV Cache Transform Coding for Compact Storage in LLM Inference | Nvidia | C | ||
| 2609 | Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders | Google Deepmind | C | ||
| 2610 | Is In-Context Learning Learning? | Microsoft | C | ||
| 2611 | Interpreting Language Models Through Concept Descriptions: A Survey | Technical University of Berlin | C | ||
| 2612 | Internal bias in reasoning models leads to overthinking | Nanjing University | C | ||
| 2613 | Internal Value Alignment in Large Language Models throughControlled Value Vector Activation | USTC | C | ||
| 2614 | Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models | Neel Nanda (Google DeepMind) / Durham University | C | ||
| 2615 | Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures | University of Chicago | C | ||
| 2616 | Improving Continual Pre-training Through Seamless Data Packing | Fudan University | C | ||
| 2617 | Identifying indicators of consciousness in Al systems | University of Oxford | C | ||
| 2618 | ICA Lens: Interpreting LLM Embeddings and Activations with Independent Component Analysis | EEEAI Lab | C | ||
| 2619 | Hyperloop Transformers :arXiv | MIT | C | ||
| 2620 | How large language models encode theory-of-mind: a study on sparse parameter patterns | Stanford University | C | ||
| 2621 | How does Alignment Enhance LLMs’ Multilingual Capabilities? A Language Neurons Perspective | Nanjing University / Microsoft Research Asia | C | ||
| 2622 | How Do Multilingual Language Models Remember Facts? | University of Copenhagen | C | ||
| 2623 | Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders (SAERL) | Tsinghua University | C | ||
| 2624 | Generalist Reward Models | Nanjing University | C | ||
| 2625 | GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs | Maxim Zhelnin, Viktor Moskvoretskii, Egor Shvetsov, Egor Ven | C | ||
| 2626 | From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning | Stanford University / LeCun | C | ||
| 2627 | From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms | Hong Kong Baptist University / South China Normal University | C | ||
| 2628 | From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence | CMU & NYU | C | ||
| 2629 | FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling | Alibaba | C | ||
| 2630 | Feeling the Strength but Not the Source: Partial Introspection in LLMs | Harvard University | C | ||
| 2631 | FLEXOLMO: Open Language Models for Flexible Data Use | Weijia Shi, Akshita Bhagia, Kevin Farhat, Niklas Muennighoff | C | ||
| 2632 | Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Models | Renmin University of China | C | ||
| 2633 | Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking | Huawei Noah's Ark Lab | C | ||
| 2634 | Exclusive Self Attention :arXiv | Apple | C | ||
| 2635 | EvoWiki: Evaluating LLMs on Evolving Knowledge | Fudan University | C | ||
| 2636 | Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory | Google DeepMind | C | ||
| 2637 | Episodic memories enable powerful algorithms for goal-directed decision making that are context-aware and explainable | Graz University of Technology | C | ||
| 2638 | End-to-End Test-Time Training for Long Context | Stanford University / NVIDIA / UC Berkeley / UC San Diego / Astera Institute | C | ||
| 2639 | Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning | C | |||
| 2640 | ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer. | C | |||
| 2641 | Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space | SEED | C | ||
| 2642 | Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates | Advanced Micro Devices | C | ||
| 2643 | Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features | Standford | C | ||
| 2644 | Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference :arXiv | Carnegie Mellon University、University of Maryland | C | ||
| 2645 | Diverse Preference Learning for Capabilities and Alignment | MIT | C | ||
| 2646 | DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging | The University of Tokyo | C | ||
| 2647 | CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis | Tongyi Lab | C | ||
| 2648 | Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss | ByteDance Seed | C | ||
| 2649 | Could Thinking Multilingually Empower LLM Reasoning? (Close Reading) | Nanjing University | C | ||
| 2650 | Copyright-Protected Language Generation via Adaptive Model Fusion | ETH Zurich | C | ||
| 2651 | Conditional Memory via Scalable Lookup:A New Axis of Sparsity for Large Language Models | DeepSeek-AI / Peking University | C | ||
| 2652 | Compression is all you need: modeling mathematics | Michael Freedman, Princeton | C | ||
| 2653 | Code-Switching and Syntax: A Large-Scale Experiment | University of Cambridge | C | ||
| 2654 | Chain-of-Thought is Not Explainability | Oxford / Google DeepMind / Mila | C | ||
| 2655 | CL-bench: A Benchmark for Context Learning | Tencent Hunyuan | C | ||
| 2656 | CL-bench Life: Can Language Models Learn from Real-Life Context? | Tencent Hunyuan LLM | C | ||
| 2657 | Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies | Institute of Automation, Chinese Academy of Sciences | C | ||
| 2658 | Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning | Qwen | C | ||
| 2659 | Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs | Anthropic | C | ||
| 2660 | BLEUBERI: BLEU is a surprisingly effective reward for instruction following | University of Maryland | C | ||
| 2661 | AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora | HKUST | C | ||
| 2662 | Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation | Fudan University | C | ||
| 2663 | Attention Residuals | Kimi | C | ||
| 2664 | Anthropic plans Claude memory update with new Memory Files :Blog | Antropic | C | ||
| 2665 | Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels | Fudan University | C | ||
| 2666 | AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models | USTC | C | ||
| 2667 | Alignment Faking in Large Language Models | Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Rog | C | ||
| 2668 | Activation Steering via Generative Causal Mediation | MIT + Stanford + Goodfire Research | C | ||
| 2669 | ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality | MIT / Stanford University / Google | C | ||
| 2670 | AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents | Harbin Institute of Technology | C | ||
| 2671 | ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning | Peking University | C | ||
| 2672 | A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility | Tübingen AI Center, University of Tübingen / University of Cambridge | C | ||
| 2673 | A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce | Salesforce AI Research, UIUC | C | ||
| 2674 | A Mechanistic Analysis of Looped Reasoning Language Models | University of Oxford | C | ||
| 2675 | A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models | University of Oregon,Adobe,Cisco | C | ||
| 2676 | A Formal Comparison Between Chain of Thought and Latent Thought : ICML26 | The University of Tokyo | C | ||
| 2677 | A Comprehensive Survey of Reward Models: Taxonomy,Applications, Challenges, and Future | Peking University / Fudan University | C | ||
| 2678 | "RAG is an algorithm, not magic." — A First-Principles Analysis of Why RAG Works | Unisound | C |