Elize Herrewijnen, Benedetta Muscato, Gizem Gezici +1 more
Ask a language model why it answered something and it will tell you, fluently. Those self-explanations look like a promising route to interpretability, since the model appears to be describing its own reasoning in language anyone can read.
The position taken here is that plausibility and faithfulness are different properties, and only one of them is being demonstrated. A model trained to produce text that reads convincingly will produce explanations that read convincingly, whether or not they correspond to whatever actually determined the output. The explanation is generated by the same process that generated the answer, which means it is another output rather than a window onto the mechanism.
The distinction in the title is the practical one. An explanation is actionable if acting on it changes something reliably, and a story that merely sounds right can be worse than no explanation, because it manufactures confidence.
Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations.Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior.However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question. In this…
Scaling Behavior Foundation Model for Humanoid Robots
arXiv (cs.LG) · July 16, 2026On-Policy Delta Distillation
arXiv (cs.CL) · July 16, 2026Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies
arXiv (cs.AI) · July 16, 2026Concept-Guided Spatial Regularization for World Models in Atari Pong