Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

可解释性

原始 PDF

“我们所看见的一切,都是由我们看不见的事物投下的影子。”——Martin Luther King, Jr.,1961 年在 Detroit Council of Churches 的布道

本章(目前仍为提纲性质)转向讨论 LLM 可解释性(interpretability),有时也称为 机制可解释性(mechanistic interpretability)。研究者在这一领域中寻找方法,帮助我们理解语言模型如何实现其行为。大量工作聚焦于揭示模型的内部表示,例如它们似乎编码了何种意义;还关注模型进行推理的路径,以及我们能否描述特定的“回路”。