Mechanistic interpretability for AI safety--a review

L Bereska, E Gavves - arXiv preprint arXiv:2404.14082, 2024 - arxiv.org
Understanding AI systems' inner workings is critical for ensuring value alignment and safety.
This review explores mechanistic interpretability: reverse engineering the computational …

Atp*: An efficient and scalable method for localizing llm behaviour to components

J Kramár, T Lieberum, R Shah, N Nanda - arXiv preprint arXiv:2403.00745, 2024 - arxiv.org
Activation Patching is a method of directly computing causal attributions of behavior to
model components. However, applying it exhaustively requires a sweep with cost scaling …

Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques

R Gupta, I Arcuschin, T Kwa… - Advances in Neural …, 2024 - proceedings.neurips.cc
Mechanistic interpretability methods aim to identify the algorithm a neural network
implements, but it is difficult to validate such methods when the true algorithm is unknown …

Unboxing the Black Box: A Survey on Mechanistic Interpretability for Algorithmic Understanding of Neural Networks

B Kowalska, H Kwaśnicka - Machine Learning, 2026 - Springer
The black box nature of deep neural networks poses a significant challenge for the
deployment of transparent and trustworthy artificial intelligence (AI) systems. With the …