Mechanistic interpretability for AI safety--a review
Understanding AI systems' inner workings is critical for ensuring value alignment and safety.
This review explores mechanistic interpretability: reverse engineering the computational …
This review explores mechanistic interpretability: reverse engineering the computational …
Atp*: An efficient and scalable method for localizing llm behaviour to components
Activation Patching is a method of directly computing causal attributions of behavior to
model components. However, applying it exhaustively requires a sweep with cost scaling …
model components. However, applying it exhaustively requires a sweep with cost scaling …
Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques
Mechanistic interpretability methods aim to identify the algorithm a neural network
implements, but it is difficult to validate such methods when the true algorithm is unknown …
implements, but it is difficult to validate such methods when the true algorithm is unknown …
Unboxing the Black Box: A Survey on Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
B Kowalska, H Kwaśnicka - Machine Learning, 2026 - Springer
The black box nature of deep neural networks poses a significant challenge for the
deployment of transparent and trustworthy artificial intelligence (AI) systems. With the …
deployment of transparent and trustworthy artificial intelligence (AI) systems. With the …