The Mechanistic Ledger
The Mechanistic Ledger
For years, compliance officers and bank examiners have leaned on a convenient fiction: that post-hoc explainability tools like LIME and SHAP could open the black box of artificial intelligence. Researchers have since shown that an adversarial wrapper can hide a model's reliance on race behind a pristine attribution chart. The obligations, meanwhile, have not moved. A lender must still state the actual reasons for an adverse decision, and the April 2026 interagency model risk guidance still demands effective challenge of a model's logic. Output monitoring cannot meet that standard; it describes what a model did, never why. Mechanistic interpretability, the mapping of internal features, sparse autoencoders and raw activations, offers the first credible record of the why. In this issue, we peel back the weights to examine what neural networks are actually doing when no one is watching.