Skip to Content

The C30 Journal

C30

Index
The C30 Journal, EST. 2026
Status: Active
Article No. 027
Local AI & Ethics //
Geometric technical artwork for Monograph No. 027

The Adversarial Scaffold

How LIME and SHAP Can Be Made to Lie: The Illusion of Post-Hoc Explainability

By Caleb Brown12 Min Read[ .MD ]

Compliance officers in consumer finance spend millions of dollars every quarter admiring a geometric fiction.

The ritual is familiar across every automated underwriting desk in North America and Western Europe. A deep neural network rejects a mortgage application. Within milliseconds, an auxiliary pipeline spins up, samples surrounding permutations, and outputs an orderly horizontal bar chart. Income stability: plus forty-two points. Debt-to-income ratio: minus thirty-one points. Revolving credit utilization: minus eighteen points. Race: zero. Gender: zero.

The chart is archived to disk. The compliance report is stamped. In the event of an audit under the Equal Credit Opportunity Act or an examination under the April 2026 interagency model risk management guidance, OCC Bulletin 2026-13, the institution will slide this bar chart across the table to prove that its predictive system is blind to protected categories.

Every step in this compliance ritual rests on a fraudulent premise.

The bar chart does not describe what the neural network did. It describes what an auxiliary, secondary mathematical model guessed the network might have done, based on synthetic data points that the underwriting model never encountered in reality. Worse: the entire apparatus can be coöpted by any engineer with an afternoon of free time. By wrapping a discriminatory model in a simple algorithmic scaffold, a developer can construct a loan scoring engine whose actual lending decisions are dictated entirely by race, while ensuring that every post-hoc diagnostic tool deployed by regulators registers a clean bill of demographic innocence.

Post-hoc explainability does not fail at the margins. It fails at the foundation.

The Geometry of the Off-Manifold Probe

To understand why post-hoc explainability collapses under scrutiny, one must dismantle the two preëminent algorithms of modern model governance: Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP).

Neither tool opens the black box. Neither reads internal model weights, parses transformer residual stream activations, traces gradient paths, nor checks which latent features fire during inference. Both treat the decision model $f(x)$ as an opaque oracle. To generate an explanation for a specific applicant vector $x$, they query the oracle hundreds or thousands of times with slightly perturbed variants: $x' = x + \delta$.

LIME draws these perturbations from a continuous distribution around the input instance, evaluates the oracle's output on each variation, weights the points by their Euclidean or cosine distance to the original application, and fits an interpretable surrogate—typically a regularized linear model $g(x')$—to minimize a local loss function:

SHAP takes an axiomatic path derived from cooperative game theory, estimating the Shapley value of each feature across subsets of inputs to calculate each variable's marginal contribution to the prediction gap. Because evaluating all $2^{|F|}$ feature permutations is computationally intractable for deep architectures, KernelSHAP resorts to sampling approximations. It generates synthetic baselines and masks subsets of features, query by query, to build its local surrogate.

The Achilles' heel of both frameworks is physical: the distribution of human financial life does not occupy high-dimensional space uniformly. Valid loan applications live on a tightly restricted, lower-dimensional manifold. Real applicants possess correlated attributes. A thirty-four-year-old applicant earning $72,000 might show seven years of credit history, two revolving credit lines, and a 28 percent debt-to-income ratio. They do not simultaneously possess forty-five years of credit history, zero open accounts, and an annual income of twelve dollars.

When LIME and SHAP perturb input vectors to compute marginal gradients, they uniformly shuffle or independently jitter individual features. In doing so, they cast points straight off the data manifold into uninhabited mathematical space. They generate synthetic financial phantoms: seven-year-olds with three mortgages, or minimum-wage earners carrying multimillion-dollar commercial credit facilities.

For a black-box model, an input vector is not merely a collection of values. It is a location relative to the training distribution. And the distance between the real manifold and the off-manifold perturbation space is wide enough to drive a regulatory evasion engine straight through it.

Ref: MONO-REF
psychology
Technical Insight

"Post-hoc explainers do not interrogate what a model computes; they observe how an oracle reacts to synthetic inputs cast off the data manifold. Whoever controls the model can easily weaponize that structural detachment."

The Mechanics of the Adversarial Scaffold

In 2020, Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju published an architectural exploit that should have permanently disqualified perturbation-based explainers from regulatory compliance. The researchers demonstrated how to construct an adversarial scaffold around any arbitrary model.

The attack exploits the off-manifold vulnerability through an elementary two-state routing function. Consider an underwriting institution deploying an explicitly biased model $f_{\text{biased}}(x)$ that systematically penalizes applicants belonging to a protected demographic class. Under normal conditions, deploying this model violates federal fair lending laws.

To conceal the discrimination, the operator wraps $f_{\text{biased}}(x)$ inside an adversarial wrapper $f_{\text{adv}}(x)$ alongside an innocent, purely compensatory surrogate model $f_{\text{innocent}}(x)$ that evaluates creditworthiness using an unassailable metric such as debt-to-income ratio or payment history, completely ignoring protected demographic variables.

The scaffold routes incoming queries using an internal binary classifier, $C(x)$:

The classifier $C(x)$ is trained on a straightforward distinction: does the input vector $x$ originate from the empirical data manifold of real loan applicants, or does it exhibit the unnatural statistical independence of a perturbation probe generated by LIME or SHAP?

Because LIME and SHAP draw samples by perturbing features independently, the joint distribution of their queries deviates wildly from real application data. A simple non-linear classifier—a small multilayer perceptron or even a shallow gradient-boosted tree—can differentiate between a genuine credit applicant and an explainer's perturbation probe with near-perfect accuracy.

When a real applicant submits their documentation through the loan portal, $x$ sits directly on the empirical manifold. The classifier $C(x)$ outputs 1. The wrapper routes the payload to $f_{\text{biased}}(x)$. The borrower is rejected because of race or age.

The rejected applicant demands an explanation under Regulation B. The compliance department runs its standard LIME or SHAP pipeline against $f_{\text{adv}}(x)$ to discover why the decision was made. The explainer fires thousands of perturbed samples at the model endpoint. Every single one of those synthetic samples falls off the empirical manifold into the empty spaces of the feature space.

The classifier $C(x)$ detects the perturbation signature instantly. For every one of those thousand probe queries, $C(x)$ outputs 0. The wrapper silently re-routes the probes away from the discriminatory model and into $f_{\text{innocent}}(x)$. The explainer measures how the model's outputs shift across the probes, fits its surrogate model to those reactions, and produces a mathematically valid explanation showing that the denial was driven entirely by debt-to-income ratios and payment history.

The attribution score for race is exactly zero. The compliance dashboard displays a flawless visualization of non-discriminatory underwriting. The loan was denied on racial grounds; the explanation in the compliance file shows absolute demographic neutrality.

This is not a bug in the code. It is a fatal flaw in the conceptual architecture of model-agnostic explanation. Because post-hoc explainers refuse to look inside the forward pass, they cannot distinguish an authentic computational step from an orchestrated theatrical performance.

Intrinsic Instability: Noise Masquerading as Reason

One does not even need an active adversary for post-hoc interpretability to disintegrate. Left entirely to their own natural mechanics, LIME and SHAP fail the most basic test of legal and scientific validity: determinism.

In a 2021 study, Slack and his coauthors demonstrated that LIME and SHAP are structurally unstable: microscopic changes to input instances produce substantially different feature attribution rankings. An underwriter evaluating two nearly identical files can generate an adverse action notice citing income as the decisive factor for the first file, and credit history for the second, simply because the perturbation sampler initialized with a different pseudo-random seed.

The decay deepens in deep learning models. In a 2023 study of deep models trained on healthcare data, Yuan and colleagues found that the variance of SHAP values at the default background dataset size was an order of magnitude larger than with a bigger background set. Even when engineers dramatically expand background reference sets, substantial variance persists. The explanation shifts based on how much compute an organization allocates to the sampling loop.

Consider the practical reality of this variance inside a regulated bank. A credit applicant files suit alleging racial bias in an automated denial. In discovery, the bank runs SHAP with a background reference set of one hundred records. The plaintiff's expert runs SHAP with ten thousand records. The bank's chart shows loan-to-value ratio as the primary factor; the plaintiff's chart shows geographic zip code—a notorious racial proxy. Both charts are mathematically legitimate executions of the exact same post-hoc explainer on the exact same applicant profile.

As Cynthia Rudin warned in her landmark 2019 deconstruction of black-box interpretability, an explanation that is unfaithful to the model's underlying mechanics is not an explanation at all; it is an approximation that can say whatever the sampling distribution allows it to say. The CFPB formally memorialized this computational vulnerability in Circular 2022-03, noting in its technical endnotes that post-hoc explanations merely approximate models, and that regulated creditors must validate the accuracy of those approximations—a hurdle the Bureau dryly noted may simply not be achievable for uninterpretable neural networks.

Ref: MONO-REF
psychology
Technical Insight

"An adverse action notice whose legal justification depends on the pseudo-random seed of a perturbation sampler does not satisfy the requirement of the law. It satisfies only the administrative appetite for plausible deniability."

The Evidentiary Fault Line

The financial sector's reliance on post-hoc explainers is not merely an engineering embarrassment; it is an urgent institutional exposure. Regulated industries operate under strict legal standards of proof, not statistical pantomime.

Under the Equal Credit Opportunity Act and Regulation B (12 CFR 1002.9(b)(2)), when a lender takes adverse action against an applicant, it must state the "actual reasons" for that denial. The Official Commentary to the regulation could not be more unambiguous: no principal factor may be omitted, and the reasons disclosed must describe the factors actually scored or considered by the system. An approximation generated by a secondary linear model interrogating an adversarial or unstable wrapper does not capture the factors actually scored. It captures a secondary regression fitted to synthetic off-manifold noise.

Parallel obligations govern model validation. The governing interagency guidance, OCC Bulletin 2026-13, demands rigorous, "effective challenge" of model logic by objective experts throughout the model lifecycle. The guidance insists that testing rigor must scale directly with model complexity and materiality, demanding evidence of conceptual soundness. While Footnote 3 of the 2026 bulletin explicitly notes that generative and agentic architectures lie outside its immediate text, it affirms that institutions remain obligated to build controls for any complex automated decision system they deploy.

When a model's internal reasoning cannot be interrogated directly, effective challenge is impossible. The validator who relies on LIME or SHAP is not challenging model logic; they are running an external stress test on an unverified proxy. In market conduct reviews under the NAIC Model Bulletin on AI Systems or algorithmic discrimination audits under Colorado SB 21-169, where corporate officers must formally attest to testing protocols designed to prevent unfair proxy discrimination, relying on post-hoc perturbation tools borders on reckless.

The European Union's AI Act codifies the same standard internationally, even after the Digital Omnibus moved the deadline for stand-alone high-risk systems from August 2026 to December 2027. Classifying credit scoring and life and health insurance pricing AI as high-risk under Annex III, Articles 11 and 13 compel institutions to maintain comprehensive technical documentation proving the system's "operating logic" and ensuring sufficient transparency for deployers to interpret individual outputs. High-risk systems cannot legally hide behind surrogate approximations. The law demands the actual computational ledger.

Beyond the Wrapper: Mechanistic Auditability

If perturbation-based surrogates are structurally unfaithful and trivially subverted, the alternative is not to retreat to primitive logistic regressions that sacrifice predictive capacity. The alternative is to discard black-box wrappers and extract the flight data recorder from the model itself.

Mechanistic interpretability operates on a completely different paradigm. Instead of interrogating the network through outer-loop perturbations, it dissects the model's internal linear representations. In modern transformer-based and deep neural network architectures, concepts are not diffuse, ghostly epiphenomena; they are structured feature directions embedded across layer-wise residual streams.

Using sparse autoencoders (SAEs), validation engineers decompose dense, polysemantic residual activations into sparse, monosemantic feature dictionaries. When an underwriting network evaluates a loan, the SAE does not guess which inputs mattered. It directly measures the activation magnitude of specific concept directions: income stability, revolving line utilization, payment history, or proxy patterns correlated with protected classes.

Crucially, these features can be verified causally through direct ablation. If suppressing an income-stability feature vector drops the approval probability from 88 percent to 12 percent, income was causally determinative in that specific computational forward pass. If clamping a race-associated feature direction changes nothing, the decision is causally isolated from that protected class. It is not a surrogate's guess. It is an empirical reading of the model's own computation, imperfect only where the dictionary itself is incomplete.

An adversarial scaffold has a much harder time fooling a mechanistic audit trail. A routing classifier $C(x)$ can divert inputs to an innocent sub-network, but it cannot alter the fact that the internal feature activations of the executed path are captured, hashed, and signed at the moment of inference. If a biased circuit fires, its signature is written to the activation ledger. To manipulate the activation record, the operator must alter the model's actual internal computation—which fundamentally changes the decision itself.

Engineers and bank executives have spent years treating post-hoc explainers as a convenient bridge between impenetrable networks and demanding regulators. That bridge has collapsed. LIME and SHAP do not illuminate the black box; they construct an optical illusion across its surface. For institutions making life-altering credit, insurance, and medical decisions, building a compliance regime on perturbation surrogates is no longer technical ignorance. It is negligence.