The Model They Can Argue With

Explainability is not what analysis gives up in exchange for power. In applied work it is how a result stays valid once it leaves the analyst’s screen.

Rob from InsightPal

A regulated product has to demonstrate compliance with a numerical standard. The testing is expensive, the measurement is noisy, and both sides of the process — the regulator and the manufacturer — have a stake in where the line gets drawn.

The obvious statistical move is a hypothesis test. Take the certification data, test whether the mean exceeds the standard, report a p-value. Defensible, publishable, and completely useless for the decision in front of the room.

What the parties actually need is a guard band: a margin between the measured result and the regulatory limit that accounts for test-to-test variability. Set it too tight and compliant products fail certification. Set it too loose and non-compliant products pass. Nobody in that room needs to know whether a difference is statistically significant. They need to know how often each of them will be wrong.

So the useful deliverable is not a test. It is a variance-components analysis that separates test-to-test variability from product-to-product variability, expressed as a pair of risk curves — one showing the probability that a compliant product fails, one showing the probability that a non-compliant product passes, both plotted against candidate guard-band widths.

That deliverable does something a p-value cannot. It lets each party locate their own exposure on a chart and argue about where the line should sit. The analysis stops announcing an answer and starts making the trade-off visible.

Nobody in that room needs to know whether a difference is statistically significant. They need to know how often each of them will be wrong.

Explainability is not a concession

There is a persistent idea in technical work that explainability is what you give up power for. That the simpler model, the cruder method, the less efficient estimator is chosen because the audience cannot handle the real thing.

That gets it backwards, at least in applied work.

The people closest to the process know it far better than the analyst does. They know that the third shift runs differently, that one of the gauges drifts, that the batch from a particular supplier behaves oddly. That knowledge does not appear in the dataset. It appears in the room, and only if the analysis is presented in a form they can interrogate.

A model the decision-maker cannot argue with is a model whose assumptions never get checked by the one person qualified to check them. That is not a communication problem. That is a validity problem.

What this looks like in practice

On an investigation into declining first-pass yield, the natural approach is to test every candidate driver of the decline and rank the results by significance. With enough observations, that ranking is close to meaningless. Sample size alone will push variables to the top of the list that moved the yield by a fraction of a percent, while a genuinely important effect measured on fewer lots sits further down.

A more useful screening framework combines two things: statistical significance, and the magnitude of the effect in the units the team actually manages. The output is a ranked list of candidate drivers with both quantities visible side by side. Engineers can look at it and say, plausibly, “that one is real, but we cannot do anything about it,” or “that one is smaller than you think because of how we sample.” Both of those conversations improve the analysis.

On a medical-device validation program, the useful deliverable is not a table of regression coefficients. It is an application that presents covariate-adjusted normative ranges — the clinician selects the patient characteristics and sees the expected range. The model underneath is the same either way. The difference is whether the person making the decision can see what the model implies for the case in front of them.

A test that holds

When a method is being chosen, one of the questions that matters is: if this analysis is wrong, will anyone notice?

An approach that produces a single number, from a procedure the decision-maker cannot follow, fails that test. If the assumptions are violated, the number will still look like a number. Nobody will catch it.

An approach that exposes the trade-off, or shows the effect in the team’s own units, or lets someone vary an input and watch the answer move, will get challenged. Sometimes the challenge is wrong and can be explained. Sometimes the challenge is right and the analysis improves. Either way the error rate drops.

There is a real cost to this. Explainable formulations are sometimes less efficient. Power, precision, or elegance is occasionally given up. In a paper that is a genuine loss. In a decision that must survive contact with a regulator, an auditor, and an operations team, an answer nobody can check is worth less than a slightly noisier answer that three informed people have pushed on and failed to break.

The point of an applied analysis is not to be correct in private. It is to be correct in a way that holds up when someone who knows the process pushes back.

That is the standard a product for applied analysis has to meet: outputs a reviewer can interrogate, assumptions that stay visible, and results expressed in the units a decision actually uses. InsightPal is built around that standard.