Open the Black Box
A black-box model gives an answer without a reason; a glass-box model shows its work. Explanations should say what the system did, what it’s doing, what comes next, and what information it’s using.
Show people why, not just what.
Explainable AI gives people a view into why a system produced a result, so they can judge it, challenge it, and decide how far to rely on it. The goal isn’t more trust, but the right amount.
Explanation in AI is older than the black box. In the early 1970s, Edward Shortliffe’s MYCIN, a Stanford research system for diagnosing blood infections, could show which of its hand-written rules led to a diagnosis. As opaque models such as deep neural networks took over, even their designers often couldn’t say why a model produced a particular result.
Explainable AI (XAI) grew into its own field in the 2010s, with methods like LIME and SHAP, DARPA’s XAI program, and the right to explanation in the EU’s GDPR. Cynthia Rudin argued that high-stakes decisions should use models that are interpretable by design, and research by Gagan Bansal, Besmira Nushi, Dan Weld and colleagues found that explanations can make people accept an AI’s wrong answers as readily as its right ones.
Explainability is how well people can understand why an AI system produced a given result. Interpretability is about how a model works overall; explainability is about this result, for this person. Good explanations let people check a decision, contest it, and calibrate how much they rely on the system.
A black-box model gives an answer without a reason; a glass-box model shows its work. Explanations should say what the system did, what it’s doing, what comes next, and what information it’s using.
People rarely want to know how a whole model works; they want to know why it said this, about them. Local methods such as LIME and SHAP show which inputs pushed a single prediction up or down.
People usually ask why this result rather than the one they expected. Tim Miller’s review of social science research found that good explanations are contrastive and pick out a few causes instead of listing every factor.
The same explanation can be too shallow for an expert and too technical for a beginner. Lead with a short reason in plain words, and let people open the details they need.
In a 2019 study of an admissions algorithm, interactive explanations helped people understand its decisions but didn’t make them trust it more. People still saw it as too rigid next to human decision-makers who can weigh exceptions and appeals.
A 2021 study found that explanations raised how often people accepted an AI’s recommendation, whether it was right or wrong. The aim is calibrated trust: enough to use the system, not so much that people stop checking.
Researchers found an image classifier that recognized horses by a copyright tag in the corner of the photos, not by the horse. Explanations help teams spot shortcuts like this before users meet them.
Cynthia Rudin argues that for high-stakes decisions, models that are interpretable by design beat black boxes with explanations added afterward, because a separate explanation can’t be fully faithful to the model it describes.
Language models write fluent explanations on request. That makes explanations cheap to produce and easy to believe, whether or not they reflect how an answer was made.
A model’s step-by-step reasoning reads like an explanation, but it isn’t guaranteed to reflect how the answer was actually produced. Don’t present it as proof; pair it with evidence people can check.
Point each claim to the source it came from, so people can verify it themselves. A citation lets people check an answer; a confident tone only asks them to believe it.