Make Clear What the System Can Do
Before anyone relies on the feature, they should be able to tell what it is for. State the capability in plain words at the point of use, not in a help page nobody opens.
A checklist for the parts of AI that face the user.
Eighteen guidelines for designing AI-infused products, grouped by when they bite: what you promise up front, how you behave during use, what happens when the system is wrong, and how it changes over time. The fourth group is the one most teams skip.
The set was published at CHI 2019 by thirteen researchers and designers: Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. Twelve of them worked at Microsoft in Redmond; the thirteenth came from the Paul G. Allen School at the University of Washington. It was not invented from scratch: the team gathered more than 150 recommendations from two decades of research and industry writing, collapsed the duplicates, and put what survived through four rounds of testing.
The last round is the part worth knowing. Forty-nine design practitioners applied the candidate guidelines to twenty AI products they used daily, judging where each one was honoured or broken. That is how eighteen survived, and why each is written as something you can look for in a screen rather than a principle you nod along to.
AI-infused products fail at the seams rather than in the model. The guidelines name those seams and order them by when they hurt: the promise you make before use, the manners you keep during it, the recovery you offer when you are wrong, and the way you change underneath someone over time.
Before anyone relies on the feature, they should be able to tell what it is for. State the capability in plain words at the point of use, not in a help page nobody opens.
Say how often the system is wrong. This is the guideline products skip most often, because admitting a hit rate feels like admitting a flaw rather than setting an expectation.
Act or interrupt according to the task in front of the person. A correct suggestion at the wrong moment is an interruption, and it is judged as one.
Display what belongs to the task at hand and nothing else. Relevance is a filter on what you show, not a ranking applied after you have shown everything.
Deliver the experience the way people expect in their social and cultural context. Tone, formality, and directness are read instantly and are not recoverable by an apology later.
The system’s language and behavior should not reinforce unfair stereotypes. Autocomplete, ranking, and defaults all carry assumptions from their training data into the interface.
Make the system easy to call when it is wanted. A capability that only appears when it decides to appear cannot be relied on for anything.
Make it easy to ignore or wave away. This is the one commercial pressure argues with, because a feature nobody can dismiss looks better in a usage chart.
Let people edit, refine, or recover when the system is wrong. Being wrong is survivable; being wrong with no way back is what turns a mistake into a complaint.
When the system does not know what was meant, it should ask or do less rather than guess confidently. Three good options beat one wrong answer delivered without a flicker of doubt.
Let people reach an explanation of the behavior. Not a model card: the reason for this route, this result, this ordering, in the place where they are looking at it.
Hold short-term memory so people can refer back instead of repeating themselves. Without it every exchange starts from nothing and the burden lands on the person.
Personalize from what people actually do rather than what they said once in a settings screen. Behavior is the more honest signal, and it keeps being given for free.
Limit disruptive change. A system that improves overnight has also moved everything the person learned, and the cost of relearning is charged to them, not to you.
Let people correct the system in small, specific ways during normal use. A thumbs-down on a whole session tells you far less than a correction on one item.
Show what a correction will change. Feedback that vanishes into the system teaches people their input does not matter, and they stop giving it.
Offer one place to customize what the system watches and how it behaves. Per-item controls are not a substitute for being able to turn the whole thing down.
Say when capabilities are added or changed. People build habits around what a system could do last month, and silent updates break those habits without explanation.
Written for recommenders and autocomplete in 2019, then handed a generation of systems that talk back. Most still hold; a few now land differently.
Saying how well the system performs was awkward for a recommender and is harder for a model that writes fluent prose at any confidence. The second guideline is now the least well served of the eighteen.
Efficient dismissal was always the guideline commercial pressure argues with. An assistant placed in every surface makes ignoring it the main interaction, and the cost of getting that wrong compounds.