Microsoft’s 18 AI Guidelines

A checklist for the parts of AI that face the user.

Eighteen guidelines for designing AI-infused products, grouped by when they bite: what you promise up front, how you behave during use, what happens when the system is wrong, and how it changes over time. The fourth group is the one most teams skip.

Origin

The set was published at CHI 2019 by thirteen researchers and designers: Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. Twelve of them worked at Microsoft in Redmond; the thirteenth came from the Paul G. Allen School at the University of Washington. It was not invented from scratch: the team gathered more than 150 recommendations from two decades of research and industry writing, collapsed the duplicates, and put what survived through four rounds of testing.

The last round is the part worth knowing. Forty-nine design practitioners applied the candidate guidelines to twenty AI products they used daily, judging where each one was honoured or broken. That is how eighteen survived, and why each is written as something you can look for in a screen rather than a principle you nod along to.

The Guidelines

AI-infused products fail at the seams rather than in the model. The guidelines name those seams and order them by when they hurt: the promise you make before use, the manners you keep during it, the recovery you offer when you are wrong, and the way you change underneath someone over time.

Published by
Saleema Amershi and colleagues (2019)
Source
Guidelines for Human-AI Interaction, CHI 2019
In practice
Audit a feature against all four phases
When to Use
How to Use

Initially

G1–G2 What the product promises before anyone relies on it.
Initially 01 / 20

Make Clear What the System Can Do

Before anyone relies on the feature, they should be able to tell what it is for. State the capability in plain words at the point of use, not in a help page nobody opens.

Best for
Onboarding and empty states
Use when
The feature is new to them
Avoid when
The capability is self-evident
Initially 02 / 20

Make Clear How Well It Does It

Say how often the system is wrong. This is the guideline products skip most often, because admitting a hit rate feels like admitting a flaw rather than setting an expectation.

Best for
Predictions and suggestions
Use when
Errors are visible to the user
Avoid when
The output is verifiable at a glance

During interaction

G3–G6 How it behaves while someone is in the middle of a task.
During interaction 03 / 20

Time Services Based on Context

Act or interrupt according to the task in front of the person. A correct suggestion at the wrong moment is an interruption, and it is judged as one.

Best for
Proactive assistance
Use when
The system can act unprompted
Avoid when
The user just asked for it
During interaction 04 / 20

Show Contextually Relevant Information

Display what belongs to the task at hand and nothing else. Relevance is a filter on what you show, not a ranking applied after you have shown everything.

Best for
Results and side panels
Use when
Context is known
Avoid when
The user is browsing deliberately
During interaction 05 / 20

Match Relevant Social Norms

Deliver the experience the way people expect in their social and cultural context. Tone, formality, and directness are read instantly and are not recoverable by an apology later.

Best for
Anything that speaks or writes
Use when
Output reaches people directly
Avoid when
Output is numeric only
During interaction 06 / 20

Mitigate Social Biases

The system’s language and behavior should not reinforce unfair stereotypes. Autocomplete, ranking, and defaults all carry assumptions from their training data into the interface.

Best for
Generated text and ranking
Use when
Output describes people
Avoid when
The domain has no human referents

When wrong

G7–G11 What it offers once it has already made a mistake.
When wrong 07 / 20

Support Efficient Invocation

Make the system easy to call when it is wanted. A capability that only appears when it decides to appear cannot be relied on for anything.

Best for
Assistants and helpers
Use when
Help is occasionally needed
Avoid when
The system is always active
When wrong 08 / 20

Support Efficient Dismissal

Make it easy to ignore or wave away. This is the one commercial pressure argues with, because a feature nobody can dismiss looks better in a usage chart.

Best for
Any volunteered suggestion
Use when
The system offers unprompted
Avoid when
Dismissal would lose real work
When wrong 09 / 20

Support Efficient Correction

Let people edit, refine, or recover when the system is wrong. Being wrong is survivable; being wrong with no way back is what turns a mistake into a complaint.

Best for
Anything generated or inferred
Use when
Errors are likely
Avoid when
The action was already reversible
When wrong 10 / 20

Scope Services When in Doubt

When the system does not know what was meant, it should ask or do less rather than guess confidently. Three good options beat one wrong answer delivered without a flicker of doubt.

Best for
Ambiguous input
Use when
Confidence is measurable
Avoid when
Guessing is free to undo
When wrong 11 / 20

Make Clear Why It Did What It Did

Let people reach an explanation of the behavior. Not a model card: the reason for this route, this result, this ordering, in the place where they are looking at it.

Best for
Ranked or filtered output
Use when
The order is not obvious
Avoid when
The reason is the input itself

Over time

G12–G18 How it changes underneath a person across months of use.
Over time 12 / 20

Remember Recent Interactions

Hold short-term memory so people can refer back instead of repeating themselves. Without it every exchange starts from nothing and the burden lands on the person.

Best for
Conversational interfaces
Use when
Tasks span several turns
Avoid when
Each session must start clean
Over time 13 / 20

Learn From User Behavior

Personalize from what people actually do rather than what they said once in a settings screen. Behavior is the more honest signal, and it keeps being given for free.

Best for
Recommenders and feeds
Use when
Use is repeated
Avoid when
There is not enough history yet
Over time 14 / 20

Update and Adapt Cautiously

Limit disruptive change. A system that improves overnight has also moved everything the person learned, and the cost of relearning is charged to them, not to you.

Best for
Shipped, settled products
Use when
Behavior changes with use
Avoid when
The current behavior is harmful
Over time 15 / 20

Encourage Granular Feedback

Let people correct the system in small, specific ways during normal use. A thumbs-down on a whole session tells you far less than a correction on one item.

Best for
Personalized systems
Use when
The model learns from use
Avoid when
Feedback changes nothing
Over time 16 / 20

Convey the Consequences of Actions

Show what a correction will change. Feedback that vanishes into the system teaches people their input does not matter, and they stop giving it.

Best for
Feedback and preference controls
Use when
Input shapes future behavior
Avoid when
The effect is immediate and visible
Over time 17 / 20

Provide Global Controls

Offer one place to customize what the system watches and how it behaves. Per-item controls are not a substitute for being able to turn the whole thing down.

Best for
Monitoring and personalization
Use when
The system observes behavior
Avoid when
Nothing is retained
Over time 18 / 20

Notify Users About Changes

Say when capabilities are added or changed. People build habits around what a system could do last month, and silent updates break those habits without explanation.

Best for
Continuously shipped products
Use when
Capabilities move
Avoid when
The change is invisible in use
✦

The Guidelines in the Age of AI

Written for recommenders and autocomplete in 2019, then handed a generation of systems that talk back. Most still hold; a few now land differently.

✦ AI Era 19 / 20

Confidence Is Now the Hard One

Saying how well the system performs was awkward for a recommender and is harder for a model that writes fluent prose at any confidence. The second guideline is now the least well served of the eighteen.

Shift
Accuracy → calibration
Use when
Output reads as authoritative
Watch for
A hedge nobody can act on
✦ AI Era 20 / 20

Dismissal Under Pressure

Efficient dismissal was always the guideline commercial pressure argues with. An assistant placed in every surface makes ignoring it the main interaction, and the cost of getting that wrong compounds.

Shift
Opt-out → unavoidable
Use when
Assistance is ambient
Watch for
Dismissal buried in settings
Further Reading