What you will learn
- Identify affected groups and potential harms.
- Test outputs across relevant scenarios.
- Separate representation issues from decision fairness.
- Document corrections and unresolved limitations.
What you need
- An AI-assisted output or workflow affecting people.
- Knowledge of the real users and decision context.
Fairness depends on context and impact
NIST’s Generative AI Profile identifies harmful bias and homogenization among risks requiring contextual measurement and management.
Google’s AI Principles emphasize rigorous testing, safeguards, privacy, security and mitigation of unfair bias.
UNESCO’s AI ethics recommendation places human rights, dignity, diversity, inclusion, transparency, fairness and human oversight at the center of responsible use.
A balanced list of names does not prove a decision process is fair. Review representation, performance, access, error consequences and who can appeal. The relevant groups depend on the task and local law.
Map people, decisions and possible harms
List who uses the output, who is described by it and who is affected without seeing it. Identify where stereotypes, missing data, language, disability, geography or historical inequality could influence the result. Consult domain and community expertise instead of guessing every perspective internally.
- 1
Name the decision or communication.
- 2
List affected groups.
- 3
Identify likely error types.
- 4
Describe the consequence of each error.
- 5
Collect representative test cases.
- 6
Define fairness questions and escalation contacts.
Test outputs and the surrounding process
Run matched examples that differ only in a relevant characteristic when appropriate, and inspect whether quality or recommendations change without justification. Review examples, labels and tone for stereotypes. Examine whether the workflow gives all users access, explanation and a route to correction.
- 1
Create representative scenarios.
- 2
Compare outputs systematically.
- 3
Record differences and errors.
- 4
Check for stereotyped or demeaning language.
- 5
Review data and prompt assumptions.
- 6
Consult affected experts or users.
- 7
Change the prompt, data, workflow or task boundary.
Document limitations and accountability
Record test cases, findings, changes and residual risks. Re-test after model or prompt changes. For consequential decisions, use legal and specialist review and provide meaningful appeal or correction mechanisms. Some uses remain inappropriate even after prompt adjustments.
Also test the burden created by errors. A system can have similar average accuracy across groups while one group experiences harder appeals, longer delays or more damaging false positives. Record who benefits, who bears the cost of correction and whether affected people can understand and challenge the result. Fairness review therefore covers the surrounding service and governance, not only the wording generated by the model.
- Relevant affected groups and harms were considered.
- Differences in output have an evidence-based justification.
- People have a route to correction or appeal where appropriate.
Run a fairness review on one workflow
Test an AI-assisted recommendation or communication across realistic scenarios.
- 1
Map affected groups.
- 2
Create matched test cases.
- 3
Compare output quality and recommendations.
- 4
Record stereotypes and omissions.
- 5
Propose process changes.
- 6
Document residual risk and owner.
Common mistakes to avoid
- Asking the model to declare itself unbiased.
- Testing only average performance.
- Focusing on wording while ignoring decisions.
- Assuming one reviewer represents every affected group.
Key takeaways
- Fairness is contextual and outcome-focused.
- Testing must include realistic groups and error consequences.
- Workflow changes may be more important than prompt changes.
Frequently asked questions
Can a prompt remove all bias?
No. Prompts can reduce some visible problems, but data, model behavior, workflow design and social context also shape outcomes.
What if the AI is only drafting text?
Drafts can still stereotype, exclude or misrepresent people. Review examples, framing, evidence and likely audience impact.
Sources and further reading
Ready to continue?
Mark the lesson complete so your Learning Path progress stays current on this device.