Summary
- Prompt tests should cover comparable users, contexts and consequences rather than isolated wording.
- Material gaps need owners, release criteria, appeal routes and evidence that remediation worked.
Sexism in a conversational system can enter through training material, labels, product rules and the way users are interpreted. A polished answer may still recommend different roles, assign different credibility or fail more often for one group. Teams should define consequential tasks, test matched scenarios across languages and identities, and examine downstream decisions—not just toxicity scores. The next useful evidence is a repeatable audit before and after a model change, with affected people able to challenge the result. Fairness is an operating discipline, not a personality claimed by the interface.


