Righteous AI Gallery
When Independent Research Exposed Hidden Bias in AI
Year: 2018
Location: United States
Principal Researchers: Joy Buolamwini and Timnit Gebru
Research Institution: MIT Media Lab / Microsoft Research
Historical Theme: AI Fairness · Accountability · Independent Testing · Equality · Human Responsibility
Historical Significance
In 2018, researchers Joy Buolamwini and Timnit Gebru published Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, a landmark study examining whether commercial facial-analysis systems performed equally across different skin types and genders. (Proceedings of Machine Learning Research)
The researchers developed a more demographically balanced facial-image benchmark and evaluated three commercial gender-classification systems.
The results revealed substantial differences in performance.
The study found that darker-skinned women were the most frequently misclassified group, with error rates reaching 34.7%, while the maximum error rate for lighter-skinned men was 0.8%. (Proceedings of Machine Learning Research)
The research transformed an abstract concern about algorithmic bias into a measurable, reproducible, and publicly documented finding.

Figure 1 — Discovering the Coded Gaze
Title: When the Machine Could Not See Everyone Equally
Description: A museum-quality historical reconstruction representing the experience that motivated Buolamwini’s investigation: a facial-analysis system interacting differently with faces of different skin tones, leading a researcher to question whether the technology was truly neutral.
The Research Question
The central question was simple but consequential:
Does an AI system perform equally well across different groups of people?
Existing evaluations could report a single overall accuracy number.
But a high overall accuracy could conceal substantial differences between groups.
Buolamwini and Gebru therefore examined performance at the intersection of gender and skin type, rather than treating each category separately. Their research introduced a more inclusive benchmark and evaluated four intersectional groups: darker females, darker males, lighter females, and lighter males. (MIT Media Lab)
This approach changed the question from:
“How accurate is the AI?”
to:
“How accurate is the AI for different people?”
Title: From Overall Accuracy to Intersectional Evaluation
Description: The table shows how the research examined AI performance across intersecting demographic and phenotypic groups rather than relying only on aggregate accuracy.
Table 1 — The Gender Shades Research Design
| Research Element | Description |
|---|---|
| AI Technology | Commercial facial-analysis gender-classification systems |
| Researchers | Joy Buolamwini and Timnit Gebru |
| Systems Evaluated | Three commercial gender-classification systems |
| Faces Evaluated | 1,270 unique individuals |
| Skin-Type Method | Fitzpatrick Skin Type classification |
| Gender Groups | Female and male labels used by the evaluated systems |
| Intersectional Groups | Darker females, darker males, lighter females, lighter males |
| Purpose | Determine whether system performance differed substantially across groups |
| Core Innovation | Intersectional evaluation of gender and skin type |
The study’s benchmark included 1,270 individuals and was designed to provide more balanced representation than the datasets examined by the researchers. (MIT Media Lab)
The Evidence
The researchers examined two existing facial-analysis benchmarks and found substantial representation imbalances.
The IJB-A and Adience datasets were overwhelmingly composed of lighter-skinned subjects: 79.6% in IJB-A and 86.2% in Adience. (Proceedings of Machine Learning Research)
The researchers then developed the Pilot Parliaments Benchmark (PPB), containing 1,270 individuals selected from three African and three European countries to create greater balance in gender and skin type representation. (MIT Media Lab)
The results demonstrated that the systems did not perform equally across the tested groups.
Title: When Aggregate Accuracy Conceals Inequality
Description: A comparison of the reported error-rate extremes identified by the Gender Shades research.
Table 2 — The Performance Gap
| Group | Reported Maximum Error Rate |
|---|---|
| Darker-Skinned Females | 34.7% |
| Lighter-Skinned Males | 0.8% |
| Difference | 33.9 percentage points |
The paper reported darker-skinned females as the most misclassified group and lighter-skinned males as the least misclassified group among the evaluated categories. (Proceedings of Machine Learning Research)
The significance of the finding was not simply that one AI system made mistakes.
It demonstrated that the probability of error could differ dramatically depending on who was being evaluated.

Figure 2 — The Evidence of Bias
Title: Four Groups, Four Different Experiences
Description: A museum visualization showing four representative demographic groups being evaluated by the same facial-analysis AI system, with visibly different error-rate indicators. The image emphasizes that a single AI system can produce unequal performance across groups.
The Righteous Choice
The importance of Gender Shades lies not only in discovering a technical problem.
It lies in the decision to measure what others might overlook.
The research demonstrated several principles central to responsible AI:
Truth
A system’s performance should be measured according to evidence rather than assumptions about technological neutrality.
Fairness
AI systems should not be evaluated only according to the groups for whom they perform best.
Accountability
Independent evaluation can reveal problems that may remain hidden when system performance is reported only in aggregate.
Inclusion
People who are underrepresented in datasets should not become invisible in the evaluation of AI systems.
Courageous Inquiry
Researchers can challenge widely accepted assumptions by testing them systematically.
Title: Why Gender Shades Is Considered for the Righteous AI Gallery
Description: The Museum’s analytical framework evaluates the historical choices and principles demonstrated by the Gender Shades research.
Table 3 — Righteousness Test
| Criterion | Historical Assessment |
|---|---|
| Conflict | Technological claims of general performance vs. evidence of unequal subgroup performance |
| Question | Does AI work equally well for different people? |
| Principle | Truth, fairness, inclusion, and accountability |
| Action | Independent empirical testing |
| Innovation | Intersectional evaluation of gender and skin type |
| Evidence | Quantitative performance differences across demographic groups |
| Public Value | Exposed a problem relevant to responsible AI development |
| Historical Impact | Became an influential landmark in discussions of algorithmic bias and AI fairness |
| Model for Humanity | Demonstrated how independent investigation can expose hidden problems and demand greater accountability |
From Personal Experience to Public Evidence
The story of Gender Shades began with a personal technological experience.
Buolamwini had encountered facial-analysis systems that failed to detect her face or classified it incorrectly. Rather than treating the experience as an isolated technical failure, she asked whether similar problems affected other people. (MIT Media Lab)
That question led to systematic research.
The progression was:
Personal Experience
↓
Research Question
↓
Independent Testing
↓
Better Benchmark
↓
Quantitative Evidence
↓
Public Accountability
This transformation is one of the most important aspects of the case.
A personal experience became a reproducible scientific investigation.
Title: How a Question Became a Movement for Accountability
Description: This table traces the transformation of a personal technological experience into a broader contribution to responsible AI research.
Table 4 — From Individual Experience to Historical Impact
| Stage | Development |
|---|---|
| 01 — Experience | Facial-analysis technology produced inconsistent results for Buolamwini. |
| 02 — Question | She questioned whether the problem was specific to her or reflected a broader pattern. |
| 03 — Investigation | Researchers systematically evaluated commercial systems. |
| 04 — Benchmark | A more balanced dataset was developed to test performance across groups. |
| 05 — Measurement | Accuracy was examined across intersecting gender and skin-type categories. |
| 06 — Evidence | Significant disparities in error rates were documented. |
| 07 — Accountability | The findings challenged developers and researchers to evaluate AI more inclusively. |
| 08 — Legacy | The work became an important reference point in the development of algorithmic fairness and responsible AI. |

Figure 3 — From Research to Accountability
Title: Making Hidden Bias Visible
Description: A museum installation representing the transformation of research evidence into public accountability: a researcher, scientific charts, facial-analysis results, and policymakers or technology professionals examining documented disparities.
Historical Impact
Gender Shades helped change how algorithmic fairness could be investigated.
The research showed that it was insufficient to ask only whether an AI system was accurate overall.
Researchers increasingly needed to ask:
Who is represented in the data?
Who is missing?
Who experiences the highest error rate?
What happens when demographic categories intersect?
The project also emphasized the importance of transparent subgroup performance reporting. MIT’s Gender Shades project describes the work as an approach to inclusive product testing for AI and emphasizes the need for intersectional evaluation. (MIT Media Lab)
Title: From Algorithmic Bias to Responsible AI Evaluation
Description: Major ideas demonstrated or advanced by the Gender Shades research and their significance for responsible AI.
Table 5 — The Historical Legacy of Gender Shades
| Area | Contribution |
|---|---|
| AI Fairness | Demonstrated measurable disparities in facial-analysis performance. |
| Algorithmic Auditing | Demonstrated the value of independent empirical testing. |
| Dataset Diversity | Highlighted the consequences of imbalanced benchmark datasets. |
| Intersectionality | Demonstrated why intersecting characteristics can reveal disparities hidden by aggregate statistics. |
| Transparency | Encouraged more detailed reporting of subgroup performance. |
| Accountability | Created evidence that technology developers could use to investigate and improve systems. |
| Responsible Innovation | Showed that identifying a technological weakness can be a form of constructive innovation. |
The Righteous Innovation
The significance of Gender Shades is not that it created a new AI product.
Its innovation was methodological and ethical.
The researchers changed the way an important class of AI systems could be evaluated.
Instead of accepting:
“The system works.”
the research demanded a more rigorous question:
“For whom does the system work, and for whom does it fail?”
This is a powerful example of righteous innovation.
Innovation does not always mean building something new.
Sometimes innovation means creating a better way to test, question, measure, and correct what already exists.

Figure 4 — Righteous Innovation
Title: A Better Way to Measure AI
Description: A conceptual museum image showing a transition from a conventional AI evaluation system using one overall accuracy score to an intersectional evaluation system examining multiple demographic groups separately.
The Human–AI Lesson
Gender Shades demonstrates that AI systems do not exist independently of human decisions.
Human beings decide:
- What data to collect
- Which people to represent
- Which benchmarks to use
- Which metrics to report
- Which errors to investigate
- Which problems to correct
- Which systems are ready for deployment
An AI system can therefore reproduce limitations embedded in the data, design, evaluation methods, and assumptions of its creators.
The lesson is not that AI is inherently unrighteous.
The lesson is that human responsibility remains essential throughout the AI lifecycle.
Title: Where Righteousness Enters AI Development
Description: The table identifies points at which human decisions can influence fairness, accuracy, transparency, and accountability in AI systems.
Table 6 — The Human–AI Responsibility Chain
| Stage | Human Decision | Righteousness Question |
|---|---|---|
| Data | Which people are represented? | Is representation sufficiently inclusive? |
| Design | What problem is being solved? | Is the purpose legitimate and beneficial? |
| Training | What data and methods are used? | Could systematic bias be introduced? |
| Testing | Who is evaluated? | Are vulnerable or underrepresented groups included? |
| Measurement | Which metrics are reported? | Can aggregate results conceal important disparities? |
| Deployment | Where is the system used? | Are the risks appropriate for the application? |
| Monitoring | What happens after deployment? | Are failures detected and corrected? |
| Accountability | Who answers for harm? | Can responsibility be identified and exercised? |
Museum Assessment
Righteous AI Gallery Classification
Primary Classification
Human–AI Righteousness
Historical Character
Pioneering Independent AI Audit
Core Principle
Truth Through Evidence
Key Virtues
Fairness · Accountability · Courageous Inquiry · Inclusion
Righteous Action
Seeing a possible injustice → questioning the assumption → measuring the evidence → making the problem visible → enabling correction
Historical Significance
Gender Shades represents a particularly important form of Human–AI righteousness:
using technical expertise to reveal a hidden problem rather than allowing a convenient assumption to remain unchallenged.
Why This Event Belongs in the Righteousness Museum
The historical importance of Gender Shades is not simply that an AI system made mistakes.
AI systems can make mistakes.
The deeper significance is that researchers looked for unequal outcomes, developed a method to measure them, documented the evidence, and brought the problem into public view.
The event therefore represents a model of responsible technological leadership:
Do not assume that technology is neutral.
Do not measure only what is convenient.
Do not hide unequal outcomes behind averages.
Test the system.
Reveal the evidence.
Improve the technology.
The Enduring Question
When an AI system appears to work well, who has the responsibility to ask whether it works equally well for everyone?
Gender Shades demonstrates that progress in artificial intelligence does not come only from making machines more capable.
It also comes from making the measurement of those machines more truthful.
Question the assumption.
Measure the evidence.
Reveal the disparity.
Demand accountability.
Build better AI.
References
Buolamwini, J., & Gebru, T. (2018). Gender Shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 77–91. https://proceedings.mlr.press/v81/buolamwini18a.html (Proceedings of Machine Learning Research)
Buolamwini, J. (2017). Gender Shades: Intersectional phenotypic and demographic evaluation of face datasets and gender classifiers [Master’s thesis, Massachusetts Institute of Technology]. MIT Media Lab. https://www.media.mit.edu/publications/full-gender-shades-thesis-17/ (MIT Media Lab)
MIT Media Lab. (2018). Gender Shades. https://www.media.mit.edu/projects/gender-shades/ (MIT Media Lab)
MIT Media Lab. (2018). Gender Shades: Publications. https://www.media.mit.edu/projects/gender-shades/publications/ (MIT Media Lab)
MIT Media Lab. (2018). Gender Shades: Frequently asked questions. https://www.media.mit.edu/projects/gender-shades/faq/ (MIT Media Lab)
Gender Shades. (2018). Gender Shades: How well do IBM, Microsoft, and Face++ AI services guess the gender of a face? https://gendershades.org/ (Gender Shades)
Curatorial Note
Exhibit is preserved as a historical study of Human–AI righteousness through independent inquiry, evidence, fairness, and accountability.
The exhibit does not claim that the researchers proved that all facial-analysis AI is inherently harmful or that every error represents intentional discrimination.
Instead, it preserves a more precise historical lesson:
When technology affects people, responsible innovation requires the courage to measure its performance honestly—including where it fails and whom those failures affect.
