Gender Shades

Righteous AI Gallery

When Independent Research Exposed Hidden Bias in AI

Year: 2018
Location: United States
Principal Researchers: Joy Buolamwini and Timnit Gebru
Research Institution: MIT Media Lab / Microsoft Research
Historical Theme: AI Fairness · Accountability · Independent Testing · Equality · Human Responsibility


Historical Significance

In 2018, researchers Joy Buolamwini and Timnit Gebru published Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, a landmark study examining whether commercial facial-analysis systems performed equally across different skin types and genders. (Proceedings of Machine Learning Research)

The researchers developed a more demographically balanced facial-image benchmark and evaluated three commercial gender-classification systems.

The results revealed substantial differences in performance.

The study found that darker-skinned women were the most frequently misclassified group, with error rates reaching 34.7%, while the maximum error rate for lighter-skinned men was 0.8%. (Proceedings of Machine Learning Research)

The research transformed an abstract concern about algorithmic bias into a measurable, reproducible, and publicly documented finding.


Figure 1 — Discovering the Coded Gaze

Title: When the Machine Could Not See Everyone Equally

Description: A museum-quality historical reconstruction representing the experience that motivated Buolamwini’s investigation: a facial-analysis system interacting differently with faces of different skin tones, leading a researcher to question whether the technology was truly neutral.


The Research Question

The central question was simple but consequential:

Does an AI system perform equally well across different groups of people?

Existing evaluations could report a single overall accuracy number.

But a high overall accuracy could conceal substantial differences between groups.

Buolamwini and Gebru therefore examined performance at the intersection of gender and skin type, rather than treating each category separately. Their research introduced a more inclusive benchmark and evaluated four intersectional groups: darker females, darker males, lighter females, and lighter males. (MIT Media Lab)

This approach changed the question from:

“How accurate is the AI?”

to:

“How accurate is the AI for different people?”


Title: From Overall Accuracy to Intersectional Evaluation

Description: The table shows how the research examined AI performance across intersecting demographic and phenotypic groups rather than relying only on aggregate accuracy.

Table 1 — The Gender Shades Research Design

Research ElementDescription
AI TechnologyCommercial facial-analysis gender-classification systems
ResearchersJoy Buolamwini and Timnit Gebru
Systems EvaluatedThree commercial gender-classification systems
Faces Evaluated1,270 unique individuals
Skin-Type MethodFitzpatrick Skin Type classification
Gender GroupsFemale and male labels used by the evaluated systems
Intersectional GroupsDarker females, darker males, lighter females, lighter males
PurposeDetermine whether system performance differed substantially across groups
Core InnovationIntersectional evaluation of gender and skin type

The study’s benchmark included 1,270 individuals and was designed to provide more balanced representation than the datasets examined by the researchers. (MIT Media Lab)


The Evidence

The researchers examined two existing facial-analysis benchmarks and found substantial representation imbalances.

The IJB-A and Adience datasets were overwhelmingly composed of lighter-skinned subjects: 79.6% in IJB-A and 86.2% in Adience. (Proceedings of Machine Learning Research)

The researchers then developed the Pilot Parliaments Benchmark (PPB), containing 1,270 individuals selected from three African and three European countries to create greater balance in gender and skin type representation. (MIT Media Lab)

The results demonstrated that the systems did not perform equally across the tested groups.


Title: When Aggregate Accuracy Conceals Inequality

Description: A comparison of the reported error-rate extremes identified by the Gender Shades research.

Table 2 — The Performance Gap

GroupReported Maximum Error Rate
Darker-Skinned Females34.7%
Lighter-Skinned Males0.8%
Difference33.9 percentage points

The paper reported darker-skinned females as the most misclassified group and lighter-skinned males as the least misclassified group among the evaluated categories. (Proceedings of Machine Learning Research)

The significance of the finding was not simply that one AI system made mistakes.

It demonstrated that the probability of error could differ dramatically depending on who was being evaluated.


Figure 2 — The Evidence of Bias

Title: Four Groups, Four Different Experiences

Description: A museum visualization showing four representative demographic groups being evaluated by the same facial-analysis AI system, with visibly different error-rate indicators. The image emphasizes that a single AI system can produce unequal performance across groups.


The Righteous Choice

The importance of Gender Shades lies not only in discovering a technical problem.

It lies in the decision to measure what others might overlook.

The research demonstrated several principles central to responsible AI:

Truth

A system’s performance should be measured according to evidence rather than assumptions about technological neutrality.

Fairness

AI systems should not be evaluated only according to the groups for whom they perform best.

Accountability

Independent evaluation can reveal problems that may remain hidden when system performance is reported only in aggregate.

Inclusion

People who are underrepresented in datasets should not become invisible in the evaluation of AI systems.

Courageous Inquiry

Researchers can challenge widely accepted assumptions by testing them systematically.


Title: Why Gender Shades Is Considered for the Righteous AI Gallery

Description: The Museum’s analytical framework evaluates the historical choices and principles demonstrated by the Gender Shades research.

Table 3 — Righteousness Test

CriterionHistorical Assessment
ConflictTechnological claims of general performance vs. evidence of unequal subgroup performance
QuestionDoes AI work equally well for different people?
PrincipleTruth, fairness, inclusion, and accountability
ActionIndependent empirical testing
InnovationIntersectional evaluation of gender and skin type
EvidenceQuantitative performance differences across demographic groups
Public ValueExposed a problem relevant to responsible AI development
Historical ImpactBecame an influential landmark in discussions of algorithmic bias and AI fairness
Model for HumanityDemonstrated how independent investigation can expose hidden problems and demand greater accountability

From Personal Experience to Public Evidence

The story of Gender Shades began with a personal technological experience.

Buolamwini had encountered facial-analysis systems that failed to detect her face or classified it incorrectly. Rather than treating the experience as an isolated technical failure, she asked whether similar problems affected other people. (MIT Media Lab)

That question led to systematic research.

The progression was:

Personal Experience

↓

Research Question

↓

Independent Testing

↓

Better Benchmark

↓

Quantitative Evidence

↓

Public Accountability

This transformation is one of the most important aspects of the case.

A personal experience became a reproducible scientific investigation.


Title: How a Question Became a Movement for Accountability

Description: This table traces the transformation of a personal technological experience into a broader contribution to responsible AI research.

Table 4 — From Individual Experience to Historical Impact

StageDevelopment
01 — ExperienceFacial-analysis technology produced inconsistent results for Buolamwini.
02 — QuestionShe questioned whether the problem was specific to her or reflected a broader pattern.
03 — InvestigationResearchers systematically evaluated commercial systems.
04 — BenchmarkA more balanced dataset was developed to test performance across groups.
05 — MeasurementAccuracy was examined across intersecting gender and skin-type categories.
06 — EvidenceSignificant disparities in error rates were documented.
07 — AccountabilityThe findings challenged developers and researchers to evaluate AI more inclusively.
08 — LegacyThe work became an important reference point in the development of algorithmic fairness and responsible AI.

Figure 3 — From Research to Accountability

Title: Making Hidden Bias Visible

Description: A museum installation representing the transformation of research evidence into public accountability: a researcher, scientific charts, facial-analysis results, and policymakers or technology professionals examining documented disparities.


Historical Impact

Gender Shades helped change how algorithmic fairness could be investigated.

The research showed that it was insufficient to ask only whether an AI system was accurate overall.

Researchers increasingly needed to ask:

Who is represented in the data?

Who is missing?

Who experiences the highest error rate?

What happens when demographic categories intersect?

The project also emphasized the importance of transparent subgroup performance reporting. MIT’s Gender Shades project describes the work as an approach to inclusive product testing for AI and emphasizes the need for intersectional evaluation. (MIT Media Lab)


Title: From Algorithmic Bias to Responsible AI Evaluation

Description: Major ideas demonstrated or advanced by the Gender Shades research and their significance for responsible AI.

Table 5 — The Historical Legacy of Gender Shades

AreaContribution
AI FairnessDemonstrated measurable disparities in facial-analysis performance.
Algorithmic AuditingDemonstrated the value of independent empirical testing.
Dataset DiversityHighlighted the consequences of imbalanced benchmark datasets.
IntersectionalityDemonstrated why intersecting characteristics can reveal disparities hidden by aggregate statistics.
TransparencyEncouraged more detailed reporting of subgroup performance.
AccountabilityCreated evidence that technology developers could use to investigate and improve systems.
Responsible InnovationShowed that identifying a technological weakness can be a form of constructive innovation.

The Righteous Innovation

The significance of Gender Shades is not that it created a new AI product.

Its innovation was methodological and ethical.

The researchers changed the way an important class of AI systems could be evaluated.

Instead of accepting:

“The system works.”

the research demanded a more rigorous question:

“For whom does the system work, and for whom does it fail?”

This is a powerful example of righteous innovation.

Innovation does not always mean building something new.

Sometimes innovation means creating a better way to test, question, measure, and correct what already exists.


Figure 4 — Righteous Innovation

Title: A Better Way to Measure AI

Description: A conceptual museum image showing a transition from a conventional AI evaluation system using one overall accuracy score to an intersectional evaluation system examining multiple demographic groups separately.


The Human–AI Lesson

Gender Shades demonstrates that AI systems do not exist independently of human decisions.

Human beings decide:

  • What data to collect
  • Which people to represent
  • Which benchmarks to use
  • Which metrics to report
  • Which errors to investigate
  • Which problems to correct
  • Which systems are ready for deployment

An AI system can therefore reproduce limitations embedded in the data, design, evaluation methods, and assumptions of its creators.

The lesson is not that AI is inherently unrighteous.

The lesson is that human responsibility remains essential throughout the AI lifecycle.


Title: Where Righteousness Enters AI Development

Description: The table identifies points at which human decisions can influence fairness, accuracy, transparency, and accountability in AI systems.

Table 6 — The Human–AI Responsibility Chain

StageHuman DecisionRighteousness Question
DataWhich people are represented?Is representation sufficiently inclusive?
DesignWhat problem is being solved?Is the purpose legitimate and beneficial?
TrainingWhat data and methods are used?Could systematic bias be introduced?
TestingWho is evaluated?Are vulnerable or underrepresented groups included?
MeasurementWhich metrics are reported?Can aggregate results conceal important disparities?
DeploymentWhere is the system used?Are the risks appropriate for the application?
MonitoringWhat happens after deployment?Are failures detected and corrected?
AccountabilityWho answers for harm?Can responsibility be identified and exercised?

Museum Assessment

Righteous AI Gallery Classification

Primary Classification

Human–AI Righteousness

Historical Character

Pioneering Independent AI Audit

Core Principle

Truth Through Evidence

Key Virtues

Fairness · Accountability · Courageous Inquiry · Inclusion

Righteous Action

Seeing a possible injustice → questioning the assumption → measuring the evidence → making the problem visible → enabling correction

Historical Significance

Gender Shades represents a particularly important form of Human–AI righteousness:

using technical expertise to reveal a hidden problem rather than allowing a convenient assumption to remain unchallenged.


Why This Event Belongs in the Righteousness Museum

The historical importance of Gender Shades is not simply that an AI system made mistakes.

AI systems can make mistakes.

The deeper significance is that researchers looked for unequal outcomes, developed a method to measure them, documented the evidence, and brought the problem into public view.

The event therefore represents a model of responsible technological leadership:

Do not assume that technology is neutral.

Do not measure only what is convenient.

Do not hide unequal outcomes behind averages.

Test the system.

Reveal the evidence.

Improve the technology.


The Enduring Question

When an AI system appears to work well, who has the responsibility to ask whether it works equally well for everyone?

Gender Shades demonstrates that progress in artificial intelligence does not come only from making machines more capable.

It also comes from making the measurement of those machines more truthful.

Question the assumption.

Measure the evidence.

Reveal the disparity.

Demand accountability.

Build better AI.


References

Buolamwini, J., & Gebru, T. (2018). Gender Shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 77–91. https://proceedings.mlr.press/v81/buolamwini18a.html (Proceedings of Machine Learning Research)

Buolamwini, J. (2017). Gender Shades: Intersectional phenotypic and demographic evaluation of face datasets and gender classifiers [Master’s thesis, Massachusetts Institute of Technology]. MIT Media Lab. https://www.media.mit.edu/publications/full-gender-shades-thesis-17/ (MIT Media Lab)

MIT Media Lab. (2018). Gender Shades. https://www.media.mit.edu/projects/gender-shades/ (MIT Media Lab)

MIT Media Lab. (2018). Gender Shades: Publications. https://www.media.mit.edu/projects/gender-shades/publications/ (MIT Media Lab)

MIT Media Lab. (2018). Gender Shades: Frequently asked questions. https://www.media.mit.edu/projects/gender-shades/faq/ (MIT Media Lab)

Gender Shades. (2018). Gender Shades: How well do IBM, Microsoft, and Face++ AI services guess the gender of a face? https://gendershades.org/ (Gender Shades)


Curatorial Note

Exhibit is preserved as a historical study of Human–AI righteousness through independent inquiry, evidence, fairness, and accountability.

The exhibit does not claim that the researchers proved that all facial-analysis AI is inherently harmful or that every error represents intentional discrimination.

Instead, it preserves a more precise historical lesson:

When technology affects people, responsible innovation requires the courage to measure its performance honestly—including where it fails and whom those failures affect.