A hospital system wants to find patients who would benefit from extra care: more nurse check-ins, quicker access to specialists, closer follow-up. It buys an algorithm that scores patients by health risk, and the algorithm looks reasonable. It predicts what it was built to predict, and the people it flags do tend to need care.
In 2019, a team led by Ziad Obermeyer published a study in Science examining an algorithm of this kind. They found that it was working as designed and still produced a troubling result. At the same risk score, Black patients were considerably sicker than White patients. The cause was the algorithm’s target. It predicted future healthcare spending as a stand-in for healthcare need, and spending is not the same as sickness. The researchers estimated that correcting the disparity would raise the share of Black patients receiving additional help from 17.7% to 46.5%.
The case raises a harder question than “is the model accurate?” If a system performs well overall but works differently for different groups, what does it mean for that system to be fair? That question is what AI fairness is about, and answering it is more involved than it first appears.
What Is AI Fairness?
AI fairness is the effort to make sure an AI system does not produce unjustified, harmful, or discriminatory outcomes for people or groups, and to judge that against a standard suited to the situation. In plain English, it asks whether a system treats the people it affects in a way that is defensible. That does not mean identical treatment or identical outcomes. Consider a loan model that approves exactly the same percentage of applicants in every group. That sounds fair until you learn the model achieved it by ignoring repayment ability. Now consider a model that treats every applicant identically but makes far more mistakes for one group than another. That also sounds fair on paper and is not.
NIST, the U.S. standards agency behind the AI Risk Management Framework, treats fairness as a cluster of related concerns: equality and equity, harmful bias, and discrimination. It also stresses that fairness standards are complex and differ across cultures and applications. NIST lists “fair, with harmful bias managed” as one of the characteristics of trustworthy AI, next to qualities like safety, security, and transparency. The OECD AI Principles likewise place fairness among the human-centered values AI systems should respect.
The common thread in these sources is context. Fairness is not a switch inside a model that is either on or off. It is a judgment about a particular system, used for a particular purpose, affecting particular people.
Why Is AI Fairness So Difficult to Define?
If fairness were simple, everyone would agree on a formula. Instead, reasonable people disagree about what a fair outcome looks like, and different applications call for different answers. Depending on the setting, fairness might mean giving equally qualified people an equal chance of being selected. It might mean making similar kinds of mistakes at similar rates across groups. It might mean selecting people at similar rates, or ensuring a score carries the same meaning no matter who receives it. It might mean treating similar individuals similarly, or avoiding unjustified discrimination of any kind. These goals sound compatible, but they are not interchangeable. Reaching one can mean giving up another.
Reuben Binns, a researcher then at the University of Oxford, made this point in a 2018 paper that connects machine learning fairness to political philosophy. Philosophers have argued for a long time about what equality requires: equal treatment, equal opportunity, or equal outcomes. Binns argued that the same disagreements sit beneath the technical definitions used in machine learning. When a team picks a fairness metric, it is taking a position in an old debate, whether it realizes it or not. The disagreement is not only philosophical. In 2016, Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan showed mathematically that several intuitive fairness conditions for risk scores cannot all hold at once, except in special cases. Those cases are when the prediction is perfect or when the groups being compared have the same underlying rate of the outcome. We return to this finding below.
The practical meaning is that choosing a fairness metric is partly a question about which kind of error or unequal treatment matters most for the application. A medical screening tool, a hiring filter, and a content recommendation system do not carry the same risks, so they should not be judged by the same yardstick.
AI Fairness vs. AI Bias: What’s the Difference?
The two terms are often used interchangeably, but they do different jobs. Bias describes patterns or influences that can distort outcomes or create unequal treatment. Bias can come from data, from design choices, from human judgment, or from the social setting in which a system is used. Fairness is the standard we use to decide whether a difference or outcome is acceptable, harmful, unjustified, or in need of action.
Put simply, bias is something you look for, and fairness is how you decide what to do about what you find. A difference between groups is not automatically proof of unfairness. Suppose a model that flags patients for a screening test flags older adults more often than younger ones. Because the condition is far more common in older people, that difference may be entirely appropriate. Now suppose a model flags one racial group less often even though patients in that group are equally sick. That difference signals a problem. The numbers might look similar in both cases, but the judgment differs because it depends on whether the difference is justified for the task.
This is also why fairness work cannot be reduced to hunting for bias. Bias has to be examined in light of a standard, and that standard has to be chosen. The data side of that story is covered in more depth in our guide to machine learning bias, but data is only one of several places where trouble can start.
Where Can Unfairness Enter an AI System?
It is tempting to picture unfairness as a flaw inside the model, as if an engineer wrote a bad line of code. In practice, unfairness often exists before any model is trained. The following is a practical way for beginners to think about it, not an official taxonomy. It can start with problem definition. Deciding what the system is for, and who it is for, shapes everything after. A tool built to cut costs will behave differently from one built to improve outcomes, even if both use the same data.
Next comes data collection and representation. If some groups are thinly represented, or the data was gathered in settings that do not reflect everyone the system will serve, performance may differ across groups. Datasets also record the history of the world that produced them, including its inequalities. Then there is measurement and proxies. Many things we care about, such as health, job performance, or creditworthiness, cannot be measured directly, so we substitute something measurable. The healthcare algorithm described earlier is a clear case: spending stood in for need, and the substitution carried unequal access to care into the score.
Labeling and human judgment matter too. Training labels often come from people, such as past hiring decisions, human reviewers, or clinicians’ notes, and those judgments can carry the inconsistencies and assumptions of the people who made them. During model development, choices about which features to use, how to define success, and how to trade off different kinds of errors can affect groups differently. Evaluation matters just as much. A system tested only on averages, or only on populations that resemble its developers’ own, can look better than it is.
Finally, deployment and feedback shape what happens in the real world. People may use a tool differently than designers expected. Decisions made with a system can change the data it later learns from. A system that works well in one hospital, region, or population may behave differently in another. NIST describes AI as socio-technical: its behavior comes from the interaction of technical components with human decisions, organizational practices, and social conditions. That framing explains why fixing the algorithm alone often does not fix the outcome. Where machine learning systems are involved, fairness is shaped by the whole process around the model, not just the model itself.
How Is AI Fairness Measured?
Here is the key principle: you cannot meaningfully measure fairness until you have decided what fairness means for the particular use case. Metrics are not neutral thermometers. Each one encodes a specific idea of what equal treatment should look like. To keep the explanations concrete, imagine a hypothetical model that screens loan applicants and predicts who will repay. We compare two groups of applicants, Group A and Group B.
Demographic parity
Demographic parity asks whether positive outcomes occur at similar rates across groups. If the model approves 40% of Group A applicants, does it approve about 40% of Group B applicants? It is easy to calculate and understand. Its weakness is that it ignores whether the groups differ in the underlying thing being predicted, so a model can satisfy it while still making worse decisions for one group.
Equal opportunity
Equal opportunity focuses on people who actually deserve the positive outcome. In our example, those are applicants who would really repay. A true positive rate is the share of those people the model correctly approves. Equal opportunity asks whether that rate is similar across groups, so that qualified applicants have a similar chance of being recognized regardless of group. The idea was formalized by Moritz Hardt, Eric Price, and Nathan Srebro in 2016.
Equalized odds
Equalized odds adds a second requirement. Besides similar true positive rates, it asks for similar false positive rates, meaning the share of people who would not repay but are approved anyway. The idea is that both kinds of mistakes, wrongly denying someone and wrongly approving someone, should be spread evenly across groups. This matters when both errors carry real costs.
Predictive parity and calibration
These ask whether a score means the same thing for everyone. Calibration means that among people given a score of, say, 70%, about 70% actually have the outcome, and this holds in each group. If a “70% likely to repay” score means 70% in Group A but only 55% in Group B, the score is not equally trustworthy for both groups. Predictive parity is a closely related idea focused on how often positive predictions turn out to be correct.
Individual fairness
Individual fairness shifts attention from groups to people. Its core idea, set out by Cynthia Dwork and colleagues in 2012, is that similar individuals should be treated similarly for the task at hand. The difficulty is defining “similar.” Deciding which differences are relevant to a loan decision and which are not is itself a value-laden choice.
Counterfactual fairness
Counterfactual fairness asks a “what if” question. Would the model’s prediction change if the person’s demographic attribute were different, while the relevant circumstances were held appropriately constant? The idea was proposed by Matt Kusner and colleagues in 2017 and relies on assumptions about how attributes and circumstances causally relate. Those assumptions can be hard to establish, so the approach is more often used in research than in routine auditing.
Even this short tour shows that fairness measurement is a set of different questions, not one test with one answer. Research on how people understand these metrics reflects the difficulty: a study by Saha and colleagues at ICML 2020 examined how non-experts comprehend fairness metrics, which is a reminder that explaining them clearly is part of using them responsibly.
Why One Fairness Metric Is Not Enough
Because each metric asks a different question, two metrics can tell different stories about the same system. In our hypothetical loan model, the approval rates might be nearly identical across groups, so demographic parity looks satisfied. Yet if the model wrongly denies far more qualified applicants in Group B, the error rates tell another story. The first metric says “balanced” while the second says “not balanced,” and neither is lying.
The same happens with risk scores. A score can be well calibrated, meaning it carries the same meaning in each group, while the groups still experience different false positive and false negative rates. Kleinberg, Mullainathan, and Raghavan showed why this is not just bad luck. They considered three conditions for a risk score: calibration within each group, equal average scores among people who actually have the outcome, and equal average scores among those who do not. They proved that, except in special circumstances, you cannot satisfy all three at once. The special circumstances are perfect prediction or equal underlying rates of the outcome across groups, and real-world data rarely offers either.
This does not mean AI cannot be fair. It means that different definitions of fairness answer different questions, and they can pull in different directions. Organizations therefore have to decide which risks and outcomes matter most for their application and be transparent about that decision. In a medical screening setting, missing a sick patient may be the harm that matters most. In another setting, wrongly flagging someone may carry the heavier cost. The right metric follows from the harm.
Why Accuracy Alone Does Not Prove AI Is Fair
Overall accuracy, the percentage of predictions a model gets right, is the number most people reach for first. It is useful, but it can hide a great deal. Imagine a model evaluated on 1,000 people, 900 from a large group and 100 from a smaller one. Suppose it is right 95% of the time for the large group and 70% of the time for the small group. Its overall accuracy would be about 92.5%, which sounds excellent, yet the smaller group experiences a much worse system. The average conceals the gap because the larger group dominates it.
Digging deeper means looking at false positives (the model says yes when the answer is no) and false negatives (the model says no when the answer is yes). These carry different human costs. A false negative in a cancer screening can mean a missed diagnosis. A false positive in a fraud system can mean a frozen account for an innocent customer. If those errors fall unevenly across groups, a high accuracy figure says little about who bears the burden.
This is why researchers recommend disaggregated evaluation, which means measuring performance separately for relevant subgroups rather than only in aggregate. The right subgroups depend on the application, and in some cases on combinations of attributes, as the next section shows.
Real-World Examples of AI Fairness Problems
Three well-documented cases show how these ideas play out.
A healthcare algorithm and the proxy problem
The Obermeyer study, published in Science in 2019, examined an algorithm used to identify patients for extra care programs. It predicted future healthcare costs, using cost as a proxy for health need. The researchers found that at the same risk score, Black patients were considerably sicker than White patients. The authors point to unequal access to care as one reason: when people face barriers to getting care, less money is spent on them even when they are equally ill, so a cost-based score underestimates their need.
The lesson is that a model can be accurate at predicting the thing it was told to predict and still fail at the thing people care about. The unfairness was introduced by the choice of target, not by any explicit use of race.
Gender Shades and intersectional error rates
In 2018, Joy Buolamwini and Timnit Gebru evaluated commercial gender-classification systems, programs that guess whether a face appears to be male or female. This is not face identification, which matches a face to a specific person. They built a benchmark balanced by skin tone and gender and tested systems from three companies. The error rates differed sharply by group, reaching up to 34.7% for darker-skinned women compared with a maximum of 0.8% for lighter-skinned men.
The study’s contribution was to show that looking at gender or skin tone alone can miss the worst gaps, which appeared at their intersection. It is also a snapshot: these figures describe the systems tested at that time and do not describe every commercial system today. What endures is the method. Test performance by subgroup, including combinations of attributes.
NIST and facial recognition
In 2019, NIST published a large evaluation of face recognition algorithms from many developers, examining how accuracy varied across demographic groups such as age, sex, and race. NIST found evidence of demographic differentials in a majority of the algorithms it tested. For one-to-one matching, false positive rates were often higher for Asian and African American faces than for Caucasian faces, by factors NIST described as ranging from about 10 to 100 depending on the algorithm. Results varied widely across algorithms, and some showed much smaller differences.
NIST was careful about causes. Its report noted that the study did not determine what produced the differentials, so it would be wrong to claim the evidence proves that unrepresentative training data explains them in every case. The lesson is partly about rigor: independent, large-scale testing can reveal that performance differs across groups even when the reasons remain open.
Each case teaches something different. The healthcare example is about choosing the right target, Gender Shades is about testing the right subgroups, and the NIST work is about independent measurement and honesty about what is not yet known.
Does Removing Race or Gender Make AI Fair?
A natural first idea is to delete sensitive attributes such as race or gender from the data so the model cannot use them. This approach is sometimes called “fairness through unawareness,” and it usually falls short.
The reason is proxies. A proxy variable is a piece of information that is correlated with a sensitive characteristic and so can carry some of the same signal. ZIP code can correlate with race or income because of how neighborhoods developed. The school someone attended, their occupation, their employment history, their location, or even their language patterns can carry information about gender, socioeconomic background, or ethnicity. A model trained without the sensitive column can often still reconstruct a good deal of it from the remaining ones.
The healthcare case shows a version of this. Race was not the input. Cost was, and cost carried the effects of unequal access. This does not mean that using any of these variables automatically creates discrimination. ZIP code, for example, may have a legitimate role in some tasks. The point is conceptual: removing a column does not remove the patterns the column reflected. In fact, deleting sensitive attributes can make unfairness harder to detect, because you lose the ability to check how the system performs across groups. Many fairness evaluations depend on having that information available for testing, even if it is not used to make decisions. Whether and how sensitive attributes can be collected or used also depends on the legal and ethical context, which is a question for qualified professionals rather than a general rule.
What Does a Fairer AI System Actually Require?
Fairness cannot be demonstrated by a single chart or a single test. A fairer system is the product of many decisions made well, and each can help reduce risk or provide evidence, though none guarantees an outcome.
It starts with a clearly defined purpose: what decision is being supported and what harm could occur if the system errs. From there, teams need to identify the people and groups who could be affected, then ask whether the data covers them appropriately. Representative coverage can help, but so does careful choice of the target and any proxies, since the healthcare example shows how a poor target can undermine an otherwise sound model. Thoughtful labeling matters because labels carry human judgment.
Evaluation then needs to include subgroup testing and fairness metrics suited to the context. Human and domain expertise is essential here, because numbers alone cannot say which differences are justified. Documentation and transparency help others understand how a system was built and what it was and was not designed for. Datasheets for datasets, proposed by Timnit Gebru and colleagues, are one example: a structured way of recording how a dataset was created, what it contains, and what uses it suits.
After deployment, monitoring can reveal problems that testing did not, and a way to investigate and correct harmful outcomes gives the whole effort accountability. This reflects the broader approach of the NIST AI Risk Management Framework, which treats managing harmful bias as an ongoing risk-management activity across the system’s life, not a one-time check.
How Can Organizations Improve AI Fairness?
Turning those requirements into a workflow looks something like this. First, define the decision and the potential harm. What does the system influence, and what happens to a person if it gets them wrong? Then identify who can be affected, including groups that may not be obvious at the outset. Next, inspect the data. Look at who is represented, how outcomes were recorded, and what the variables actually measure. Then decide what fairness means in context, ideally with input from domain experts and the people affected, and choose evaluation measures that match. A hiring tool and a medical triage tool will not share the same answer.
With those choices made, test performance across relevant groups. When gaps appear, investigate causes rather than simply adjusting the metric until it looks acceptable. A gap might come from data coverage, a flawed proxy, inconsistent labels, or the way the tool is used, and the right fix depends on which. Changing a threshold to equalize one number can leave the real problem untouched, or create a new one.
Document the decisions and trade-offs so that others can review them later. Monitor the system after deployment, since real-world conditions change. Finally, create a process for responding to problems when they surface, including who is responsible and how affected people can raise concerns. For a more detailed look at specific techniques, see our guide on how to reduce AI bias. This general guidance is not legal advice, and legal requirements for fairness testing vary by jurisdiction and industry, so organizations should consult qualified counsel for compliance questions.
AI Fairness in Generative AI and Language Models
Fairness concerns do not stop at systems that score or classify people. They also apply to systems that generate text, images, and code. Modern generative AI can reproduce stereotypes found in its training material, or represent some identities and groups less richly than others. Quality can vary as well: a model may respond less accurately to certain dialects or to languages with less training data. Behavior around safety can also differ, such as refusing some requests more readily when they mention particular groups, or handling similar prompts inconsistently depending on the identities involved.
Many of these issues trace back to training-data patterns, but not all of them do. Choices about how language models are fine-tuned, evaluated, and moderated also shape what users experience. The same questions from earlier still apply, though they are harder to answer. There is rarely a single correct output to compare against, so evaluating fairness often means examining patterns across many responses rather than checking individual answers.
A Simple Way to Think About AI Fairness
When you meet any claim that an AI system is “fair” or “unfair,” four questions can help you evaluate it. This is a practical way of thinking offered by LegacyVia, not an academic framework.
Fair for whom? Every system affects different people in different ways. Start by naming the groups and individuals involved, including those who are easy to overlook, such as people with less data about them or fewer ways to object.
Fair according to what standard? Equal selection rates, equal error rates, equally meaningful scores, and similar treatment of similar people are different standards. A claim of fairness means little until it says which one it is using and why.
Fair in which context? The same model can be reasonable in one setting and harmful in another. The stakes, the alternatives, and the way people use the output all matter.
Fair based on what evidence? Look for testing across relevant groups, documentation of decisions, and ongoing monitoring. A claim without evidence is a promise, not a finding.
Common Misconceptions About AI Fairness
Fair AI means everyone gets the same result
Identical outcomes are rarely the goal and are sometimes the wrong one. Fairness usually concerns whether differences are justified and whether people face similar chances or similar kinds of mistakes, depending on the setting.
Removing race or gender from the dataset removes bias
Other variables can carry the same information, so the pattern can survive the deletion. Removing the attribute can also make it harder to check whether outcomes differ.
High accuracy means an AI system is fair
Accuracy averages over everyone. A system can score well overall while making far more mistakes for a smaller group, as the hypothetical above and the Gender Shades study illustrate.
Fairness is only a machine-learning problem
Many sources of unfairness sit outside the model: how the problem was framed, what was measured, who labeled the data, and how people use the results. That is why NIST treats AI as socio-technical.
One fairness metric can prove a model is fair
Each metric answers a single question, and different metrics can disagree. A good result on one is evidence about one aspect of fairness, not a verdict on the whole system.
If the model is fair, the entire AI system is fair
A well-evaluated model can still sit inside an unfair process. If people apply its outputs inconsistently, if affected individuals have no way to contest a decision, or if the surrounding policy is unjust, a good model does not fix that.
Frequently Asked Questions About AI Fairness
What is AI fairness in simple terms?
AI fairness is the effort to ensure AI systems do not create unjustified or harmful differences in how people are treated, judged against a standard that fits the specific use. It is not the same as giving everyone identical results.
Why is AI fairness important?
AI systems increasingly influence decisions about health care, lending, hiring, and public services. If they work worse for some groups, the harm can scale quickly and be hard to notice. Fairness work helps surface those problems and gives organizations a basis for addressing them.
What is the difference between AI fairness and AI bias?
Bias refers to patterns or influences that can distort outcomes or produce unequal treatment. Fairness is the standard used to judge whether a given difference is acceptable or harmful. Bias is what you look for, and fairness helps decide what to do about it.
How is AI fairness measured?
By choosing a fairness criterion that fits the application and then testing the system against it, usually by comparing results across groups. Measurement comes after deciding what fairness means for the case at hand.
What are common AI fairness metrics?
Common ones include demographic parity, equal opportunity, equalized odds, and calibration or predictive parity. Researchers also discuss individual fairness and counterfactual fairness, which focus on individuals and “what if” comparisons.
What is demographic parity?
Demographic parity asks whether positive outcomes, such as approvals or selections, occur at similar rates across groups. It is simple to compute, but it does not account for whether the groups differ in the underlying outcome being predicted.
What is equal opportunity in AI?
Equal opportunity asks whether people who truly qualify for a positive outcome have a similar chance of receiving it across groups. Technically, it compares true positive rates.
Can AI ever be completely fair?
There is no single definition of fairness that every system can satisfy completely, and some reasonable definitions cannot all hold at once. That does not mean fairness is out of reach. Systems can be made fairer by choosing suitable standards, testing carefully, and correcting problems over time.
Can removing race or gender eliminate AI bias?
No. Other variables can act as proxies and carry similar information, so removing the attribute does not remove the underlying patterns. It can also make unfairness harder to detect.
Why can two fairness metrics disagree?
Each metric measures a different idea, such as equal selection rates or equal error rates. Research by Kleinberg, Mullainathan, and Raghavan shows that some common criteria cannot all be met at once except in special cases, so improving one can worsen another.
Conclusion
The healthcare algorithm at the start of this article was not malicious, and it was not inaccurate in the narrow sense. It did what it was asked to do. The problem was that what it was asked to do did not match what mattered. That is the heart of AI fairness. It is not a single score, and it is not a promise that every group will receive identical outcomes. It is a process of deciding what fairness means for a specific system, examining who is affected, identifying meaningful sources of unequal treatment, selecting appropriate evidence, and continuing to evaluate the system as it is used.A fair AI system is not one that has passed a test once; it is one whose builders keep asking who it works for, who it fails, and what they will do about it.
Sources
Click
- NIST AI Risk Management Framework: AI Risks and Trustworthiness — https://airc.nist.gov/airmf-resources/airmf/3-sec-characteristics/
- NIST Face Recognition Vendor Test (FRVT) Part 3: Demographic Effects — https://www.nist.gov/publications/face-recognition-vendor-test-part-3-demographic-effects
- OECD AI Principles — https://www.oecd.org/en/topics/ai-principles.html
- Obermeyer et al., “Dissecting racial bias in an algorithm used to manage the health of populations,” Science (2019) — https://doi.org/10.1126/science.aax2342
- Buolamwini and Gebru, “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,” PMLR 81 (2018) — https://proceedings.mlr.press/v81/buolamwini18a.html
- Binns, “Fairness in Machine Learning: Lessons from Political Philosophy,” PMLR 81 (2018) — https://proceedings.mlr.press/v81/binns18a
- Kleinberg, Mullainathan, and Raghavan, “Inherent Trade-Offs in the Fair Determination of Risk Scores” (2016) — https://arxiv.org/abs/1609.05807
- Hardt, Price, and Srebro, “Equality of Opportunity in Supervised Learning” (2016) — https://arxiv.org/abs/1610.02413
- Dwork et al., “Fairness Through Awareness” (2012) — https://arxiv.org/abs/1104.3913
- Kusner et al., “Counterfactual Fairness” (2017) — https://arxiv.org/abs/1703.06856
- Saha et al., “Measuring Non-Expert Comprehension of Machine Learning Fairness Metrics,” PMLR 119 (2020) — https://proceedings.mlr.press/v119/saha20c.html
- Gebru et al., “Datasheets for Datasets” — https://www.microsoft.com/en-us/research/?p=563964

