Race, classification, and unequal treatment
A fictional recruitment office notices that applicants assigned to two demographic categories receive interviews at different rates. One manager calls the disparity proof of biased screening. Another says it merely reflects differences in qualifications. A third argues that the categories are too crude to be useful. Each response identifies a possible issue, but none has yet examined the process. We need to know how the categories were produced, what the outcome measures and what comparison could distinguish explanations.
Racial classification and unequal treatment are related subjects, but they are not the same analytical task. We can investigate how institutions define groups, how people identify themselves, how others perceive them and how a decision responds to a perceived category. Keeping these questions distinct makes it possible to study consequential inequalities without treating the classifications as timeless descriptions of human essence.
Categories have institutional histories
The US National Archives' description of the 1790 census lists categories separating free white males by age, free white females, other free persons and enslaved people. The enumeration combined racialized labels, sex and age with legal condition. We are reading the Archives' institutional description of the questions, not inspecting an original household manuscript or treating those historical labels as our own categories of worth. 1790 census records.
The document shows that an administrative count is organized by decisions about what distinctions to record. It does not simply discover a single set of natural boxes waiting in the population. A category's presence can reveal what officials sought to administer or distinguish, while its absence can make other experiences difficult to reconstruct from that record alone.
This is not a reason to conclude that all records are useless. It is a reason to read them historically. We can ask who designed the form, who supplied the information, who assigned the category and what uses followed. A count may be accurate within its rules while remaining inadequate for a different question. The rules belong in the interpretation of the number.
Comparisons across time therefore require more than matching labels. A label may change meaning, a form may allow more than one response, or the method of assignment may change. An apparent increase in a group's size can combine population change with changed classification. Before explaining a trend, establish whether the units being compared were defined consistently enough for the claim.
Self-identification and perceived classification answer different questions
Imagine a fictional survey that asks applicants how they identify, while recruiters see only names and application materials. The survey measures a self-reported category. The hiring decision may respond to a perceived category inferred from a signal. These measures can be related without being identical. A mismatch is not necessarily a dishonest response; the two observers are answering different questions with different information.
This distinction affects study design. To investigate experiences associated with self-identification, a survey response may be relevant. To test whether a recruiter changes behavior in response to a perceived signal, researchers need a design that varies or observes that signal. A demographic label in a database cannot tell us exactly what a decision-maker perceived at the moment of choice.
Multiple classifications can also apply in different settings. Someone may be read one way by a stranger, identify differently in a survey and encounter a third administrative classification. The resulting experiences need not be uniform. An analysis should name the classification relevant to the mechanism instead of merging every measure into a single supposedly obvious fact about the person.
Socially produced categories can have material consequences. A boundary does not need to be biologically fixed to affect access, treatment or recognition. The relevant causal chain runs through practices and institutions that act on the classification. Establishing that chain requires evidence of what people and organizations do, not an assumption that a category itself acts independently of them.
A disparity is an outcome to explain
Suppose our fictional office receives one hundred applications in each of two recorded groups and invites twenty people from one and ten from the other. The invitation rates are twenty percent and ten percent. The difference is ten percentage points, and the higher rate is twice the lower. These calculations establish a disparity in this applicant pool under the recorded classifications.
They do not identify its cause. The pool may contain differences in relevant experience, the screening process may respond differently to equivalent materials, earlier barriers may have shaped who applied, or several mechanisms may operate together. Listing these possibilities does not show that they are equally likely. It identifies what further evidence should investigate.
The population also matters. Applicants are not all people who might have wanted the job. A requirement, referral system or expectation about treatment could affect who enters the pool. A fair comparison within the pool would not necessarily establish equal access to applying. Conversely, a population disparity cannot automatically be attributed to the final recruiter if earlier stages differ.
It helps to draw the sequence: awareness, application, eligibility screening, interview, offer and acceptance. A gap can emerge or change at each stage. The relevant denominator changes accordingly. An offer rate among interviewees answers a different question from an offer rate among all applicants. A careful account follows people through the sequence and explains which stage a comparison addresses.
A correspondence experiment creates a narrower comparison
Bertrand and Mullainathan's published 2004 study sent 4,870 fictitious résumés to Boston and Chicago newspaper job advertisements during 2001–2002. Names associated with perceived racial categories were randomly assigned within the résumé design. The reported callback rates were 9.65 percent for White-sounding names and 6.45 percent for African-American-sounding names, with 2,435 résumés in each condition. The study examined callbacks in selected occupations and cities, not job offers or lifetime earnings. Original article, methods and Table 1.
Random assignment is the crucial difference from simply comparing two existing groups of applicants. In a correspondence experiment, researchers construct application materials and vary a signal according to a planned assignment. This can make the signal less entangled with the other attributes that differ across real applicants. The resulting comparison targets how the application is treated under the tested conditions.
Names can carry more than one association. The authors explicitly examine concerns about social-background signals; their checks do not make every possible interpretation of a name disappear. The causal object is the assigned name signal in the study's design. It is not an experimentally assigned intrinsic racial property of a person, and it cannot by itself disclose every decision-maker's conscious motive.
The distinction is essential for interpreting the evidence honestly. An experiment can demonstrate unequal responses to a signal while leaving questions about perception and mechanism. It can also provide strong evidence within a selected setting without becoming a timeless estimate for every occupation, medium or employer. Limitations define the claim; they do not erase the result.
Read the numerical result without changing its meaning
Subtracting the reported rates gives 3.20 percentage points. Dividing 9.65 by 6.45 gives about 1.496, so the higher callback rate is about 49.6 percent larger relative to the lower rate. These are two descriptions of the same comparison, using different scales. Calling the difference “3.2 percent” would be ambiguous, while calling it a 49.6-percentage-point gap would be wrong.
The baseline matters. A relative increase of about one half sounds large, but the absolute rates also show that most applications in both conditions received no callback. Neither observation cancels the other. Reporting both absolute and relative differences helps the reader understand the size of the outcome without relying on a single rhetorically convenient scale.
Do not convert this result into a claim that one group must submit an exact number of applications to obtain a job. A callback is not a job; applications to different employers need not be independent; and repeated applications by an actual person may differ from the study's materials. An inverse-rate calculation can be a mathematical illustration under assumptions, but it would not establish a real person's required search effort.
The study also does not authorize a reader to repeat it casually by sending deceptive applications. Our exercise uses the published results. The learning objective is to understand assignment, measurement and scope, not to conduct an unreviewed field experiment that consumes employers' time or involves unsuspecting participants.
Unequal treatment can operate through rules and routines
Consider another fictional office that recruits only through employees' personal referrals. Every referred applicant receives identical formal treatment. If employees' networks reach some groups more readily than others, the recruitment channel can still distribute access unevenly. The mechanism here is the structure of information and entry, rather than a different response to two otherwise equivalent applications at the final screen.
This distinction helps separate interpersonal treatment from institutional processes. A rule can produce unequal opportunities without requiring every participant to hold the same intention. But the existence of an unequal outcome does not establish which routine produced it. We need evidence connecting the rule to who heard about the opening, who was referred and who applied.
Suppose the office adds a public advertisement and a structured assessment. If the applicant pool changes, that is evidence about access to the recruitment channel. If interview decisions change among comparable applicants, that addresses another stage. A single before-and-after total can mix both changes, along with changes in the labor market or the jobs offered. The evaluation should track the intended pathways separately.
Rules can also interact with earlier inequalities. Requiring unpaid availability at short notice may select for people with financial buffers and flexible obligations. Whether the requirement is necessary for the work, whether alternatives exist and how its burden is distributed are empirical and evaluative questions. A neutral-looking sentence on a form is not a complete account of the conditions under which people must meet it.
Statistical adjustment needs a causal question
Suppose a researcher adjusts a hiring comparison for education and experience. The resulting estimate describes a difference conditional on those measured variables and the model used. It does not automatically become the one true measure of inequality. Education and experience may be relevant qualifications, but they can also reflect earlier access and treatment that a broader question seeks to understand.
If the question concerns the final screening decision among applicants with similar measured preparation, adjustment may help. If the question concerns the entire pathway from early opportunities to employment, holding preparation fixed can remove part of that pathway from the comparison. Neither question is illegitimate. The mistake is to answer one and announce that the other has been settled.
Measurement quality matters as well. A broad credential category may not capture specific skills; years of experience may conceal different tasks. An unexplained residual can include omitted variables and measurement error as well as unequal treatment. Conversely, explaining a gap statistically through a variable does not establish that the process producing that variable was fair.
This is why a residual should not be casually renamed discrimination, and a smaller adjusted gap should not be casually renamed proof of equal opportunity. A stronger argument combines a clearly specified question with evidence suited to the proposed mechanism. Experiments, institutional records, historical analysis and interviews can illuminate different parts of the process.
Variation within groups is part of the explanation
A group average can conceal differences by occupation, age, location, gender or other circumstances. That variation matters when a mechanism operates only in a particular setting. The relevant question is not whether a category has one uniform experience, but how classification interacts with situations and institutions. A study should define its population closely enough to make the comparison interpretable.
At the same time, dividing a sample into many small subgroups can create unstable estimates. An apparently dramatic difference based on a few observations may have substantial uncertainty. Searching every possible subgroup and reporting only the most striking one increases the risk of mistaking chance variation for a stable pattern. Good analysis states how comparisons were chosen and what the data can support.
Individual counterexamples do not settle population questions. A successful applicant from a disadvantaged group does not prove the absence of a barrier; an unsuccessful applicant does not establish its presence. A barrier can alter probabilities while allowing exceptions. The task is to estimate and explain a pattern, then recognize the limits of applying that pattern to a particular person's biography.
Missing classifications require attention too. If an office reports rates only for applicants who supplied demographic information, the result covers that subset. Nonresponse may vary with trust, the application channel or other circumstances. It would be misleading to assign everyone with a missing response to a residual group and interpret that group as a coherent identity. Report the missing share, explain how it is handled and consider whether the observed subset supports the intended comparison.
What a responsible conclusion sounds like
Return to the recruitment office. A defensible first statement reports the categories, applicant population, stage and rates. A second statement identifies plausible mechanisms and the evidence currently available. A third distinguishes what has been established from what remains open. This sequence allows a strong conclusion where the design supports one without making every observed gap carry more causal information than it contains.
For the historical census, the conclusion concerns how an institution classified people. For the correspondence experiment, it concerns differential responses to randomized name signals in a specific hiring context. For the fictional referral system, it concerns a proposed access mechanism that would require investigation. These examples belong together because each makes the production and consequences of categories visible, while preserving the difference between a record, an experiment and an illustration.
The moral evaluation remains important. We may object to unequal treatment, restricted access or inherited disadvantage for stated reasons. But ethical urgency does not remove the need to identify the process accurately. Understanding where a barrier operates is part of designing a response capable of changing it.
Application
Prepare a 400-word evidence note comparing the fictional recruitment disparity with the published correspondence experiment. State the population, classification, outcome, denominator and causal comparison for each. Calculate the experiment's absolute and relative callback differences. Then propose evidence that would test the fictional referral mechanism without submitting deceptive applications or collecting private information.
Check your understanding: Does a 3.20-percentage-point callback difference equal a 3.20-percent relative difference, and does a remaining gap after adjusting for education automatically identify discrimination at the final hiring stage?
Expected answer: No. The reported callback rates differ by about 49.6 percent relative to the lower rate, while their absolute difference is 3.20 percentage points. An adjusted residual depends on the variables, measurement and model; it does not alone identify a mechanism. A causal claim needs a comparison suited to the stage and question being investigated.