enlumn.
Practical Judgment and Intellectual Honesty

Did this actually cause that?

A fictional workshop introduces an optional mentoring session. At the end of the month, participants who attended mentoring complete more projects than those who did not. The organizers announce that mentoring improved performance. A skeptical volunteer replies that motivated people were more likely to attend in the first place.

Both statements concern the same pattern, but they explain it differently. The organizers imagine an effect of mentoring. The volunteer imagines a difference between the people who selected it. Both processes could operate together. A table showing the association does not yet tell us how much work each explanation does.

Causal reasoning asks what would change if we changed something. That makes it essential to practical judgment. We do not merely want to know which people attend mentoring. We want to know whether offering, improving or requiring it will help. To answer, we need a comparison that represents what would have happened otherwise.

The comparison we want cannot be observed directly

For a particular participant, imagine two versions of the same month. In one, the person receives mentoring. In the other, everything relevant is the same except that they do not. The difference in completed projects would reveal the effect for that person under those conditions.

We cannot observe both versions of that person's actual month. Once one happens, the other remains counterfactual: an alternative outcome used in reasoning, rather than an observed event. The problem is not solved by comparing the person with themselves a month earlier, because other things may have changed. Nor is it solved merely by comparing them with someone else, because the two people may differ.

This does not make causal knowledge impossible. It explains why the design of a comparison matters. We seek evidence that makes one observed group or period a credible guide to the unobserved alternative. Each design earns credibility in particular ways and remains vulnerable in others.

Start by stating the intervention and outcome. “Mentoring helps” is underspecified. Are we asking about an invitation, actual attendance, a particular mentor, or a requirement? Are we measuring completed projects, quality, confidence, retention or safe practice? Different questions can have different answers without contradiction.

A third variable can explain an association

Suppose experienced participants both seek mentoring and complete more projects. Experience can influence attendance and performance. It therefore offers a route by which the two measured variables become associated even if mentoring adds nothing. In this setting, experience is a possible confounder.

The objection does not require an accusation that the organizers manipulated the data. A sincere observation can support a misleading inference. People often enter programs for reasons connected with the outcome the program aims to change. Those reasons have to be considered rather than wished away.

Now suppose less experienced participants are especially likely to seek help. Mentoring could genuinely improve their performance while they still complete fewer projects than experienced nonparticipants. The raw comparison might hide a benefit. Confounding does not always exaggerate an effect; it can also obscure one or change its apparent direction.

The practical lesson is therefore stronger than “correlation is not causation.” That slogan names a problem but does not investigate it. Ask which factors could influence both participation and the outcome, in which direction, and whether the available comparison addresses them. NIST's introduction to experimental design makes the central distinction between observing relationships and designing changes to study causes. NIST: Experiments and Experimental Design.

Work through a misleading total

Consider this invented record from two levels of workshop experience. Completion means finishing an agreed project within the month. The counts are chosen to make the arithmetic inspectable.

Experience Mentoring: completed / total No mentoring: completed / total
New participants 18 / 30 4 / 10
Experienced participants 9 / 10 24 / 30
All participants 27 / 40 28 / 40

Among new participants, completion is 60 percent with mentoring and 40 percent without. Among experienced participants, it is 90 percent with mentoring and 80 percent without. Yet the combined rate is 67.5 percent with mentoring and 70 percent without. The aggregate comparison appears to favor no mentoring.

The composition explains the reversal. Most mentored participants are new, and most unmentored participants are experienced. Combining the groups mixes a comparison about mentoring with a difference in experience. The aggregate is not arithmetically wrong. It answers a question whose interpretation requires knowing who is included.

Do the within-experience comparisons prove that mentoring causes improvement? No. They address one specified difference, but other differences may remain. Perhaps the new participants who attend mentoring have more free time. Perhaps mentors select projects that are easier to finish. Breaking down the table improves the inquiry without magically completing it.

This is how to use a statistical complication constructively. Identify the distortion, calculate the relevant comparison, and state what the correction does and does not establish. Invoking a famous paradox without doing the accounting would add vocabulary while leaving the decision unchanged.

What random assignment contributes

Suppose the workshop has more volunteers requesting a new mentoring format than available places. With their informed agreement and a fair allocation procedure, it randomly assigns some to an immediate offer and others to a later offer. Random allocation can make the groups comparable without requiring the organizers to identify and measure every prior difference.

The reasoning is about the assignment process. If assignment does not depend on motivation, experience or anticipated success, those factors do not systematically determine who receives the offer. Chance imbalances remain possible, especially in a small group. Randomization is a method for controlling a source of bias, not a guarantee that every realized group is identical.

The study still needs a clearly measured outcome and follow-up of both groups. If disappointed participants disappear from the records, a comparison of those who remain can become misleading. If observers know the assignments and judge quality differently, the outcome assessment can be biased. The word “randomized” does not replace inspection of the whole procedure.

It also matters whether we compare assignment or attendance. Some people offered mentoring may not attend. Some people in the later group may seek help elsewhere. Comparing everyone by the original assignment estimates the effect of the offer under those conditions. Comparing only those who chose to attend can reintroduce selection. There is no universal substitute for stating precisely which effect was measured.

When an experiment is unavailable

Many important questions cannot be answered by an appropriate experiment. We cannot randomly assign a city's historical migration or a person's childhood. Some interventions would be unacceptable to impose. We then need other evidence and a candid account of its limitations.

One option is to compare changes over time in a group exposed to a change with changes in a relevant comparison group. If both groups had previously followed similar patterns, their divergence may be informative. But the interpretation depends on whether some other change affected one group differently. The comparison must be defended historically or substantively, not only calculated.

Another option is to examine a rule that creates different treatment for otherwise similar cases near a threshold. Such a design can be useful when the rule is genuinely followed and people cannot readily manipulate their position. Its result may apply most directly near that threshold rather than to everyone. The attraction of the method does not remove the need to understand how the rule actually operates.

You need not become a specialist in every design to ask productive questions. What supplies the comparison? Why should it represent what would otherwise have happened? What could make that resemblance fail? Which people, places and periods does the result actually concern? These questions identify the structure of the inference beneath the label.

A mechanism helps, but a good story is not enough

Mentoring might improve performance by catching errors early, providing feedback or helping participants choose manageable projects. These mechanisms make the proposed effect intelligible. They suggest intermediate observations worth collecting, such as whether participants revise plans after feedback.

Yet a plausible mechanism does not establish that a benefit occurred. Feedback could be poor, attendance too brief, or the project measure inappropriate. We can explain how a bridge is intended to carry weight without proving that this bridge was built correctly. Similarly, an intervention's theory of operation is one kind of evidence, not a substitute for outcomes.

The reverse is also possible: a credible comparison may show an effect before its mechanism is fully understood. That does not make the comparison worthless. It means the account is incomplete in a different way. We should distinguish confidence that an effect occurred from confidence about the pathway that produced it.

Mechanisms become especially useful when considering transfer. If mentoring works because it provides equipment unavailable elsewhere, the effect may shrink in a workshop where everyone already has that equipment. If it works through skilled feedback, a larger program that cannot recruit suitable mentors may disappoint. Understanding the process helps identify which conditions matter when a result moves to a new setting.

Look for what disappeared from the sample

Suppose the organizers survey only people still attending at the end of the month. Those who found mentoring discouraging may have left. Their absence can make the remaining group look unusually successful. A result based on survivors needs an account of who left and why.

Now suppose the workshop publicizes its three most impressive graduates. Their stories may be true and worth hearing. They do not reveal the typical outcome or the proportion who benefited. Selection can occur when people enter a program, when records are retained, and when results are displayed. Each stage changes what the audience is allowed to see.

This applies to your own learning as well. Remembering the occasions when a confident intuition succeeded while forgetting the occasions when it failed makes intuition look better than the full record would. You do not need a theory of universal human bias to see the logical problem: the selected cases do not represent the cases needed for the conclusion.

A useful practice is to ask for the denominator. Out of how many participants did these successful cases emerge? How many began, completed, declined or disappeared? What happened to the people who received the alternative? Counts cannot answer every question about quality, but they can expose a story whose apparent strength depends on missing failures.

Return to a decision, with a narrower claim

The workshop may decide to continue mentoring while improving evaluation. That decision can be reasonable even if the initial causal claim was too strong. Participants may value the support; the cost may be manageable; a better comparison may be feasible next term. Uncertainty about a precise effect does not require abandoning every useful activity.

But the public explanation should match the evidence: “Mentored participants completed more projects in the initial records, although selection makes the size of the effect uncertain. We will compare clearer outcomes and track all participants next term.” If the within-group table points elsewhere, report that instead. A method is valuable only if it is allowed to change the story.

The chapter's central skill is to make the alternative outcome visible. Whenever someone says a program, habit, policy or product caused an improvement, ask what comparison supports the word “caused.” Then examine the comparison rather than arguing from your enthusiasm for the proposed cause. The next chapter asks how to rely on other people's answers when you cannot perform the whole investigation yourself.

Application

A school reports that students who use its optional study room receive higher marks. It concludes that requiring every student to use the room will raise everyone's marks. Identify the observational claim, the causal claim and the additional claim about changing the policy.

An explained answer: The observed association concerns students who chose the room. The causal claim attributes their better marks to room use rather than prior motivation, available time or other differences. The policy claim assumes a requirement will have similar effects, even though attendance, crowding and the mix of students may change. Better evidence needs a credible comparison and attention to the difference between voluntary use and a universal requirement.

Numerical check: In the workshop table, explain the aggregate reversal without calling either set of figures false. The mentoring group contains more beginners, whose baseline completion is lower. The totals combine different mixtures; examining experience changes the comparison, while leaving possible unmeasured differences unresolved.

Next chapter →