How much does the evidence tell us?
A repair cooperative buys a device that flags defective components. The manufacturer reports that it catches nine out of ten defective components and wrongly flags one out of ten sound components. A component triggers the warning. How likely is it to be defective?
“Ninety percent” is a tempting answer. It is also incomplete. The answer depends on how common defects were among the components being tested. A warning can be informative without making the warned-about condition likely. We will work through this hypothetical example carefully, because the distinction matters far beyond testing machines.
The purpose is not to attach numbers to every uncertain thought. It is to learn how evidence changes a comparison between possibilities. The numbers make that comparison visible. When exact probabilities are unavailable, the same questions can still improve your reasoning: What was plausible before this observation? How expected would the observation be under each explanation? What relevant information have I left out?
Conditional questions point in different directions
Begin with the manufacturer's first statement. Among defective components, nine out of ten trigger a warning. That answers a question about the device's response when the defect is already specified. It does not directly answer the question about the component when the warning is specified.
Compare two ordinary sentences: most bicycles have wheels; most things with wheels are bicycles. The first can be true while the second is false. Reversing the direction of a relationship changes the question. Probability statements have the same problem, even when the two directions sound almost identical in conversation.
Write the distinction in words before using notation. “Probability of a warning given a defect” differs from “probability of a defect given a warning.” The word “given” identifies the information with which we begin. In the first question, the component is defective and we ask about the device. In the second, the device has sounded and we ask about the component.
The false-warning rate matters because sound components can also produce the observation. If there are many more sound components than defective ones, even a relatively small fraction of sound components can contribute a large share of all warnings. We need to count both routes to the observed result.
Build a population you can count
Assume that exactly 100 of a batch's 1,000 components are defective. These are stipulated teaching numbers, not measured claims about a real product. Apply the device's rates to this batch: it flags 90 defective components and misses 10. Of the 900 sound components, it wrongly flags 90 and leaves 810 unflagged.
| Component condition | Warning | No warning | Total |
|---|---|---|---|
| Defective | 90 | 10 | 100 |
| Sound | 90 | 810 | 900 |
| Total | 180 | 820 | 1,000 |
Now select a component from the warning group. That group contains 180 components, of which 90 are defective. The proportion defective is 90 divided by 180: one half. Under the stated assumptions, a warning raises the probability of a defect from 10 percent to 50 percent.
This is a substantial change. Calling the warning “only fifty-fifty” can obscure how much information it supplied. Before the test, sound components outnumbered defective ones nine to one. After a warning, the two possibilities are equally represented. Evidence need not establish a conclusion beyond doubt to matter a great deal.
It also follows that a warning does not justify describing the component as certainly defective. A sensible response might be inspection, retesting with a different method, or temporary separation from the usable stock. Which response is justified depends on the consequences and costs. The probability calculation is an input to that decision, not the decision itself.
Change the starting frequency
Suppose a different supplier sends 1,000 components, of which 500 are defective. Keep the same test behavior. The device flags 450 defective components and 50 sound ones. Now 450 of the 500 warnings identify defects: 90 percent.
Nothing about the device improved. The population changed. Defects were much more common before testing, so they account for a larger share of warnings afterward. The starting frequency is often called the base rate. Ignoring it can make the same evidence appear to have a fixed meaning in every setting.
Move in the other direction. Suppose only 10 of 1,000 components are defective. The device flags nine of those and wrongly flags 99 of the 990 sound components. Nine of the 108 warnings identify a defect, approximately 8.3 percent. Most warnings are false in this population, even though the device still catches nine out of ten defects.
The result does not mean that testing is necessarily useless. It means that the warning must be interpreted with the population and purpose in view. A cheap follow-up inspection may be worthwhile. Discarding every flagged component could be wasteful. Using the test as a final verdict without considering its false warnings would be a different decision again.
Base rates are not a license to ignore individual evidence. The test is individual evidence, and it changes the probability. The lesson is that starting information and new information must be combined. Neither should simply erase the other.
A compact way to express the comparison
The warning is nine times as likely for a defective component as for a sound one: 0.9 divided by 0.1 equals nine. This ratio is a likelihood ratio. It measures how strongly this observation favors one specified possibility over the other.
In the first batch, prior odds of a defect were 1 to 9. Multiply those odds by nine and the posterior odds become 1 to 1. Odds of 1 to 1 correspond to a probability of one half. This is the odds form of Bayes' rule: posterior odds equal prior odds multiplied by the likelihood ratio of the evidence.
You need not memorize notation to retain the reasoning. A warning is more expected if the component is defective, so it favors that explanation. But the strength of that favor must be combined with the fact that defects were initially uncommon. The table and the formula describe the same accounting.
Be careful about what the ratio compares. It does not say that a defect is nine times as likely after every imaginable piece of information. It compares the probability of this warning under these two conditions, using the stipulated rates. A different device, a different defect or a different operating environment might require different rates.
The formula is exact only when its inputs and definitions fit the question. Elegant arithmetic cannot rescue an irrelevant population or a manufacturer whose performance figures do not apply to your use. Good numerical reasoning includes examining where the numbers came from.
Two reports may contain only one piece of evidence
The cooperative tests a flagged component again. The same device, using the same measurement, produces the same warning. Can you multiply the odds by nine again? Only if the relevant assumptions about the relationship between the tests are justified.
Suppose both warnings arise from a harmless surface mark that consistently confuses the device. The second warning adds little: once the first warning and its underlying cause are known, repetition is expected whether the component is defective or sound. Treating the two warnings as independent would exaggerate the information.
For a simpler case, imagine that you photocopy the first test report and place the copy beside the original. You now have two documents but one observation. No one would reasonably double their confidence because the folder became thicker. Repeated measurements can be more useful than photocopies, but their value depends on whether they bring genuinely additional information.
The same issue arises with testimony. Five articles may all rely on one unnamed witness. Three friends may have heard the story from the same person. Agreement across reports is more informative when the routes to the conclusion are meaningfully independent. Counting endorsements without tracing their dependence can turn repetition into apparent corroboration.
Independence is not all or nothing in everyday inquiry. A second specialist may inspect the same evidence with a different method. That can add something even if both depend on the original sample. State what is shared and what is new, rather than invoking “multiple sources” as a guarantee.
Absence can be evidence, but only against an expectation
What does no warning tell us in the first batch? There are 820 unflagged components, of which ten are defective. The defective proportion is approximately 1.22 percent. No warning lowers the probability of a defect substantially but does not make it zero.
The absence matters because the device usually warns when a defect is present. If the device were unplugged, silence would tell us nothing about the component. The contrast explains a general rule: failing to observe something counts against a possibility when the observation was reasonably expected if that possibility were true.
Suppose someone says a meeting occurred, but no minutes exist. The missing record matters differently in an organization that reliably records every meeting and one that rarely records anything. You must examine the process that could have produced the evidence. “There is no record” is not a complete argument until we know whether a record should exist and whether the search could find it.
This also limits what you can conclude from a field visit. Not seeing a bird does not establish that the species is absent from the area. Time of day, season, habitat and observation method affect the chance of detection. You do not need to quantify every factor to recognize the missing step between “I did not observe it” and “it is not there.”
Avoid precision you have not earned
Outside a teaching example, the defect rate may be estimated from a small sample. Test performance may vary. The relevant population may be disputed. Reporting a posterior probability to several decimal places could suggest a certainty the inputs do not support.
One response is sensitivity analysis: examine how the conclusion changes under several plausible inputs. We already did this by changing the defect rate. If your recommended action stays the same across a defensible range, you have learned something useful. If a small change reverses the recommendation, the uncertain input deserves attention.
Another response is to keep the judgment qualitative. You might say that an observation moderately favors one explanation, while acknowledging that you cannot estimate the strength accurately. That can be more honest than inventing a number. Quantification should clarify the evidence, not decorate uncertainty with an appearance of measurement.
There is a further limit. A calculation can compare only the possibilities represented in it. If the cooperative considers merely “defective component” and “sound component under ordinary conditions,” but the device is malfunctioning, the model is missing an important explanation. An unexpected pattern across many tests might give a reason to inspect the device itself.
Good updating therefore includes two tasks: changing confidence within an existing set of explanations and reconsidering whether that set is adequate. Neither endless model revision nor blind attachment to the first model is a substitute for attention to the actual problem.
Probability does not determine what someone is owed
Return to the decision about a flagged component. A 50 percent chance of a costly failure may justify a cheap inspection. A 50 percent chance of a harmless cosmetic flaw may not justify throwing away an expensive part. Different consequences can support different actions at the same probability.
When people are involved, additional constraints matter. A suspicion about a colleague is not permission to announce an accusation as fact. The need to investigate can coexist with duties of fairness, privacy and accurate description. A numerical estimate cannot cancel those duties merely by being written as a percentage.
The intellectual achievement of this chapter is to distinguish an observation's informativeness from certainty, and certainty from the threshold for action. Next we will examine a different challenge: even when two things reliably occur together, what would justify saying that one caused the other?
Application
In a hypothetical batch of 1,000 components, 200 are defective. A device flags 80 percent of defective components and 10 percent of sound components. Construct the table and calculate the proportion defective among flagged components.
An explained answer: The device flags 160 defective components and 80 sound ones. There are 240 warnings, and 160/240 equals two thirds, approximately 66.7 percent. The 80 percent figure describes warnings among defects; it is not the answer to defects among warnings. A second identical warning cannot automatically be treated as independent evidence.
Transfer question: Three websites report the same allegation and all cite one interview. What should you investigate before treating their agreement as strong corroboration? Identify their source dependence, whether any made independent checks, and whether the interview itself supplies evidence for the actual allegation.