QUICK-REFERENCE BOX — Don't Misread the p-value
- The p-value is the probability of data this extreme if the null hypothesis were true — not the probability the null is true.
- Type I error (α): false positive — saying there's an effect when there isn't.
- Type II error (β): false negative — missing a real effect. Power = 1 − β.
- A 95% CI that excludes the null value (1 for ratios, 0 for differences) means p < 0.05.
- Statistical significance ≠ clinical importance — big studies make trivial effects "significant."
1. Overview
The Core Idea
Hypothesis testing asks whether observed data are compatible with a null hypothesis of "no effect." The p-value measures that compatibility, the confidence interval shows the range of plausible effects (and their precision), and Type I/II errors and power tell you how likely the study was to reach a correct conclusion.
2. The p-value
What It Is (and Isn't)
- It is the probability of results as extreme as (or more than) observed, assuming the null hypothesis is true.
- It is not the probability that the null hypothesis is true, nor the probability the finding is due to chance in the everyday sense.
- The α threshold (commonly 0.05) is a convention for the acceptable false-positive rate — arbitrary, and best interpreted in context.
3. Type I & Type II Errors
| Truth: no effect | Truth: real effect |
| Study says "effect" | Type I error (α, false positive) | Correct (true positive) |
| Study says "no effect" | Correct (true negative) | Type II error (β, false negative) |
4. Statistical Power
Detecting a True Effect
- Power = 1 − β — the probability of detecting an effect that truly exists.
- Power rises with larger sample size, larger effect size, lower variability, and a higher α.
- An underpowered "negative" study cannot distinguish "no effect" from "not enough data."
5. Confidence Intervals
Range & Precision
- A 95% CI gives a range of plausible values for the true effect; a wider CI means less precision (often a smaller sample).
- If a 95% CI excludes the null (1 for RR/OR, 0 for a difference), the result is significant at p < 0.05.
- CIs convey both significance and clinical magnitude — more informative than a p-value alone.
6. Statistically Significant vs Clinically Important
The Distinction
A very large study can make a clinically trivial difference statistically significant. Conversely, an important effect can miss significance in a small, underpowered trial. Always judge the size of the effect and its confidence interval, not just whether p crossed 0.05.
7. Key Pearls
High-Value Points
- p-value = probability of the data given the null, not the probability of the null.
- Type I = false positive (α); Type II = false negative (β); Power = 1 − β.
- Increase power mainly by increasing sample size.
- A CI excluding the null = significant; CI width = precision.
- Significant ≠ important — weigh effect size and CI.
8. Common Mistakes to Avoid
Misreadings & Better Practice
Frequent errors with p-values and CIs.
| Mistake | Why it's wrong | Better reading |
| "p = 0.04 means 4% chance the null is true." | Misdefines the p-value. | It's the probability of the data given the null. |
| "Non-significant = no effect." | May be underpowered. | Check power and the CI. |
| Reporting only p-values. | Hides magnitude/precision. | Report effect size with a CI. |
9. Board-Style High-Yield Summary
Key Takeaways
- p-value = probability of data this extreme if the null is true.
- Type I (α) false positive; Type II (β) false negative; Power = 1 − β (raise with sample size).
- A 95% CI excluding the null → p < 0.05; CI width shows precision.
- CIs beat p-values for conveying magnitude and precision.
- Statistical significance is not clinical importance.
10. References
- 1.Wasserstein RL, Lazar NA. The ASA statement on p-values: context, process, and purpose. Am Stat. 2016.
- 2.Gardner MJ, Altman DG. Confidence intervals rather than P values. BMJ. 1986.
- 3.Guyatt G, et al. Users' Guides to the Medical Literature.
Back to Clinical Guideline Hubs