Home / Clinical Guideline Hubs / Chapter 15.4
Section 15 — Biostatistics & Evidence-Based Medicine Board-prep reference v1.0 · July 2026

Chapter 15.4 — Hypothesis Testing, p-values & Confidence Intervals

What "significant" really means · A board-prep reference

Educational reference. Statistical concepts for exam preparation, not patient-specific clinical instructions.
QUICK-REFERENCE BOX — Don't Misread the p-value

1. Overview

The Core Idea

Hypothesis testing asks whether observed data are compatible with a null hypothesis of "no effect." The p-value measures that compatibility, the confidence interval shows the range of plausible effects (and their precision), and Type I/II errors and power tell you how likely the study was to reach a correct conclusion.

2. The p-value

What It Is (and Isn't)
  • It is the probability of results as extreme as (or more than) observed, assuming the null hypothesis is true.
  • It is not the probability that the null hypothesis is true, nor the probability the finding is due to chance in the everyday sense.
  • The α threshold (commonly 0.05) is a convention for the acceptable false-positive rate — arbitrary, and best interpreted in context.

3. Type I & Type II Errors

Truth: no effectTruth: real effect
Study says "effect"Type I error (α, false positive)Correct (true positive)
Study says "no effect"Correct (true negative)Type II error (β, false negative)

4. Statistical Power

Detecting a True Effect
  • Power = 1 − β — the probability of detecting an effect that truly exists.
  • Power rises with larger sample size, larger effect size, lower variability, and a higher α.
  • An underpowered "negative" study cannot distinguish "no effect" from "not enough data."

5. Confidence Intervals

Range & Precision
  • A 95% CI gives a range of plausible values for the true effect; a wider CI means less precision (often a smaller sample).
  • If a 95% CI excludes the null (1 for RR/OR, 0 for a difference), the result is significant at p < 0.05.
  • CIs convey both significance and clinical magnitude — more informative than a p-value alone.

6. Statistically Significant vs Clinically Important

The Distinction

A very large study can make a clinically trivial difference statistically significant. Conversely, an important effect can miss significance in a small, underpowered trial. Always judge the size of the effect and its confidence interval, not just whether p crossed 0.05.

7. Key Pearls

High-Value Points
  • p-value = probability of the data given the null, not the probability of the null.
  • Type I = false positive (α); Type II = false negative (β); Power = 1 − β.
  • Increase power mainly by increasing sample size.
  • A CI excluding the null = significant; CI width = precision.
  • Significant ≠ important — weigh effect size and CI.

8. Common Mistakes to Avoid

Misreadings & Better Practice

Frequent errors with p-values and CIs.

MistakeWhy it's wrongBetter reading
"p = 0.04 means 4% chance the null is true."Misdefines the p-value.It's the probability of the data given the null.
"Non-significant = no effect."May be underpowered.Check power and the CI.
Reporting only p-values.Hides magnitude/precision.Report effect size with a CI.

9. Board-Style High-Yield Summary

Key Takeaways
  • p-value = probability of data this extreme if the null is true.
  • Type I (α) false positive; Type II (β) false negative; Power = 1 − β (raise with sample size).
  • A 95% CI excluding the null → p < 0.05; CI width shows precision.
  • CIs beat p-values for conveying magnitude and precision.
  • Statistical significance is not clinical importance.

10. References

Back to Clinical Guideline Hubs