Evidence-Based Medicine, Hierarchy of Evidence and Types of Research
Evidence-based medicine (EBM) is the integration of the best available research evidence with clinical expertise and patient values and preferences to make informed healthcare decisions.
EBM does not mean following published evidence blindly. High-quality care requires interpretation of research in the context of the individual patient, the clinical setting, available resources, potential harms and expected benefits.
Evidence-based medicine is not “cookbook medicine” — evidence must be critically appraised and applied to the right patient.
The Three Components of Evidence-Based Medicine
| Component | Meaning |
|---|---|
| Best available evidence | High-quality and clinically relevant research from appropriate study designs. |
| Clinical expertise | The clinician's ability to diagnose, assess risk, interpret evidence and individualise treatment. |
| Patient values and preferences | Patient goals, expectations, risk tolerance, circumstances and preferences. |
Steps of Evidence-Based Practice
A commonly used framework is the 5 A's of evidence-based practice.
- Ask: formulate a focused and answerable clinical question.
- Acquire: search for the best available evidence.
- Appraise: critically evaluate validity, importance and applicability.
- Apply: integrate the evidence with clinical expertise and patient preferences.
- Assess: evaluate the outcome and the effectiveness of the decision-making process.
Formulating a Clinical Question – PICO
Clinical questions are often structured using the PICO framework.
| Letter | Meaning | Example |
|---|---|---|
| P | Patient / Population / Problem | Elderly patients with intertrochanteric fracture |
| I | Intervention | Intramedullary nail |
| C | Comparison | Sliding hip screw |
| O | Outcome | Reoperation, fixation failure or functional outcome |
Time or study design may sometimes be added, giving frameworks such as PICOT or PICOS.
Types of Clinical Questions
| Question Type | Commonly Preferred Evidence |
|---|---|
| Therapy / Intervention | Randomised controlled trial / systematic review |
| Diagnosis | Diagnostic accuracy study compared with reference standard |
| Prognosis | Prospective cohort study |
| Harm / Risk | Cohort or case-control study |
| Etiology | Cohort or case-control study |
| Patient experience | Qualitative research |
The “best” study design depends on the clinical question being asked.
Broad Types of Research
Clinical research can broadly be divided into:
- Primary research: investigators collect or analyse original patient-level or experimental data.
- Secondary research: investigators synthesise or interpret existing studies.
- Quantitative research: uses numerical measurements and statistical analysis.
- Qualitative research: explores experiences, beliefs, perceptions and behaviours.
- Mixed-methods research: combines quantitative and qualitative approaches.
Observational versus Experimental Research
| Feature | Observational Study | Experimental Study |
|---|---|---|
| Intervention assigned by investigator | No | Yes |
| Examples | Cohort, case-control, cross-sectional | Randomised controlled trial |
| Major limitation | Confounding and selection bias | Cost, feasibility, ethics and external validity |
Descriptive Studies
Descriptive studies describe the occurrence, presentation or characteristics of disease without necessarily testing a causal hypothesis.
Case Report
A detailed description of an individual patient with an unusual diagnosis, presentation, treatment response or complication.
Case reports are useful for generating hypotheses and identifying rare or unexpected events but cannot establish treatment effectiveness or causality.
Case Series
A case series describes a group of patients with a similar disease, intervention or outcome. There is usually no comparison group.
Case series are particularly common in surgical research but are highly vulnerable to selection bias and confounding.
Cross-Sectional Study
A cross-sectional study measures exposure and outcome in a population at a particular point or short period in time.
It is commonly used to estimate prevalence.
Advantages
- Relatively quick.
- Usually inexpensive.
- Useful for describing disease burden.
- Can evaluate multiple variables simultaneously.
Limitations
- Temporal relationship between exposure and outcome may be unclear.
- Generally cannot establish causality.
Case-Control Study
A case-control study begins with the outcome.
Patients with the disease or outcome of interest are identified as cases, and patients without the outcome are selected as controls. Previous exposure is then compared between the two groups.
Case-control study: start with disease → look backward for exposure.
Best Uses
- Rare diseases.
- Outcomes with long latency periods.
- Investigation of multiple potential exposures.
Major Limitations
- Recall bias.
- Selection bias.
- Confounding.
- Direct incidence cannot usually be calculated.
Odds Ratio
The common measure of association in a case-control study is the odds ratio (OR).
| Disease | No Disease | |
|---|---|---|
| Exposed | a | b |
| Not exposed | c | d |
Conceptually: Odds ratio = ad / bc
- OR = 1: no association.
- OR > 1: exposure associated with greater odds of outcome.
- OR < 1: exposure may be protective.
Cohort Study
A cohort study begins with exposure status.
Exposed and unexposed groups are followed or reconstructed over time to compare the occurrence of an outcome.
Cohort study: start with exposure → observe subsequent outcome.
Types
- Prospective cohort: subjects are followed forward from the time of enrolment.
- Retrospective cohort: existing records are used to reconstruct exposure and outcomes that have already occurred.
Advantages
- Can calculate incidence.
- Temporal sequence is usually clearer than in cross-sectional studies.
- Can study multiple outcomes from one exposure.
Limitations
- Confounding.
- Loss to follow-up.
- Can be expensive and time consuming if prospective.
- Inefficient for very rare outcomes.
Relative Risk
The relative risk (RR), or risk ratio, compares the risk of an outcome between an exposed or treated group and a comparison group.
Conceptually:
RR = Risk in exposed group / Risk in unexposed group
- RR = 1: no difference in risk.
- RR > 1: higher risk in the exposed group.
- RR < 1: lower risk in the exposed group.
Randomised Controlled Trial
In a randomised controlled trial (RCT), eligible participants are randomly allocated to different interventions or control groups and outcomes are compared.
Randomisation aims to distribute known and unknown prognostic factors between treatment groups, thereby reducing confounding.
For therapeutic interventions, a well-conducted RCT is generally the strongest individual primary study design for estimating causal treatment effects.
Randomisation and Allocation Concealment
Randomisation
Randomisation refers to assigning participants to intervention groups using an unpredictable random sequence.
Allocation Concealment
Allocation concealment prevents the person enrolling participants from knowing the upcoming treatment assignment.
It protects against selection bias.
Allocation concealment is different from blinding.
Blinding
Blinding means keeping one or more individuals unaware of treatment allocation after participants have been assigned.
Depending on the study, blinding may involve:
- Participants.
- Treating clinicians.
- Outcome assessors.
- Data analysts.
Blinding is often difficult or impossible in orthopaedic surgical trials, but blinded outcome assessment may still reduce bias.
Intention-to-Treat Analysis
In an intention-to-treat (ITT) analysis, participants are analysed according to the group to which they were originally randomised, regardless of crossover, non-compliance or protocol deviation.
ITT helps preserve the benefits of randomisation and usually provides an estimate of the effect of assigning a treatment strategy in real clinical practice.
Common Types of Randomised Trials
| Trial Type | Principle |
|---|---|
| Parallel-group trial | Participants remain in one allocated treatment group. |
| Crossover trial | Participants receive multiple interventions sequentially with appropriate washout. |
| Cluster randomised trial | Groups such as hospitals or clinics rather than individuals are randomised. |
| Factorial trial | Evaluates two or more interventions within the same trial. |
Superiority, Non-Inferiority and Equivalence Trials
| Design | Purpose |
|---|---|
| Superiority | Determine whether one intervention is better than another. |
| Non-inferiority | Determine whether a new intervention is not unacceptably worse than the standard by more than a prespecified margin. |
| Equivalence | Determine whether two interventions have effects sufficiently similar within predefined margins. |
Diagnostic Accuracy Studies
Diagnostic studies evaluate the ability of an index test to identify a target condition compared with an accepted reference standard.
| Disease Present | Disease Absent | |
|---|---|---|
| Test Positive | True positive | False positive |
| Test Negative | False negative | True negative |
Sensitivity and Specificity
Sensitivity
Sensitivity is the proportion of patients with the disease who have a positive test.
Sensitivity = TP / (TP + FN)
Specificity
Specificity is the proportion of patients without the disease who have a negative test.
Specificity = TN / (TN + FP)
Highly sensitive tests are useful for reducing missed disease when negative, while highly specific tests are useful for supporting a diagnosis when positive, although interpretation should consider the entire clinical context and likelihood ratios.
Positive and Negative Predictive Values
Positive Predictive Value
Probability that a patient with a positive test actually has the disease.
PPV = TP / (TP + FP)
Negative Predictive Value
Probability that a patient with a negative test truly does not have the disease.
NPV = TN / (TN + FN)
Unlike sensitivity and specificity, predictive values are strongly influenced by disease prevalence in the tested population.
Likelihood Ratios
Likelihood ratios describe how much a test result changes the probability of disease.
Positive likelihood ratio: Sensitivity / (1 − Specificity)
Negative likelihood ratio: (1 − Sensitivity) / Specificity
Larger positive likelihood ratios and smaller negative likelihood ratios provide greater diagnostic information.
Qualitative Research
Qualitative research explores experiences, perceptions, beliefs and behaviour rather than primarily measuring numerical outcomes.
Common methods include:
- Individual interviews.
- Focus groups.
- Participant observation.
- Document or thematic analysis.
Qualitative research can be particularly valuable for understanding patient experiences, barriers to treatment, rehabilitation adherence and shared decision-making.
Systematic Review
A systematic review uses a predefined and reproducible methodology to identify, select, critically appraise and synthesise all relevant studies addressing a focused research question.
Important components include:
- Focused research question.
- Predefined inclusion and exclusion criteria.
- Comprehensive literature search.
- Transparent study selection.
- Risk-of-bias assessment.
- Structured evidence synthesis.
A systematic review may or may not include a meta-analysis.
Meta-Analysis
Meta-analysis is a statistical technique used to combine quantitative results from multiple sufficiently similar studies to estimate an overall treatment or exposure effect.
Potential advantages include increased statistical precision and improved estimation of treatment effects.
However, the reliability of a meta-analysis depends on the quality and comparability of the included studies.
A meta-analysis of poor-quality studies does not automatically become high-quality evidence.
Heterogeneity in Meta-Analysis
Heterogeneity refers to differences between studies included in a systematic review.
Clinical Heterogeneity
Differences in patients, interventions, outcomes or settings.
Methodological Heterogeneity
Differences in study design, risk of bias or measurement methods.
Statistical Heterogeneity
Variation in effect estimates beyond that expected from sampling error alone.
The I² statistic is commonly used to describe statistical heterogeneity, but it should not be interpreted without considering the clinical and methodological differences between studies.
Forest Plot
A forest plot graphically displays the results of individual studies and the pooled estimate in a meta-analysis.
- The square usually represents the effect estimate from an individual study.
- The size of the square generally reflects the study's statistical weight.
- The horizontal line represents the confidence interval.
- The diamond commonly represents the pooled estimate.
- The vertical line represents the line of no effect.
Publication Bias and Funnel Plot
Publication bias occurs when the probability of a study being published is influenced by the nature or direction of its results, often favouring statistically significant or apparently positive findings.
A funnel plot may be used in meta-analysis to explore possible small-study effects and publication bias when a sufficient number of studies are available.
Funnel-plot asymmetry is not specific for publication bias and may have other explanations.
Narrative Review versus Systematic Review
| Feature | Narrative Review | Systematic Review |
|---|---|---|
| Question | Often broad | Usually focused |
| Search strategy | May not be systematic | Predefined and reproducible |
| Study selection | Potentially subjective | Explicit inclusion/exclusion criteria |
| Risk of bias assessment | Not always performed | Usually formal |
Hierarchy of Evidence
A hierarchy of evidence ranks study designs according to their usual ability to minimise bias and support reliable clinical conclusions.
| Approximate Level | Study Type |
|---|---|
| Highest | High-quality systematic reviews and meta-analyses of appropriate high-quality studies |
| Well-conducted randomised controlled trials | |
| Prospective cohort studies | |
| Retrospective cohort studies | |
| Case-control studies | |
| Cross-sectional studies | |
| Case series | |
| Case reports | |
| Lowest | Expert opinion / mechanistic reasoning without direct clinical evidence |
This hierarchy is a useful general guide, but it is not absolute. The appropriate hierarchy differs according to whether the question concerns therapy, diagnosis, prognosis, harm or patient experience.
Study design alone does not determine evidence quality — a poorly performed RCT may provide less reliable evidence than a well-conducted observational study.
Levels of Evidence
Several organisations use formal levels-of-evidence systems. Exact definitions differ between organisations and according to the clinical question.
A simplified therapeutic hierarchy often resembles:
| Level | Typical Evidence |
|---|---|
| Level I | High-quality RCTs or systematic reviews of high-quality RCTs |
| Level II | Lower-quality RCTs or prospective comparative studies |
| Level III | Retrospective comparative studies / case-control studies |
| Level IV | Case series |
| Level V | Expert opinion / mechanistic reasoning |
The exact scheme being used should always be stated because one organisation's “Level II” may not be identical to another's.
Clinical Practice Guidelines
Clinical practice guidelines are systematically developed recommendations intended to assist clinicians and patients in making decisions about appropriate healthcare.
A high-quality guideline should:
- Use a systematic evidence search.
- Evaluate the quality or certainty of evidence.
- Consider benefits and harms.
- Consider patient preferences and feasibility.
- Declare conflicts of interest.
- Clearly link recommendations to evidence.
GRADE – Certainty of Evidence
The GRADE approach is widely used to assess the certainty of evidence for specific outcomes and to support clinical guideline recommendations.
Certainty is commonly classified as:
- High.
- Moderate.
- Low.
- Very low.
Evidence may be downgraded because of:
- Risk of bias.
- Inconsistency.
- Indirectness.
- Imprecision.
- Publication bias.
Bias
Bias is a systematic error that causes an estimated association or treatment effect to differ from the truth.
| Bias | Example |
|---|---|
| Selection bias | Comparison groups differ systematically because of how participants were selected. |
| Recall bias | Cases remember previous exposure differently from controls. |
| Observer / detection bias | Outcome assessment is influenced by knowledge of treatment group. |
| Performance bias | Groups receive different care apart from the intervention being studied. |
| Attrition bias | Differential loss to follow-up affects results. |
| Publication bias | Positive studies are preferentially published. |
Confounding
A confounder is a factor associated with both the exposure and the outcome that can distort the apparent relationship between them.
For example, if patients receiving one surgical procedure are systematically younger and healthier than those receiving another procedure, age and general health may confound the relationship between procedure and outcome.
Methods used to reduce confounding include:
- Randomisation.
- Restriction.
- Matching.
- Stratification.
- Multivariable statistical adjustment.
- Propensity-score methods in selected observational studies.
Bias versus Random Error
| Feature | Bias | Random Error |
|---|---|---|
| Nature | Systematic error | Chance variation |
| Solved simply by increasing sample size? | No | Often reduced |
| Main concern | Validity | Precision |
A very large study can still give a highly precise but biased answer.
Internal and External Validity
Internal Validity
Internal validity refers to whether the study's result is likely to be correct for the participants who were actually studied.
External Validity
External validity, or generalisability, refers to whether the findings can reasonably be applied to other patients, populations and clinical settings.
A tightly controlled trial may have excellent internal validity while having limited generalisability to routine clinical practice.
P Value
The p value is the probability, under the assumptions of the statistical model and null hypothesis, of observing data at least as incompatible with the null hypothesis as the data actually observed.
Conventionally, a threshold such as p < 0.05 is often used to define statistical significance, although this threshold is arbitrary and should not be interpreted as proof of clinical importance or truth.
Statistical significance does not necessarily mean clinical significance.
Confidence Interval
A confidence interval provides a range of values compatible with the observed data under the statistical model and indicates the precision of an effect estimate.
Narrow confidence intervals indicate greater precision, while wide intervals indicate greater uncertainty.
For relative measures such as RR or OR, a confidence interval crossing 1 includes the null value. For an absolute difference, the null value is usually 0.
Absolute Risk Reduction and Number Needed to Treat
Absolute effects are often more clinically intuitive than relative effects.
Absolute Risk Reduction (ARR) is the difference in event risk between control and treatment groups.
ARR = Control Event Rate − Experimental Event Rate
Number Needed to Treat (NNT) = 1 / ARR
ARR must be expressed as a proportion rather than a percentage when calculating NNT.
Simple Example of NNT
Suppose fixation failure occurs in:
- 20% of the control group.
- 10% of the treatment group.
ARR = 0.20 − 0.10 = 0.10
NNT = 1 / 0.10 = 10
Therefore, approximately 10 patients would need to receive the treatment rather than control to prevent one additional fixation failure over the specified follow-up period, assuming the estimate is valid and applicable.
Number Needed to Harm
When an intervention increases the absolute risk of an adverse event, the reciprocal of the absolute risk increase can be expressed as the Number Needed to Harm (NNH).
NNT and NNH should always be interpreted together with the follow-up period and baseline risk.
Statistical versus Clinical Significance
A difference can be statistically significant while being too small to matter to the patient.
Conversely, a potentially clinically important difference may fail to achieve statistical significance if the study is too small or imprecise.
Clinical interpretation should therefore consider:
- Magnitude of effect.
- Confidence interval.
- Patient-important outcomes.
- Baseline risk.
- Potential harms.
- Cost and burden of intervention.
Minimal Clinically Important Difference
The minimal clinically important difference (MCID) refers broadly to the smallest change in an outcome considered meaningful to patients or clinically important.
MCID is particularly useful when interpreting patient-reported outcome measures, pain scales and functional scores.
A statistically significant improvement smaller than the clinically meaningful threshold may have limited practical relevance.
Type I and Type II Errors
| Error | Meaning | Common Term |
|---|---|---|
| Type I error | Rejecting a true null hypothesis | False positive |
| Type II error | Failing to reject a false null hypothesis | False negative |
- Alpha (α) relates to the tolerated probability of a Type I error.
- Beta (β) relates to the probability of a Type II error.
- Power = 1 − β.
Statistical Power
Statistical power is the probability that a study will detect a specified true effect when that effect exists, under the assumptions used for the calculation.
Power generally increases with:
- Larger sample size.
- Larger true effect size.
- Lower variability.
- Higher accepted Type I error threshold.
An underpowered study may fail to identify a genuine clinically important difference.
Prospective versus Retrospective Research
| Feature | Prospective | Retrospective |
|---|---|---|
| Data collection | Defined before future outcomes occur | Uses previously recorded data/events |
| Control over measurements | Usually greater | Limited by available records |
| Time/cost | Often greater | Often lower |
| Missing data | Can be planned prospectively | Often a significant limitation |
Registry and Database Studies
Registries and large administrative databases allow investigators to study outcomes in large real-world patient populations.
Advantages
- Large sample sizes.
- Ability to study uncommon outcomes.
- Good representation of routine clinical practice.
Limitations
- Confounding.
- Missing data.
- Variable data quality.
- Coding errors.
- Lack of detailed clinical variables.
Basic Science and Translational Research
Basic science research investigates fundamental biological or mechanical processes and may include laboratory, cellular, biomechanical and animal studies.
Translational research attempts to move discoveries from fundamental science toward clinically useful applications and, conversely, use clinical observations to generate scientific questions.
Basic science is essential for understanding mechanisms, but mechanistic plausibility alone does not establish that a treatment improves clinically important patient outcomes.
Clinical Outcomes and Surrogate Outcomes
Patient-Important Outcomes
- Pain.
- Function.
- Quality of life.
- Return to work or sport.
- Reoperation.
- Complications.
- Mortality.
Surrogate Outcomes
Surrogate outcomes are substitute measurements expected to predict clinically meaningful outcomes. Examples in orthopaedics may include certain radiographic measurements, biomarkers or intermediate measures of healing.
Surrogate improvement does not always translate into better patient-important outcomes.
How to Critically Appraise a Research Paper
Critical appraisal can be simplified into three major questions:
1. Is the study valid?
Assess study design, patient selection, randomisation, allocation concealment, blinding, follow-up, measurement methods, confounding and risk of bias.
2. What are the results?
Determine the magnitude of effect, precision, confidence intervals, statistical significance and clinically important outcomes.
3. Can I apply the results?
Consider whether the patient population, intervention, expertise, risks, benefits and clinical setting resemble your own practice.
Questions to Ask When Reading an RCT
- Was the randomisation truly random?
- Was allocation adequately concealed?
- Were important baseline characteristics similar?
- Were participants and outcome assessors blinded where feasible?
- Was follow-up sufficiently complete?
- Were participants analysed in their assigned groups?
- Were co-interventions similar between groups?
- Were clinically important outcomes measured?
- What was the magnitude of benefit or harm?
- How precise was the estimate?
- Are the results applicable to my patient?
Questions to Ask When Reading an Observational Study
- How were participants selected?
- Were exposure and outcome measured reliably?
- Was the comparison group appropriate?
- Could selection bias explain the findings?
- What important confounders were present?
- How were confounders adjusted?
- Was follow-up adequate?
- Was there missing data?
- Is the effect large enough to be clinically meaningful?
- Can causality reasonably be inferred?
Association Does Not Automatically Mean Causation
Observational research can demonstrate an association between an exposure and an outcome, but the observed relationship may be due to:
- True causation.
- Chance.
- Bias.
- Confounding.
- Reverse causation.
Stronger causal inference requires consideration of study design, temporality, consistency, biological plausibility and alternative explanations.
Challenges of Evidence-Based Research in Orthopaedics
Orthopaedic research presents several specific difficulties.
- Blinding surgeons is often impossible.
- Surgical expertise varies between operators.
- Implant technology changes rapidly.
- Fracture patterns and patient populations are heterogeneous.
- Surgeon treatment preference can cause selection bias.
- Learning curves may affect outcomes.
- Rare complications require very large sample sizes.
- Crossovers between treatment groups may occur.
- Radiographic outcomes may not correlate perfectly with patient-reported function.
These limitations make careful interpretation of both randomised and high-quality observational studies especially important in orthopaedic practice.
High-Yield Comparison of Study Designs
| Design | Starting Point | Useful Measure / Purpose | Major Problem |
|---|---|---|---|
| Case report | One patient | Novel observation | No comparator |
| Case series | Group with condition/treatment | Description | No control group |
| Cross-sectional | Population at one time | Prevalence | Temporality unclear |
| Case-control | Outcome | Odds ratio | Recall/selection bias |
| Cohort | Exposure | Incidence, relative risk | Confounding |
| RCT | Random allocation | Treatment effect | Feasibility, cost, applicability |
| Systematic review | Multiple studies | Evidence synthesis | Depends on included studies |
| Meta-analysis | Quantitative study estimates | Pooled effect | Heterogeneity / publication bias |
Exam Pearls
- Evidence-based medicine combines research evidence, clinical expertise and patient preferences.
- PICO = Patient, Intervention, Comparison, Outcome.
- Case-control studies begin with disease/outcome and look for previous exposure.
- Cohort studies begin with exposure and evaluate subsequent outcomes.
- Odds ratio is classically associated with case-control studies.
- Relative risk can be directly calculated in cohort studies and RCTs when risks are observed.
- Cross-sectional studies are particularly useful for estimating prevalence.
- Randomisation primarily helps reduce confounding by distributing prognostic factors between groups.
- Allocation concealment prevents foreknowledge of upcoming treatment assignment and reduces selection bias.
- Blinding occurs after allocation and helps reduce performance and detection biases.
- Intention-to-treat analysis preserves the main advantages of randomisation.
- Sensitivity = TP / (TP + FN).
- Specificity = TN / (TN + FP).
- PPV and NPV depend strongly on disease prevalence.
- A systematic review does not necessarily contain a meta-analysis.
- A meta-analysis is a statistical pooling technique.
- A poor-quality RCT does not automatically provide better evidence than a high-quality observational study.
- Type I error = false-positive conclusion.
- Type II error = false-negative conclusion.
- Statistical significance does not necessarily equal clinical significance.
- NNT = 1 / absolute risk reduction.
- Internal validity concerns whether the result is valid within the study; external validity concerns generalisability.
- Bias is systematic error and cannot simply be eliminated by increasing sample size.
- Association alone does not prove causation.
Common Viva Questions
What is evidence-based medicine?
The integration of the best available evidence with clinical expertise and patient values and preferences.
What does PICO stand for?
Patient or Population, Intervention, Comparison and Outcome.
What is the difference between a cohort and case-control study?
A cohort study begins with exposure and assesses outcome, whereas a case-control study begins with outcome and evaluates previous exposure.
What is the main measure of association in a case-control study?
Odds ratio.
What is the purpose of randomisation?
To create comparable treatment groups by distributing known and unknown prognostic factors by chance and thereby reduce confounding.
What is allocation concealment?
Prevention of foreknowledge of the upcoming treatment allocation during participant enrolment.
What is intention-to-treat analysis?
Analysing participants according to their original randomised group regardless of protocol deviations or crossover.
What is a systematic review?
A structured synthesis of research using a predefined, comprehensive and reproducible methodology.
What is meta-analysis?
Statistical pooling of results from multiple sufficiently comparable studies.
What is confounding?
Distortion of the exposure-outcome relationship by another factor related to both the exposure and outcome.
What is the difference between statistical and clinical significance?
Statistical significance concerns compatibility of data with a statistical hypothesis, whereas clinical significance concerns whether the magnitude of difference matters to patients or clinical practice.
Take-Home Approach
- Ask the right question: structure the clinical uncertainty using PICO.
- Choose the right evidence: the ideal study design depends on whether the question concerns therapy, diagnosis, prognosis or harm.
- Understand the hierarchy: systematic reviews and RCTs often provide strong therapeutic evidence, but study quality remains crucial.
- Recognise observational designs: case-control studies start with outcome; cohort studies start with exposure.
- Look beyond the p value: examine effect size, confidence interval, absolute risk and clinical relevance.
- Search for bias and confounding: these can substantially distort apparently convincing results.
- Assess applicability: determine whether the participants, intervention and clinical setting resemble the patient in front of you.
- Integrate evidence rather than obey it: research findings must be combined with clinical expertise and patient preferences.
The purpose of evidence-based medicine is not simply to find the highest-level paper. It is to identify the most trustworthy evidence for the clinical question, understand its limitations and apply it appropriately to the individual patient.