The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Trials

Reading a trial's population section before its results

Solved
Solved by t.abubakar in post #8
Everything in post #4 holds. The case it does not cover is the one I have. A treatment-policy estimand asks what happens to people assigned to a strategy, including those who abandon it. A hypothetical estimand asks what would have happened had everyone continued. Both are legitimate and they give different numbers.…

Jump to the accepted answer →

W
WickramasingheTL2Member4 May 2026#1

On the subject in the title: Reading a trial's population section before its results Working notes rather than a conclusion.

Reading STEP 2 (Lancet, 2021) for the population rather than the effect, which I have not done properly before.

The baseline table is more restrictive than the way the trial gets discussed here. Several of the questions in this category come from people who would not have been enrolled.

What is the honest way to describe what the trial says to somebody outside its population?

50 likes 3mo
ZO
z.onwukaTL25 May 2026#2

Worth separating two things that the opening post runs together.

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

Worth reading the earlier posts in this thread before acting on mine.

0 likes 3mo
CD
c.dahlbergTL26 May 2026#3
Wickramasinghe, post #1: On the subject in the title: Reading a trial's population section before its results Working notes rather than a conclusion. Reading STEP 2 ( Lancet , 2021) for the population rather than the effect, which I have not done properly before. The baseline table is more restrictive than the way the trial gets discussed here. Several of the… Go to post

Trial duration determines what can be observed. A weight-change trajectory at 40 weeks and at 72 weeks are different observations and both get quoted as the result.

2 likes in reply to #1 3mo
JB
j.baptistaTL27 May 2026 · edited#4

Composite endpoints should be read component by component. A composite driven entirely by its softest component is a different finding from one where the components move together.

8 likes 3mo
VK
v.krastevTL27 May 2026#5

Taking post #2 at face value and following it one step further.

The first question about any trial is what it set out to estimate, not what it found. Once the estimand is on the table the rest of the discussion is tractable.

A single observation, in a thread that deserves better than single observations.

20 likes 3mo
MA
m.achebeTL28 May 2026#6

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

I would put the burden of proof on the interesting explanation, not the dull one.

0 likes 3mo
P
PSkarbekTL3Regular8 May 2026#7
c.dahlberg, post #3: Trial duration determines what can be observed. A weight-change trajectory at 40 weeks and at 72 weeks are different observations and both get quoted as the result. Go to post

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

0 likes in reply to #3 3mo
TA
t.abubakarTL2 Solution9 May 2026#8

Everything in post #4 holds. The case it does not cover is the one I have.

A treatment-policy estimand asks what happens to people assigned to a strategy, including those who abandon it. A hypothetical estimand asks what would have happened had everyone continued. Both are legitimate and they give different numbers.

That is what the documentation says. What happens in practice is usually close.

12 likes 3mo
SO
s.ostergaardTL29 May 2026#9

Reading the supplementary appendix is where most of the real information is, and it is where almost nobody goes. The baseline table alone answers half the generalisability questions asked here.

The reasoning is more useful than the number, which is why I have shown it.

14 likes 3mo
BV
bias_varianceTL4Biostatistician10 May 2026#10
j.baptista, post #4: Composite endpoints should be read component by component. A composite driven entirely by its softest component is a different finding from one where the components move together. Go to post

Safety findings from a trial powered for efficacy are underpowered by construction. Absence of a signal in that setting is weak evidence of absence.

28 likes in reply to #4 3mo
AS
a.salcedoTL3Regular10 May 2026#11

Post #7 describes the usual case. This is about the unusual one.

A trial that answers a slightly different question from the one you have is the normal situation rather than a failure of the trial. The skill is describing the gap precisely.

19 likes 3mo
KH
ka.haddadTL211 May 2026#12

Adding the measurement that post #11 says would settle it.

Absolute and relative effects answer different questions. Write down the event rate in each arm and the difference between them; everything quotable is derived from those two numbers.

Marking that as an opinion rather than a finding.

8 likes 3mo
GI
g.ibarraTL211 May 2026 · edited#13
v.krastev, post #5: Taking post #2 at face value and following it one step further. The first question about any trial is what it set out to estimate, not what it found. Once the estimand is on the table the rest of the discussion is tractable. A single observation, in a thread that deserves better than single observations. Go to post

Funding and trial conduct should be stated and are a weak predictor of anything on their own. Design quality is the stronger signal and it is checkable.

That is the honest state of it as of this week.

0 likes in reply to #5 3mo
MY
m.yilmazTL212 May 2026#14

Registration before enrolment, with the primary endpoint declared, is what makes outcome switching detectable. Checking the registry against the paper takes five minutes and is worth doing.

This is where my knowledge stops and I would rather mark the edge than blur it.

0 likes 3mo
GH
g.haalandTL3Regular12 May 2026#15

Effect sizes in a trial population reflect adherence achieved under trial conditions, which is generally better than adherence outside them.

26 likes 3mo
ID
il.dumitruTL212 May 2026#16
s.ostergaard, post #9: Reading the supplementary appendix is where most of the real information is, and it is where almost nobody goes. The baseline table alone answers half the generalisability questions asked here. The reasoning is more useful than the number, which is why I have shown it. Go to post

Picking up post #15: that is the part I would want checked first.

Nothing in a trial report is medical advice about an individual, and the gap between a population estimate and a person is exactly where clinical judgement lives.

12 likes in reply to #9 3mo
FA
f.abrahamsenTL2Member13 May 2026#17
j.baptista, post #4: Composite endpoints should be read component by component. A composite driven entirely by its softest component is a different finding from one where the components move together. Go to post

I had written a reply contradicting post #15 and deleted it. Here is what survived.

A treatment-policy estimand asks what happens to people assigned to a strategy, including those who abandon it. A hypothetical estimand asks what would have happened had everyone continued. Both are legitimate and they give different numbers.

The short answer was in the first line; everything after is the working.

2 likes in reply to #4 3mo
EC
e.coelhoTL213 May 2026#18

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

The right answer here may simply be that it has not been measured.

0 likes 3mo
SK
s.kimaniTL214 May 2026#19

Entry criteria, run-in periods and the self-selection of people willing to enter a multi-year trial all narrow the population. That is how internal validity is bought and it constrains generalisation.

The step people skip is the one I have spelled out.

8 likes 2mo
EP
e.piresTL214 May 2026#20

Discontinuation handling is the methodological detail that most changes a result and gets the least attention. Read how missing data was imputed before reading the effect size.

2 likes 2mo
SC
s.chowdhuryTL3Regular14 May 2026#21

Post #20 answers the question as asked. The question underneath it is different.

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

I would put this at better than even and not much better.

12 likes 2mo
I
IRenaudinTL2Member15 May 2026#22

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

Adding this to the thread rather than to the wiki, because I am not confident enough for the wiki.

25 likes 2mo
NK
n.kaufmannTL215 May 2026#23

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

The general answer and the answer for your case may diverge here.

0 likes 2mo
AK
a.kwiatkowskiTL2Member15 May 2026 · edited#24
a.salcedo, post #11: Post #7 describes the usual case. This is about the unusual one. A trial that answers a slightly different question from the one you have is the normal situation rather than a failure of the trial. The skill is describing the gap precisely. Go to post

Worth separating two things that post #22 runs together.

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

I have no interest in any supplier named above.

1 like in reply to #11 2mo
NB
n.boatengTL216 May 2026#25

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

7 likes 2mo
CR
crossover_reviewTL3Regular16 May 2026#26

Saving this. It is the version I will quote when the question comes round again.

18 likes 2mo
KB
k.batistaTL217 May 2026#27
s.ostergaard, post #9: Reading the supplementary appendix is where most of the real information is, and it is where almost nobody goes. The baseline table alone answers half the generalisability questions asked here. The reasoning is more useful than the number, which is why I have shown it. Go to post

Picking up post #24: that is the part I would want checked first.

An open-label trial is not worthless and its subjective endpoints deserve more scepticism than its objective ones. That is a graded judgement rather than a verdict.

I have separated what I observed from what I concluded, which does not always happen.

0 likes in reply to #9 2mo
MM
methods_marginTL3Regular17 May 2026#28
g.haaland, post #15: Effect sizes in a trial population reflect adherence achieved under trial conditions, which is generally better than adherence outside them. Go to post

On post #27 — agreed on the reasoning, with one qualification.

Absolute and relative effects answer different questions. Write down the event rate in each arm and the difference between them; everything quotable is derived from those two numbers.

I would treat the number as indicative rather than as a measurement.

0 likes in reply to #15 2mo
RN
r.nakamuraTL217 May 2026#29

Discontinuation handling is the methodological detail that most changes a result and gets the least attention. Read how missing data was imputed before reading the effect size.

24 likes 2mo
DT
dexa_twice_yearlyTL3Regular18 May 2026#30

Entry criteria, run-in periods and the self-selection of people willing to enter a multi-year trial all narrow the population. That is how internal validity is bought and it constrains generalisation.

I would want to see it done twice before believing it once.

0 likes 2mo