Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.
Coming back to: Sample size calculations, read backwards from the published number posts 61–89
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
Adding a reference point for Sample size calculations. Mine is a single case, collected without controls, and I am posting the method alongside it so it can be discounted appropriately.
Post #60 answers the question as asked. The question underneath it is different.
Composite endpoints should be read component by component. A composite driven entirely by its softest component is a different finding from one where the components move together.
That matches what I have seen, for whatever a single anecdote is worth.
I have no financial interest in anything named in this thread and I want to say so before I comment on Sample size calculations, because it is the sort of subject where it matters.
On post #67 — agreed on the reasoning, with one qualification.
Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.
Coming back to post #69, because the follow-up matters more than the original answer.
Discontinuation handling is the methodological detail that most changes a result and gets the least attention. Read how missing data was imputed before reading the effect size.
I am describing what is, rather than arguing for what should be.
Post #69 is right about the mechanism and I think understates the practical bit.
Sample size calculations looks different depending on whether you are reading the primary literature or the summaries of it, and the difference is not in our favour.
Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.
Anyone with a larger sample, please post it.
Where I would push back on the Sample size calculations consensus is the confidence, not the direction. The direction looks right. The confidence is borrowed.
Collapsed as off-topic by two members at trust level 3 or above
I read post #73 twice before replying, because I had assumed the opposite.
Having read the whole Sample size calculations thread before replying: the question in the first post has not actually been answered yet, and three of us have answered a nearby one instead.
Absolute and relative effects answer different questions. Write down the event rate in each arm and the difference between them; everything quotable is derived from those two numbers.
Taking post #76 at face value and following it one step further.
An honest declaration on Sample size calculations: I have a prior here and it is strong enough that you should weight what I say downward. Stating it rather than hiding it.
Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.
Scoping that to what I have actually seen rather than what I have read.
The failure mode on Sample size calculations is boring rather than dramatic. It is almost always the step everyone assumes was done correctly because it is too simple to get wrong.
Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.
On post #80 — agreed on the reasoning, with one qualification.
A treatment-policy estimand asks what happens to people assigned to a strategy, including those who abandon it. A hypothetical estimand asks what would have happened had everyone continued. Both are legitimate and they give different numbers.
Post #82 is the version of this I will quote in future. One addition.
I would call the community position on Sample size calculations likely rather than established, and I would be comfortable defending that hedge.
Composite endpoints should be read component by component. A composite driven entirely by its softest component is a different finding from one where the components move together.
Caveat: everything above assumes the paperwork is what it says it is.
That is clearer than the version I had in my head. Thank you.
One more thing on Sample size calculations that took me far too long to see: the two figures people quote are not measuring the same quantity. Once you notice that, the apparent contradiction disappears.
Building on post #86 rather than restating it.
Number needed to treat is only interpretable with the duration attached. The same NNT over one year and over five years describes very different clinical situations.
This topic was referenced in
- How to read a forest plot, properly, from scratchEvidence › Trials · 81 replies
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
[2026 update] Reading a supplementary appendix and finding the interesting part
On the subject in the title: Reading a supplementary appendix and finding the interesting part Working notes rather than a conclusion. Collecting what is known about Reading a supplementary appendix in one…
|
+39 | 44 | 8.2k | 5mo |
|
How to read a forest plot, properly, from scratch
The question in the title: How to read a forest plot, properly, from scratch I will give what I have already checked below so nobody repeats it. Reading STEP 4 ( JAMA , 2021) for the population rather than…
|
+73 | 81 | 17k | 7mo |
|
Follow-up: What a statistical analysis plan adds that the paper does not
What a statistical analysis plan adds that the paper does not I have a specific reason for asking rather than idle curiosity, and the context is below. A follow-up question about statistical analysis plan…
|
2 | 7.2k | 3mo | |
|
Composite endpoints and the component doing the work
Composite endpoints and the component doing the work Writing it up because I had to work it out twice and would rather nobody else did. Collecting what is known about composite endpoints in one place, because…
|
+3 | 7 | 33k | 13h |
|
Interim analyses and stopping rules — does this still hold?
Interim analyses and stopping rules — does this still hold? — that is the question, and I have not found it answered plainly anywhere I have looked. Interim analyses and stopping rules, from the point of view…
|
+30 | 34 | 29k | 13mo |
Related topics — sharing the tags estimand, discontinuation & dropout, surrogate endpoints
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Second pass at: Reading a combination trial: attributing effect to components
Second pass at: Reading a combination trial: attributing effect to components — setting out what I have, and where I think it stops being reliable. The question about Reading a combination trial that I…
|
+39 | 43 | 7.3k | 12mo |
|
Journal club: STEP 8 and the fairness of the comparator dose — a second dataset
Posting this under the heading it deserves: Journal club: STEP 8 and the fairness of the comparator dose — a second dataset Everything below is what sits behind that. Posting a small dataset on STEP 8. It is…
|
+18 | 22 | 656 | 13h |
|
[2026 update] Journal club: FLOW and the renal composite, component by component
Journal club: FLOW and the renal composite, component by component Writing it up because I had to work it out twice and would rather nobody else did. A question about what the numbers mean rather than what…
|
2 | 21k | 9d | |
|
Journal club: SURMOUNT-4 and continuation versus withdrawal — a second dataset
Posting this under the heading it deserves: Journal club: SURMOUNT-4 and continuation versus withdrawal — a second dataset Everything below is what sits behind that. Posting a small dataset on SURMOUNT-4 and…
|
+58 | 64 | 17k | 13mo |
|
Primary endpoint hierarchies and why order matters
Primary endpoint hierarchies and why order matters Writing it up because I had to work it out twice and would rather nobody else did. I was wrong about primary endpoint hierarchies in a thread last spring and…
|
4 | 308 | 3mo |