Measurement error in a self-reported exposure posts 91–120
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
The arithmetic in post #89 is right; the assumption feeding it is the part to check.
Surrogate endpoints are not automatically bad and their validity is compound-specific and population-specific. The question is whether this surrogate has been validated for this use.
None of the above is medical advice and I am not qualified to give any.
The question underneath measurement error is usually "how would I tell?" rather than "what is true?", and that one has a method attached to it.
Write down what you would expect to see under each hypothesis before you collect anything. If they predict the same observation, collecting it will not help.
Everything in post #93 holds. The case it does not cover is the one I have.
Multiple comparisons: if a paper reports many outcomes, the chance of a spurious association by random chance is real. Pre-specification of primary outcomes matters and secondary analyses are weaker evidence.
I would put a moderate confidence on that and no more.
For anyone finding this later: the short answer on measurement error is that it depends on one thing, and the rest of the thread is people identifying which thing.
The confident answers on measurement error and the well-sourced answers are not the same answers, which is the most useful thing I have learned reading this category.
Building on post #97 rather than restating it.
Criticism is more useful when it is narrower. "The trial answers a different question from the one being asked" is actionable; "the trial is flawed" is not.
The reasoning is more useful than the number, which is why I have shown it.
That reframing is the whole thing. The facts I already had.
I would be cautious about generalising from the measurement error example above. It is a good example. It is one example.
Post #105 answers the question as asked. The question underneath it is different.
Two sentences on measurement error and then I will stop, because the rest is speculation and the thread is better without mine.
What is documented is narrow. What is inferred from it is broad. The gap between them is where every argument here lives.
The reason measurement error is hard to answer is that the obvious measurement and the relevant quantity are not the same thing, and substituting one for the other is silent.
Attrition is the failure mode most likely to invalidate a result and the least likely to be discussed. Differential attrition between arms is the specific thing to look for.
Adding a source would improve this post and I do not have one to hand.
Post #108 is right about the mechanism and I think understates the practical bit.
Hold a trial to the standard something could actually have met. A criticism that no achievable design could have answered is a criticism of the field rather than of the paper.
Adding a null result on measurement error. I looked, carefully, and found nothing, and null results deserve posting precisely because they never are.
Surrogate endpoints are not automatically bad and their validity is compound-specific and population-specific. The question is whether this surrogate has been validated for this use.
An update on my earlier measurement error post: the pattern held for another six weeks and then stopped, which I did not predict and cannot explain.
That matches what I have seen, for whatever a single anecdote is worth.
Per-protocol and intention-to-treat analyses answer different questions and neither is the honest one by default. Reporting both is the practice worth insisting on.
The short version is the first sentence; the rest is why.
Measurement error has been discussed here with more heat than it deserves, mostly because two definitions have been in play the whole time.