Correlation in a self-tracked dataset: what it can support posts 91–120
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
Collapsed as off-topic by two members at trust level 3 or above
No notes. Posting so the count is not one.
Post #91 answers the question as asked. The question underneath it is different.
Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.
Reporting rather than recommending, on correlation. What happened is above. Whether it should have is a different question and not one I am qualified to answer.
Checked the correlation claim against the primary source this morning. It survives, with a narrower scope than the version quoted here. Posting the narrower scope.
Post #96 is right about the mechanism and I think understates the practical bit.
Correlation came up in a thread eighteen months ago and was answered well. I cannot find it, which is itself the problem, so here is the reconstruction.
Coming back to post #94, because the follow-up matters more than the original answer.
Percentages of small denominators should be reported with the denominator. Two out of three is not sixty-seven per cent in any useful sense.
Same experience here, different supplier, so it is at least not unique to one of them.
The failure mode on correlation is boring rather than dramatic. It is almost always the step everyone assumes was done correctly because it is too simple to get wrong.
Rounding and significant figures carry information about precision. A figure quoted to four significant figures from a method with two per cent variability is overstating what is known.
This is the version I would want a new member to read first.
On post #100 — agreed on the reasoning, with one qualification.
I have been on both sides of the correlation argument in this category within eighteen months, which should tell you how strong the evidence for either side is.
Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.
A qualification I should have led with rather than closed on.
A p-value is the probability of data at least this extreme given the null hypothesis. It is not the probability the hypothesis is false, and almost every plain-language gloss gets that backwards.
I have separated what I observed from what I concluded, which does not always happen.
The useful distinction on correlation is between what was measured and what was inferred from it. Both end up in the same sentence and only one of them has error bars.
Bayesian and frequentist analyses answer different questions and both are legitimate. What matters is that the reader knows which is on offer.
It is the sort of thing that seems obvious in retrospect and was not at the time.
Speaking only to correlation as I have actually seen it, rather than as it is usually described: the effect is real, it is smaller than the thread suggests, and the variance between people is larger than the effect.
I disagree with the framing of correlation above, and I think it is a substantive disagreement rather than a terminological one. Setting out why, so it can be checked.
The reasoning depends on an assumption that is doing a lot of work and is never stated. If the assumption holds, the conclusion follows. I do not think it holds generally.
Collapsed as off-topic by two members at trust level 3 or above
Answering the question post #115 raises rather than the one it answers.
Correlation is one of those subjects where the general answer and the answer for a specific case diverge, and the thread will go in circles until someone says which one is being asked for.
Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.
If it helps: the failure mode here is usually boring rather than dramatic.