Correlation in a self-tracked dataset: what it can support posts 61–90
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
This is the sort of exchange that makes the archive worth searching.
Collapsed as off-topic by two members at trust level 3 or above
I read post #63 twice before replying, because I had assumed the opposite.
Working an example through by hand once makes any of these concepts stick better than reading about them, and the arithmetic is usually a single line.
If the premise is wrong, everything after it is decoration.
Post #63 answers the question as asked. The question underneath it is different.
I read the earlier replies on correlation twice before writing this, because I had assumed the opposite and wanted to be sure I was disagreeing with what was said rather than what I expected.
The honest answer on correlation is that it depends, and the useful part is the list of what it depends on. Four items, in rough order of how much they matter.
Most people get the first two right and then argue about the fourth.
Rounding and significant figures carry information about precision. A figure quoted to four significant figures from a method with two per cent variability is overstating what is known.
The claim is narrower than it sounds, and deliberately so.
My position on correlation is current rather than settled. I have revised it once already and I expect to again, so treat it accordingly.
Collapsed as off-topic by two members at trust level 3 or above
Post #70 is the version of this I will quote in future. One addition.
Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.
That holds under the stated conditions and I have stated them.
The version of correlation that I was taught turned out to be a teaching simplification. Useful, and not true in the way I had assumed it was.
One caution on correlation: everything above assumes the underlying documentation is what it claims to be. That assumption is doing real work and is rarely stated.
Answering the question post #72 raises rather than the one it answers.
Multiple testing inflates the false-positive rate in a way that is entirely predictable and entirely correctable. The correction should be declared in advance.
Noting that I have skin in this question and have tried to discount for it.
What I want from this correlation thread is the list of things that would need to be true for the claim to hold. If we can write that list, we can check it.
The thing about correlation that took me longest to accept is that a plausible mechanism is not evidence of an effect. It is a reason to look, not a result.
Narrowing post #74, because the general version has more than one answer.
Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.
The honest answer is that it depends, and here is what it depends on.
Everything in post #76 holds. The case it does not cover is the one I have.
Baseline imbalance in a randomised trial is expected by chance and adjusting for it post hoc is a choice that should have been pre-specified.
That has held every time I have looked, which is not the same as always.
Collapsed as off-topic by two members at trust level 3 or above
Coming back to post #76, because the follow-up matters more than the original answer.
Two questions I would want answered before drawing anything from the correlation data above: how were the cases selected, and what happened to the ones that dropped out.
Post #77 put the caveat in the right place and I want to underline it.
Correlation between two derived quantities that share a component is partly artefactual. It is a common trap in analyses of ratios.
The variance between people here is larger than the effect being discussed.
Time-to-event analysis handles differing follow-up in a way a simple proportion cannot, which is why event rates and Kaplan-Meier estimates can differ.
Speaking for myself and not for anyone else who has posted here.
Before the thread moves on from correlation — what is the sample size behind the claim? I am not being difficult; I have seen the same figure quoted from an n of four and from an n of four hundred.
Adding the boring version of correlation, because the interesting version keeps getting posted and the boring one is usually right.
Check the ordinary explanations, in order, and stop when one of them accounts for what you are seeing. Most of the time the second one does.
P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".
A weak preference rather than a position.
Source for the correlation figure, since it was asked for. It is in the discussion rather than the abstract, which is why the version circulating is stronger than the paper is.
Reading the surrounding paragraph is worth the two minutes. The authors are more careful than their summarisers.
Where an analysis was changed after seeing the data, the honest thing is to report both and say which was pre-specified.
That is where I would start, not where I would stop.
Confirming post #89 from a second method, which matters more than confirming it from a second person.
What would change my mind on correlation is a second dataset collected by someone with no stake in the first. Until then I hold it loosely and I would rather say so than pretend to more.