Correlation in a self-tracked dataset: what it can support posts 121–150
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
I would call the community position on correlation likely rather than established, and I would be comfortable defending that hedge.
Post #121 answers the question as asked. The question underneath it is different.
Summarising the correlation thread so far, since it is long and the answer is buried: the first reply has the method, the fourth has the correction to it, and the rest is people agreeing at length.
That reframing is the whole thing. The facts I already had.
Baseline imbalance in a randomised trial is expected by chance and adjusting for it post hoc is a choice that should have been pre-specified.
A p-value is the probability of data at least this extreme given the null hypothesis. It is not the probability the hypothesis is false, and almost every plain-language gloss gets that backwards.
Not a strong opinion, just a consistent one.
Picking up post #128: that is the part I would want checked first.
Percentages of small denominators should be reported with the denominator. Two out of three is not sixty-seven per cent in any useful sense.
On post #126 — agreed on the reasoning, with one qualification.
The arithmetic on correlation is the easy part and it is where the errors are, which is an uncomfortable combination. Show your working and someone will catch it.
Post #129 is the version of this I will quote in future. One addition.
A distribution shown is worth ten summary statistics. Where a paper shows individual data points, read those first.
Someone will know this better than I do and I hope they say so.
Post #129 describes the usual case. This is about the unusual one.
I changed my mind about correlation after someone here asked me for the source and I could not produce one. That is worth saying out loud because it is the ordinary way it happens.
Two claims get bundled together under correlation and they need separating. The descriptive one — this is what was observed — is usually well supported. The causal one — this is why — usually is not.
Almost every disagreement in threads like this one dissolves once you say which of the two you are making.
Where I have landed on correlation, having got it wrong once in public: the direction is clear, the magnitude is not, and anyone quoting a precise magnitude has borrowed it from somewhere that did not measure it.
Worth separating correlation as a question about the compound from correlation as a question about the documentation. They get answered by different people and only one of them is answerable here.
P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".
Two sentences on correlation and then I will stop, because the rest is speculation and the thread is better without mine.
What is documented is narrow. What is inferred from it is broad. The gap between them is where every argument here lives.
Survivorship in a self-reporting population biases every aggregate produced from it, and the bias is in the flattering direction.
The step people skip is the one I have spelled out.
Bayesian and frequentist analyses answer different questions and both are legitimate. What matters is that the reader knows which is on offer.
Where I would look next, rather than where I would stop.
Collapsed as off-topic by two members at trust level 3 or above
Coming back to post #144, because the follow-up matters more than the original answer.
I would rather this thread reach "we do not know" about correlation than reach a confident answer that nobody can support when asked.
That matches what I have seen, for whatever a single anecdote is worth.
Post #149 put the caveat in the right place and I want to underline it.
I have three months of notes on correlation and the honest summary is that the trend is real and the week-to-week numbers are noise. I nearly drew the opposite conclusion from the first fortnight.