The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Research Methods · Statistics · continued

Correlation in a self-tracked dataset: what it can support posts 121–150

This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.

JV
j.vogelTL225 Aug 2024#121

Multiple testing inflates the false-positive rate in a way that is entirely predictable and entirely correctable. The correction should be declared in advance.

I would rather post the uncertainty than round it away.

23 likes 23mo
TH
TL4_HalvorsenTL4Leader · Journal club25 Aug 2024#122
buffer_margin, post #44: Bayesian and frequentist analyses answer different questions and both are legitimate. What matters is that the reader knows which is on offer. Anyone who has looked at this more carefully, please correct the record. Go to post

I would call the community position on correlation likely rather than established, and I would be comfortable defending that hedge.

0 likes in reply to #44 23mo
HL
h.lindqvistTL225 Aug 2024#123

Post #121 answers the question as asked. The question underneath it is different.

Summarising the correlation thread so far, since it is long and the answer is buried: the first reply has the method, the fourth has the correction to it, and the rest is people agreeing at length.

1 like 23mo
AD
appeals_deskTL3Regular25 Aug 2024#124

That reframing is the whole thing. The facts I already had.

6 likes 23mo
AK
a.kirchnerTL225 Aug 2024 · edited#125
e.kuipers, post #4: Following this. I have the same question and no better information than the first post. Go to post

Baseline imbalance in a randomised trial is expected by chance and adjusting for it post hoc is a choice that should have been pre-specified.

30 likes in reply to #4 23mo
LI
l.ibarraTL2Regular25 Aug 2024#126
policy_reader, post #32: Where I part company with post #28, and it is a narrow parting. The reason correlation is hard to answer is that the obvious measurement and the relevant quantity are not the same thing, and substituting one for the other is silent. Go to post

A p-value is the probability of data at least this extreme given the null hypothesis. It is not the probability the hypothesis is false, and almost every plain-language gloss gets that backwards.

Not a strong opinion, just a consistent one.

0 likes in reply to #32 23mo
AD
a.delgadoTL225 Aug 2024#127

A note on how correlation gets discussed rather than on correlation itself: the confident posts get the replies and the careful ones get ignored, and the careful ones have been right more often.

3 likes 23mo
MH
m.haddadTL2Regular25 Aug 2024#128

Everything in post #126 holds. The case it does not cover is the one I have.

Worth stating the null on correlation before we explain it: the observation may be nothing. That possibility deserves a sentence and usually does not get one.

10 likes 23mo
KM
k.marchandTL225 Aug 2024 · edited#129
k.radich, post #119: I read post #115 twice before replying, because I had assumed the opposite. Where the correlation reasoning breaks down for me is the step from the group result to the individual case. That step is almost never argued for. Go to post

Picking up post #128: that is the part I would want checked first.

Percentages of small denominators should be reported with the denominator. Two out of three is not sixty-seven per cent in any useful sense.

0 likes in reply to #119 23mo
FT
fr.translation_moTL2Translator · FR25 Aug 2024#130

On post #126 — agreed on the reasoning, with one qualification.

The arithmetic on correlation is the easy part and it is where the errors are, which is an uncomfortable combination. Show your working and someone will catch it.

1 like 23mo
PN
p.novakTL225 Aug 2024 · edited#131

A request rather than an answer: could whoever has the primary source for correlation post it? I have seen the claim three times this month and each version had lost a qualifier.

15 likes 23mo
TK
t.kulkarniTL3Regular25 Aug 2024#132

Post #129 is the version of this I will quote in future. One addition.

A distribution shown is worth ten summary statistics. Where a paper shows individual data points, read those first.

Someone will know this better than I do and I hope they say so.

5 likes 23mo
AV
a.vermeulenTL225 Aug 2024#133

Post #129 describes the usual case. This is about the unusual one.

I changed my mind about correlation after someone here asked me for the source and I could not produce one. That is worth saying out loud because it is the ordinary way it happens.

1 like 23mo
RJ
r.jhannsdttirTL3Regular25 Aug 2024#134
f.rasmussen, post #31: Post #30 is the version of this I will quote in future. One addition. Correlation: I would want to see the raw numbers rather than the summary before agreeing. Summaries lose exactly the information that would settle this. Go to post

Two claims get bundled together under correlation and they need separating. The descriptive one — this is what was observed — is usually well supported. The causal one — this is why — usually is not.

Almost every disagreement in threads like this one dissolves once you say which of the two you are making.

0 likes in reply to #31 23mo
HK
h.kimaniTL225 Aug 2024#135

Right — I had this wrong and I am glad to have read it before it mattered.

21 likes 23mo
B
BDraganovTL2Member25 Aug 2024#136

Building on post #133 rather than restating it.

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

9 likes 23mo
JP
j.palaciosTL225 Aug 2024#137

Where I have landed on correlation, having got it wrong once in public: the direction is clear, the magnitude is not, and anyone quoting a precise magnitude has borrowed it from somewhere that did not measure it.

2 likes 23mo
HA
h.almeidaTL2Member25 Aug 2024#138
PSkarbek, post #46: Acknowledging rather than arguing. The reasoning holds as far as I can follow it. Go to post

Worth separating correlation as a question about the compound from correlation as a question about the documentation. They get answered by different people and only one of them is answerable here.

0 likes in reply to #46 23mo
NC
n.chowdhuryTL225 Aug 2024#139
a.ibarra, post #110: I read post #108 twice before replying, because I had assumed the opposite. Correlation is a question about a distribution, not about a value, and treating it as a value is what produces the confident wrong answers. Go to post

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

28 likes in reply to #110 23mo
IT
integrator_traceTL2Member25 Aug 2024#140

Two sentences on correlation and then I will stop, because the rest is speculation and the thread is better without mine.

What is documented is narrow. What is inferred from it is broad. The gap between them is where every argument here lives.

14 likes 23mo
ID
i.dumitruTL225 Aug 2024#141

A note on scope: what I am saying about correlation applies to the case in the first post and I would not extend it further without checking.

20 likes 23mo
EL
endpoint_lineTL3Regular25 Aug 2024#142
isotonic_sheet, post #56: Coming back to post #54, because the follow-up matters more than the original answer. Two things can be true about correlation at once: the mechanism is plausible and the evidence for the size of the effect is thin. Most of the argument here is people defending the first against attacks on the second. Go to post

Survivorship in a self-reporting population biases every aggregate produced from it, and the bias is in the flattering direction.

The step people skip is the one I have spelled out.

0 likes in reply to #56 23mo
SD
s.demirTL225 Aug 2024#143

Picking up post #142: that is the part I would want checked first.

Adding a null result on correlation. I looked, carefully, and found nothing, and null results deserve posting precisely because they never are.

2 likes 23mo
OP
o.pasqualeTL1Member25 Aug 2024#144

Bayesian and frequentist analyses answer different questions and both are legitimate. What matters is that the reader knows which is on offer.

Where I would look next, rather than where I would stop.

9 likes 23mo
PO
p.onwukaTL225 Aug 2024 · edited#145

The most useful thing anyone has posted about correlation in this category was a table of what had been measured and by whom. That is what I would want again.

14 likes 23mo
LA
l.aaltonenTL325 Aug 2024#146
KK
k.karlsenTL225 Aug 2024#147

Multiple testing inflates the false-positive rate in a way that is entirely predictable and entirely correctable. The correction should be declared in advance.

The conclusion is tentative; the arithmetic underneath it is not.

0 likes 23mo
HN
h.nicolaidesTL3Regular25 Aug 2024#148

That matches what I have seen, for whatever a single anecdote is worth.

5 likes 23mo
WM
w.moreauTL226 Aug 2024#149

Building on post #147 rather than restating it.

The bit of correlation that nobody enjoys is that the answer changes depending on what you are trying to decide with it. Say what the decision is and the thread will converge.

0 likes 23mo
Z
ZieglerTL3Regular26 Aug 2024#150

Post #149 put the caveat in the right place and I want to underline it.

I have three months of notes on correlation and the honest summary is that the trend is real and the week-to-week numbers are noise. I nearly drew the opposite conclusion from the first fortnight.

0 likes 23mo