The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Research Methods · Statistics · continued

Correlation in a self-tracked dataset: what it can support posts 61–90

This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.

YA
y.adebayoTL222 Aug 2024#61

A p-value is the probability of data at least this extreme given the null hypothesis. It is not the probability the hypothesis is false, and almost every plain-language gloss gets that backwards.

Take it as a starting point and not as a specification.

3 likes 23mo
MH
ms_hollowayTL4Mass spectrometrist22 Aug 2024#62
policy_reader, post #32: Where I part company with post #28, and it is a narrow parting. The reason correlation is hard to answer is that the obvious measurement and the relevant quantity are not the same thing, and substituting one for the other is silent. Go to post

This is the sort of exchange that makes the archive worth searching.

0 likes in reply to #32 23mo
MI
m.ibarraTL222 Aug 2024 · edited#63

On correlation I would separate what is worth knowing from what is worth acting on. The first list is long and the second is short, and conflating them is how threads get heated.

23 likes 23mo
SL
s.leclercTL4 Moderator22 Aug 2024#64

The most useful reply I ever got about correlation was a request to state my units. It sounds like pedantry and it has saved me twice.

11 likes 23mo
CA
c.amankwahTL222 Aug 2024#65
TH
TL4_HalvorsenTL4Leader · Journal club22 Aug 2024#66
owen.brady, post #40: A distribution shown is worth ten summary statistics. Where a paper shows individual data points, read those first. Adding it in case it saves somebody the afternoon it cost me. Go to post

Post #63 answers the question as asked. The question underneath it is different.

I read the earlier replies on correlation twice before writing this, because I had assumed the opposite and wanted to be sure I was disagreeing with what was said rather than what I expected.

1 like in reply to #40 23mo
RE
r.ekstromTL223 Aug 2024#67
NHuddleston, post #26: Building on post #23 rather than restating it. A methods point on correlation rather than a substantive one: if the comparison is not like for like, the difference you are measuring is the difference in method. Go to post

The honest answer on correlation is that it depends, and the useful part is the list of what it depends on. Four items, in rough order of how much they matter.

Most people get the first two right and then argue about the fourth.

31 likes in reply to #26 23mo
EF
endo_fellow_rkTL3Endocrinology fellow23 Aug 2024#68

Rounding and significant figures carry information about precision. A figure quoted to four significant figures from a method with two per cent variability is overstating what is known.

The claim is narrower than it sounds, and deliberately so.

16 likes 23mo
LV
l.vukovicTL223 Aug 2024#69

Reading back through, this was answered upthread and I missed it. My fault.

10 likes 23mo
SC
s.chowdhuryTL3Regular23 Aug 2024#70

My position on correlation is current rather than settled. I have revised it once already and I expect to again, so treat it accordingly.

3 likes 23mo
JD
j.delacroixTL323 Aug 2024#71
SO
sa.okonkwoTL223 Aug 2024#72

The version of correlation that I was taught turned out to be a teaching simplification. Useful, and not true in the way I had assumed it was.

0 likes 23mo
AW
a.westergaardTL3Regular23 Aug 2024#73

One caution on correlation: everything above assumes the underlying documentation is what it claims to be. That assumption is doing real work and is rarely stated.

4 likes 23mo
MM
m.malinowskiTL223 Aug 2024#74
sharps_bin, post #16: Worth separating two things that post #15 runs together. What I can speak to on correlation is narrow, so I will keep it narrow rather than generalising from it. Beyond that boundary I do not know. Go to post

Answering the question post #72 raises rather than the one it answers.

Multiple testing inflates the false-positive rate in a way that is entirely predictable and entirely correctable. The correction should be declared in advance.

Noting that I have skin in this question and have tried to discount for it.

12 likes in reply to #16 23mo
ML
m.lindqvistTL223 Aug 2024 · edited#75

What I want from this correlation thread is the list of things that would need to be true for the claim to hold. If we can write that list, we can check it.

19 likes 23mo
RR
r.restrepoTL223 Aug 2024#76

The thing about correlation that took me longest to accept is that a plausible mechanism is not evidence of an effect. It is a reason to look, not a result.

0 likes 23mo
CC
c.correiaTL223 Aug 2024#77
n.kuusela, post #8: Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p Go to post

Narrowing post #74, because the general version has more than one answer.

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

The honest answer is that it depends, and here is what it depends on.

2 likes in reply to #8 23mo
ME
me.eriksenTL223 Aug 2024#78
m.guerrero, post #25: Two people in this thread mean different things by correlation and are disagreeing about the definition while believing they are disagreeing about the facts. Worth pausing to define it. Go to post

Everything in post #76 holds. The case it does not cover is the one I have.

Baseline imbalance in a randomised trial is expected by chance and adjusting for it post hoc is a choice that should have been pre-specified.

That has held every time I have looked, which is not the same as always.

8 likes in reply to #25 23mo
KS
k.salinasTL223 Aug 2024#79

Careful with the language on correlation. "Not detected" and "not present" are different findings and the first is a statement about the method.

4 likes 23mo
DF
d.fontaineTL223 Aug 2024#80
C
chromatogramTL4Analytical chemist23 Aug 2024#81
m.onwuka, post #39: Percentages of small denominators should be reported with the denominator. Two out of three is not sixty-seven per cent in any useful sense. For what it is worth, the same held on the two occasions I checked. Go to post

Post #77 put the caveat in the right place and I want to underline it.

Correlation between two derived quantities that share a component is partly artefactual. It is a common trap in analyses of ratios.

The variance between people here is larger than the effect being discussed.

2 likes in reply to #39 23mo
MI
m.ibarraTL223 Aug 2024#82
v.malinowski, post #21: Where I part company with post #19, and it is a narrow parting. Adding a reference point for correlation. Mine is a single case, collected without controls, and I am posting the method alongside it so it can be discounted appropriately. Go to post

Time-to-event analysis handles differing follow-up in a way a simple proportion cannot, which is why event rates and Kaplan-Meier estimates can differ.

Speaking for myself and not for anyone else who has posted here.

0 likes in reply to #21 23mo
EF
endo_fellow_rkTL3Endocrinology fellow23 Aug 2024#83

Before the thread moves on from correlation — what is the sample size behind the claim? I am not being difficult; I have seen the same figure quoted from an n of four and from an n of four hundred.

26 likes 23mo
TD
t.dumitruTL223 Aug 2024 · edited#84

Adding the boring version of correlation, because the interesting version keeps getting posted and the boring one is usually right.

Check the ordinary explanations, in order, and stop when one of them accounts for what you are seeing. Most of the time the second one does.

13 likes 23mo
CB
c.bakkerTL223 Aug 2024#85
k.okafor, post #49: P-values and significance: p Same conclusion as the reply above, reached differently, which is mildly reassuring. Go to post

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

A weak preference rather than a position.

4 likes in reply to #49 23mo
JN
j.nascimentoTL223 Aug 2024#86

Source for the correlation figure, since it was asked for. It is in the discussion rather than the abstract, which is why the version circulating is stronger than the paper is.

Reading the surrounding paragraph is worth the two minutes. The authors are more careful than their summarisers.

0 likes 23mo
SL
s.leclercTL4 Moderator23 Aug 2024#87

Helpful, and easy to find again, which is half of what a good reply is.

0 likes 23mo
MB
m.brobergTL223 Aug 2024#88

Post #85 answers the question as asked. The question underneath it is different.

Practical note on correlation: write down what you expect before you look. The number of times I have found what I went looking for is higher than chance would allow.

18 likes 23mo
AF
a.friskTL223 Aug 2024#89
y.adebayo, post #61: A p-value is the probability of data at least this extreme given the null hypothesis. It is not the probability the hypothesis is false, and almost every plain-language gloss gets that backwards. Take it as a starting point and not as a specification. Go to post

Where an analysis was changed after seeing the data, the honest thing is to report both and say which was pre-specified.

That is where I would start, not where I would stop.

20 likes in reply to #61 23mo
TN
t.ndiayeTL223 Aug 2024#90

Confirming post #89 from a second method, which matters more than confirming it from a second person.

What would change my mind on correlation is a second dataset collected by someone with no stake in the first. Until then I hold it loosely and I would rather say so than pretend to more.

9 likes 23mo