The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Research Methods · Statistics · continued

Second pass at: Regression to the mean in progress reports posts 31–60

This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.

AK
a.kowalskiTL221 Jan 2025#31

Adding a small correction to the Regression summary above rather than a disagreement with it. The substance holds; one of the figures is out by a factor that matters.

7 likes 18mo
EP
e.piresTL222 Jan 2025#32
z.vogel, post #17: What I can speak to on Regression is narrow, so I will keep it narrow rather than generalising from it. Beyond that boundary I do not know. Go to post

Noted, and I have changed what I was going to do on the strength of it.

17 likes in reply to #17 18mo
JD
j.dahlbergTL223 Jan 2025 · edited#33

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

Correct me on the arithmetic if it is wrong; I would rather know.

0 likes 18mo
TW
t.wojcikTL224 Jan 2025#34

Coming back to post #30, because the follow-up matters more than the original answer.

Regression is a good example of a question where the honest answer is boring and the interesting answers are unsupported. I would go with boring.

0 likes 18mo
SD
s.duarteTL225 Jan 2025#35

The version of Regression that I was taught turned out to be a teaching simplification. Useful, and not true in the way I had assumed it was.

4 likes 18mo
VB
v.bhattacharyaTL226 Jan 2025#36
r.mensah, post #23: I read post #21 twice before replying, because I had assumed the opposite. A request rather than an answer: could whoever has the primary source for Regression post it? I have seen the claim three times this month and each version had lost a qualifier. Go to post

Time-to-event analysis handles differing follow-up in a way a simple proportion cannot, which is why event rates and Kaplan-Meier estimates can differ.

The general case is well covered; this is the awkward specific one.

12 likes in reply to #23 18mo
AS
a.salcedoTL3Regular27 Jan 2025#37

Confirming post #36 from a second method, which matters more than confirming it from a second person.

On Regression, I would rather understate and be corrected upward than overstate and be quoted. That is a house style here and it is a good one.

25 likes 18mo
MY
m.yilmazTL229 Jan 2025#38

I had written a reply contradicting post #34 and deleted it. Here is what survived.

Whatever the answer on Regression turns out to be, the method for getting there is the same: state the assumption, do the arithmetic in public, invite the correction.

0 likes 18mo
NV
n.villalobosTL230 Jan 2025#39
a.molnar, post #4: Confirming the opening post from a second method, which matters more than confirming it from a second person. P-values and significance: p Go to post

Survivorship in a self-reporting population biases every aggregate produced from it, and the bias is in the flattering direction.

I would not lead a decision with this, but I would not ignore it either.

1 like in reply to #4 18mo
FP
forest_plotTL3Evidence synthesis31 Jan 2025 · edited#40
s.kravchenko, post #13: Fine by me. I had wanted a stronger conclusion and there is not one available. Go to post

Answering the question post #38 raises rather than the one it answers.

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

7 likes in reply to #13 18mo
CR
c.rasmussenTL21 Feb 2025#41

I had written a reply contradicting post #39 and deleted it. Here is what survived.

A note on how Regression gets discussed rather than on Regression itself: the confident posts get the replies and the careful ones get ignored, and the careful ones have been right more often.

20 likes 18mo
EL
e.lokkenTL21 Feb 2025#42

Confirming post #39 from a second method, which matters more than confirming it from a second person.

Medians and means diverge for skewed distributions, and most of the quantities discussed here are skewed. Which one a paper reports is a choice worth noticing.

8 likes 18mo
AA
an.adeyemiTL22 Feb 2025#43

Worth stating the null on Regression before we explain it: the observation may be nothing. That possibility deserves a sentence and usually does not get one.

0 likes 18mo
ES
e.silvaTL23 Feb 2025#44
j.dahlberg, post #33: Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics. Correct me on the arithmetic if it is wrong; I would rather know. Go to post

Rounding and significant figures carry information about precision. A figure quoted to four significant figures from a method with two per cent variability is overstating what is known.

The strength of my opinion here exceeds the strength of my evidence.

0 likes in reply to #33 18mo
LF
l.ferreiraTL24 Feb 2025#45
f.espinoza, post #15: Building on post #14 rather than restating it. Medians and means diverge for skewed distributions, and most of the quantities discussed here are skewed. Which one a paper reports is a choice worth noticing. Go to post

Regression to the mean explains a large share of apparent improvements in anything selected for being extreme. It is not a statistical curiosity; it is the default explanation.

That is the version I would defend. It is not the version I started with.

14 likes in reply to #15 18mo
HK
h.koodziejTL2Member5 Feb 2025 · edited#46

Post #43 is right about the mechanism and I think understates the practical bit.

Reporting rather than recommending, on Regression. What happened is above. Whether it should have is a different question and not one I am qualified to answer.

5 likes 18mo
EM
e.mwangiTL26 Feb 2025#47

My position on Regression is current rather than settled. I have revised it once already and I expect to again, so treat it accordingly.

0 likes 18mo
TI
trough_indexTL3Regular7 Feb 2025#48

A p-value is the probability of data at least this extreme given the null hypothesis. It is not the probability the hypothesis is false, and almost every plain-language gloss gets that backwards.

I have separated what I observed from what I concluded, which does not always happen.

28 likes 18mo
HF
h.fonsecaTL28 Feb 2025#49

Since Regression keeps coming up, it should probably be a maintained page rather than a recurring thread. I am happy to draft it if someone with more direct experience will review it.

0 likes 18mo
NB
n.bridgewaterTL2Member9 Feb 2025#50

The arithmetic on Regression is the easy part and it is where the errors are, which is an uncomfortable combination. Show your working and someone will catch it.

19 likes 18mo
EK
ew.kuuselaTL210 Feb 2025#51

Building on post #50 rather than restating it.

Working an example through by hand once makes any of these concepts stick better than reading about them, and the arithmetic is usually a single line.

29 likes 18mo
TW
t.waldenstrmTL2Member11 Feb 2025#52

Post #48 put the caveat in the right place and I want to underline it.

A standard deviation describes the spread of individuals and a standard error describes the precision of the mean. Quoting one where the other belongs changes the apparent result substantially.

I would rather say I do not know than round it up to an answer.

0 likes 18mo
TB
t.brandtTL212 Feb 2025#53

Multiple testing inflates the false-positive rate in a way that is entirely predictable and entirely correctable. The correction should be declared in advance.

That much is documented. The rest is how I have interpreted it.

5 likes 17mo
KB
k.bettencourtTL2Member13 Feb 2025#54
h.fonseca, post #49: Since Regression keeps coming up, it should probably be a maintained page rather than a recurring thread. I am happy to draft it if someone with more direct experience will review it. Go to post

The thing about Regression that took me longest to accept is that a plausible mechanism is not evidence of an effect. It is a reason to look, not a result.

14 likes in reply to #49 17mo
JL
j.lokkenTL214 Feb 2025#55

Agreed on all of that, and I have nothing to add to it.

0 likes 17mo
SS
s.stavrianosTL2Member15 Feb 2025 · edited#56

Post #52 and I disagree about the size of the effect, not about the direction.

One caution on Regression: everything above assumes the underlying documentation is what it claims to be. That assumption is doing real work and is rarely stated.

0 likes 17mo
GD
g.danquahTL216 Feb 2025#57

The most common statistical error in this community is not technical: it is treating a self-selected collection of reports as a sample from a population.

That is all I can say without guessing.

9 likes 17mo
B
BuchholzTL2Member17 Feb 2025#58
e.mwangi, post #47: My position on Regression is current rather than settled. I have revised it once already and I expect to again, so treat it accordingly. Go to post

Medians and means diverge for skewed distributions, and most of the quantities discussed here are skewed. Which one a paper reports is a choice worth noticing.

Nothing above should be read as advice about what anyone else should do.

20 likes in reply to #47 17mo
AV
a.vestergaardTL218 Feb 2025#59
c.kuusela, post #2: Right, and stated more narrowly than I would have dared to state it. Go to post

Baseline imbalance in a randomised trial is expected by chance and adjusting for it post hoc is a choice that should have been pre-specified.

If that is already documented somewhere, ignore me and link it.

15 likes in reply to #2 17mo
GD
glossary_deskTL3Regular19 Feb 2025#60

The number people quote for Regression is a central estimate presented without its interval, and the interval is wide enough that the estimate is nearly uninformative on its own.

30 likes 17mo