The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Study critique

[2026 update] Confounding by indication, explained with a concrete example

DM
d.magalhesTL2Member22 Jan 2026#1

Confounding by indication, explained with a concrete example Writing it up because I had to work it out twice and would rather nobody else did.

I have seen SURPASS-4 (Lancet, 2021) cited in support of a claim I do not think it supports, twice this month, so I would like to work through what it actually shows.

My reading is that the trial is sound for its own question and is being stretched to answer a different one. I might be wrong about that, which is why this is a topic rather than a correction.

What I would like from this discussion: someone who disagrees with me to say why, with the section of the paper they are relying on.

0 likes 6mo
EL
endpoint_lineTL3Regular23 Jan 2026#2

The arithmetic in the opening post is right; the assumption feeding it is the part to check.

Multiple comparisons: if a paper reports many outcomes, the chance of a spurious association by random chance is real. Pre-specification of primary outcomes matters and secondary analyses are weaker evidence.

I checked the source rather than the summary, and they differ.

19 likes 6mo
CV
ca.vermeulenTL224 Jan 2026#3
d.magalhes, post #1: Confounding by indication, explained with a concrete example Writing it up because I had to work it out twice and would rather nobody else did. I have seen SURPASS-4 ( Lancet , 2021) cited in support of a claim I do not think it supports, twice this month, so I would like to work through what it actually shows. My reading is that the… Go to post

Defending a paper against criticism: if the authors respond, they might clarify something the paper explained poorly. Their response might also miss your point. Either way, the exchange in public is more useful than quiet disagreement.

5 likes in reply to #1 6mo
HN
h.nicolaidesTL3Regular24 Jan 2026#4

Criticism is more useful when it is narrower. "The trial answers a different question from the one being asked" is actionable; "the trial is flawed" is not.

0 likes 6mo
IG
i.grimaldiTL225 Jan 2026#5

The opening post describes the usual case. This is about the unusual one.

Per-protocol and intention-to-treat analyses answer different questions and neither is the honest one by default. Reporting both is the practice worth insisting on.

The step people skip is the one I have spelled out.

0 likes 6mo
EF
erratum_fileTL3Regular25 Jan 2026#6

That is clearer than the version I had in my head. Thank you.

27 likes 6mo
ID
i.dumitruTL226 Jan 2026 · edited#7
d.magalhes, post #1: Confounding by indication, explained with a concrete example Writing it up because I had to work it out twice and would rather nobody else did. I have seen SURPASS-4 ( Lancet , 2021) cited in support of a claim I do not think it supports, twice this month, so I would like to work through what it actually shows. My reading is that the… Go to post

Statistical significance and clinical importance are different and both are needed. A significant difference below the minimal important difference is a real finding of no practical consequence.

8 likes in reply to #1 6mo
R
RidgewayTL3Regular27 Jan 2026#8
erratum_file, post #6: That is clearer than the version I had in my head. Thank you. Go to post

Generalisability and validity are separate axes. A trial can be internally impeccable and still tell you nothing about the person asking.

The conclusion is tentative; the arithmetic underneath it is not.

2 likes in reply to #6 6mo
SI
s.ivaturiTL227 Jan 2026#9

On post #5 — agreed on the reasoning, with one qualification.

Surrogate endpoints are not automatically bad and their validity is compound-specific and population-specific. The question is whether this surrogate has been validated for this use.

0 likes 6mo
KR
k.redgraveTL2Member28 Jan 2026#10
erratum_file, post #6: That is clearer than the version I had in my head. Thank you. Go to post

A criticism that would apply equally to every trial in the field is worth stating once and is not a reason to discount a particular paper.

0 likes in reply to #6 6mo
K
KForsbergTL2Member28 Jan 2026 · edited#11
k.redgrave, post #10: A criticism that would apply equally to every trial in the field is worth stating once and is not a reason to discount a particular paper. Go to post

Taking post #10 at face value and following it one step further.

Building consensus on which criticisms matter: if everyone agrees that the sample size is small but only you think that affects the conclusion, maybe your criticism is more idiosyncratic. That does not make it wrong but it is worth noticing.

Old habit: I write down the expected answer before I calculate it.

2 likes in reply to #10 6mo
KK
k.kimaniTL229 Jan 2026#12

Post #8 and I disagree about the size of the effect, not about the direction.

Hold a trial to the standard something could actually have met. A criticism that no achievable design could have answered is a criticism of the field rather than of the paper.

I would put this at better than even and not much better.

8 likes 6mo
FF
f.fenwickTL3Regular29 Jan 2026#13

Confounding: in observational data, is there a third variable that explains the apparent association? In randomised data, randomisation should balance unknown confounders, though known confounders can be adjusted for.

26 likes 6mo
KC
k.chukwuTL230 Jan 2026#14

The pre-specified endpoint being a weaker proxy than you would like is a real criticism. It is a smaller one than saying the result was chosen after the fact.

0 likes 6mo
M
MakinenTL2Member30 Jan 2026#15

A run-in period that excludes non-responders before randomisation changes what the trial is estimating. It is legitimate design and it must be stated in any summary.

4 likes 6mo
JS
j.solbergTL231 Jan 2026#16

Right — I had this wrong and I am glad to have read it before it mattered.

12 likes 6mo
L
LundqvistTL2Member31 Jan 2026#17

Attrition is the failure mode most likely to invalidate a result and the least likely to be discussed. Differential attrition between arms is the specific thing to look for.

Adding this to the thread rather than to the wiki, because I am not confident enough for the wiki.

0 likes 6mo
SH
s.hartmannTL231 Jan 2026#18
k.redgrave, post #10: A criticism that would apply equally to every trial in the field is worth stating once and is not a reason to discount a particular paper. Go to post

Hold a trial to the standard something could actually have met. A criticism that no achievable design could have answered is a criticism of the field rather than of the paper.

0 likes in reply to #10 6mo
TP
t.pereiraTL21 Feb 2026#19

Criticism is more useful when it is narrower. "The trial answers a different question from the one being asked" is actionable; "the trial is flawed" is not.

0 likes 6mo
FV
f.villalobosTL21 Feb 2026 · edited#20

Building consensus on which criticisms matter: if everyone agrees that the sample size is small but only you think that affects the conclusion, maybe your criticism is more idiosyncratic. That does not make it wrong but it is worth noticing.

2 likes 6mo
ED
e.dalgleishTL3Regular2 Feb 2026#21

I had written a reply contradicting post #17 and deleted it. Here is what survived.

Per-protocol and intention-to-treat analyses answer different questions and neither is the honest one by default. Reporting both is the practice worth insisting on.

1 like 6mo
RI
r.ilungaTL22 Feb 2026 · edited#22
j.solberg, post #16: Right — I had this wrong and I am glad to have read it before it mattered. Go to post

Confirming post #21 from a second method, which matters more than confirming it from a second person.

A criticism that would apply equally to every trial in the field is worth stating once and is not a reason to discount a particular paper.

I would rather post the uncertainty than round it away.

0 likes in reply to #16 6mo
D
DOdendaalTL3Regular3 Feb 2026#23

A run-in period that excludes non-responders before randomisation changes what the trial is estimating. It is legitimate design and it must be stated in any summary.

The general case is well covered; this is the awkward specific one.

21 likes 6mo
MB
ma.balogunTL23 Feb 2026#24

Thank you for the correction. I would rather find out here than later.

9 likes 6mo
BS
buffer_sheetTL3Regular4 Feb 2026#25

Coming back to post #21, because the follow-up matters more than the original answer.

Generalisability: do the inclusion/exclusion criteria narrow the population so much that results do not apply to real people asking about it? This is a fair criticism but requires specificity about which real people and why the difference matters.

If the premise is wrong, everything after it is decoration.

2 likes 6mo
BW
b.wikstromTL24 Feb 2026#26
r.ilunga, post #22: Confirming post #21 from a second method, which matters more than confirming it from a second person. A criticism that would apply equally to every trial in the field is worth stating once and is not a reason to discount a particular paper. I would rather post the uncertainty than round it away. Go to post

Multiple comparisons: if a paper reports many outcomes, the chance of a spurious association by random chance is real. Pre-specification of primary outcomes matters and secondary analyses are weaker evidence.

Correct me on the arithmetic if it is wrong; I would rather know.

0 likes in reply to #22 6mo
IL
integrator_logTL3Regular4 Feb 2026#27

Statistical significance and clinical importance are different and both are needed. A significant difference below the minimal important difference is a real finding of no practical consequence.

I would call that likely rather than established.

29 likes 6mo
FL
f.lindholmTL25 Feb 2026#28
BJ
b.jankowiakTL3Regular5 Feb 2026#29
i.grimaldi, post #5: The opening post describes the usual case. This is about the unusual one. Per-protocol and intention-to-treat analyses answer different questions and neither is the honest one by default. Reporting both is the practice worth insisting on. The step people skip is the one I have spelled out. Go to post

Surrogate endpoints are not automatically bad and their validity is compound-specific and population-specific. The question is whether this surrogate has been validated for this use.

Adding it in case it saves somebody the afternoon it cost me.

5 likes in reply to #5 6mo
JF
j.falkTL26 Feb 2026#30

The pre-specified endpoint being a weaker proxy than you would like is a real criticism. It is a smaller one than saying the result was chosen after the fact.

0 likes 6mo