The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Study critique

Confounding by indication, explained with a concrete example

TV
to.vargaTL229 Aug 2025#1

On the subject in the title: Confounding by indication, explained with a concrete example Working notes rather than a conclusion.

Comparing SURMOUNT-1 (N Engl J Med, 2022) with SURPASS-2 (N Engl J Med, 2021) and finding the comparison harder than it looks.

Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.

Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?

0 likes 11mo
AR
a.reyesTL4 Admin30 Aug 2025#2

Noted, and I have changed what I was going to do on the strength of it.

1 like 11mo
BD
b.demirTL230 Aug 2025#3

The arithmetic in the opening post is right; the assumption feeding it is the part to check.

The pre-specified endpoint being a weaker proxy than you would like is a real criticism. It is a smaller one than saying the result was chosen after the fact.

6 likes 11mo
SB
s.bruunTL230 Aug 2025#4

Hold a trial to the standard something could actually have met. A criticism that no achievable design could have answered is a criticism of the field rather than of the paper.

17 likes 11mo
CD
c.delgadoTL230 Aug 2025 · edited#5

Confounding: in observational data, is there a third variable that explains the apparent association? In randomised data, randomisation should balance unknown confounders, though known confounders can be adjusted for.

If the premise is wrong, everything after it is decoration.

32 likes 11mo
AL
a.lindholmTL231 Aug 2025#6
to.varga, post #1: On the subject in the title: Confounding by indication, explained with a concrete example Working notes rather than a conclusion. Comparing SURMOUNT-1 ( N Engl J Med , 2022) with SURPASS-2 ( N Engl J Med , 2021) and finding the comparison harder than it looks. Different populations, different durations, different endpoints defined… Go to post

Where I part company with post #4, and it is a narrow parting.

Attrition is the failure mode most likely to invalidate a result and the least likely to be discussed. Differential attrition between arms is the specific thing to look for.

The claim is narrower than it sounds, and deliberately so.

0 likes in reply to #1 11mo
ME
m.eriksenTL231 Aug 2025#7

Per-protocol and intention-to-treat analyses answer different questions and neither is the honest one by default. Reporting both is the practice worth insisting on.

3 likes 11mo
ME
m.ekstromTL231 Aug 2025#8

Surrogate endpoints are not automatically bad and their validity is compound-specific and population-specific. The question is whether this surrogate has been validated for this use.

For what it is worth, the same held on the two occasions I checked.

11 likes 11mo
T
ThibodeauTL3Regular31 Aug 2025#9

Confirming post #6 from a second method, which matters more than confirming it from a second person.

Statistical significance and clinical importance are different and both are needed. A significant difference below the minimal important difference is a real finding of no practical consequence.

Written quickly, so the reasoning may be tighter than the wording.

1 like 11mo
HA
h.amankwahTL21 Sep 2025#10
m.ekstrom, post #8: Surrogate endpoints are not automatically bad and their validity is compound-specific and population-specific. The question is whether this surrogate has been validated for this use. For what it is worth, the same held on the two occasions I checked. Go to post

I had written a reply contradicting post #8 and deleted it. Here is what survived.

Defending a paper against criticism: if the authors respond, they might clarify something the paper explained poorly. Their response might also miss your point. Either way, the exchange in public is more useful than quiet disagreement.

6 likes in reply to #8 11mo
IR
i.rasmussenTL21 Sep 2025#11

Coming back to post #7, because the follow-up matters more than the original answer.

Generalisability and validity are separate axes. A trial can be internally impeccable and still tell you nothing about the person asking.

That much is documented. The rest is how I have interpreted it.

4 likes 11mo
M
MSaarinenTL3Regular1 Sep 2025#12
s.bruun, post #4: Hold a trial to the standard something could actually have met. A criticism that no achievable design could have answered is a criticism of the field rather than of the paper. Go to post

I will take the caveat as seriously as the claim, which is the point of putting it there.

0 likes in reply to #4 11mo
IB
i.brobergTL21 Sep 2025#13
to.varga, post #1: On the subject in the title: Confounding by indication, explained with a concrete example Working notes rather than a conclusion. Comparing SURMOUNT-1 ( N Engl J Med , 2022) with SURPASS-2 ( N Engl J Med , 2021) and finding the comparison harder than it looks. Different populations, different durations, different endpoints defined… Go to post

Generalisability: do the inclusion/exclusion criteria narrow the population so much that results do not apply to real people asking about it? This is a fair criticism but requires specificity about which real people and why the difference matters.

The conclusion is tentative; the arithmetic underneath it is not.

26 likes in reply to #1 11mo
V
VThorvaldsenTL3Regular1 Sep 2025#14

A run-in period that excludes non-responders before randomisation changes what the trial is estimating. It is legitimate design and it must be stated in any summary.

The step people skip is the one I have spelled out.

12 likes 11mo
ET
e.tammTL21 Sep 2025 · edited#15

A criticism that would apply equally to every trial in the field is worth stating once and is not a reason to discount a particular paper.

7 likes 11mo
EM
e.mikkelsenTL2Member2 Sep 2025#16
s.bruun, post #4: Hold a trial to the standard something could actually have met. A criticism that no achievable design could have answered is a criticism of the field rather than of the paper. Go to post

Post #15 answers the question as asked. The question underneath it is different.

Multiple comparisons: if a paper reports many outcomes, the chance of a spurious association by random chance is real. Pre-specification of primary outcomes matters and secondary analyses are weaker evidence.

1 like in reply to #4 11mo
EK
e.krastevTL22 Sep 2025#17

Worth separating two things that post #15 runs together.

Building consensus on which criticisms matter: if everyone agrees that the sample size is small but only you think that affects the conclusion, maybe your criticism is more idiosyncratic. That does not make it wrong but it is worth noticing.

It is the kind of thing that is obvious once and never again.

0 likes 11mo
SP
s.poulsenTL32 Sep 2025#18
RO
r.oyelaranTL22 Sep 2025#19
m.eriksen, post #7: Per-protocol and intention-to-treat analyses answer different questions and neither is the honest one by default. Reporting both is the practice worth insisting on. Go to post

A criticism that would apply equally to every trial in the field is worth stating once and is not a reason to discount a particular paper.

Marking that as an opinion rather than a finding.

12 likes in reply to #7 11mo
KF
k.farrugiaTL3Regular2 Sep 2025#20
b.demir, post #3: The arithmetic in the opening post is right; the assumption feeding it is the part to check. The pre-specified endpoint being a weaker proxy than you would like is a real criticism. It is a smaller one than saying the result was chosen after the fact. Go to post

Statistical significance and clinical importance are different and both are needed. A significant difference below the minimal important difference is a real finding of no practical consequence.

4 likes in reply to #3 11mo
SE
septum_entryTL2Member2 Sep 2025#21

Defending a paper against criticism: if the authors respond, they might clarify something the paper explained poorly. Their response might also miss your point. Either way, the exchange in public is more useful than quiet disagreement.

I would want to see it done twice before believing it once.

0 likes 11mo
CF
c.falkTL23 Sep 2025 · edited#22

Hold a trial to the standard something could actually have met. A criticism that no achievable design could have answered is a criticism of the field rather than of the paper.

A partial answer, offered because a partial answer beats none.

2 likes 11mo
AT
apostille_traceTL1Member3 Sep 2025#23

Post #20 answers the question as asked. The question underneath it is different.

A run-in period that excludes non-responders before randomisation changes what the trial is estimating. It is legitimate design and it must be stated in any summary.

9 likes 11mo
KC
k.chukwuTL23 Sep 2025#24
b.demir, post #3: The arithmetic in the opening post is right; the assumption feeding it is the part to check. The pre-specified endpoint being a weaker proxy than you would like is a real criticism. It is a smaller one than saying the result was chosen after the fact. Go to post

Attrition is the failure mode most likely to invalidate a result and the least likely to be discussed. Differential attrition between arms is the specific thing to look for.

21 likes in reply to #3 11mo
KB
k.bettencourtTL2Member3 Sep 2025#25

That is a fair summary of where the discussion has got to.

0 likes 11mo
JS
j.solbergTL23 Sep 2025#26

Worth separating two things that post #22 runs together.

Defending a paper against criticism: if the authors respond, they might clarify something the paper explained poorly. Their response might also miss your point. Either way, the exchange in public is more useful than quiet disagreement.

I have said this before in a thread nobody could find, so it is worth repeating.

1 like 11mo

Suggested topics

TopicParticipantsRepliesViewsActivity
Immortal time bias in a claims-database study
Immortal time bias in a claims-database study — setting out what I have, and where I think it stops being reliable. Asking about immortal time bias directly, because I have read four threads on it and each…
MACDSOIDNN+29 34 28k 11mo
A well-designed study with a badly written abstract
Posting this under the heading it deserves: A well-designed study with a badly written abstract Everything below is what sits behind that. Reading back through what has been written here about well-designed…
RCMICBMRPN+117 128 7.2k 14mo
Coming back to: A critique that turned out to be unfair, retracted by its author
A critique that turned out to be unfair, retracted by its author — setting out what I have, and where I think it stops being reliable. Reading back through what has been written here about critique, three…
IAQLJB 2 1.3k 14mo
A critique that turned out to be unfair, retracted by its author
Posting this under the heading it deserves: A critique that turned out to be unfair, retracted by its author Everything below is what sits behind that. Critique, and specifically the version of it that the…
MSASFIONS+34 38 20k 17h
Criticising the method without criticising the authors — one year on
On the subject in the title: Criticising the method without criticising the authors — one year on Working notes rather than a conclusion. A methods question rather than a substantive one, about Criticising…
KOSKJNAKSC+16 20 61k 11mo

Related topics — sharing the tags confounding, observational data, disputed

TopicParticipantsRepliesViewsActivity
Why plausible mechanism is not evidence of effect
Asking directly, because I could not find a straight answer: Why plausible mechanism is not evidence of effect What changes if the standard account of plausible mechanism is wrong? I ask because I have been…
FYGFSPCNAK+14 18 57k 13mo
IGF-1 as a surrogate endpoint and its limitations — the long version
On the subject in the title: IGF-1 as a surrogate endpoint and its limitations — the long version Working notes rather than a conclusion. A narrow question about IGF-1, deliberately narrow, because the broad…
NOMGCRKBGP+14 18 53k 12mo
Reading a rodent study on a secretagogue without over-extrapolating — the long version
Reading a rodent study on a secretagogue without over-extrapolating — the long version — setting out what I have, and where I think it stops being reliable. A narrow question about Reading a rodent study,…
CDDBANKBMA+59 66 30k 4h
Why this subcategory is stricter about sourcing than most
Asking directly, because I could not find a straight answer: Why this subcategory is stricter about sourcing than most A comparison question rather than a question about one compound. Two things in the same…
DOCAEFRESC+48 52 617 21d
Ipamorelin selectivity claims and where they came from — a second dataset
Posting this under the heading it deserves: Ipamorelin selectivity claims and where they came from — a second dataset Everything below is what sits behind that. Posting a small dataset on Ipamorelin…
TIDCM 2 27k 14mo