Revisiting: Journal club: SURMOUNT-1, read without the press release posts 31–60
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
Adding a small correction to the SURMOUNT-1 summary above rather than a disagreement with it. The substance holds; one of the figures is out by a factor that matters.
On post #29 — agreed on the reasoning, with one qualification.
Posting my SURMOUNT-1 numbers with the method attached so they can be discounted properly. Uncontrolled, unblinded, and collected by someone who wanted a particular answer.
The most common failure mode is spending fifty minutes on the effect size and ten on the population. Reversing that ratio would improve most sessions.
I would not lead a decision with this, but I would not ignore it either.
One caution on SURMOUNT-1: everything above assumes the underlying documentation is what it claims to be. That assumption is doing real work and is rarely stated.
Post #33 describes the usual case. This is about the unusual one.
SURMOUNT-1 is a good example of a question where the honest answer is boring and the interesting answers are unsupported. I would go with boring.
Adding the measurement that post #37 says would settle it.
The baseline characteristics table is the single most useful page for the questions that get asked here, because it tells you who the result applies to.
The claim is narrower than it sounds, and deliberately so.
Building on post #37 rather than restating it.
The number people quote for SURMOUNT-1 is a central estimate presented without its interval, and the interval is wide enough that the estimate is nearly uninformative on its own.
Where the group cannot agree, record the disagreement rather than resolving it by seniority. The recorded disagreement is more honest and more useful later.
The figures usually contain the finding and the text usually contains the interpretation. Separating them for the first twenty minutes is a discipline worth keeping.
This follows post #42 rather than contradicting it.
The honest answer on SURMOUNT-1 is that it depends, and the useful part is the list of what it depends on. Four items, in rough order of how much they matter.
Most people get the first two right and then argue about the fourth.
Helpful, and short, which on this subject is harder than long.
Sessions on negative or null results are consistently the most instructive and the hardest to get anyone to attend.
I read post #45 twice before replying, because I had assumed the opposite.
A note on how SURMOUNT-1 gets discussed rather than on SURMOUNT-1 itself: the confident posts get the replies and the careful ones get ignored, and the careful ones have been right more often.
Taking post #46 at face value and following it one step further.
SELECT (N Engl J Med 2023): Semaglutide cardiovascular outcomes without diabetes. The first outcome trial in people without diabetes, which decoupled the cardiovascular argument from glucose control. Read the absolute numbers, not just the relative reduction.
Reading it again, the caveat matters more than the finding.
Post #46 is the version of this I will quote in future. One addition.
I would call the community position on SURMOUNT-1 likely rather than established, and I would be comfortable defending that hedge.
The most common failure mode is spending fifty minutes on the effect size and ten on the population. Reversing that ratio would improve most sessions.
I would want a second opinion before relying on that.
The confident answers on SURMOUNT-1 and the well-sourced answers are not the same answers, which is the most useful thing I have learned reading this category.
SURMOUNT-1 (N Engl J Med 2022): Tirzepatide obesity trial. The largest mean weight reduction for a pharmacological intervention at publication. Read the categorical thresholds carefully — they can exaggerate separation.
Before the thread moves on from SURMOUNT-1 — what is the sample size behind the claim? I am not being difficult; I have seen the same figure quoted from an n of four and from an n of four hundred.
Picking up post #51: that is the part I would want checked first.
Two things can be true about SURMOUNT-1 at once: the mechanism is plausible and the evidence for the size of the effect is thin. Most of the argument here is people defending the first against attacks on the second.
I read post #51 twice before replying, because I had assumed the opposite.
Reading order that works for these sessions: registry entry, methods, baseline table, primary result, then abstract last. Reading the abstract first anchors everything that follows.
I am confident about the direction and much less about the magnitude.
The baseline characteristics table is the single most useful page for the questions that get asked here, because it tells you who the result applies to.
I disagree with the framing of SURMOUNT-1 above, and I think it is a substantive disagreement rather than a terminological one. Setting out why, so it can be checked.
The reasoning depends on an assumption that is doing a lot of work and is never stated. If the assumption holds, the conclusion follows. I do not think it holds generally.
Post #59 is the version of this I will quote in future. One addition.
Source for the SURMOUNT-1 figure, since it was asked for. It is in the discussion rather than the abstract, which is why the version circulating is stronger than the paper is.
Reading the surrounding paragraph is worth the two minutes. The authors are more careful than their summarisers.