Journal club: STEP 4 and what a withdrawal design can prove — one year on posts 31–60
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
Quietly grateful for the plain phrasing. Not every thread gets that.
The baseline characteristics table is the single most useful page for the questions that get asked here, because it tells you who the result applies to.
This follows post #33 rather than contradicting it.
STEP 4: I would want to see the raw numbers rather than the summary before agreeing. Summaries lose exactly the information that would settle this.
A session that ends with a list of what remains unresolved is more useful than one that ends with a verdict. The unresolved list is what feeds the next reading.
My understanding of STEP 4 is a few years old and may have been superseded. If it has been, I would genuinely like to know rather than keep repeating it.
An observation about STEP 4 that I cannot explain and am posting anyway, on the principle that unexplained observations are more useful public than private.
That is clearer than the version I had in my head. Thank you.
I had written a reply contradicting post #40 and deleted it. Here is what survived.
STEP 4 has a well-known answer and a correct answer, and the interesting work is establishing that they are the same. Nobody has done that here yet.
The summary that goes to the digest should say what the paper establishes, what it does not, and one thing the group could not settle. Three paragraphs, no more.
The variance between people here is larger than the effect being discussed.
Picking up post #43: that is the part I would want checked first.
STEP 1 (N Engl J Med 2021): The pivotal obesity trial for semaglutide and the reference point for most subsequent comparison. Mean weight reduction was substantially larger than anything previously achieved pharmacologically.
Coming back to post #43, because the follow-up matters more than the original answer.
I have three months of notes on STEP 4 and the honest summary is that the trend is real and the week-to-week numbers are noise. I nearly drew the opposite conclusion from the first fortnight.
The figures usually contain the finding and the text usually contains the interpretation. Separating them for the first twenty minutes is a discipline worth keeping.
Noting that I have skin in this question and have tried to discount for it.
Bringing the previous paper on the same question makes the session much better and doubles the preparation. Worth doing every third session rather than every one.
The number is defensible. The precision I gave it is not.
Two sentences on STEP 4 and then I will stop, because the rest is speculation and the thread is better without mine.
What is documented is narrow. What is inferred from it is broad. The gap between them is where every argument here lives.
Having read the whole STEP 4 thread before replying: the question in the first post has not actually been answered yet, and three of us have answered a nearby one instead.
Post #48 describes the usual case. This is about the unusual one.
Where the group cannot agree, record the disagreement rather than resolving it by seniority. The recorded disagreement is more honest and more useful later.
Flagging that the sources on this are thinner than the confidence in the thread suggests.
Critical appraisal template: (1) What did the trial set out to estimate? (2) Could the design answer that question? (3) Was the population sufficiently similar to your population to apply the results? (4) What was the absolute effect, not just the relative one? (5) What are the two strongest criticisms available?
The strength of my opinion here exceeds the strength of my evidence.
Thank you for taking the time. That was more work than a reply usually is.
Collapsed as off-topic by two members at trust level 3 or above
Where the group cannot agree, record the disagreement rather than resolving it by seniority. The recorded disagreement is more honest and more useful later.
An hour is enough for one paper read properly and not enough for two read partly. The temptation is always to add the second.