The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Trials

Revisiting: Primary endpoint hierarchies and why order matters

AH
a.hartmannTL23 May 2026#1

Posting this under the heading it deserves: Revisiting: Primary endpoint hierarchies and why order matters Everything below is what sits behind that.

I was wrong about Primary endpoint hierarchies in a thread last spring and I would like to correct it publicly rather than quietly.

The error was in the units, which changed the conclusion by an order of magnitude. Setting out the corrected version, and the way I now check for that class of mistake.

43 likes 3mo
G
GDashwoodTL3Regular4 May 2026 · edited#2

Second this, and I would have said it less carefully.

0 likes 3mo
JT
j.teixeiraTL24 May 2026#3

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

3 likes 3mo
CV
c.vermeulenTL25 May 2026#4

Subgroup analyses are hypothesis-generating unless pre-specified and adequately powered, and almost none are the second. The interaction test matters more than the subgroup point estimate.

That is the version I use. It may not be the version that is correct.

11 likes 3mo
MS
m.silvaTL25 May 2026#5

Post #4 is the version of this I will quote in future. One addition.

I would call the community position on Primary endpoint hierarchies likely rather than established, and I would be comfortable defending that hedge.

17 likes 3mo
VS
v.salgadoTL26 May 2026 · edited#6
a.hartmann, post #1: Posting this under the heading it deserves: Revisiting: Primary endpoint hierarchies and why order matters Everything below is what sits behind that. I was wrong about Primary endpoint hierarchies in a thread last spring and I would like to correct it publicly rather than quietly. The error was in the units, which changed the conclusion… Go to post

Where I part company with post #3, and it is a narrow parting.

A treatment-policy estimand asks what happens to people assigned to a strategy, including those who abandon it. A hypothetical estimand asks what would have happened had everyone continued. Both are legitimate and they give different numbers.

I would be interested in a counterexample if anyone has one.

32 likes in reply to #1 3mo
LC
lu.cabreraTL26 May 2026#7

Registration before enrolment, with the primary endpoint declared, is what makes outcome switching detectable. Checking the registry against the paper takes five minutes and is worth doing.

1 like 3mo
NS
n.stanescuTL27 May 2026#8

Primary endpoint hierarchies is a question about a distribution, not about a value, and treating it as a value is what produces the confident wrong answers.

7 likes 3mo
LC
l.cabreraTL27 May 2026#9

Agreed on all of that, and I have nothing to add to it.

11 likes 3mo
AS
a.stephanopoulosTL3Regular7 May 2026#10

Worth separating two things that post #6 runs together.

Pre-specification is the property that makes a primary endpoint trustworthy. An endpoint chosen after seeing the data can be the best endpoint in the world and it no longer carries the same guarantee.

Marking that as an opinion rather than a finding.

24 likes 3mo
EF
endo_fellow_rkTL3Endocrinology fellow8 May 2026#11

Primary endpoint hierarchies came up in a thread eighteen months ago and was answered well. I cannot find it, which is itself the problem, so here is the reconstruction.

0 likes 3mo
YA
y.adebayoTL28 May 2026#12

I read the earlier replies on Primary endpoint hierarchies twice before writing this, because I had assumed the opposite and wanted to be sure I was disagreeing with what was said rather than what I expected.

32 likes 3mo
MH
ms_hollowayTL4Mass spectrometrist8 May 2026#13
lu.cabrera, post #7: Registration before enrolment, with the primary endpoint declared, is what makes outcome switching detectable. Checking the registry against the paper takes five minutes and is worth doing. Go to post

Post #10 describes the usual case. This is about the unusual one.

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

If that is already documented somewhere, ignore me and link it.

16 likes in reply to #7 3mo
MI
m.ibarraTL29 May 2026#14

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

6 likes 3mo
FN
formulary_notesTL3Regular9 May 2026#15

Absolute and relative effects answer different questions. Write down the event rate in each arm and the difference between them; everything quotable is derived from those two numbers.

Caveat: everything above assumes the paperwork is what it says it is.

0 likes 3mo
CA
c.amankwahTL29 May 2026 · edited#16

Seconded. It reads as careful rather than confident, which is the right register.

24 likes 3mo
TH
TL4_HalvorsenTL4Leader · Journal club10 May 2026#17
n.stanescu, post #8: Primary endpoint hierarchies is a question about a distribution, not about a value, and treating it as a value is what produces the confident wrong answers. Go to post

On post #13 — agreed on the reasoning, with one qualification.

If you are new and reading this thread for the answer to Primary endpoint hierarchies: the answer is conditional, the conditions are in the third reply, and the rest of the thread is worth skipping.

11 likes in reply to #8 3mo
CC
ch.correiaTL210 May 2026#18
m.silva, post #5: Post #4 is the version of this I will quote in future. One addition. I would call the community position on Primary endpoint hierarchies likely rather than established, and I would be comfortable defending that hedge. Go to post

Picking up post #17: that is the part I would want checked first.

The figure that circulates in coverage is almost always whichever estimand gives the larger effect. That is not fraud; it is selection, and it is why the paper matters more than the summary.

Nothing above should be read as advice about what anyone else should do.

3 likes in reply to #5 3mo
NM
n.moreauTL210 May 2026#19

Coming back to post #17, because the follow-up matters more than the original answer.

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

It took me longer than it should have to see that.

3 likes 3mo
PM
p.mbekiTL210 May 2026#20
lu.cabrera, post #7: Registration before enrolment, with the primary endpoint declared, is what makes outcome switching detectable. Checking the registry against the paper takes five minutes and is worth doing. Go to post

Post #17 is right about the mechanism and I think understates the practical bit.

What I can speak to on Primary endpoint hierarchies is narrow, so I will keep it narrow rather than generalising from it. Beyond that boundary I do not know.

0 likes in reply to #7 3mo
S
SHermansenTL2Member11 May 2026#21

Adding a small correction to the Primary endpoint hierarchies summary above rather than a disagreement with it. The substance holds; one of the figures is out by a factor that matters.

29 likes 3mo
NR
n.ramosTL211 May 2026#22
lu.cabrera, post #7: Registration before enrolment, with the primary endpoint declared, is what makes outcome switching detectable. Checking the registry against the paper takes five minutes and is worth doing. Go to post

Worth separating two things that post #20 runs together.

Pre-specification is the property that makes a primary endpoint trustworthy. An endpoint chosen after seeing the data can be the best endpoint in the world and it no longer carries the same guarantee.

0 likes in reply to #7 3mo
R
RodriguesTL3Regular11 May 2026#23

Post #22 answers the question as asked. The question underneath it is different.

The thing about Primary endpoint hierarchies that took me longest to accept is that a plausible mechanism is not evidence of an effect. It is a reason to look, not a result.

20 likes 3mo
AZ
an.zamoraTL212 May 2026#24
IS
isotonic_sheetTL3Regular12 May 2026#25
an.zamora, post #24: Posting my Primary endpoint hierarchies numbers with the method attached so they can be discounted properly. Uncontrolled, unblinded, and collected by someone who wanted a particular answer. Go to post

Adding a note of thanks rather than an opinion. I did not know most of that.

21 likes in reply to #24 3mo
PT
p.trevinoTL212 May 2026#26
v.salgado, post #6: Where I part company with post #3, and it is a narrow parting. A treatment-policy estimand asks what happens to people assigned to a strategy, including those who abandon it. A hypothetical estimand asks what would have happened had everyone continued. Both are legitimate and they give different numbers. I would be interested in a… Go to post

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

0 likes in reply to #6 3mo
Promoted into the documentation commons. The content of this topic is maintained at PIONEER 6 — trial digest, with named maintainers and a review date. The promotion was discussed in doc review. Corrections are best raised against the document, which is the version that gets kept current.

Suggested topics

TopicParticipantsRepliesViewsActivity
Run-in periods and the population they select — the long version
Run-in periods and the population they select — the long version Writing it up because I had to work it out twice and would rather nobody else did. A narrow question about Run-in periods, deliberately narrow,…
ALFREK 2 8.9k 4mo
About the Trials category
Individual randomised trials: design, population, endpoints, results, and limitations. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its…
JMBNSSZOEM+6 10 27k 12mo
Non-inferiority margins: how they are chosen and how they are abused
Posting this under the heading it deserves: Non-inferiority margins: how they are chosen and how they are abused Everything below is what sits behind that. A question about non-inferiority margins that I…
NTKCTDOPA+42 46 10k 5mo
Interim analyses and stopping rules
Interim analyses and stopping rules Writing it up because I had to work it out twice and would rather nobody else did. Asking about interim analyses and stopping rules on behalf of the question I keep seeing…
TKHJRIGEL+25 29 35k 16d
Coming back to: Sample size calculations, read backwards from the published number
On the subject in the title: Sample size calculations, read backwards from the published number Working notes rather than a conclusion. Sample size calculations — I have the observation and I do not trust my…
PRIEBMMTV+82 88 1.8k 11mo

Related topics — sharing the tags estimand, surrogate endpoints, discontinuation & dropout

TopicParticipantsRepliesViewsActivity
Adjudicated events and why the definition matters — the long version
Adjudicated events and why the definition matters — the long version Writing it up because I had to work it out twice and would rather nobody else did. A narrow question about Adjudicated events, deliberately…
JSEFIBSSRE+125 146 13k 15mo
Follow-up: Open-label extensions: what survives and what does not
Open-label extensions: what survives and what does not — setting out what I have, and where I think it stops being reliable. I have spent a fortnight trying to pin Open-label extensions down and I want to set…
MATPBLCMA+22 26 397 16mo
Interim analyses and stopping rules — does this still hold?
Interim analyses and stopping rules — does this still hold? — that is the question, and I have not found it answered plainly anywhere I have looked. Interim analyses and stopping rules, from the point of view…
JDMHKMFTHL+30 34 29k 13mo
Journal club: the CagriSema phase 2 combination paper
On the subject in the title: Journal club: the CagriSema phase 2 combination paper Working notes rather than a conclusion. CagriSema phase 2 combination paper — I have the observation and I do not trust my…
PFKMM 2 329 21h
About the Journal club category
Our recurring session. One named paper per topic, methods first, conclusions last. Everyone reads before posting. This post is a community wiki: any member at trust level 3 or above can edit it, and every…
KHSDEMSCA+2 6 12k 16mo