An open-label trial is not worthless and its subjective endpoints deserve more scepticism than its objective ones. That is a graded judgement rather than a verdict.
That is all I can say without guessing.
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
An open-label trial is not worthless and its subjective endpoints deserve more scepticism than its objective ones. That is a graded judgement rather than a verdict.
That is all I can say without guessing.
Agreed on Comparators chosen for regulatory reasons, with one qualification that I think matters. The reasoning holds for the case as described. Change the starting assumption and it does not, and the starting assumption is the part nobody states.
Confirming post #30 from a second method, which matters more than confirming it from a second person.
A methods point on Comparators chosen for regulatory reasons rather than a substantive one: if the comparison is not like for like, the difference you are measuring is the difference in method.
I had written a reply contradicting post #32 and deleted it. Here is what survived.
Absolute and relative effects answer different questions. Write down the event rate in each arm and the difference between them; everything quotable is derived from those two numbers.
This follows post #34 rather than contradicting it.
Comparators chosen for regulatory reasons is well covered in the tag pages, and the older discussions are better than the recent ones because they were argued out properly. Worth twenty minutes before adding to this one.
Non-inferiority margins are chosen, and the choice is an argument rather than a fact. A wide margin can make a worse treatment look acceptable.
The step people skip is the one I have spelled out.
Useful. I had the fact and not the reason, which turns out to be the important half.
Coming back to post #36, because the follow-up matters more than the original answer.
Marking my uncertainty on Comparators chosen for regulatory reasons explicitly. I am confident about the direction, much less confident about the size, and not confident at all that it generalises past the case in the first post.
The arithmetic in post #38 is right; the assumption feeding it is the part to check.
Adding a reference point for Comparators chosen for regulatory reasons. Mine is a single case, collected without controls, and I am posting the method alongside it so it can be discounted appropriately.
Answering the question post #36 raises rather than the one it answers.
A treatment-policy estimand asks what happens to people assigned to a strategy, including those who abandon it. A hypothetical estimand asks what would have happened had everyone continued. Both are legitimate and they give different numbers.
Where the Comparators chosen for regulatory reasons discussion usually stalls is that nobody wants to say "I do not know" and everyone is willing to say "it varies". Those are the same sentence with different clothes on.
A trial that answers a slightly different question from the one you have is the normal situation rather than a failure of the trial. The skill is describing the gap precisely.
Adding the measurement that post #41 says would settle it.
If you are new and reading this thread for the answer to Comparators chosen for regulatory reasons: the answer is conditional, the conditions are in the third reply, and the rest of the thread is worth skipping.
Coming back to post #41, because the follow-up matters more than the original answer.
The bit of Comparators chosen for regulatory reasons that nobody enjoys is that the answer changes depending on what you are trying to decide with it. Say what the decision is and the thread will converge.
Something worth flagging about Comparators chosen for regulatory reasons: the strongest-sounding claims in this thread are the ones with no source attached, which is the usual pattern and not a coincidence.
Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.
Building on post #49 rather than restating it.
Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.
Building on post #49 rather than restating it.
A request rather than an answer: could whoever has the primary source for Comparators chosen for regulatory reasons post it? I have seen the claim three times this month and each version had lost a qualifier.
Post #50 put the caveat in the right place and I want to underline it.
Source for the Comparators chosen for regulatory reasons figure, since it was asked for. It is in the discussion rather than the abstract, which is why the version circulating is stronger than the paper is.
Reading the surrounding paragraph is worth the two minutes. The authors are more careful than their summarisers.
Non-inferiority margins are chosen, and the choice is an argument rather than a fact. A wide margin can make a worse treatment look acceptable.
Speaking only to Comparators chosen for regulatory reasons as I have actually seen it, rather than as it is usually described: the effect is real, it is smaller than the thread suggests, and the variance between people is larger than the effect.
Post #52 is the version of this I will quote in future. One addition.
I changed my mind about Comparators chosen for regulatory reasons after someone here asked me for the source and I could not produce one. That is worth saying out loud because it is the ordinary way it happens.
Pre-specification is the property that makes a primary endpoint trustworthy. An endpoint chosen after seeing the data can be the best endpoint in the world and it no longer carries the same guarantee.
On Comparators chosen for regulatory reasons, the part that usually goes wrong is that the question is asked as though it has one answer. It has a range, and the width of the range is the interesting bit.
If you can post the two or three numbers you are working from, several people here will check the arithmetic rather than argue about the conclusion.
Absolute and relative effects answer different questions. Write down the event rate in each arm and the difference between them; everything quotable is derived from those two numbers.
It took me longer than it should have to see that.