Measurement error in a self-reported exposure posts 31–60
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
Adding a small correction to the measurement error summary above rather than a disagreement with it. The substance holds; one of the figures is out by a factor that matters.
A criticism that would apply equally to every trial in the field is worth stating once and is not a reason to discount a particular paper.
I had written a reply contradicting post #32 and deleted it. Here is what survived.
Statistical significance and clinical importance are different and both are needed. A significant difference below the minimal important difference is a real finding of no practical consequence.
Two people can read the same figure differently here and both be reasonable.
The version of measurement error that I was taught turned out to be a teaching simplification. Useful, and not true in the way I had assumed it was.
The thing about measurement error that took me longest to accept is that a plausible mechanism is not evidence of an effect. It is a reason to look, not a result.
Generalisability and validity are separate axes. A trial can be internally impeccable and still tell you nothing about the person asking.
Everything in post #35 holds. The case it does not cover is the one I have.
I would keep measurement error and the decision it usually gets used for separate in this thread. They are related and they are not the same question, and merging them is why the last one went badly.
Useful. I had the fact and not the reason, which turns out to be the important half.
Measurement error looks different depending on whether you are reading the primary literature or the summaries of it, and the difference is not in our favour.
Attrition is the failure mode most likely to invalidate a result and the least likely to be discussed. Differential attrition between arms is the specific thing to look for.
That is the honest state of it as of this week.
Everything in post #42 holds. The case it does not cover is the one I have.
Small methodological point on measurement error: repeating a measurement is cheap and resolves most of what is being argued about here at no cost to anyone.
Taking post #44 at face value and following it one step further.
Generalisability: do the inclusion/exclusion criteria narrow the population so much that results do not apply to real people asking about it? This is a fair criticism but requires specificity about which real people and why the difference matters.
On measurement error the community has more anecdote than the confidence in this thread implies, and I include my own contribution in that.
Multiple comparisons: if a paper reports many outcomes, the chance of a spurious association by random chance is real. Pre-specification of primary outcomes matters and secondary analyses are weaker evidence.
Marking that as an opinion rather than a finding.
I read post #46 twice before replying, because I had assumed the opposite.
Careful with the language on measurement error. "Not detected" and "not present" are different findings and the first is a statement about the method.
Collapsed as off-topic by two members at trust level 3 or above
Building consensus on which criticisms matter: if everyone agrees that the sample size is small but only you think that affects the conclusion, maybe your criticism is more idiosyncratic. That does not make it wrong but it is worth noticing.
I would want to see it done twice before believing it once.
This is the answer, and the reason it is the answer is the more useful part.
Criticism is more useful when it is narrower. "The trial answers a different question from the one being asked" is actionable; "the trial is flawed" is not.
It is a small point and it changes the answer, which is an awkward combination.
Post #51 put the caveat in the right place and I want to underline it.
I would call the community position on measurement error likely rather than established, and I would be comfortable defending that hedge.
Hold a trial to the standard something could actually have met. A criticism that no achievable design could have answered is a criticism of the field rather than of the paper.
Reading it again, the caveat matters more than the finding.
Summarising the measurement error thread so far, since it is long and the answer is buried: the first reply has the method, the fourth has the correction to it, and the rest is people agreeing at length.
Where I part company with post #53, and it is a narrow parting.
Since measurement error keeps coming up, it should probably be a maintained page rather than a recurring thread. I am happy to draft it if someone with more direct experience will review it.
Post #57 is the version of this I will quote in future. One addition.
Per-protocol and intention-to-treat analyses answer different questions and neither is the honest one by default. Reporting both is the practice worth insisting on.
A partial answer, offered because a partial answer beats none.
The most useful reply I ever got about measurement error was a request to state my units. It sounds like pedantry and it has saved me twice.