An ABAB design with a data table and honest limitations — a second dataset posts 61–90
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
I had written a reply contradicting post #58 and deleted it. Here is what survived.
Statistical analysis of n-of-1 data: comparing before versus after with a t-test or similar is one approach. Plotting the data visually is another. Both are valid.
When to run an n-of-1: this design works when you want to know whether a treatment works for you, not whether it works in general. For that purpose, it is efficient.
Adding the caveat now so it does not have to be extracted later.
Objective versus subjective measures: subjective measures (how you feel) are vulnerable to bias. Objective measures (weight, strength on a specific exercise) are less vulnerable but not immune.
Noting that I have skin in this question and have tried to discount for it.
Right, and stated more narrowly than I would have dared to state it.
Withdrawal and reintroduction is the strongest design available to an individual, and it only works for effects that reverse on a timescale you can observe.
Reporting the observation and leaving the explanation open deliberately.
Picking up post #64: that is the part I would want checked first.
Source for the ABAB design figure, since it was asked for. It is in the discussion rather than the abstract, which is why the version circulating is stronger than the paper is.
Reading the surrounding paragraph is worth the two minutes. The authors are more careful than their summarisers.
I disagree with the framing of ABAB design above, and I think it is a substantive disagreement rather than a terminological one. Setting out why, so it can be checked.
The reasoning depends on an assumption that is doing a lot of work and is never stated. If the assumption holds, the conclusion follows. I do not think it holds generally.
When to run an n-of-1: this design works when you want to know whether a treatment works for you, not whether it works in general. For that purpose, it is efficient.
A rolling mean over several days is far more informative than any single reading for anything that varies day to day, which is nearly everything.
Written in the hope of being told what I have missed.
Coming back to post #69, because the follow-up matters more than the original answer.
What I want from this ABAB design thread is the list of things that would need to be true for the claim to hold. If we can write that list, we can check it.
ABAB design is a good example of a question where the honest answer is boring and the interesting answers are unsupported. I would go with boring.
Collapsed as off-topic by two members at trust level 3 or above
Stopping rules: decide in advance when you will stop measuring (after a defined duration, after a defined number of measurements, or after a defined condition is met). Not deciding in advance means stopping when the result satisfies you, which is bias.
The honest answer is that it depends, and here is what it depends on.
I had written a reply contradicting post #74 and deleted it. Here is what survived.
Designing a personal experiment that could actually change your mind: that is the standard for an n-of-1 design. An experiment designed so that any result confirms what you already believed has not changed anything.
That holds under the stated conditions and I have stated them.
Confirming post #74 from a second method, which matters more than confirming it from a second person.
The best thing about this subcategory is that people post their protocols before their results. That order is what keeps it honest.
It does not tell you about anybody else, which is why aggregating these accounts does not produce evidence of the kind people want it to.
That is what I would do. It may not be what is correct.
Posting my ABAB design numbers with the method attached so they can be discounted properly. Uncontrolled, unblinded, and collected by someone who wanted a particular answer.
Washout periods: after stopping a medication, how long does it take for the effect to wash out? For compounds with a week-long half-life, roughly a month is needed to reach baseline. Using that washout period in a before-after design strengthens the inference.
That is one dataset and I would not build a rule on it.
The useful distinction on ABAB design is between what was measured and what was inferred from it. Both end up in the same sentence and only one of them has error bars.
Post #82 answers the question as asked. The question underneath it is different.
Reporting the whole series rather than the interesting segment is the discipline that makes personal data worth reading. Selective reporting is the default without effort.
Measure the same thing the same way at the same time of day. Most of the noise in personal data is measurement protocol rather than biology.
The short answer was in the first line; everything after is the working.
Narrowing post #86, because the general version has more than one answer.
ABAB design is one of those subjects where the general answer and the answer for a specific case diverge, and the thread will go in circles until someone says which one is being asked for.
For anyone finding this later: the short answer on ABAB design is that it depends on one thing, and the rest of the thread is people identifying which thing.
Building on post #86 rather than restating it.
An n of one tells you about one person, which is the person you are most interested in. That is the whole value and it is not nothing.
Post #88 put the caveat in the right place and I want to underline it.
Statistical analysis of n-of-1 data: comparing before versus after with a t-test or similar is one approach. Plotting the data visually is another. Both are valid.