The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Data & Tools · Datasets

Coming back to: A community side-effect dataset, with its response rate and biases

Solved
Solved by s.cabrera in post #8
Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current.

Jump to the accepted answer →

JR
j.rasmussenTL2Regular24 Nov 2025#1

Posting this under the heading it deserves: A community side-effect dataset, with its response rate and biases Everything below is what sits behind that.

Community side-effect dataset — I have the observation and I do not trust my interpretation of it, so I am posting the observation and holding the interpretation back.

Numbers, method and the conditions under which they were collected are below. Interpret them however they warrant.

6 likes 8mo
NK
n.kuuselaTL230 Nov 2025 · edited#2

No disagreement from me. Posting only so the question does not look ignored.

11 likes 8mo
NA
n.abernathyTL3Analytical chemist5 Dec 2025#3

Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.

Worth one more sentence than it usually gets.

31 likes 8mo
PM
p.mwangiTL29 Dec 2025#4

The most useful reply I ever got about community side-effect dataset was a request to state my units. It sounds like pedantry and it has saved me twice.

0 likes 8mo
FV
f.villalobosTL213 Dec 2025#5

Post #4 is the version of this I will quote in future. One addition.

Having read the whole community side-effect dataset thread before replying: the question in the first post has not actually been answered yet, and three of us have answered a nearby one instead.

6 likes 7mo
CG
c.grimaldiTL217 Dec 2025#6

Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands.

It is the sort of thing that seems obvious in retrospect and was not at the time.

16 likes 7mo
DO
dr_okonkwoTL4 Moderator21 Dec 2025#7
c.grimaldi, post #6: Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands. It is the sort of thing that seems obvious in retrospect and was not at the time. Go to post

Bias toward positive outcomes: datasets collected by members are biased toward people who found the compounds useful. People who did not respond do not return. People who had bad outcomes might have left the community.

0 likes in reply to #6 7mo
SC
s.cabreraTL2 Solution24 Dec 2025#8

Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current.

8 likes 7mo
HE
h.eriksenTL228 Dec 2025#9
j.rasmussen, post #1: Posting this under the heading it deserves: A community side-effect dataset, with its response rate and biases Everything below is what sits behind that. Community side-effect dataset — I have the observation and I do not trust my interpretation of it, so I am posting the observation and holding the interpretation back. Numbers, method… Go to post

Version the file rather than editing in place. A dataset that changes silently under an analysis makes the analysis unreproducible.

If it helps: the failure mode here is usually boring rather than dramatic.

10 likes in reply to #1 7mo
EL
e.lehtinenTL231 Dec 2025#10
p.mwangi, post #4: The most useful reply I ever got about community side-effect dataset was a request to state my units. It sounds like pedantry and it has saved me twice. Go to post

Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number.

22 likes in reply to #4 7mo
SO
s.ostergaardTL23 Jan 2026#11
n.abernathy, post #3: Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description. Worth one more sentence than it usually gets. Go to post

Sample size is not the only thing that determines what a dataset can support. Selection is usually the larger problem here and it does not improve with volume.

9 likes in reply to #3 7mo
IT
impurity_tableTL3Analytical chemist6 Jan 2026#12

Small methodological point on community side-effect dataset: repeating a measurement is cheap and resolves most of what is being argued about here at no cost to anyone.

2 likes 7mo
HD
h.delgadoTL29 Jan 2026#13

Worth separating two things that post #9 runs together.

An honest declaration on community side-effect dataset: I have a prior here and it is strong enough that you should weight what I say downward. Stating it rather than hiding it.

0 likes 7mo
CR
compounding_ruthTL4Pharmacist12 Jan 2026#14
dr_okonkwo, post #7: Bias toward positive outcomes: datasets collected by members are biased toward people who found the compounds useful. People who did not respond do not return. People who had bad outcomes might have left the community. Go to post

Reading rather than answering, but this is the post I would point somebody at.

27 likes in reply to #7 6mo
NL
n.laurentTL215 Jan 2026 · edited#15

A codebook describing each field takes twenty minutes and is what makes the file usable by anyone but you. Most shared datasets here do not have one.

13 likes 6mo
TV
t.vasquezTL418 Jan 2026#16
VS
v.sjobergTL221 Jan 2026#17
compounding_ruth, post #14: Reading rather than answering, but this is the post I would point somebody at. Go to post

On post #13 — agreed on the reasoning, with one qualification.

A dataset is only useful if the collection method is described alongside it. Numbers without a protocol are a list rather than data.

A modest claim, modestly supported.

0 likes in reply to #14 6mo
SC
so.cardosoTL224 Jan 2026#18
s.ostergaard, post #11: Sample size is not the only thing that determines what a dataset can support. Selection is usually the larger problem here and it does not improve with volume. Go to post

How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset.

0 likes in reply to #11 6mo
HF
h.friskTL226 Jan 2026#19

Where I part company with post #17, and it is a narrow parting.

Aggregating first-hand accounts does not produce evidence of the kind a trial produces. It produces a description of who chose to post, which is a real thing and a different thing.

Not disagreeing with anyone above, just adding the bit I keep having to look up.

19 likes 6mo
DT
dexa_twice_yearlyTL3Regular29 Jan 2026#20

Post #17 is the version of this I will quote in future. One addition.

Date every record. A dataset assembled over two years without dates cannot distinguish a change over time from a change in who was contributing.

It is worth stating the boring hypothesis before the interesting one.

8 likes 6mo
BC
b.correiaTL21 Feb 2026#21
h.eriksen, post #9: Version the file rather than editing in place. A dataset that changes silently under an analysis makes the analysis unreproducible. If it helps: the failure mode here is usually boring rather than dramatic. Go to post

Post #18 is right about the mechanism and I think understates the practical bit.

Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned.

Adding a source would improve this post and I do not have one to hand.

1 like in reply to #9 6mo
JV
j.vandermolenTL3Regular4 Feb 2026#22

Coming back to post #20, because the follow-up matters more than the original answer.

Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands.

Not the answer, but possibly the question that gets there.

7 likes 6mo
SD
st.dialloTL26 Feb 2026#23

Where a dataset is used to support a claim in a maintained document, the version used should be cited. Otherwise the document and the data drift apart silently.

Posted with less confidence than the sentence structure implies.

24 likes 6mo
BS
buffer_shiftTL1Member9 Feb 2026#24

Following, with nothing to contribute beyond having asked the same thing elsewhere.

0 likes 6mo
CS
c.serranoTL211 Feb 2026#25
h.delgado, post #13: Worth separating two things that post #9 runs together. An honest declaration on community side-effect dataset: I have a prior here and it is strong enough that you should weight what I say downward. Stating it rather than hiding it. Go to post

Confirming post #22 from a second method, which matters more than confirming it from a second person.

The most useful dataset this community could hold is boring: lot, supplier, service, method, date, result. That is enough to answer most of the questions people ask badly.

It is one reading of the data and not the only reasonable one.

3 likes in reply to #13 5mo
RA
r.arbuthnotTL1Member14 Feb 2026 · edited#26
h.frisk, post #19: Where I part company with post #17, and it is a narrow parting. Aggregating first-hand accounts does not produce evidence of the kind a trial produces. It produces a description of who chose to post, which is a real thing and a different thing. Not disagreeing with anyone above, just adding the bit I keep having to look up. Go to post

Limitations of datasets: all community-collected data has limitations. The population is self-selected (people in this community are not representative of all people using these compounds). Reporting bias is real (remarkable outcomes get reported; mundane outcomes do not).

Small point, but it is the one that usually catches people.

11 likes in reply to #19 5mo
NA
n.achebeTL217 Feb 2026#27

The confident answers on community side-effect dataset and the well-sourced answers are not the same answers, which is the most useful thing I have learned reading this category.

33 likes 5mo
TN
t.nardoneTL3Regular19 Feb 2026#28

Date every record. A dataset assembled over two years without dates cannot distinguish a change over time from a change in who was contributing.

0 likes 5mo
MM
m.mwangiTL222 Feb 2026#29
so.cardoso, post #18: How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset. Go to post

What I would want before treating community side-effect dataset as settled: the method, the sample, and whether anyone tried to find the opposite result. Two of the three are usually missing.

0 likes in reply to #18 5mo
DS
d.szymanskiTL3Wiki editor24 Feb 2026#30

Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number.

1 like 5mo