The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Data & Tools · Datasets

Coming back to: A community side-effect dataset, with its response rate and biases

AA
a.amankwahTL28 Jul 2025#1

On the subject in the title: A community side-effect dataset, with its response rate and biases Working notes rather than a conclusion.

Reading back through what has been written here about community side-effect dataset, three questions come up every time and only one has ever been answered properly.

Listing all three, with what I think the state of the answer is for each, so the thread can start further along than the last one did.

0 likes 13mo
TS
t.steenkampTL2Member14 Jul 2025#2

Everything in the opening post holds. The case it does not cover is the one I have.

Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number.

1 like 12mo
CN
c.nybergTL219 Jul 2025#3

Practical experience of community side-effect dataset, offered as one case with the conditions stated, not as a general finding. Conditions first, because they are what make it interpretable.

6 likes 12mo
P
PWendelboeTL1Member23 Jul 2025#4

Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned.

A modest claim, modestly supported.

17 likes 12mo
MA
mi.amankwahTL227 Jul 2025#5

A dataset is only useful if the collection method is described alongside it. Numbers without a protocol are a list rather than data.

I have kept the units in throughout, for the obvious reason.

6 likes 12mo
L
LJankowiakTL3Regular31 Jul 2025#6

Clear enough that I do not think I have a follow-up, which is unusual.

16 likes 12mo
AK
ak.kravchenkoTL24 Aug 2025 · edited#7

Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current.

31 likes 12mo
AD
ambient_draftTL3Regular7 Aug 2025#8
c.nyberg, post #3: Practical experience of community side-effect dataset, offered as one case with the conditions stated, not as a general finding. Conditions first, because they are what make it interpretable. Go to post

How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset.

Small point, but it is the one that usually catches people.

23 likes in reply to #3 12mo
RS
r.sobczakTL211 Aug 2025#9

Reframing community side-effect dataset slightly, because I think the disagreement is about the question rather than the answer. If the question is "does it happen", yes. If it is "how often", nobody here knows.

10 likes 12mo
EA
e.almeidaTL2Member14 Aug 2025#10

Bias toward positive outcomes: datasets collected by members are biased toward people who found the compounds useful. People who did not respond do not return. People who had bad outcomes might have left the community.

Not disagreeing with anyone above, just adding the bit I keep having to look up.

22 likes 11mo
KM
k.marchandTL217 Aug 2025 · edited#11

Sample size is not the only thing that determines what a dataset can support. Selection is usually the larger problem here and it does not improve with volume.

5 likes 11mo
MH
m.haddadTL2Regular20 Aug 2025#12

Building on post #11 rather than restating it.

Where a dataset is used to support a claim in a maintained document, the version used should be cited. Otherwise the document and the data drift apart silently.

It is worth stating the boring hypothesis before the interesting one.

0 likes 11mo
EN
e.nilsenTL223 Aug 2025#13
PWendelboe, post #4: Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned. A modest claim, modestly supported. Go to post

Everything in post #11 holds. The case it does not cover is the one I have.

Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands.

Not the answer, but possibly the question that gets there.

28 likes in reply to #4 11mo
FT
fr.translation_moTL2Translator · FR26 Aug 2025#14
t.steenkamp, post #2: Everything in the opening post holds. The case it does not cover is the one I have. Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number. Go to post

The most useful dataset this community could hold is boring: lot, supplier, service, method, date, result. That is enough to answer most of the questions people ask badly.

Adding a source would improve this post and I do not have one to hand.

14 likes in reply to #2 11mo
VS
v.stanescuTL229 Aug 2025#15

Post #11 and I disagree about the size of the effect, not about the direction.

Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.

I am confident about the direction and much less about the magnitude.

2 likes 11mo
AL
aliquot_lineTL3Regular1 Sep 2025#16

Reading rather than answering, but this is the post I would point somebody at.

0 likes 11mo
FP
f.piresTL24 Sep 2025#17
ak.kravchenko, post #7: Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current. Go to post

Date every record. A dataset assembled over two years without dates cannot distinguish a change over time from a change in who was contributing.

The strength of my opinion here exceeds the strength of my evidence.

21 likes in reply to #7 11mo
N
NicolaidesTL3Regular6 Sep 2025#18
PWendelboe, post #4: Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned. A modest claim, modestly supported. Go to post

Limitations of datasets: all community-collected data has limitations. The population is self-selected (people in this community are not representative of all people using these compounds). Reporting bias is real (remarkable outcomes get reported; mundane outcomes do not).

9 likes in reply to #4 11mo
LK
l.krastevTL29 Sep 2025#19

Bias toward positive outcomes: datasets collected by members are biased toward people who found the compounds useful. People who did not respond do not return. People who had bad outcomes might have left the community.

It reads as pedantry until the day it does not.

13 likes 11mo
GD
glossary_deskTL3Regular12 Sep 2025 · edited#20
r.sobczak, post #9: Reframing community side-effect dataset slightly, because I think the disagreement is about the question rather than the answer. If the question is "does it happen", yes. If it is "how often", nobody here knows. Go to post

Offering a way to settle community side-effect dataset rather than another opinion about it. Two measurements, taken the same way, a fortnight apart. If the difference is within the noise, the question was not answerable at this precision.

4 likes in reply to #9 10mo
DT
d.tammTL215 Sep 2025#21

Picking up post #20: that is the part I would want checked first.

How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset.

This is the version I would want a new member to read first.

2 likes 10mo
AN
a.nwosuTL217 Sep 2025#22

On post #18 — agreed on the reasoning, with one qualification.

I would put moderate confidence on the mainstream reading of community side-effect dataset and no more. That is not scepticism for its own sake; it is where the sourcing actually stops.

8 likes 10mo
AB
a.batistaTL220 Sep 2025#23
l.krastev, post #19: Bias toward positive outcomes: datasets collected by members are biased toward people who found the compounds useful. People who did not respond do not return. People who had bad outcomes might have left the community. It reads as pedantry until the day it does not. Go to post

Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.

19 likes in reply to #19 10mo
AN
a.novakTL223 Sep 2025#24
t.steenkamp, post #2: Everything in the opening post holds. The case it does not cover is the one I have. Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number. Go to post

Date every record. A dataset assembled over two years without dates cannot distinguish a change over time from a change in who was contributing.

0 likes in reply to #2 10mo
BN
bench_notesTL4 Moderator25 Sep 2025 · edited#25

A codebook describing each field takes twenty minutes and is what makes the file usable by anyone but you. Most shared datasets here do not have one.

4 likes 10mo
EV
e.vargaTL228 Sep 2025#26

Worth separating two things that post #22 runs together.

Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current.

I would rather be precise about what I do not know than vague about what I do.

12 likes 10mo
RA
r.aldana_pharmdTL4Pharmacist30 Sep 2025#27

A dataset is only useful if the collection method is described alongside it. Numbers without a protocol are a list rather than data.

Reporting the observation and leaving the explanation open deliberately.

26 likes 10mo
SO
s.okaforTL23 Oct 2025#28
e.varga, post #26: Worth separating two things that post #22 runs together. Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current. I would rather be precise about what I do not know than vague about what I do. Go to post

Second this, and I would have said it less carefully.

0 likes in reply to #26 10mo
BJ
b.jansenTL25 Oct 2025#29

Where a dataset is used to support a claim in a maintained document, the version used should be cited. Otherwise the document and the data drift apart silently.

7 likes 10mo
BP
baseline_peakTL2Member8 Oct 2025#30

The most useful dataset this community could hold is boring: lot, supplier, service, method, date, result. That is enough to answer most of the questions people ask badly.

18 likes 10mo