The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Data & Tools · Datasets · continued

Coming back to: A community side-effect dataset, with its response rate and biases posts 61–90

This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.

NL
ne.laurentTL26 May 2026 · edited#61

Limitations of datasets: all community-collected data has limitations. The population is self-selected (people in this community are not representative of all people using these compounds). Reporting bias is real (remarkable outcomes get reported; mundane outcomes do not).

Speaking for myself and not for anyone else who has posted here.

12 likes 3mo
NB
n.bridgewaterTL2Member8 May 2026#62

Bias toward positive outcomes: datasets collected by members are biased toward people who found the compounds useful. People who did not respond do not return. People who had bad outcomes might have left the community.

26 likes 3mo
SL
s.lundgrenTL210 May 2026#63
e.nilsen, post #53: Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description. Go to post

Taking post #62 at face value and following it one step further.

The bit of community side-effect dataset that nobody enjoys is that the answer changes depending on what you are trying to decide with it. Say what the decision is and the thread will converge.

0 likes in reply to #53 3mo
VM
v.milanoviTL312 May 2026#64
NC
n.chowdhuryTL214 May 2026#65

Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number.

Noting that the question and the thing people usually mean by it are different.

8 likes 2mo
AD
ambient_draftTL3Regular16 May 2026#66
n.achebe, post #27: The confident answers on community side-effect dataset and the well-sourced answers are not the same answers, which is the most useful thing I have learned reading this category. Go to post

Reframing community side-effect dataset slightly, because I think the disagreement is about the question rather than the answer. If the question is "does it happen", yes. If it is "how often", nobody here knows.

19 likes in reply to #27 2mo
NK
n.kirchnerTL218 May 2026#67
endo_fellow_rk, post #49: Sensible. I would want the same detail before I acted on it either. Go to post

Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned.

One of those cases where knowing the mechanism does not help the decision.

0 likes in reply to #49 2mo
IT
integrator_traceTL2Member20 May 2026#68

Fair, and the limits you put on it are the part I will remember.

2 likes 2mo
HK
h.kimaniTL222 May 2026#69

Appreciated. The plain phrasing does more work here than a longer post would.

25 likes 2mo
AS
a.schaefferTL2Member24 May 2026#70

A codebook describing each field takes twenty minutes and is what makes the file usable by anyone but you. Most shared datasets here do not have one.

I would treat that as a working assumption and revisit it.

0 likes 2mo
GB
g.bakkenTL226 May 2026 · edited#71

How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset.

Written in the hope of being told what I have missed.

7 likes 2mo
TN
t.nguyen_newTL1Member29 May 2026#72

Sample size is not the only thing that determines what a dataset can support. Selection is usually the larger problem here and it does not improve with volume.

The variance between people here is larger than the effect being discussed.

1 like 2mo
SL
s.lindqvistTL231 May 2026#73
sa.rasmussen, post #44: A dataset is only useful if the collection method is described alongside it. Numbers without a protocol are a list rather than data. I would put the burden of proof on the interesting explanation, not the dull one. Go to post

My experience of community side-effect dataset contradicts the reply above. I am posting it as a data point rather than as a refutation, because one person's experience is exactly that.

0 likes in reply to #44 2mo
ST
sterile_tableTL3Regular2 Jun 2026#74

Building on post #71 rather than restating it.

A dataset is only useful if the collection method is described alongside it. Numbers without a protocol are a list rather than data.

18 likes 2mo
JM
j.marchettiTL24 Jun 2026#75

Aggregating first-hand accounts does not produce evidence of the kind a trial produces. It produces a description of who chose to post, which is a real thing and a different thing.

One more caveat and then I will stop qualifying: the sample selected itself.

11 likes 2mo
FR
figure_reviewTL2Member6 Jun 2026#76

Taking community side-effect dataset seriously for a moment rather than deflecting: the honest position is that the community has observations and no controlled comparison, and those two things support very different sentences.

3 likes 2mo
AA
a.amankwahTL28 Jun 2026#77
l.krastev, post #55: Post #53 describes the usual case. This is about the unusual one. Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands. This is the sort of thing the wiki should carry and currently does not. Go to post

Where I part company with post #75, and it is a narrow parting.

Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.

0 likes in reply to #55 2mo
RM
r.marsdenTL3Regular10 Jun 2026#78
Ziegler, post #38: Summarising the community side-effect dataset thread so far, since it is long and the answer is buried: the first reply has the method, the fourth has the correction to it, and the rest is people agreeing at length. Go to post

Post #75 is the version of this I will quote in future. One addition.

Date every record. A dataset assembled over two years without dates cannot distinguish a change over time from a change in who was contributing.

24 likes in reply to #38 2mo
MY
m.yildizTL212 Jun 2026#79

Small correction to my own earlier position on community side-effect dataset. I had the units the wrong way round, which changes the conclusion by an order of magnitude and therefore changes it entirely.

17 likes 2mo
LM
lyophil_marginTL3Regular14 Jun 2026#80
Nicolaides, post #58: Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current. Go to post

Post #79 answers the question as asked. The question underneath it is different.

Worth separating community side-effect dataset as a question about the compound from community side-effect dataset as a question about the documentation. They get answered by different people and only one of them is answerable here.

7 likes in reply to #58 1mo
I
IsaksenTL3Regular16 Jun 2026#81
h.eriksen, post #9: Version the file rather than editing in place. A dataset that changes silently under an analysis makes the analysis unreproducible. If it helps: the failure mode here is usually boring rather than dramatic. Go to post

Version the file rather than editing in place. A dataset that changes silently under an analysis makes the analysis unreproducible.

Not a strong opinion, just a consistent one.

1 like in reply to #9 1mo
RC
r.coelhoTL218 Jun 2026#82
t.nardone, post #28: Date every record. A dataset assembled over two years without dates cannot distinguish a change over time from a change in who was contributing. Go to post

Answering the question post #78 raises rather than the one it answers.

Sample size is not the only thing that determines what a dataset can support. Selection is usually the larger problem here and it does not improve with volume.

That is the shape of it. The detail is where I would expect to be corrected.

6 likes in reply to #28 1mo
BP
bench_peakTL3Regular20 Jun 2026#83

Practical note on community side-effect dataset: write down what you expect before you look. The number of times I have found what I went looking for is higher than chance would allow.

15 likes 1mo
MN
m.ndiayeTL222 Jun 2026#84

Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands.

31 likes 1mo
FE
footnote_entryTL3Regular24 Jun 2026#85
h.frisk, post #19: Where I part company with post #17, and it is a narrow parting. Aggregating first-hand accounts does not produce evidence of the kind a trial produces. It produces a description of who chose to post, which is a real thing and a different thing. Not disagreeing with anyone above, just adding the bit I keep having to look up. Go to post

Adding the measurement that post #84 says would settle it.

Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.

If anyone can point at the primary source I would be grateful.

3 likes in reply to #19 1mo
HC
h.castellanosTL226 Jun 2026 · edited#86

Post #82 describes the usual case. This is about the unusual one.

Date every record. A dataset assembled over two years without dates cannot distinguish a change over time from a change in who was contributing.

10 likes 1mo
K
KStephanopoulosTL3Regular28 Jun 2026#87

Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned.

I would call that likely rather than established.

22 likes 30d
SV
s.vogelTL229 Jun 2026#88

Noted, and thank you for writing it out rather than summarising it.

0 likes 28d
B
BirkelandTL3Regular1 Jul 2026#89

Picking up post #87: that is the part I would want checked first.

The question underneath community side-effect dataset is usually "how would I tell?" rather than "what is true?", and that one has a method attached to it.

Write down what you would expect to see under each hypothesis before you collect anything. If they predict the same observation, collecting it will not help.

0 likes 26d
PF
p.fontaineTL23 Jul 2026#90
endo_fellow_rk, post #49: Sensible. I would want the same detail before I acted on it either. Go to post

I would put moderate confidence on the mainstream reading of community side-effect dataset and no more. That is not scepticism for its own sake; it is where the sourcing actually stops.

1 like in reply to #49 25d