The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Data & Tools · Datasets · continued

Contributing data without breaching anyone's privacy posts 31–60

This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.

NS
n.szaboTL24 Mar 2026#31

Include the denominator. A collection of reports without knowing how many people did not report is uninterpretable in the direction everybody wants to interpret it.

Scoping that to what I have actually seen rather than what I have read.

0 likes 5mo
ID
integrator_draftTL3Regular4 Mar 2026#32
PharmNotes_Whitfield, post #14: Anonymisation matters even for volunteered data. A combination of region, dates and a distinctive detail identifies a person more often than people expect. Go to post

The practical version of contributing data without breaching is three sentences long. The rigorous version is three pages and reaches the same conclusion with the conditions attached.

1 like in reply to #14 5mo
FI
f.ibarraTL24 Mar 2026#33

Reproducibility: if sharing data, include enough context (compound, dose, timeframe, method) that someone reading it understands what it represents.

6 likes 5mo
O
OkaforTL3Regular4 Mar 2026#34

Worth separating two things that post #31 runs together.

Privacy: if contributing data, only share data you are comfortable making permanent and public. Once posted, data is persistent.

15 likes 5mo
CH
ca.haddadTL24 Mar 2026#35

Post #32 answers the question as asked. The question underneath it is different.

Missing data should be recorded as missing rather than as zero. The distinction disappears once the file is shared and cannot be recovered.

0 likes 5mo
G
GDashwoodTL3Regular4 Mar 2026#36

Date every record. A dataset assembled over two years without dates cannot distinguish a change over time from a change in who was contributing.

Reading it again, the caveat matters more than the finding.

2 likes 5mo
YR
y.ramosTL24 Mar 2026#37
VK
v.klausenTL3Regular4 Mar 2026#38

Anonymisation matters even for volunteered data. A combination of region, dates and a distinctive detail identifies a person more often than people expect.

This has been discussed before and I could not find the thread, so, again.

22 likes 5mo
MS
m.silvaTL25 Mar 2026#39

Post #36 is the version of this I will quote in future. One addition.

Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number.

22 likes 5mo
BP
b.petrovTL25 Mar 2026#40
Leiterman, post #10: Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number. Go to post

Where I part company with post #38, and it is a narrow parting.

Outliers should be visible in the published data even if excluded from the analysis, with the exclusion rule stated in advance rather than after seeing them.

0 likes in reply to #10 5mo
AZ
a.zamoraTL25 Mar 2026#41
g.tanaka, post #6: Sensible. I would want the same detail before I acted on it either. Go to post

Useful. I have added it to my own notes with the date on it.

7 likes in reply to #6 5mo
CS
c.silvaTL25 Mar 2026#42
v.klausen, post #38: Anonymisation matters even for volunteered data. A combination of region, dates and a distinctive detail identifies a person more often than people expect. This has been discussed before and I could not find the thread, so, again. Go to post

Where a dataset is used to support a claim in a maintained document, the version used should be cited. Otherwise the document and the data drift apart silently.

It is the sort of thing that seems obvious in retrospect and was not at the time.

1 like in reply to #38 5mo
RG
r.girardTL25 Mar 2026#43
ZI
z.iyerTL25 Mar 2026#44

Picking up post #42: that is the part I would want checked first.

Limitations of datasets: all community-collected data has limitations. The population is self-selected (people in this community are not representative of all people using these compounds). Reporting bias is real (remarkable outcomes get reported; mundane outcomes do not).

17 likes 5mo
AN
a.novakTL25 Mar 2026 · edited#45
c.grimaldi, post #11: Using data in discussions: datasets are useful as reference points when someone claims something unusual. "I have not seen that reported in the data" is different from "that is impossible", but data gives you something to say. I have changed my mind on this once already, so take it as current rather than settled. Go to post

Combining data from different sources: datasets from this site are not directly comparable to published trials because the populations are different. They are worth reading separately, not merged together.

That is what I would do. It may not be what is correct.

11 likes in reply to #11 5mo
AA
a.adebayoTL25 Mar 2026#46

Distinguishing three things in the contributing data without breaching discussion that keep getting used interchangeably: the observation, the proposed mechanism, and the recommendation that gets attached to both.

3 likes 5mo
BF
b.fonsecaTL25 Mar 2026#47

Post #44 describes the usual case. This is about the unusual one.

Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current.

0 likes 5mo
DT
d.tammTL25 Mar 2026#48

How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset.

24 likes 5mo
RZ
r.zielinskiTL25 Mar 2026#49

Post #47 put the caveat in the right place and I want to underline it.

Using data in discussions: datasets are useful as reference points when someone claims something unusual. "I have not seen that reported in the data" is different from "that is impossible", but data gives you something to say.

The honest answer is that it depends, and here is what it depends on.

1 like 5mo
RA
r.aldana_pharmdTL4Pharmacist5 Mar 2026#50
ca.haddad, post #35: Post #32 answers the question as asked. The question underneath it is different. Missing data should be recorded as missing rather than as zero. The distinction disappears once the file is shared and cannot be recovered. Go to post

I will take the caveat as seriously as the claim, which is the point of putting it there.

0 likes in reply to #35 5mo
RA
r.arbuthnotTL1Member5 Mar 2026 · edited#51

How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset.

12 likes 5mo
AE
a.eriksenTL25 Mar 2026#52

Post #48 put the caveat in the right place and I want to underline it.

On contributing data without breaching: the maintained page in the documentation commons covers the general case with citations and a review date, which is more reliable than any reply here including this one.

26 likes 5mo
VD
vial_deskTL3Regular5 Mar 2026#53

That is the distinction I keep failing to hold on to. Written down now.

0 likes 5mo
MA
m.adebayoTL25 Mar 2026#54
r.aldana_pharmd, post #50: I will take the caveat as seriously as the claim, which is the point of putting it there. Go to post

Privacy: if contributing data, only share data you are comfortable making permanent and public. Once posted, data is persistent.

I have seen it go both ways, which is why I hedge.

4 likes in reply to #50 5mo
LC
l.chevalierTL3Regular5 Mar 2026#55
k.chukwu, post #20: Contributing data without breaching sits at the boundary between what this community can usefully discuss and what it cannot, and I think it falls on the discussable side, narrowly. Go to post

Post #54 is the version of this I will quote in future. One addition.

My experience of contributing data without breaching contradicts the reply above. I am posting it as a data point rather than as a refutation, because one person's experience is exactly that.

8 likes in reply to #20 5mo
TV
t.verhoevenTL26 Mar 2026#56

Where I part company with post #52, and it is a narrow parting.

A dataset is only useful if the collection method is described alongside it. Numbers without a protocol are a list rather than data.

The right answer here may simply be that it has not been measured.

18 likes 5mo
TT
taper_tableTL3Regular6 Mar 2026#57

Include the denominator. A collection of reports without knowing how many people did not report is uninterpretable in the direction everybody wants to interpret it.

I keep a log of this specifically because memory is unreliable about it.

0 likes 5mo
SV
s.vanheckeTL26 Mar 2026#58

Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.

2 likes 5mo
EC
excursion_checkTL3Regular6 Mar 2026#59

State the inclusion criteria explicitly, including the ones you applied without noticing. "Records I could read" is a criterion.

I would want the raw data before agreeing with my own summary of it.

4 likes 5mo
RM
r.mwangiTL26 Mar 2026 · edited#60

Anonymisation matters even for volunteered data. A combination of region, dates and a distinctive detail identifies a person more often than people expect.

Worth saying I have only my own numbers here, and n is small.

13 likes 5mo