The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Data & Tools · Datasets

Whether an aggregate is worth publishing at all: a disputed topic

PE
ppm_errorTL3Analytical chemist10 Jun 2025#1

On the subject in the title: Whether an aggregate is worth publishing at all: a disputed topic Working notes rather than a conclusion.

An honest uncertainty about whether an aggregate rather than a disguised assertion.

I do not know the answer and I have not been able to find one. What I have is the shape of the question, which may be worth more than my guess at the answer.

1 like 14mo
GL
glossary_lineTL1Member12 Jun 2025#2

No notes. Posting so the count is not one.

4 likes 14mo
IG
in.guerreroTL213 Jun 2025#3

Practical note on whether an aggregate: write down what you expect before you look. The number of times I have found what I went looking for is higher than chance would allow.

13 likes 13mo
HN
h.nicolaidesTL3Regular15 Jun 2025#4

Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned.

27 likes 13mo
MD
m.dumitruTL216 Jun 2025#5
Z
ZieglerTL3Regular17 Jun 2025 · edited#6

Worth separating two things that post #4 runs together.

The useful distinction on whether an aggregate is between what was measured and what was inferred from it. Both end up in the same sentence and only one of them has error bars.

8 likes 13mo
ZV
z.vogelTL218 Jun 2025#7

Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.

A qualification I should have led with rather than closed on.

19 likes 13mo
DW
diluent_watchTL2Member19 Jun 2025#8
glossary_line, post #2: No notes. Posting so the count is not one. Go to post

A request rather than an answer: could whoever has the primary source for whether an aggregate post it? I have seen the claim three times this month and each version had lost a qualifier.

0 likes in reply to #2 13mo
AP
ar.petrovTL220 Jun 2025#9
z.vogel, post #7: Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description. A qualification I should have led with rather than closed on. Go to post

Limitations of datasets: all community-collected data has limitations. The population is self-selected (people in this community are not representative of all people using these compounds). Reporting bias is real (remarkable outcomes get reported; mundane outcomes do not).

0 likes in reply to #7 13mo
FF
f.fenwickTL3Regular21 Jun 2025#10
ppm_error, post #1: On the subject in the title: Whether an aggregate is worth publishing at all: a disputed topic Working notes rather than a conclusion. An honest uncertainty about whether an aggregate rather than a disguised assertion. I do not know the answer and I have not been able to find one. What I have is the shape of the question, which may be… Go to post

The confident answers on whether an aggregate and the well-sourced answers are not the same answers, which is the most useful thing I have learned reading this category.

0 likes in reply to #1 13mo
RB
r.bruunTL222 Jun 2025#11
z.vogel, post #7: Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description. A qualification I should have led with rather than closed on. Go to post

Speaking only to whether an aggregate as I have actually seen it, rather than as it is usually described: the effect is real, it is smaller than the thread suggests, and the variance between people is larger than the effect.

28 likes in reply to #7 13mo
SC
sourced_claimsTL3Regular23 Jun 2025 · edited#12

A structured format with one observation per row is worth the initial effort. Wide formats are easier to enter and much harder to analyse.

The number is defensible. The precision I gave it is not.

13 likes 13mo
MR
m.radichTL224 Jun 2025#13

I disagree with the framing of whether an aggregate above, and I think it is a substantive disagreement rather than a terminological one. Setting out why, so it can be checked.

The reasoning depends on an assumption that is doing a lot of work and is never stated. If the assumption holds, the conclusion follows. I do not think it holds generally.

2 likes 13mo
HO
h.oyelowoTL2Regular25 Jun 2025#14

Narrowing post #11, because the general version has more than one answer.

Units in the column header, always, and the same units down the whole column. Mixed units in one field is the commonest defect in shared spreadsheets here.

0 likes 13mo
AA
a.adeyemiTL226 Jun 2025#15
r.bruun, post #11: Speaking only to whether an aggregate as I have actually seen it, rather than as it is usually described: the effect is real, it is smaller than the thread suggests, and the variance between people is larger than the effect. Go to post

Adding a data point of agreement rather than a data point.

20 likes in reply to #11 13mo
SC
s.chowdhuryTL3Regular27 Jun 2025#16
ar.petrov, post #9: Limitations of datasets: all community-collected data has limitations. The population is self-selected (people in this community are not representative of all people using these compounds). Reporting bias is real (remarkable outcomes get reported; mundane outcomes do not). Go to post

On whether an aggregate, the part that usually goes wrong is that the question is asked as though it has one answer. It has a range, and the width of the range is the interesting bit.

If you can post the two or three numbers you are working from, several people here will check the arithmetic rather than argue about the conclusion.

9 likes in reply to #9 13mo
AN
a.nybergTL228 Jun 2025#17

I read post #16 twice before replying, because I had assumed the opposite.

Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands.

0 likes 13mo
QL
quiet_lurkerTL229 Jun 2025#18
KD
k.dahlbergTL230 Jun 2025 · edited#19

Whether an aggregate has been discussed here with more heat than it deserves, mostly because two definitions have been in play the whole time.

14 likes 13mo
AR
a.reyesTL4 Admin1 Jul 2025#20
m.radich, post #13: I disagree with the framing of whether an aggregate above, and I think it is a substantive disagreement rather than a terminological one. Setting out why, so it can be checked. The reasoning depends on an assumption that is doing a lot of work and is never stated. If the assumption holds, the conclusion follows. I do not think it holds… Go to post

This follows post #19 rather than contradicting it.

A codebook describing each field takes twenty minutes and is what makes the file usable by anyone but you. Most shared datasets here do not have one.

5 likes in reply to #13 13mo
RS
r.sobczakTL21 Jul 2025#21

The arithmetic on whether an aggregate is the easy part and it is where the errors are, which is an uncomfortable combination. Show your working and someone will catch it.

19 likes 13mo
EA
e.almeidaTL2Member2 Jul 2025#22
m.radich, post #13: I disagree with the framing of whether an aggregate above, and I think it is a substantive disagreement rather than a terminological one. Setting out why, so it can be checked. The reasoning depends on an assumption that is doing a lot of work and is never stated. If the assumption holds, the conclusion follows. I do not think it holds… Go to post

Post #19 put the caveat in the right place and I want to underline it.

Where the whether an aggregate discussion usually stalls is that nobody wants to say "I do not know" and everyone is willing to say "it varies". Those are the same sentence with different clothes on.

0 likes in reply to #13 13mo
NR
n.ramosTL23 Jul 2025#23

Combining data from different sources: datasets from this site are not directly comparable to published trials because the populations are different. They are worth reading separately, not merged together.

Written in the hope of being told what I have missed.

0 likes 13mo
I
IHollingworthTL2Member4 Jul 2025#24

Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.

I am reporting what happened, not recommending it.

4 likes 13mo
MM
m.marchettiTL25 Jul 2025#25

That is the distinction I keep failing to hold on to. Written down now.

26 likes 13mo

Suggested topics

TopicParticipantsRepliesViewsActivity
Whether an aggregate is worth publishing at all: a disputed topic — does this still hold?
Whether an aggregate is worth publishing at all: a disputed topic — does this still hold? I have a specific reason for asking rather than idle curiosity, and the context is below. An honest uncertainty about…
BDHFBWIBOF+14 18 3.4k 2mo
Revisiting: Sample size in a voluntary survey: the selection problem
Revisiting: Sample size in a voluntary survey: the selection problem — setting out what I have, and where I think it stops being reliable. I have spent a fortnight trying to pin Sample size down and I want to…
RAZAR 2 56k 9mo
About the Datasets category
Community-collected datasets, their collection methods, and their limitations. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its author and a…
BNOFSVPSAA+6 11 42k 12mo
Coming back to: A community side-effect dataset, with its response rate and biases
On the subject in the title: A community side-effect dataset, with its response rate and biases Working notes rather than a conclusion. Reading back through what has been written here about community…
AATSCNPMA+27 31 46k 9mo
A community side-effect dataset, with its response rate and biases
A community side-effect dataset, with its response rate and biases — setting out what I have, and where I think it stops being reliable. Community side-effect dataset: setting out the arithmetic in full,…
SBSPGC 2 5.1k 19h

Related topics — sharing the tags confounding, site feedback, observational data

TopicParticipantsRepliesViewsActivity
TB-500 and thymosin beta-4: the fragment versus the protein
TB-500 and thymosin beta-4: the fragment versus the protein Writing it up because I had to work it out twice and would rather nobody else did. TB-500 and thymosin beta-4: what I expected, what I found, and…
COGDJLDVS+107 118 22k 11mo
Sample size intuition for a personal experiment
On the subject in the title: Sample size intuition for a personal experiment Working notes rather than a conclusion. A follow-up question about sample size intuition that I did not know to ask the first time.…
RKHECAHTT 4 62k 18mo
Second pass at: The mass-error checker and its intended scope
Posting this under the heading it deserves: Second pass at: The mass-error checker and its intended scope Everything below is what sits behind that. Posting a small dataset on mass-error checker. It is mine,…
HROLNVPNBV+70 74 1.7k 8mo
Trust level 3 requirements: a clarification — does this still hold?
Asking directly, because I could not find a straight answer: Trust level 3 requirements: a clarification — does this still hold? Two things I would like separated before anyone answers on Trust level 3…
ASTVIABVGR+77 90 57k 9h
What a confidence interval means, from scratch
What a confidence interval means, from scratch — that is the question, and I have not found it answered plainly anywhere I have looked. Asking about confidence interval directly, because I have read four…
MHSNO 2 2.6k 2mo