The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Data & Tools · Datasets · continued

Aggregated purity results across four services, with caveats — a second dataset posts 121–150

This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.

SO
s.ostergaardTL210 Sep 2025#121

Units in the column header, always, and the same units down the whole column. Mixed units in one field is the commonest defect in shared spreadsheets here.

The literature is thinner on this than the confidence in the thread implies.

0 likes 11mo
BV
bias_varianceTL4Biostatistician10 Sep 2025#122

I would rather this thread reach "we do not know" about Aggregated purity results across four than reach a confident answer that nobody can support when asked.

1 like 11mo
DF
d.ferreiraTL211 Sep 2025#123
t.dumitru, post #7: Aggregating first-hand accounts does not produce evidence of the kind a trial produces. It produces a description of who chose to post, which is a real thing and a different thing. Go to post

That is a cleaner way of putting what I was circling around.

6 likes in reply to #7 11mo
B
batchlogTL3Regular11 Sep 2025 · edited#124

On post #122 — agreed on the reasoning, with one qualification.

Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands.

It took me longer than it should have to see that.

15 likes 11mo
CC
c.chowdhuryTL211 Sep 2025#125

Post #124 answers the question as asked. The question underneath it is different.

I changed my mind about Aggregated purity results across four after someone here asked me for the source and I could not produce one. That is worth saying out loud because it is the ordinary way it happens.

0 likes 11mo
TY
two_year_lineTL3Regular11 Sep 2025#126
r.sobczak, post #78: Saving this. It is the version I will quote when the question comes round again. Go to post

What would change my mind on Aggregated purity results across four is a second dataset collected by someone with no stake in the first. Until then I hold it loosely and I would rather say so than pretend to more.

3 likes in reply to #78 11mo
IB
i.balogunTL211 Sep 2025#127
s.rasmussen, post #120: Two sentences on Aggregated purity results across four and then I will stop, because the rest is speculation and the thread is better without mine. What is documented is narrow. What is inferred from it is broad. The gap between them is where every argument here lives. Go to post

Units in the column header, always, and the same units down the whole column. Mixed units in one field is the commonest defect in shared spreadsheets here.

10 likes in reply to #120 11mo
JR
j.rasmussenTL2Regular11 Sep 2025#128

Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned.

22 likes 11mo
YA
y.adeyemiTL211 Sep 2025#129

Sample size is not the only thing that determines what a dataset can support. Selection is usually the larger problem here and it does not improve with volume.

Noting that the question and the thing people usually mean by it are different.

0 likes 11mo
AT
a.thorneTL2Wiki editor11 Sep 2025#130
ambient_review, post #67: Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands. If it helps: the failure mode here is usually boring rather than dramatic. Go to post

Where I part company with post #126, and it is a narrow parting.

The most useful dataset this community could hold is boring: lot, supplier, service, method, date, result. That is enough to answer most of the questions people ask badly.

Written quickly, so the reasoning may be tighter than the wording.

5 likes in reply to #67 11mo
AL
a.lindqvistTL211 Sep 2025#131

Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current.

29 likes 11mo
GV
g.valckenaereTL3Regular11 Sep 2025 · edited#132

Narrowing post #129, because the general version has more than one answer.

The number people quote for Aggregated purity results across four is a central estimate presented without its interval, and the interval is wide enough that the estimate is nearly uninformative on its own.

14 likes 11mo
SD
st.dialloTL211 Sep 2025#133
a.adebayo, post #109: Aggregated purity results across four has a well-known answer and a correct answer, and the interesting work is establishing that they are the same. Nobody has done that here yet. Go to post

Post #129 put the caveat in the right place and I want to underline it.

Whatever the answer on Aggregated purity results across four turns out to be, the method for getting there is the same: state the assumption, do the arithmetic in public, invite the correction.

5 likes in reply to #109 11mo
H
HHidalgoTL2Member11 Sep 2025#134

How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset.

The variance between people here is larger than the effect being discussed.

0 likes 11mo
CS
c.serranoTL211 Sep 2025#135

I read post #133 twice before replying, because I had assumed the opposite.

Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number.

Speaking for myself and not for anyone else who has posted here.

0 likes 11mo
TN
t.nardoneTL3Regular11 Sep 2025#136

Fair, and the limits you put on it are the part I will remember.

20 likes 11mo
IA
i.amankwahTL211 Sep 2025#137
two_year_line, post #126: What would change my mind on Aggregated purity results across four is a second dataset collected by someone with no stake in the first. Until then I hold it loosely and I would rather say so than pretend to more. Go to post

The version of Aggregated purity results across four that I was taught turned out to be a teaching simplification. Useful, and not true in the way I had assumed it was.

9 likes in reply to #126 11mo
JV
j.vandermolenTL3Regular11 Sep 2025#138
f.danquah, post #97: Appreciated. The plain phrasing does more work here than a longer post would. Go to post

One caution on Aggregated purity results across four: everything above assumes the underlying documentation is what it claims to be. That assumption is doing real work and is rarely stated.

2 likes in reply to #97 11mo
EM
e.mensaTL211 Sep 2025 · edited#139

A codebook describing each field takes twenty minutes and is what makes the file usable by anyone but you. Most shared datasets here do not have one.

The right answer here may simply be that it has not been measured.

0 likes 11mo
VD
vial_deskTL3Regular11 Sep 2025#140

Combining results from different testing services into one column loses information that matters. Keep the service as a field.

28 likes 11mo
SA
s.antonsenTL211 Sep 2025#141

A definition problem is doing most of the work in this Aggregated purity results across four discussion. Once the term is pinned down I suspect the disagreement mostly goes away and what is left is small.

0 likes 11mo
QZ
q.zhao_qaTL3Quality assurance11 Sep 2025#142

What I would check first on Aggregated purity results across four is whether the thing being measured moved or whether the way of measuring it moved. Those look identical in a graph.

0 likes 11mo
NV
n.villalobosTL211 Sep 2025#143
vial_desk, post #140: Combining results from different testing services into one column loses information that matters. Keep the service as a field. Go to post

Post #142 is the version of this I will quote in future. One addition.

A dataset is only useful if the collection method is described alongside it. Numbers without a protocol are a list rather than data.

That is the version I use. It may not be the version that is correct.

8 likes in reply to #140 11mo
FP
forest_plotTL3Evidence synthesis11 Sep 2025#144

Publishing the raw records alongside the summary is what makes a dataset checkable. A summary alone asks for trust that nobody has earned.

I am not the right person to answer the follow-up to this.

19 likes 11mo
ZC
z.cardosoTL211 Sep 2025#145
EN
electrolyte_notesTL2Regular11 Sep 2025#146

Post #144 describes the usual case. This is about the unusual one.

Independent test results are the most valuable data this community collects, and they are only comparable when the method is captured alongside the number.

I would be glad to be shown a cleaner way of putting this.

0 likes 11mo
RL
r.lundgrenTL211 Sep 2025#147
r.szabo, post #14: How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset. Happy to expand any of that if it is the useful part. Go to post

Helpful, and short, which on this subject is harder than long.

5 likes in reply to #14 11mo
DM
d.moreauTL2Regular11 Sep 2025 · edited#148

Careful with the language on Aggregated purity results across four. "Not detected" and "not present" are different findings and the first is a statement about the method.

13 likes 11mo
VB
v.bruunTL211 Sep 2025#149

Post #146 answers the question as asked. The question underneath it is different.

Where a dataset is used to support a claim in a maintained document, the version used should be cited. Otherwise the document and the data drift apart silently.

Marking that as an opinion rather than a finding.

0 likes 11mo
KR
k.redgraveTL2Member11 Sep 2025#150

I read post #148 twice before replying, because I had assumed the opposite.

An honest declaration on Aggregated purity results across four: I have a prior here and it is strong enough that you should weight what I say downward. Stating it rather than hiding it.

4 likes 11mo