Contributing data without breaching anyone's privacy posts 31–60
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
The practical version of contributing data without breaching is three sentences long. The rigorous version is three pages and reaches the same conclusion with the conditions attached.
Where I part company with post #38, and it is a narrow parting.
Outliers should be visible in the published data even if excluded from the analysis, with the exclusion rule stated in advance rather than after seeing them.
Useful. I have added it to my own notes with the date on it.
Where a dataset is used to support a claim in a maintained document, the version used should be cited. Otherwise the document and the data drift apart silently.
It is the sort of thing that seems obvious in retrospect and was not at the time.
Collapsed as off-topic by two members at trust level 3 or above
On post #39 — agreed on the reasoning, with one qualification.
State the inclusion criteria explicitly, including the ones you applied without noticing. "Records I could read" is a criterion.
That is where I would start, not where I would stop.
Picking up post #42: that is the part I would want checked first.
Limitations of datasets: all community-collected data has limitations. The population is self-selected (people in this community are not representative of all people using these compounds). Reporting bias is real (remarkable outcomes get reported; mundane outcomes do not).
Combining data from different sources: datasets from this site are not directly comparable to published trials because the populations are different. They are worth reading separately, not merged together.
That is what I would do. It may not be what is correct.
Post #47 put the caveat in the right place and I want to underline it.
Using data in discussions: datasets are useful as reference points when someone claims something unusual. "I have not seen that reported in the data" is different from "that is impossible", but data gives you something to say.
The honest answer is that it depends, and here is what it depends on.
I will take the caveat as seriously as the claim, which is the point of putting it there.
How to contribute: if you have longitudinal data you want to add, the format is simple: date, measurement, context. Contact the maintainer of the specific dataset.
Post #48 put the caveat in the right place and I want to underline it.
On contributing data without breaching: the maintained page in the documentation commons covers the general case with citations and a review date, which is more reliable than any reply here including this one.
Privacy: if contributing data, only share data you are comfortable making permanent and public. Once posted, data is persistent.
I have seen it go both ways, which is why I hedge.
Post #54 is the version of this I will quote in future. One addition.
My experience of contributing data without breaching contradicts the reply above. I am posting it as a data point rather than as a refutation, because one person's experience is exactly that.
Where I part company with post #52, and it is a narrow parting.
A dataset is only useful if the collection method is described alongside it. Numbers without a protocol are a list rather than data.
The right answer here may simply be that it has not been measured.
Include the denominator. A collection of reports without knowing how many people did not report is uninterpretable in the direction everybody wants to interpret it.
I keep a log of this specifically because memory is unreliable about it.
Self-reported, unblinded, self-selected data has known biases and is still worth collecting, provided every one of those words appears in the description.
State the inclusion criteria explicitly, including the ones you applied without noticing. "Records I could read" is a criterion.
I would want the raw data before agreeing with my own summary of it.