Whether an aggregate is worth publishing at all: a disputed topic — does this still hold?
Confirming post #5 from a second method, which matters more than confirming it from a second person.
A codebook describing each field takes twenty minutes and is what makes the file usable by anyone but you. Most shared datasets here do not have one.
The literature is thinner on this than the confidence in the thread implies.
Marking my uncertainty on Whether an aggregate explicitly. I am confident about the direction, much less confident about the size, and not confident at all that it generalises past the case in the first post.
Coming back to post #8, because the follow-up matters more than the original answer.
Where a value is derived rather than measured, mark it. Derived columns get treated as observations the moment the file leaves your hands.
I would not lead a decision with this, but I would not ignore it either.
I read post #12 twice before replying, because I had assumed the opposite.
Aggregating first-hand accounts does not produce evidence of the kind a trial produces. It produces a description of who chose to post, which is a real thing and a different thing.
Read the full topic (19 posts)
This topic was referenced in
- Contributing data without breaching anyone's privacy — what changed sinceData & Tools › Datasets · 2 replies
- Cleaning a self-reported dataset and what you throw away — what changed sinceData & Tools › Datasets · 2 replies
- Revisiting: Sample size in a voluntary survey: the selection problemData & Tools › Datasets · 2 replies
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Sample size in a voluntary survey: the selection problem
Posting this under the heading it deserves: Sample size in a voluntary survey: the selection problem Everything below is what sits behind that. Sample size: setting out the arithmetic in full, because I have…
|
+1 | 5 | 12k | 15mo |
|
Contributing data without breaching anyone's privacy — what changed since
On the subject in the title: Contributing data without breaching anyone's privacy — what changed since Working notes rather than a conclusion. What changes if the standard account of Contributing data without…
|
2 | 58k | 17mo | |
|
Coming back to: A community side-effect dataset, with its response rate and biases
Posting this under the heading it deserves: A community side-effect dataset, with its response rate and biases Everything below is what sits behind that. Community side-effect dataset — I have the observation…
|
+96 | 101 | 33k | 2d |
|
About the Datasets category
Community-collected datasets, their collection methods, and their limitations. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its author and a…
|
+6 | 11 | 42k | 12mo |
|
A community side-effect dataset, with its response rate and biases
A community side-effect dataset, with its response rate and biases — setting out what I have, and where I think it stops being reliable. Community side-effect dataset: setting out the arithmetic in full,…
|
2 | 5.1k | 19h |
Related topics — sharing the tags data table, confounding, observational data
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Deamidation and the close-eluting pair it produces
Deamidation and the close-eluting pair it produces Writing it up because I had to work it out twice and would rather nobody else did. Something about deamidation does not reconcile and I would like a second…
|
+33 | 37 | 46k | 15mo |
|
What a confidence interval means, from scratch — the long version
The question in the title: What a confidence interval means, from scratch — the long version I will give what I have already checked below so nobody repeats it. Asking about confidence interval directly,…
|
+16 | 20 | 2.3k | 8mo |
|
A claim built entirely on a subgroup analysis — what changed since
On the subject in the title: A claim built entirely on a subgroup analysis — what changed since Working notes rather than a conclusion. What changes if the standard account of claim built is wrong? I ask…
|
2 | 2k | 18mo | |
|
Coming back to: Why the cheapest option is often not the cheapest
Why the cheapest option is often not the cheapest — that is the question, and I have not found it answered plainly anywhere I have looked. Reading back through what has been written here about cheapest…
|
+21 | 25 | 55k | 13mo |
|
About the Datasets category
Community-collected datasets, their collection methods, and their limitations. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its author and a…
|
+6 | 11 | 42k | 12mo |