Arguments about author position on a paper are zero-sum: one first-author slot, several people who deserve it.
A published dataset is not. It is a separate citable output with its own author list, and the people who built it can lead it without displacing anyone from the paper.
Why publish data separately
- It creates a citable output for people whose contribution was collection and curation.
- It is often required by funders and increasingly by journals.
- It makes the work reusable and extends its life.
- It relieves pressure on the paper author list.
Who should author a dataset
- Those who designed the collection instrument.
- Those who collected and cleaned the data.
- Those who curated and documented it.
- This list is legitimately different from the paper author list.
What a usable dataset needs
- A persistent identifier from a proper repository.
- Documentation sufficient for someone with no contact with the team.
- An explicit licence.
- A version number and a stated citation format.
- Clear statement of what has been anonymised and how.
Software as an output
- Tag releases and give each one an identifier.
- Include a citation file so users cite it correctly.
- State the licence and the level of support offered.
- Analysis code accompanying a paper should be published alongside it.
Access restrictions where needed
- Sensitive data can be published under controlled access rather than not at all.
- Publish the metadata even when the data itself is restricted.
- State the access process clearly so requests are possible.
- Restricted access is still a citable output.
One thing worth remembering
A dataset publication does not compete with anyone for a position on the paper.
That makes it the easiest genuine credit to give — no negotiation, no displacement, and it lands in exactly the places that count: databases, identifiers, and citation records that follow the person for the rest of their career.
Câu hỏi thường gặp
Why publish data as a separate output?
Because it creates a citable output for those whose contribution was collection and curation, is increasingly required, makes work reusable, and relieves pressure on the paper author list.
Who should author a dataset?
Those who designed the collection instrument, collected and cleaned the data, and curated and documented it — a list legitimately different from the paper author list.
What does a usable dataset need?
A persistent identifier from a proper repository, documentation sufficient for someone with no contact with the team, an explicit licence, a version number and citation format, and a clear anonymisation statement.
How should software be published?
Tag releases and give each an identifier, include a citation file, state the licence and level of support, and publish analysis code alongside the paper it accompanies.
What about sensitive data?
It can be published under controlled access rather than not at all — publish the metadata even when the data is restricted, and state the access process clearly.