VietnamScholar.comVietnam Scholar
+84 926 138 138
Data and Ethics

Datasets and code as outputs: the fastest way to give partners something that counts

A published dataset with its own authorship is a citable output that does not compete with anyone's paper.

Datasets and code as outputs: the fastest way to give partners something that counts

Arguments about author position on a paper are zero-sum: one first-author slot, several people who deserve it.

A published dataset is not. It is a separate citable output with its own author list, and the people who built it can lead it without displacing anyone from the paper.

Why publish data separately

  • It creates a citable output for people whose contribution was collection and curation.
  • It is often required by funders and increasingly by journals.
  • It makes the work reusable and extends its life.
  • It relieves pressure on the paper author list.

Who should author a dataset

  • Those who designed the collection instrument.
  • Those who collected and cleaned the data.
  • Those who curated and documented it.
  • This list is legitimately different from the paper author list.

What a usable dataset needs

  • A persistent identifier from a proper repository.
  • Documentation sufficient for someone with no contact with the team.
  • An explicit licence.
  • A version number and a stated citation format.
  • Clear statement of what has been anonymised and how.

Software as an output

  • Tag releases and give each one an identifier.
  • Include a citation file so users cite it correctly.
  • State the licence and the level of support offered.
  • Analysis code accompanying a paper should be published alongside it.

Access restrictions where needed

  • Sensitive data can be published under controlled access rather than not at all.
  • Publish the metadata even when the data itself is restricted.
  • State the access process clearly so requests are possible.
  • Restricted access is still a citable output.

One thing worth remembering

A dataset publication does not compete with anyone for a position on the paper.

That makes it the easiest genuine credit to give — no negotiation, no displacement, and it lands in exactly the places that count: databases, identifiers, and citation records that follow the person for the rest of their career.

Câu hỏi thường gặp

Why publish data as a separate output?

Because it creates a citable output for those whose contribution was collection and curation, is increasingly required, makes work reusable, and relieves pressure on the paper author list.

Who should author a dataset?

Those who designed the collection instrument, collected and cleaned the data, and curated and documented it — a list legitimately different from the paper author list.

What does a usable dataset need?

A persistent identifier from a proper repository, documentation sufficient for someone with no contact with the team, an explicit licence, a version number and citation format, and a clear anonymisation statement.

How should software be published?

Tag releases and give each an identifier, include a citation file, state the licence and level of support, and publish analysis code alongside the paper it accompanies.

What about sensitive data?

It can be published under controlled access rather than not at all — publish the metadata even when the data is restricted, and state the access process clearly.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
Call now +84 926 138 138