Skip to main content

Finding research datasets

Finding a dataset is easy. Finding one you can actually use takes five specific checks, and doing them first saves a week.

How do you find research datasets?

Search a cross-repository index first — DataCite Commons, Google Dataset Search, or re3data to locate the right repository — then the domain repository for your field, which has better metadata than any general index. Also check the data availability statement of papers reporting the measurements you want, since that is where deposits are actually named.

Three ways in

RouteUse whenTools
Cross-repository searchYou do not know where the data would liveDataCite Commons, Google Dataset Search, OpenAIRE, B2FIND
Domain repositoryYou know the field and want good metadataGEO, SRA, PDB, GenBank, ICPSR, PANGAEA, UK Data Service, and dozens more
From the paperYou want the data behind a specific published resultData availability statement, supplementary files, the accession number in the methods

The third route is the most reliable and the most overlooked. A paper reporting exactly the measurement you need will name its deposit in the data availability statement or give an accession number in the methods, and that is a direct pointer no keyword search would have produced.

re3data is a registry of repositories rather than of datasets. Use it to answer where would data like this be deposited, then search there — domain repositories carry structured, field-specific metadata that a general index flattens away.

General repositories

  • Zenodo — CERN-operated, takes anything, issues DOIs, integrates with GitHub for code releases.
  • Dryad — curated, mostly data underlying published papers, strong in ecology and evolution.
  • figshare — very broad, including figures, posters and supplementary material.
  • Open Science Framework — project-oriented, good where data, code, protocol and preregistration belong together.
  • Harvard Dataverse — widely used in the social sciences, with versioning and access controls.

A general repository is where data goes when no domain repository fits. If one does fit, the domain repository is almost always the better place to look: its metadata describes the measurements rather than just the file.

The five checks before you commit

Do these before downloading anything large. Each one has ended a dataset evaluation for someone who did it in the wrong order.

  1. Licence. Look for an explicit CC0, CC BY or equivalent. No licence is not permission — it is the default of all rights reserved, and it will block publication later even if the download works now.
  2. Documentation. Is there a codebook, data dictionary or README that defines every variable, its units and its missing-value codes? Undocumented columns are unusable at any size.
  3. Format and completeness. Open formats, and the actual files present rather than a landing page promising them. Check whether what is deposited is the raw data or only derived summaries.
  4. Provenance and version. Which version is this, what changed, and is there a DOI that resolves to this exact version rather than to the latest one?
  5. Ethics and restrictions. Human-subject data frequently carries consent limits, a data use agreement, or an application process. Find that out before building anything on it.

Citing data

Cite the dataset itself, not only the paper that used it: creators, year, title, repository, version, and the DOI. Most repositories publish a preferred citation on the landing page — use it.

Cite the version you used. Datasets are revised, and a DOI that always resolves to the latest version will silently misdescribe your analysis once the data changes. Where a version-specific DOI exists, that is the one to cite.

When there is no dataset

If a paper reports the measurements you need but deposits nothing, email the corresponding author. Data sharing on request is still the stated policy of many journals, and a specific request naming the variables and the intended use is answered far more often than a general one.

Check the supplementary material first, though. A surprising amount of usable data is buried in a supplementary spreadsheet that no dataset index has ever seen.

Questions

Is Google Dataset Search good enough on its own?
As a starting point, yes. It indexes repositories that publish schema.org dataset markup, which is a wide net. It is weak on field-specific filtering — assay type, organism, instrument, geography — so a domain repository will usually beat it once you know where to look.
What does FAIR data mean?
Findable, Accessible, Interoperable, Reusable: a set of principles for how research data should be published. In practice it means a persistent identifier, rich metadata, an open format, and an explicit licence. FAIR does not mean open — data can be FAIR and still access-controlled.
Can I use a dataset with no licence?
Assume not. Absent an explicit licence, default copyright applies and you have no permission to redistribute or publish derived work. Ask the depositor for a clear licence; most will add one when asked.
How do I find the data behind a specific paper?
Read the data availability statement, then the methods for an accession number, then the supplementary files. If all three are silent, search the corresponding author name in Zenodo, Dryad and the relevant domain repository before emailing them.

Last updated 2026-08-25. Part of the litscout literature search guides.

We use only the cookies needed to run the service — signing you in and keeping the session secure. No analytics, no advertising, no third-party trackers. Privacy Policy