Skip to main content
Discover RIMS See it live
See the platform live

← Back to blog

Insight

SDG mapping of research output: what universities get wrong

By Discover RIMS Admin · September 28, 2026

Almost every research information system now offers SDG mapping, and almost every university reports SDG figures somewhere. What is rarely said out loud is that two systems analysing the same publications will disagree with each other — often by more than they agree.

The uncomfortable finding

This has been measured repeatedly, and the results are consistent.

  • Armitage, Lorenz and Mikki (2020) compared the Bergen and Elsevier SDG queries in Web of Science. Using the action-based approach, the overlap between the two sets was below 25% for every SDG they tested. Even their best case — SDG 3 with the topic approach — shared 1.19 million publications while leaving roughly 492,000 and 425,000 unique to one approach. Country rankings changed depending on which approach was used.
  • Purnell (2022) compared four methods for SDG 13 alone. The largest pairwise overlap was 29.2%; the smallest was 13.4%.
  • Ottaviani and Stahlschmidt (2024), in a preprint, DOI-matched more than 15 million publications across Web of Science, Scopus and OpenAlex. The three-way overlap ranged from 1.3% to 7.2% depending on the goal.

So an SDG count is not an observation about the world. It is the output of a specific classifier, applied to a specific database, on a specific date.

Why the classifiers disagree

They are built differently, and the differences are documented.

  • Elsevier SDG Research Mapping combines Boolean Scopus queries with a machine-learning layer.
  • Aurora Universities Network publishes pure Boolean queries, built from the 169 SDG targets rather than the 17 goals. Aurora's own evaluation reports average precision around 70% and recall around 14% — deliberately conservative, because the queries use keyword combinations to suppress false positives.
  • Clarivate's InCites schema is not keyword-based at all: SDGs are assembled from citation-clustered Micro Citation Topics plus manual curation.
  • Dimensions uses supervised machine learning trained on curated data.
  • OpenAlex uses multilingual mBERT models trained on Aurora query output.

A keyword method and a citation-clustering method are answering different questions. It should not surprise anyone that they return different sets.

Why SDG 3 always looks large

Health research dominates SDG classifications, and the effect is measurable: one study found Scopus returned 124% more SDG 3 publications than Dimensions for the same period. If your institution's SDG profile is a spike at SDG 3 and a flat line elsewhere, check the classifier before concluding anything about your research strategy.

The related trap is vocabulary. Terms like "poverty" and "hunger" appear across SDGs 1, 2 and 3, which is exactly why Aurora uses keyword combinations rather than single terms.

Versions move, and nobody publishes a crosswalk

SDG mappings are revised. Elsevier's 2021 mapping captured roughly twice as many articles as its 2020 version; the 2022 revision added COVID-related terms to SDG 3; the 2023 revision changed SDGs 1, 4, 5, 7 and 14; the 2025 mapping, published in April 2025, kept the same query and algorithm but still reports slightly different publication counts.

This means a year-on-year SDG trend built across mapping versions is not a like-for-like comparison, and no vendor publishes a crosswalk that would let you correct for it. If your slide shows SDG output rising, make sure the rise is not a mapping update.

What this means for rankings submissions

Times Higher Education's Impact Rankings combine self-submitted institutional evidence with bibliometrics. THE's methodology states that a specific query is created for each SDG to narrow the metric to relevant papers, supplemented by publications identified by artificial intelligence, drawing on Elsevier's SDG mapping. Research accounts for 27% of each SDG score.

The practical consequence: the bibliometric part of your score is produced by someone else's classifier, not yours. Your internal SDG numbers will not match, and they are not supposed to. What you control is the evidence you submit and the consistency of your own reporting.

What a research office should actually do

  • Record the provenance of every SDG figure: which classifier, which version or year, which database, and the extraction date. Without those four, the number cannot be reproduced — including by you, next year.
  • Pick one classifier and stay with it for internal trends. Consistency matters more than choosing the theoretically best method.
  • Never mix sources in one chart. A Scopus-derived SDG count and a Dimensions-derived one do not belong on the same axis.
  • Treat SDG counts as indicative in any document that influences funding decisions, and say so in the footnote. The disagreement between methods is larger than most of the differences leadership will want to act on.
  • Ask your vendor which mapping they use and when it was last updated. If a system cannot tell you, its SDG figures are not auditable.

Frequently asked questions

Which SDG classification is the most accurate?

There is no established winner. Benchmarking work across multiple methods found no clear best performer, and noted that each system tends to score best on its own validation data. The honest framing is fitness for purpose: a conservative keyword method suits auditable reporting, while a machine-learning method may suit discovery.

Why do our SDG numbers differ from the rankings?

Because the ranking uses its own bibliometric classifier over its own database. Unless you replicate the exact mapping version and source, the figures will differ. That is expected, not an error on either side.

Can we compare SDG output year on year?

Only if the classifier and its version stayed the same across those years. Mappings have been revised repeatedly, and revisions change counts without any change in research activity.

Does Discover RIMS classify publications against the SDGs?

Yes, every publication is classified across the 17 goals automatically. If you are using the figures for reporting, ask us which mapping and version we apply so you can record it alongside the numbers — reproducibility depends on that, whichever system you buy.

Related reading

Sources

Checked September 2026.

Related articles