NIH Data Sharing Index Challenge · Phase 2 Finalist
SHAREscore
SHARE · S-Index · S-Impact
Make data the currency of science.
76M
datasets scored
140K
researchers ranked
12
live extensions
9
repositories
Built, live, and self-funded, today. · sharescore.org
sharescore.org · ConductScience Foundation · conductscience.com
About: open, shareable, FAIR data by default
ConductScience.com builds scientific tools
ConductScience.org maintains the standards
Shuhan He, MD
Founder, ConductScience
Faculty, Department of Emergency Medicine & Laboratory of Computer ScienceMass General Hospital · Harvard Medical School
conductscience.com
1,200+
institutions served
280
verifiable methodology citations across 184 journals
NIH Replication Prize Winner
2026
Cited in the methods of
CellNeuronNature CommunicationsScience AdvanceseLifeCell ReportsiScience+ 177 more
sharescore.org · ConductScience Foundation · conductscience.com
The vision
SHAREscore · core to the ConductScience mission
Already building the instruments & software that produce open, FAIR data at the source · SHAREscore = the missing layer that measures and rewards how well it's shared · turns a mandate into momentum.
Why this is the mission, not a side project
Instrumentation outputs FAIR, open data with provenance receipts at the source
Already ship ~2,900 instruments producing open, FAIR data at the source, at scale
SHAREscore = how we measure and reward it · sharing finally counts
sharescore.org · ConductScience Foundation · conductscience.com
How it works
Three linked metrics, built from standards the field already agrees on · SHARE grades the dataset · S-Index ranks the researcher · S-Impact weights it by real reuse.
SHARE grades the dataset
Every dataset scored 0–100 from 25 machine-checked signals · quality becomes visible, automated, comparable.
SHARE
the dataset's score
S-Index ranks the researcher
max(k): k datasets each scoring ≥ k · the h-index for data · producing high-quality data finally counts.
S-Index
the h-index for data quality
S-Impact weights it by reuse
Citations earned by a researcher's S-Index core · tracks real reuse.
S-Impact
the reuse measure
sharescore.org · ConductScience Foundation · conductscience.com
Design principles
First principle: FAIR, open data deserves a FAIR, open standard · not one any company or consortium can own or gatekeep.
01
Measure what the researcher controls
Scored: documenting, formatting, licensing, linking at upload
Scored: facts a computer can verify
Not scored: citation luck, field size, topic popularity
02
Every signal earns its place with evidence
Scored: signals that empirically predict reuse across 76M+ objects
Not scored: anything a committee likes but data doesn't back
Rule: future changes must prove predictive validity
03
Governed so no one can capture it
Council: external-majority 6:1, owns the methodology
Open: signals, code, pledges fully forkable
Recourse: disagree and you can fork it
04
Grows and adapts, permissionlessly
Anyone submits: any community member can score a dataset — or a whole repository — not just its owner
Self-onboard: a repository joins with a one-PR pledge · no approval queue
New object types: code, software, protocols score on the same 25 signals, no rule change
sharescore.org · ConductScience Foundation · conductscience.com
Where the 25 signals come from
Grounded in four standards chosen on explicit criteria, authoritative · cross-domain · machine-readable · adopted at scale
DataCite 4.6
2024 · the DOI standard
authoritative registry every research DOI already uses
Dublin Core
DCMI · ISO 15836
the cross-domain metadata foundation, ISO-ratified
schema.org/Dataset
web · Google
what powers Google Dataset Search, adoption at web scale
RDA FAIR Maturity
2020 · the community ref
the community's own yardstick for measuring FAIRness
64candidate fields
↓Standards convergence · in ≥ 2 of the 4 standards
38
↓FAIR mapping · maps to a FAIR sub-principle
33
↓Measurability · a computer can check it anywhere
29
↓Discriminative generalizability · separates good sharing from bad
25signals · 5×5
Five per dimension · not designed, what survived the funnel.
sharescore.org · ConductScience Foundation · conductscience.com
SHARE grades the dataset
SHARE = signals present / 25 × 100 · five dimensions, five machine-checked signals each · one big letter per dimension:
SStewardship
keywords
contributors
subjects
geo_context
temporal_context
HHarmonization
description
methods
references
contributor_pids
org_pids
AAccess
access_status
license_clarity
license_permissive
no_embargo
open_format
RReuse
discovery
access_events
citations
derivatives
community
EEngagement
related_pubs
related_data
funding
versioning
standards_ref
Every signal a binary check, present or absent, computed from public metadata. Each of the 25 worth 4 points: 25 × 4 = 100-point scale.
sharescore.org · ConductScience Foundation · conductscience.com
SHARE maps to FAIR at the signal level
All 25 signals map to FAIR, each under the letter it supports · also a map of data quality.
FFindable8
S1 · Geographic context F2
S2 · Temporal context F2
S3 · Contributor roles F2
H2 · Contributor PID F1
H5 · Description quality F2
S4 · Subject classification F2
R1 · Discovery / views outcome
R5 · Community engagement outcome
F3 · metadata name the data's ID archive-set
F4 · indexed in a searchable resource archive-set
AAccessible3
A1 · Open-access status A1
A4 · No embargo A1
R2 · Access / downloads outcome
A2 · metadata persist after data archive-set
Accessibility is mostly the archive's job (hosting, resolvable IDs), so few signals are depositor-controlled — which is exactly what keeps the score fair. License & format signals also count under Reusable and Interoperable.
IInteroperable7
S4 · Subject classification I2 · controlled vocab
H2 · Contributor PID I2
H3 · Organization PID I2
A5 · Open format I1
H4 · Bibliographic refs I3 · qualified
E1 · Related publication I3 · qualified
E2 · Related data I3 · qualified
RReusable10
A2 · License clarity R1.1
A3 · License permissive R1.1
H1 · Methods documentation R1.2
S3 · Contributor roles R1.2 · provenance
S5 · Ethical transparency R1.2
E3 · Funding information R1.2
E4 · Version tracking R1.2
E5 · Community-standard ref R1.3
R3 · Formal citations outcome
R4 · Derivative works outcome
All 15 FAIR sub-principles, bucketed. We score the 12 a researcher controls; the 3 greyed ones (F3, F4, A2) are set by the archive · scoring them would grade the repository, not the dataset. Violet = reuse outcomes, held out to validate the score.
sharescore.org · ConductScience Foundation · conductscience.com
From SHARE, two career metrics
Every dataset gets a SHARE score · from a researcher's scored datasets, two metrics follow.
SHARE
every dataset scored 0–100
S-Index · the researcher
How many high-scoring datasets
max(k): k datasets each scoring ≥ k on SHARE · the h-index, for data
e.g. 90 72 55 41 12 6 · 4 · 2 → S-Index = 6
S-Impact · the reuse
How much that core is cited
Σ citations of the S-Index core · empty deposits earn nothing
S-Impact = total citations of those 6 · 90 72 55 41 12 6
sharescore.org · ConductScience Foundation · conductscience.com
Section 3
The proof
Eight questions, answered on 183,872 datasets · and NIH's own registry. Next slide: how we tested it.
01Does SHARE at deposit predict reuse?
05Does it hold across fields?
02Does it hold every year?
06Does the S-Index reward quality?
03Does structure beat volume?
07Does S-Impact track influence?
04Is there a real dose-response?
08How much room to improve?
sharescore.org · ConductScience Foundation · conductscience.com
Predictor. SHARE at deposit, S + H + A + E only — the Reuse (R) bucket is held out.
3
Outcome (what we predict). Whether the data gets reused after deposit — a “derivative” (someone builds a new dataset on it) or a citation. Measured later than the score, so it can’t leak in. 407 of the 183,872 (0.22%) gained a derivative.
4
Model. Logistic regression, odds ratio per +10 points, adjusted for year, domain, and license.
Why it's clean
The reuse signals never enter the score.
We score the data at deposit, then watch what actually gets reused later. Because reuse is never part of that score, it can't secretly predict itself — no circularity, no leakage.
sharescore.org · ConductScience Foundation · conductscience.com
The proof · question 1
Does SHARE at deposit predict reuse?
5.73×
the odds of being reused
for every +10 points of SHARE score at deposit · 95% CI 4.97–6.61 · p < 0.001
183,872 Zenodo datasets deposited 2017–2022 · the outcome (a later derivative) is read years after the score
Answer
YES.
Better-documented data at deposit is far more likely to be reused later.
Citations move the same way: 3.02× per +10
p < 0.001 · Zenodo 2016 cohort, 8-year follow-up
Corroborated on NIH's own registry: ClinicalTrials.gov intent-to-share deposits score +17.2 points.
sharescore.org · ConductScience Foundation · conductscience.com
The proof · question 2
Is the effect consistent year over year?
Derivative odds ratio (per +10 SHARE points), by deposit year
Answer
Yes — every year.
Near 5.5× per +10 SHARE points every deposit year tested, 2019 through 2021. Temporally stable, not one unusual year.
sharescore.org · ConductScience Foundation · conductscience.com
The proof · question 3
Does organizing the 25 signals beat just counting how many of the 25 are present?
0.930 AUC
five dimensions, each with its own weight
vs 0.848 if you sum all 25 into one score, or 0.846 for a plain count of how many are present. Letting the five dimensions carry separate weights is the whole +0.08 gain. AUC = how reliably the score separates reused from un-reused data (0.5 = chance, 1.0 = perfect)
Answer
Structure wins.
Counting asks how many of the 25 boxes are ticked. Structure asks which kinds — the five dimensions entered as separate signals. Two datasets with the same count but different profiles score differently, and that difference is what predicts reuse.
sharescore.org · ConductScience Foundation · conductscience.com
Is there a real dose-response?
New datasets built on the original (“derivatives”) per 1,000 deposits, by SHARE bandper 1,000 so bands of different sizes compare fairly · derivatives counted at the 2026 harvest, 4–9 years after deposit
Answer
Yes — a smooth curve.
Reuse climbs across every SHARE band, 32× more derivatives at the top than the bottom — a real dose-response that rises the whole way up.
A derivative is a new dataset someone built on the original — the strongest, hardest-to-fake form of reuse, well beyond a view or a download. That's why we measure it, not raw traffic.
sharescore.org · ConductScience Foundation · conductscience.com
The proof · question 5
Does it hold across fields?
Derivative odds ratio, ecology vs the rest
Answer
Yes — every field.
Ecology is the strongest field (17.3×) but only 5,983 datasets — 3% of the cohort. Drop that 3%, and the other 177,889 (97%) still show 4.9×. Field-general.
We singled out ecology because its data-rich deposits are reused most — the toughest place to rule out a one-field artifact. Removing that small, strongest slice barely dents the effect, and the 4.9× still rests on 177,889 datasets, large and well-powered.
sharescore.org · ConductScience Foundation · conductscience.com
The proof · question 6
Does the S-Index reward quality, or just volume?
S-Index = max(k)
datasets sorted by SHARE · k is where the bars still clear the k×k line · solid bars count, faded ones don't
2,517
most datasets one depositor has
vs
56
highest S-Index observed · avg SHARE 53
Answer
Quality.
The S-Index is max(k): the largest k where a researcher has k datasets each scoring at least k. It's the h-index logic applied to data quality, so 2,000 sloppy deposits can't inflate it.
Complementary, not redundant: pre-registered to correlate with the H-Index; the reuse-weighted S-Impact then shows it tracks real influence — next slide.
sharescore.org · ConductScience Foundation · conductscience.com
The proof · question 7
Does S-Impact track real influence?
S-Impact = total citations of a researcher's S-Index core datasets
1.60×
the odds of senior authorship
on highly-cited papers · for researchers with S-Impact ≥ 5 · p = 0.003
Answer
Yes.
S-Impact counts how often a researcher's best data is actually cited. The independent check: those with S-Impact ≥ 5 are 1.60× more likely to be senior authors on highly-cited papers.
Senior authorship is an outside marker of scientific leadership — nothing to do with our score — so the link isn't circular. The people whose data gets reused are the ones leading influential work.
sharescore.org · ConductScience Foundation · conductscience.com
The proof · question 8
How much room is there to improve?
Answer
A lot.
Scores are widely dispersed — a 56-point spread, 29 to 85 across repositories, mean about 45, so most data sits far below what's achievable. And the spread is a curation choice, not a budget one: SRA is at 85, NASA at 29, both large and well-funded — so any repository can move up.
Average SHARE by repository · DB-authoritative, v20
SRA
85
GEO
74
EDI
56
OpenAIRE
54
ClinicalTrials.gov
50
OpenNeuro
50
Dryad
45
Zenodo
44
NASA
29
Why it's curation, not budget
Same resources, different scores
SRA (85) and NASA (29) are both large, well-funded repositories — the 56-point gap is metadata policy, not money.
A choice, not a constraint
requiring richer metadata at deposit is a policy any repository can adopt — which is exactly what the nudges push.
Mid-scale = the target
the breadth aggregators (44–56) have the most room to move.
sharescore.org · ConductScience Foundation · conductscience.com
Section 4
The live platform
Not a proposal — a running platform, already built and live at sharescore.org.
76M+datasets scored
140K+researchers ranked
500+institutions, 179 countries
APIopen, no-auth, over all 76M+
12live extensions
Self-serveany repo can pledge & be scored
The next slides walk the live surfaces — the global map, the leaderboards, the API, the extensions.
sharescore.org · ConductScience Foundation · conductscience.com
Global reach · /map
179 countries scored, and the score is unbiased. Country researcher-count barely predicts a depositor's score (r = 0.003); Global South 44.7 vs 44.6 worldwide; widest continental gap 1.7 points; dataset-volume barely predicts score (r = +0.02). Quality wins, not resources or size.
sharescore.org · ConductScience Foundation · conductscience.com
Researcher leaderboard · /researchers
Every researcher ranked by S-Index · the h-index for data, made personal · a policy becomes a scoreboard people care about.
sharescore.org · ConductScience Foundation · conductscience.com
Individual researcher profile
Every researcher: a public, explainable profile. S-Index, S-Impact, deposit-time scores, a SHARE-dimension radar, toggleable extensions. Not an abstraction; a page you can visit today.
sharescore.org · ConductScience Foundation · conductscience.com
Institutional leaderboard · /institutions
Rolls up to institutions, too · 500+ ranked by average S-Index, SHARE score, or data-sharing volume, across 179 countries · the same permissionless score, aggregated to where funding and hiring decisions are made.
sharescore.org · ConductScience Foundation · conductscience.com
Live accountability dashboard · /metrics
A real-time scoreboard for the 2023 data-sharing mandate · a program officer sees whether the policy is producing reusable data today, not in a year-end PDF.
What the Metrics tab shows
Mean SHARE score across all 76M+
Datasets scored & repositories pledged
Score trend, up or down over time
Score distribution + SHARE-dimension breakdown
Fully automated rollout
Weekly new-repository scoring
Daily automated tracking
Quarterly public reports
sharescore.org · ConductScience Foundation · conductscience.com
Open API · api.sharescore.org
API live today, what's possible? A no-auth REST API over all 76M+ scored objects. Three things anyone can already do:
Any publisher
drop SHARE (dataset), S-Index (researcher) & S-Impact onto any page, or expose their own repository's SHARE scores
Any repository
be scored end-to-end, giving every researcher in it a live score
Any researcher
look up any other researcher's S-Index, right now
Because it's open, others build on it — how a standard spreads:
A journal · pre-submission SHARE check · authors fix weak metadata before peer review
A university · S-Index in promotion & tenure packets
A funder · applicants ranked by data-sharing track record
A grant reviewer · proposals triaged by past data quality
An AI research agent · datasets ranked by SHARE, reusable data surfaced first
A lab · dashboard benchmarked against peers
sharescore.org · ConductScience Foundation · conductscience.com
Extensions · /extensions, 12 live tools
Twelve lenses ship live · each toggleable in a researcher's own workflow
SHARE Star, overall rating
Originality Lens, anti-duplicate
Compliance Checker, policy check
Data Substance Validator, ghost-data
Pre-submission Feedback, pre-deposit fix
Badges, embeddable proof
Badge Generator, shareable badge
Trend Tracker, over time
Deposit Rate Monitor, activity
SHARE-T, temporal trends
SHARE-C, peer benchmarking
Field Normalizer, field percentile
Example, the Originality Lens (anti-gaming)
Worked example, a 60-dataset profile: S-Index holds at 52 → 52 · all 60 contribute unique concepts, no series inflation · near-duplicate dumps would collapse the unique-concept score, genuine breadth doesn't.
Integrity, in the workflow: several lenses harden it directly · Originality (anti-duplicate) · Data Substance Validator (ghost-data detection) · Compliance Checker (policy & transparency).
sharescore.org · ConductScience Foundation · conductscience.com
Section 5
Implementation
How the standard actually runs, four ways:
1A competitive mechanism that drives adoption
2Governed so no one can capture it
3Changed only by an open process
4Funded to outlast any grant
sharescore.org · ConductScience Foundation · conductscience.com
Compliance becomes a competition
A mandate sets a floor · publishing every repository's SHARE ceiling turns compliance into a race to the top · adoption compounds on its own, no new enforcement.
A competitive ecosystem · the ceiling mechanism
60
Repository A
15/25 signals
80
Repository B
20/25 signals
100
Repository C
25/25 → deposits flow here
SHARE_max = signals ÷ 25 × 100 · each new signal a repository supports raises its depositors' ceiling +4 points · repositories compete to host the best data. (numbers illustrative)
sharescore.org · ConductScience Foundation · conductscience.com
Adoption mechanism · both paths work
Anyone can get a repository scored.
Repository pledge
The repository itself
An official maintainer maps the repository's metadata fields to the 25 SHARE signals and we verify them.
→ Confirmed score
Community pledge
Any researcher or third party
Anyone can propose the mapping in one pull request; it scores preliminary immediately and is verified later.
→ Preliminary score, then confirmed
The add-repo wizard · live at sharescore.org/add-repo, no JSON required
1
Repository info
2
API configuration
3
Signal mappings · 0 of 25
4
Review & verify
sharescore.org · ConductScience Foundation · conductscience.com
ConductScience Foundation · open governance and the COI firewall
1 · External-majority control
Council owns all methodology at 6 external : 1 internal · founder the only ConductScience seat · no single party, including us, can move a score.
2 · No revenue depends on scores
$0 of ConductScience revenue tied to any SHARE score or repository relationship · no incentive to game the metric · commercial support capped at 50%.
3 · Fully open & forkable
Rubric CC-BY, code Apache-2.0, all pledges public, anyone can fork it · fiscal sponsor: HTWB 501(c)(3) (EIN 88-3335809), independent of ConductScience.
sharescore.org · ConductScience Foundation · conductscience.com
How changes happen, one open, evidence-based process
1
Proposal
Anyone opens a GitHub proposal: rationale + expected impact on scores.
2
60-day comment
Public for a minimum 60-day comment period.
3
Evidence
Must show improved accuracy, better reuse prediction or fewer gaming vectors.
4
Council review
External-majority Council weighs proposal, comments, evidence.
5
Supermajority
Minor = simple majority; signal or denominator changes = ⅔ + Board ratification.
6
Public changelog
Approved changes ship with rationale, effective date, and migration notes.
Worked examples (NIH updates the DMS Policy, a new gaming vector, a community-proposed signal) and the ratchet & migration rules for existing scores: in the appendix.
sharescore.org · ConductScience Foundation · conductscience.com
Next Steps · the growth loop & roadmap to one billion, mostly live Monday
1B+objects · by end of 2026 · re3data 3,000+ vs our 9
Next datasets in the pipeline Figshare · Mendeley Data · Harvard Dataverse · PANGAEA · ICPSR · UK Data Service · OSF · biomedical tier: GEO · SRA · dbGaP · Europe PMC
Q3 2026
9 → 20 repositories
Submit to ODSS as a DMS-Policy reference metric
Ship FAIR-mapping report; nudge top-100 repos
Q4 2026
20 → 35 repos
First journal to require a SHARE score at data submission
Begin post-incentive re-validation of 5.73×
Q1 2027 · flagship
50+ repos · 1 billion objects
Publish the SHAREscore paper
Align with DMS-Policy reporting (the ask)
Success measures · composition-adjusted:
repos & researchers notified
correction / claim conversion
composition-adjusted SHARE lift
NIH-relevant datasets improved
sharescore.org · ConductScience Foundation · conductscience.com
Dissemination partners · getting SHAREscore into curricula and clinics
Where the next generation learns to value shared data.
~30K
AMSA
American Medical Student Association · SHAREscore built into the teaching curriculum for medical students.
~540
FUN
Faculty for Undergraduate Neuroscience · ~540 faculty carrying it into undergraduate neuroscience programs.
MGB
Biobank
Mass General Brigham Biobank · biomedical integration, letter of support on file.
Plus our own reach at conductscience.com (~22,000 sessions/month) · additional letters of support on file from Dr. Jonathan Rosand and Google for Health.
sharescore.org · ConductScience Foundation · conductscience.com
The ask
We don't need funding. We need recognition.
Already built, live, self-funded · scoring 76M+ datasets across 9 repositories, including NIH's own GEO and SRA · what accelerates it: NIH putting its weight behind a standard, not money.
RECOGNIZE
the NIH Office of Data Science Strategy (ODSS) names the S-Index as a reference metric aligned with the 2023 Data Management & Sharing (DMS) Policy
ANNOUNCE
co-announce it to the research community, press & guidance reach
POINT
direct NIH-repository grantees (GEO, SRA, already scored) to their live SHARE scores
Runs with or without NIH · Apache-2.0, forkable, reproducible from public code · recognition accelerates a flywheel that already turns.
Zero cost. Zero risk. · sharescore.org · Shuhan He, MD
sharescore.org · ConductScience Foundation · conductscience.com
Thank you
The team that builds and runs it, eleven people across medicine, computer science, bioinformatics, math, and machine learning.
Shuhan He, MD
Informaticist · Lead
Yijian Henry He, PhD
Computer Science · CTO
Louise Corscadden, PhD
Community Engagement
Allison Goff, PhD
Bioinformatics
Boyu Peng, MS
Computer Science
Santosh Adhikari, MS
Machine Learning
Lawrence Jiang, BS
Computer Science
Shivani Pimparkar, MS
Bioinformatics
Pedram Safari, PhD
Mathematics
Yaning "Abby" Zheng
Scientific Forecasting
Xinhui Qian, MS
Machine Learning
Operational team, not owners. The ConductScience Foundation 501(c)(3) = a separate governance council controlling the methodology · framework, code, and pledges all open.
sharescore.org · ConductScience Foundation · conductscience.com
Section 8
Appendix
sharescore.org · ConductScience Foundation · conductscience.com · Appendix
S·H·A·E come from the deposit · reuse held out to validate · you can't buy your score.
One rubric, everywhere · a ratchet, never a penalty
25 binary signals on any repository · a score never falls · improving is pure upside.
It shows its work
Every score traces to the exact signals behind it.
Covers what NIH asked for
NIH's criteria
SHARE
FAIR adherence
SHA
Timeliness of sharing
AE
Annotation & completeness
SH
Frequency of reuse
R
Downstream influence
RE
sharescore.org · ConductScience Foundation · conductscience.com · Appendix
Appendix · why the 5.73× holds, and what we've pre-registered
The derivative-reuse effect survives a full robustness battery · the checks we do not yet have are pre-registered, with a published-negatives commitment · honest by design.
It's not an artifact
Dose–response (per 1,000)0 → 15.06
…top vs bottom bin32×
Temporal 2019 / 2020 / 20215.49 / 5.37 / 5.83
Ecology17.30×
Non-ecology4.91×
Monotonic, temporally stable, and field-general, not a threshold, a fluke year, or one data-rich field.
It's not just prolific authors
Derivative rate, S-Index 11–20 bin37.4% vs 10.4%
…significanceχ²=97.4
Holding researcher productivity fixed (same S-Index bin), consistent sharers still create far more derivatives · not a productivity proxy.
Triangulated across 3 outcomes
Derivative (IsSourceOf)5.73×
Any relatedIdentifier10.04×
Citation > 03.58×
2017–2022 cohort (n=183,872); citation OR here is "any citation > 0," distinct from the 2016 8-yr-follow-up cohort's 3.0×. All predictors VIF < 2.5, no collinearity concern.
Pre-registered next, what we don't yet have
Author/lab-covariate model. Published OR adjusts for year, domain, license only · will add prior productivity, dataset size/richness, funding as covariates (fixed-effects / matched).
Pre-registered bar: success = OR ≥ 3.0; negative results published either way. Q4 2026 – Q1 2027.
sharescore.org · ConductScience Foundation · conductscience.com · Appendix
Appendix · when the rubric changes, the ratchet & migration rules
The natural follow-up to the change process: what happens to a score you already earned when the rubric evolves? Four rules, and one deliberate exception.
The rules, stability you can trust
1
Every score is version-stamped. Always tagged with the rubric it was computed under (e.g. v20) · a 2026 score stays comparable to a 2036 score · never apples-to-oranges.
2
An honest score never falls, the ratchet. New signals only raise ceilings · improving your deposit is pure upside, never a penalty.
3
Old & new stay comparable. Every version ships published migration notes · records dual-scored across a transition window · trend lines don't break.
4
Recomputation is deterministic. Scoring a pure function of public metadata · anyone can re-run any version and reproduce the number.
The math. Shown score = max(old-rubric score, new-rubric score) · a new signal can raise the ceiling 84 → 86 · a stricter definition computing 80 leaves the earned 84 intact · no dataset's SHARE ever falls · so the S-Index (largest k with k datasets at SHARE ≥ k) is monotonic non-decreasing across every version · only exception below: a closed exploit.
The one exception
Closing an exploit re-scores, and flags the records that used it.
If a gaming vector is found, we don't grandfather the scores that exploited it · those records re-scored under the fix and flagged, not quietly lowered · in that single case, integrity beats stability · fully transparent in the changelog.
sharescore.org · ConductScience Foundation · conductscience.com · Appendix
Appendix · built to scale, 76M today, 1B+ addressable
Why it scales, the architecture
Public metadata only. Harvest via OAI-PMH, DataCite, OpenAIRE, and open repository APIs (GBIF, NASA CMR, EBI/NCBI), no repository integration, no private access, no opt-in required.
Scoring is O(1) per record. 25 binary signals via a stateless function, embarrassingly parallel and incremental. 76M+ objects scored on commodity infrastructure.
Onboarding a repo is a bounded schema map. ~1–2 person-days each, mostly reused across repositories on the same standard (DataCite, Dublin Core).
No partner has to sign first. Coverage grows unilaterally from public data; pledges only ever raise ceilings.
The addressable ceiling
76M frozen for this competition (~96% OpenAIRE) · all genuine data repositories, not paper indexes · the ceiling depends on the counting unit:
Dataset level · tens of millions
OpenAIRE dataset subset
~93M
DataCite dataset DOIs (subset)
tens of M
Figshare · Dryad · Dataverse · PANGAEA
~10M+
Record / observation level · crosses 1B
GBIF species occurrences
~3B
NASA EarthData granules
1B+
EBI ENA sequence records
billions
1B+
scorable objects from public metadata. Tens of millions at the dataset level; GBIF and NASA EarthData each clear a billion on their own at the record level.
sharescore.org · ConductScience Foundation · conductscience.com · Appendix
Appendix · the live platform, far more than a leaderboard
Datasets browser76M+ records, every one scored & searchable by DOI, repo, reuse
S-Impactthe citation-outcome twin of the S-Index, reuse, not just deposit
Conversational interfacequery mappings & propose pledges in plain language
Pledge registryevery repository's public signal commitment, auditable
Add a repositorypermissionless onboarding, one pull request, no gatekeeper
Live coveragewhat's scored across 9 repositories, updated continuously
sharescore.org · ConductScience Foundation · conductscience.com · Appendix
Appendix · anti-gaming, why the score survives optimization
The gaming move
Why it fails
Mark signals "not applicable"
Fixed /25 denominator. Skipping lowers the score.
Blame the repository
Unsupported scores 0. No free pass.
Repository-shop for a soft grader
Same metadata, same score everywhere.
Hide in a low-scoring venue
A 60 is a 60 anywhere. Ceilings published.
Get an insider deal
Self-service pledge, community-validated.
Claim signals you don't satisfy
Attestation + PR pledge. Checked against metadata.
sharescore.org · ConductScience Foundation · conductscience.com
FAQ · anticipated questions · scope & data types
Does it work for anything besides datasets — code, software, models, protocols?
Yes. SHARE scores any research object with public metadata; the 25 signals are format-agnostic (documentation, licensing, linking, identifiers), so software, models, and protocols score the same way.
How do you score controlled-access data (dbGaP) you can't see?
SHARE scores the metadata record, never the underlying data, so gated datasets are scored like open ones — on how well the deposit is documented, licensed, and linked, which stays visible even when the data isn't.
Does it work across languages and non-English metadata?
The signals are structural (is a license present, a method documented, an ORCID linked), not linguistic, so they compute regardless of language. Identifier and controlled-vocabulary signals are language-independent by construction.
The through-line
One rubric, computed from public metadata on any object type, in any language, open or gated. Nothing here needs the data itself — only its metadata record.
sharescore.org · ConductScience Foundation · conductscience.com · FAQ
FAQ · anticipated questions · fairness & equity
Does SHARE penalize small or null-result datasets?
No. It grades documentation quality, not size or outcome. A small, well-documented null-result dataset can score 100; a huge, poorly-documented one scores low. It rewards stewardship, not volume or "positive" results.
Does the S-Index penalize early-career researchers?
The S-Index is max(k) — breadth of quality capped by count — so it grows with career stage like the h-index. We show it alongside a researcher's best and median SHARE, so a few excellent early deposits still stand out.
Some fields legally can't share raw data (clinical, qualitative). Is SHARE fair to them?
It scores the deposit that is made, including metadata-only and controlled-access records, so a field that shares rich metadata about restricted data still scores well. It measures sharing quality given what can be shared.
Does it favor large, well-resourced institutions?
No. SRA (85) and NASA (29) are both large and well-funded, so the spread is curation policy, not budget. Scoring is from public metadata with no fee or opt-in, so a depositor at any institution is scored identically.
sharescore.org · ConductScience Foundation · conductscience.com · FAQ
FAQ · anticipated questions · method & rigor
Is the score reproducible — would two people get the same number?
Yes. Every signal is a deterministic, binary, machine-checked test on public metadata, computed by open code — no human rating step. Two people, or the public API, always get the same score. Inter-rater reliability is 1.0 by construction.
Is this correlation or causation?
We claim leakage-controlled prediction, not causation: predictor at deposit, outcome later, reuse bucket held out. The pre-registered post-incentive re-validation (roadmap Q4 2026) is the causal test — does nudging a repo up actually raise later reuse.
How good a predictor is it, really?
The four-component deposit-time model reaches AUC 0.930 (0.5 = chance, 1.0 = perfect), out-of-sample and cross-validated. The full ROC and per-band rates are in the robustness appendix.
Why 25 signals — why not more or fewer?
25 is what survived a four-stage funnel from 64 candidate fields (in ≥2 of 4 standards → maps to a FAIR sub-principle → machine-checkable everywhere → discriminative). It wasn't designed to be 25; governance can add or retire signals as evidence warrants.
sharescore.org · ConductScience Foundation · conductscience.com · FAQ
Who pays after the grant — and why is a company running it?
Under $10K a year, self-funded by ConductScience. There's a structural reason it isn't grant-run: NIH's Research Tools Policy encourages dissemination but funds no upkeep, and 2 CFR §200 makes universities poor hosts for standing shared infrastructure. So we run a commercial-in-structure, infrastructure-in-mission hybrid that replaces the grant cliff with ongoing funding for maintenance and version control.
How often are scores updated? What about retractions?
Re-harvested as metadata changes; a ratchet keeps an honest score from falling. The one exception: a retraction or a closed exploit re-scores and flags the affected records rather than grandfathering them.
Can you legally score and publish others' metadata?
Yes — SHARE uses public metadata repositories already expose under open terms. The rubric is CC-BY and the code Apache-2.0, so the whole method is forkable and auditable.
How is SHARE different from F-UJI / FAIRshake / RDA maturity indicators?
Those assess FAIRness against a checklist. SHARE adds three things: it's validated against real reuse (5.73×), it separates deposit-time quality from outcomes to resist gaming, and it rolls up to researcher metrics (S-Index, S-Impact). It complements them, not replaces them.
sharescore.org · ConductScience Foundation · conductscience.com · FAQ
FAQ · anticipated questions
What if everyone games it? The Goodhart & LLM-metadata question
Once a score becomes a target, does it stop meaning anything? Our 5.73× was measured on deposits made before any incentive existed — not a gamed number. Against LLM-mass-produced but empty metadata, three layers. Live today in our extensions store: Data Substance Validator runs 8 checks incl. ghost-dataset detection · Originality Lens flags near-duplicate deposits. Being built now: a semantic check, real science vs boilerplate. The real backstop: SHARE validated against actual reuse · metadata that never earns reuse fails over time. A later rule closing a loophole → records re-scored, not grandfathered.
sharescore.org · ConductScience Foundation · conductscience.com · FAQ
Appendix · why the dataset is the unit of credit
Why now · AI can write infinite prose, but humans generate the real-world data · the paper is becoming a thin wrapper around the dataset underneath.
Our view of science, the paper is becoming a wrapper
Fig 1 · the dataset
DOI 10.5281/zenodo · SHARE 92
LLMs already generate the prose · the paper: a thin, auto-generated wrapper · the durable asset: the dataset underneath · SHAREscore makes it measurable & rewarded.
sharescore.org · ConductScience Foundation · conductscience.com · Appendix
Appendix · roadmap · SHARE 2.0
A proposed next version: move realized reuse out of the score, so all 25 signals are deposit-time and machine-checked — what you control, not citation luck.
The one change: the R bucket
1.0 R = Reuse (outcomes)
views · downloads · citations · derivatives · community — post-deposit, popularity-driven, not controllable
maps to FAIR R1 / R1.1 / R1.2 / R1.3 · candidates pending the same 64→25 derivation funnel
The proof still holds
Our headline result — 5.73× per +10, AUC 0.930, 32× dose-response — is the 4-component deposit-time model. It contains no R at all; realized reuse is already held out as the outcome. Moving R out of the score changes the presentation, not the proof.
Re-runs on existing data only: recompute the composite · re-bin the dose-response · validate the genuinely-new signals · re-run S-Impact on the deposit-time core.
Removes the Principle-1 tension and the S-Impact circularity in one move.
sharescore.org · ConductScience Foundation · conductscience.com · Appendix