Known challenges

As of 9 August 2026.

What this corpus cannot do yet, measured and dated. These are work in progress, not concessions.

1

Our current limitations

What our corpus does not yet do well, exactly what we know about it, and what it would take to fix.

This page is public for a simple reason: a knowledge infrastructure that does not publish its defects cannot be audited. A figure whose margin of error is unknown cannot be verified — it can only be believed or dismissed. We would rather it were verified.

Each challenge has three parts: what we know, with the status of the measurement · why it is not fixed, without evasion · what it would take to fix it, and who could do it. The third part is why this page exists. Several of these challenges need no funding — they need a reader, and that reader probably exists somewhere.

How to read the figures on this page

COUNTED
a query over the whole set concerned. Exact at its date.
ESTIMATED
a random sample, re-read, with its confidence interval.
READ
someone read excerpts and judged. No third party checked. Weakest status, flagged wherever it applies.
NOT MEASURED
the question exists and we have not answered it.

A “read” figure and a “counted” figure cannot be cited the same way. Without this column, twenty-five re-read excerpts and a count over 63,227 rows would carry the same apparent authority.

The corpus, counted: 72,932 observations extracted from 9,199 publications — 63,227 validated, 9,549 pending review, 156 disputed. The disease map shows 4,043 of them: prevalences and incidences, on a disease, with a country.

The challenges, heaviest first

1.Values that are not of the quantity announced

What we knowREAD

"most regions in Somalia were at least 70% likely to be below 5% prevalence"
                                              -> value kept: 70

That 70 is a probability displayed as a prevalence. Twenty-five observations re-read by hand, seven carried this defect: order of magnitude ~190 observations on the map. It is the most serious defect we know of, despite its modest extent: a reader has no way to spot it. A misattached value at least shows the right quantity; this one does not.

Why it is not fixed. Telling a prevalence from a probability, an attributable fraction or a share of cases requires understanding the sentence, not matching a pattern. Our lexical detectors fail: the one we built for a neighbouring defect reached 53% precision, and a flag that is wrong one time in two stops being read.

What it would take. An annotated set of 300 observations, labelled by quantity. The sampling protocol exists and has been used three times; the calibration costs under a dollar. What is missing is the annotator: an epidemiologist or biostatistician, about two days’ work. With that set, a classifier becomes measurable — and if it does not reach sufficient precision, we will say so rather than ship it.

2.Values that do not measure the disease they are filed under

What we knowESTIMATED

12.1% to 14.1% of the map    95% CI [5.7 – 21.0]    roughly 490 to 570 observations

A prevalence of neuropathy among diabetics is filed under “diabetes”. The value is real and correctly extracted; its attachment is wrong. One hundred observations sampled at random, classified by a model whose measured precision is 71–82%, then re-read by one person, with no third-party check.

Why it is not fixed. Automatic flagging would be wrong about one time in four. And a false judgement flag is indistinguishable from a true one without re-reading — it is therefore worse than no flag, which at least keeps the reader alert.

What it would take. A clinician re-reading 500 observations and deciding, for each, whether the displayed disease is the subject of the measurement or the population studied. We have the protocol, the sample and the classifier that prepares the work — the reader is what is missing. Five hundred judgements, at roughly ten seconds each, is under two hours for someone in the field.

3.Values that have lost the population they apply to

What we knowREAD

"Malaria prevalence was higher in males (57.7%) than females (47.2%)"
        -> 57.7 is displayed as THE prevalence of malaria

The value is true; its qualification has disappeared — a subgroup presented as the whole. About 490 observations, from the same twenty-five readings.

Why it is not fixed. Unlike challenge 1, the value is not wrong: it is incomplete. Repairing it means recovering the qualifier — “among men”, “in rural areas”, “in 2019” — which exists in the sentence but was not extracted. That is a change to the extraction schema, not a data fix.

What it would take. Add a stratum field to the schema, then re-extract the affected subset. Compute cost is small — our passes over 40,000 observations cost around ten dollars. The prior work is defining a stratum vocabulary that does not become a free-text field: sex, setting, age band, period.

4.Two thirds of locations do not come from the text

What we knowCOUNTED

46,539   observations carry a country
34,035   of which the country is INFERRED, not written in the sentence
11,848   of which it comes from the authors’ affiliation

A paper signed from Lagos does not necessarily study Nigeria. The inherited field flags this row by row, and the affected observations can be excluded with one query.

Why it is not fixed. This is not a defect but a trade-off: without the inference, two observations in three would have no country and the map would be empty. The problem is not the inference — it is that we do not know how often it is wrong. We only know the error rate exceeds 6%.

What it would take. Two hundred observations whose country is inferred, checked against the paper’s actual study site. This is abstract reading — no clinical expertise required. An intern, a librarian, a master’s student can do it in a day. Of every challenge on this page, this has the best value-to-effort ratio: two hundred readings turn an unknown into a confidence interval, on a field that affects 34,035 observations.

5.Measure type is determined only on percentages

What we knowCOUNTED

42,886   observations carry a measure type    (unit = %)
26,810   count-based        -> none
 3,118   rate-based         -> none

Without a measure type, an observation cannot reach the map, whatever it contains. This is not a quality filter: it is the scope of a treatment that has not been extended.

Why it is not fixed. Because extending the classifier would bring a known defect with it. Among count-based observations, roughly 290 carry several values for one kept — “120 children, 55 had durable viral suppression, 65 failed”, where the value kept is the sample size.

What it would take. In order: revisit extraction for multi-value claims, then extend the classifier. The second costs about $0.30 of compute and a day of calibration. The first requires a model decision — can one observation carry several values? — and it is that decision, not the compute, that blocks.

6.One observation in nine covers several diseases at once

What we knowCOUNTED

When a paper measures a co-infection — “prevalence of HIV/HBV/HCV triple infection: 0.64%” — that value appears under each of the three diseases: 11.1% of map observations, 16.4% of country-profile ones. These rows are accurate but do not add up.

Why it is not fixed. Two very different cases look alike: a co-infection measured jointly, where the value holds for each disease, and a duplication, where only one disease is measured and the other is context. Telling them apart requires the same judgement as challenge 2.

What it would take. Nothing new: the work of challenge 2 solves this one for free. One reader unblocks both.

7.The graph does not connect institutions, and barely connects researchers

What we knowCOUNTED

   0 / 80   researchers linked to an institution in the database
22.4%       of observations have a path to a researcher

The corpus’s 43 institutions do not appear in the graph. We removed them rather than display 34 isolated points: name matching reached 9 correspondences out of 80.

Why it is not fixed. The researcher → institution link does not exist in our data, and inferring it from a free-text affiliation field would produce a wrong attachment nine times in ten. We do not publish a link we cannot justify.

What it would take. Linking 80 researchers to their institution by hand, using ROR identifiers — which we already hold. Eighty decisions: one afternoon. It is the shortest challenge on this page, and it opens up a whole reading of the corpus.

8.Concepts do not all have the same granularity

What we knowDOCUMENTED LIMITATION

Some concepts separate their severe form into a distinct concept — pneumonia and severe pneumonia — while others absorb it: malaria covers severe malaria. Two rates are therefore not always comparable across concepts. Some concepts are deliberately imprecise, and that is not a defect: of thirty re-read observations of “diabetes (unspecified)”, one allowed type 1 to be told from type 2. The imprecision is in the papers, not in our vocabulary.

Why it is not fixed. Unifying the convention touches malaria, the corpus’s first concept. Splitting it without measuring what that moves would be exactly the risk this project exists to avoid — we did it once for pneumonia, at small scale, and that was enough to measure the cost.

What it would take. An explicit convention — is severity a concept or an attribute? — then a split measured before it is applied. The work is ours; what would help is a clinical opinion on the cases where the distinction changes how a figure reads.

What we have not measured

NOT MEASUREDby definition — which is why this section exists. A silent limitation is more dangerous than a quantified one.

  1. Extraction recall. We do not know how many numerical claims in a paper were NOT extracted. Every rate on this page concerns what was extracted, never what was missed.
  2. The error rate of inferred locations. Known to exceed 6%, never measured.
  3. The 9,549 observations pending review. No deadline is promised.
  4. Geographic and thematic coverage. 39 countries appear on the map; we have not measured whether the absence of others reflects a gap in the corpus or a gap in the research.
  5. The precision of extraction itself. No human annotation campaign has been run on a random sample of the whole corpus.
  6. The transportability of the first three challenges. They are measured on 4,043 observations — 6.4% of the published corpus. Nothing establishes they hold for the rest.
  7. Co-authorship between researchers. Author lists exist on 9,189 of 9,199 publications and are not exploited.
Appendix — two dated corrections, and why some figures went down
2

What we do not know

The gaps grid: what has been measured nowhere, over a declared scope. An empty cell says “no observation in this corpus”, never “no data in the world”.

The full grid, by section and country: open the gaps grid

It stays on its own screen: a 12-country by 3-section table that is read by scanning it, not by summarising it.

3

What comes next

What is under way, and what turns the above into work rather than a verdict.

This section sets out what the project will do, and what each piece of work requires — in hours, skills and access. The four are independent: none waits on the others.

Durations are derived from rates already observed on this corpus, not estimated by eye. Where no rate exists, it says so.

1.Expert review, and what it changes in the corpus

9,705 observations await human review — 9,549 pending, 156 disputed. None of the 63,227 published observations has been read by a human: validation is mechanical, it certifies form and plausibility, never truth.

At the rate given in the contribution table — 500 observations in two hours for someone in the field — the whole set represents about 39 hours of clinical or epidemiological expertise. Spread across ten reviewers, that is a day each.

  • A measured error rate. Today our best approximation comes from an internal check: fewer than 3.9% of claims cannot be found in their source text. That is a bound, not a measurement.
  • The project’s first recall figure. We know what extraction produced; we do not know what it missed.
  • Worked examples for extraction. Every reviewer/machine disagreement documents an error pattern, and patterns are fixed upstream, not row by row.
hours
COUNTED~39 h of review, at the observed rate of 500 observations in two hours. Analysing disagreements, ~15 h, is estimated.
skills
clinician, epidemiologist, or public-health resident
access
none — the interface and the dataset are public

2.Extension beyond health

The chain — collection, chunking, extraction, normalisation, validation, publication — contains nothing medical. What is domain-specific sits in two objects: the reference table — 487 diseases, 74 indicators, 32 anthropometric measures, 32 risk factors, 247 context rules — and the two extraction prompts, one per language.

Everything else — pagination, validation computation, location inheritance, the graph, the deposit — is indifferent to subject matter.

We have not done it, so we do not claim it: portability is inferred from the structure of the code, never tested on a second domain. The first attempt will reveal what in the chain was medical without our noticing. That is the reason to do it.

hours
ESTIMATED · NO OBSERVED RATE~3 weeks for a first domain, most of it building the reference table. Ours took eleven curation passes over four months, but it was built alongside the chain.
skills
an expert in the target domain, plus a developer for the prompts
access
a corpus of reachable publications and a licence that permits extraction

3.Autonomous orchestration

Today every step is launched by hand. The repository holds 25 commands; a full run chains eight of them. Between each, a human reads an output and decides to continue.

This is not an engineering failing: it was a choice that held while the corpus doubled monthly and every pass revealed a new error pattern. An autonomous chain that propagates a normalisation error across 70,000 observations costs more than eight typed commands.

What changed: the patterns have settled, the checks are now mechanical, and each step can say whether it succeeded. Orchestration becomes possible because the checks exist, not the other way round.

hours
ESTIMATED · NO OBSERVED RATE~2 weeks, half of it reworking checks so they block rather than warn.
skills
a developer familiar with job queues and failure recovery
access
a scheduler, and a model-call budget bounded per run

4.The 9,189 publications whose authors are not extracted

Author lists exist on 9,189 publications out of 9,199: 108,402 author mentions, 11.8 per publication on average. They are stored as a single stringAdegnika AA, Verweij JJ, Agnandji ST, … — never split into people.

Splitting the string is mechanical. The difficulty is elsewhere: Adegnika AA and Adegnika Ayola A. are the same person, and naive name matching would merge homonyms. A raw count gives 45,412 distinct names; the number of real people is lower, and we do not know by how much.

So this work is not “parse 9,189 strings”. It is: split, then match against ORCID and OpenAlex where an identifier exists, then leave unmerged whatever is not established — a duplicated researcher is repairable, two researchers wrongly merged can no longer be detected.

hours
ESTIMATED~1 week for splitting and identifier matching. The correct-merge rate is unmeasured, and it is what will decide the rest.
skills
a developer, plus a bibliometric opinion on the merge threshold
access
the ORCID and OpenAlex APIs, both public and free

Contributing

This project is published under CC BY 4.0, corpus and method alike — 10.5281/zenodo.21794559.

What would help most, in order:

what we needwho can do iteffort
Link 80 researchers to their institutionanyone comfortable with RORone afternoon
Check 200 inferred locationsabstract reading, no clinical expertiseone day
Re-read 500 observations, subject vs populationclinician, epidemiologist, public-health residenttwo hours
Annotate 300 observations by quantityepidemiologist or biostatisticiantwo days
A clinical opinion on concept granularityclinicianone conversation

Reporting an error is a contribution too, and the fastest one: an observation whose value, country or disease looks wrong to you interests us more than general agreement. Every observation carries an identifier and a link to its source publication.

contact@e-shepha.com

We answer reports, and we correct in public: the figures on this page change, and their measurement date is written beside them.