Skip to main content
Research

 

Research

 

From a shelf of paper to a dataset you can trust

 

Our research runs the length of a single pipeline: read the document, structure what it says, join it to the same person elsewhere, and then — the step most often skipped — work out what the machine's mistakes do to the answers. Five themes, each a genuine research problem rather than an engineering task.

1. Transcription at population scale

Historical documents resist automation in ways modern ones do not. Pages are degraded, ruled tables drift out of alignment, columns are overwritten, hands change mid-volume, and the same clerk will abbreviate a word four different ways in a single register. Off-the-shelf OCR fails on all of it.

We develop methods for automatic identification and segmentation of tables and table structures in degraded documents, and for recognition of handwritten names, dates, ages, occupations, mortality counts, birth weights and incomes. The emphasis is on scale: a method that works on a hundred pages under supervision is not the same object as one that runs unattended across a national archive and reports honestly on its own uncertainty.

Recent work moves from single-purpose recognisers to ensembles of vision–language models that read layout, typescript and handwriting in one pass, with agreement between independent models used as an intrinsic quality signal and disagreements routed to review.

That word independent is doing real work, and we have measured what happens when it fails. On the HCNC corpus, majority voting across three vision–language engines produced worse data than the best engine alone, because two of the three descended from a common model family and made the same mistakes in the same places. As foundation models converge on shared lineages, "run several models and take the majority" is becoming less safe rather than more — and the failure is invisible without ground truth to measure against. Cross-engine agreement is now used as confidence metadata rather than as a value-chooser.

2. Record linkage

A transcribed record is useful; a linked record is transformative. Linkage joins transcribed historical individuals to one another and to the Danish administrative registers, so that a single life can be followed from a birth record through schooling and hospital admissions to adult earnings and, ultimately, cause of death.

Danish sources before 1968 carry no personal identification number, so historical linkage is inherently probabilistic: it rests on names, dates and places, all of which are exactly the fields handwriting recognition finds hardest. The research problem is therefore joint — transcription quality and linkage quality cannot be optimised separately, because the error in one determines the feasible accuracy of the other.

This theme is why the unit is a partner in HisPeR, the Danish historical population register now designated national research infrastructure. HisPeR carries the CPR register back to 1645 and will eventually hold around 100 million personal registrations; it is the spine that a transcribed archive gets joined to, and without something of that kind a newly digitised collection remains an island.

3. Standardisation and classification

Historical records describe the world in free text. A census taker writes "journeyman smith, Vejle" and a modern analysis needs an occupational code that means the same thing in Denmark, Sweden, the Netherlands and the United States. Doing that by hand is slow, tedious and inconsistent between coders; doing it badly makes cross-country comparison meaningless.

We build and release systems that do this automatically — OccCANINE for occupational descriptions into HISCO, and related work on converting historical occupational descriptions into comparable outcome scores. The same problem recurs for causes of death, diagnoses, addresses and administrative categories, and we treat it as one problem rather than a series of one-off cleaning exercises.

4. Inference from machine-generated data

This is the theme that distinguishes the unit, and the one most often neglected elsewhere.

When a sample is produced by an automated transcription or classification system, the measurement error it contains is not random noise. It is systematic: a model misreads certain hands, certain name forms, certain occupations more than others — and those are rarely distributed evenly across the population being studied. Treating machine-generated variables as if they were correctly measured biases estimates in a direction that is predictable in principle but invisible in practice.

We develop estimation and inference theory for exactly this setting, and practical guidance on detecting and correcting prediction bias in historical data. It is an econometric problem, not an engineering one, and it is the reason a unit like this belongs in a department of economics.

The clearest illustration is our work on the historical US censuses, where the transcriptions were made by people rather than machines but the logic is identical: the records hardest to read belong disproportionately to the immigrant, the poor and the mobile, so a linked sample silently over-represents everybody else — and every estimate computed on it inherits that selection.

5. Space and geography

Records are located as well as dated. Geocoding historical addresses, parishes and enumeration districts turns a linked life course into a linked trajectory — where someone was born, where they moved, what kind of place they were living in when a policy reached them. It also opens the long-run questions of economic geography: agglomeration, structural change, market access and the spatial concentration of industry.

This theme expands substantially with Alexander Klein joining the unit, and connects directly to the Geolinking UK Census project, which is building a geocoded, linked version of the English and Welsh censuses from 1841 to 1921.

Space is not only a historical concern. In the Normalization project, where and when public housing replaced institutions for people with disabilities is the variation that identifies the reform's effects — so establishing who was exposed to whom, at what distance and for how long, is the empirical problem itself rather than a descriptive layer on top of it.

 

How the themes fit together

Themes 1–3 build the data. Theme 4 says how far you can trust it. Theme 5 places it. The applications — early-life conditions, health and human capital, social mobility, economic geography — are where the unit's work meets the rest of HEDG and the wider department. See Projects for what this looks like in practice.

 

Last Updated 14.09.2026