Collecting and harvesting big data
Denmark's archives hold centuries of individual-level records - censuses, parish registers, police registers, school and hospital records, tax books, death certificates. Almost all of it is handwritten, and almost none of it is machine-readable. BDAD builds the methods that turn those documents into data, and uses the resulting data to study health, human capital and social mobility over the life course and across generations.
BDAD is an interdisciplinary research unit at the Department of Economics, University of Southern Denmark, led by Professor Christian Møller Dahl. It is interdisciplinary by construction: econometrics and machine learning on one side, economic history, historical demography and epidemiology on the other. We work directly with the institutions that hold the records — the Danish National Archives, Statistics Denmark, The National Archives (UK) and the Swedish SwedPop infrastructure — and we release what we build, so that it is usable well beyond the projects that paid for it.
3 |
1.1 m |
14 m |
4 |
|
open systems released — OccCANINE, DARE, HANA |
handwritten name images published in HANA |
occupation–code pairs behind OccCANINE, in 13 languages |
national archives and statistical agencies we build with |
The problem we work on
A historical record only becomes evidence once it is machine-readable, structured and joinable to other records. That last step is the hard one. Transcribing a census by hand is a decade of work; transcribing it badly is worse than not transcribing it at all, because the errors are not random and they propagate into every estimate built on top. Our work sits at exactly that junction — building transcription and linkage systems that operate at population scale, and the statistical theory needed to use their output honestly.
The pay-off is the life course. Once a birth record, a school record, a hospital admission and an adult tax return can be joined to the same person, questions that were previously unanswerable become ordinary empirical work: how early-life conditions shape adult health and earnings, how a policy reform changed the trajectory of the children exposed to it, how advantage and disadvantage travel between generations.
What we do
-
Research
Transcription at population scale, record linkage, standardisation of free-text historical descriptions, and inference from machine-transcribed samples.
-
Projects
Work in Denmark, Norway, Sweden, the United Kingdom and the United States — from census geolinking to Danish trade statistics and Norwegian student yearbooks.
-
Tools & Data
OccCANINE, DARE and HANA — all released openly, with code, models and training data, and used internationally.
-
Work with us
For archives with holdings to digitise, researchers who want to use our tools, and prospective PhD students.
Selected results
The methods are not the point on their own. These are studies that could not have been done without the data the unit produced:
- A universal toddler-health programme, evaluated at scale. A large government trial assessed on a cohort transcribed and curated entirely by machine from archival nurse records. The Economic Journal, 2026.
- The Copenhagen Infant Health Nurse Records cohort. An archive of handwritten records turned into a documented, governed research cohort that others can apply to use. International Journal of Epidemiology, 2023.
- School closures in the 1918 pandemic. 500,000 archival death certificates and closure records for 2,100 Swedish school districts, used to identify both the mortality effect and the long-run human-capital effect of the policy.
- What makes an artist. The clustering of creative activity in the United States since 1850, built entirely from digitised historical census records — the same infrastructure serving a humanities question.
Partnerships and advisory roles
Alongside its own projects, the unit contributes to national infrastructure and to research programmes led elsewhere.
HisPeR — the Danish historical population register
HisPeR extends Denmark's CPR register backwards to 1645. Complete, it will hold some 100 million personal registrations — enough to follow families across as many as fifteen generations — and it has been designated part of Denmark's national research infrastructure roadmap by the Ministry of Higher Education and Science, with funding of nearly DKK 16 million for 2026–2030. It is led by the Danish National Archives, and SDU is one of thirteen partner institutions.
For this unit, HisPeR is the destination as much as the partnership: it is the spine that transcribed historical records are ultimately joined to, and the reason record linkage is a research theme here rather than a technical afterthought. Announcement from Rigsarkivet →
PID-scapes — PandemiX Center, Roskilde University
PID-scapes studies how infectious diseases interact at population level, and how the Russian (1889–1892) and Spanish (1918–1920) influenza pandemics reshaped the ordinary circulation of other diseases. It rests on roughly 80,000 individual patient journals from the Blegdam and Øresund hospitals, held at the Copenhagen City Archives and never transcribed.
Christian Møller Dahl is one of three external collaborators on the project, which adopts this unit's published transcription methods for the patient journals. PandemiX is a Danish centre of excellence at Roskilde University. PandemiX projects →
The Danish Brain Collection
BDAD is a partner to the Danish Brain Collection, held in Region Syddanmark, which makes brain material and associated data available to researchers on application. The collection →
Using our toolsOccCANINE, DARE and HANA are open and documented, and we would rather they were used than admired. If you are working with handwritten sources — in any language, not only Danish — and want help getting started, write to cmd@sam.sdu.dk. |