Skip to main content
Research

 

US Census Transcriptions

 

Reading the names the transcribers could not, so that more people can be linked — and a fairer sample of them

 Paper  Improving Historical Census Transcriptions: A Machine Learning Approach
Authors  Torben S. D. Johansen, Christian Møller Dahl, Sam Il Myoung Hwang, Munir Squires
 Partners  Vancouver School of Economics, University of British Columbia
 Status  Working paper
 Funding  

 

The historical censuses of the United States are among the most heavily used datasets in economics and social history. An enormous amount of what we know about American migration, mobility, immigration and inequality over the last two centuries rests on linking individuals from one decennial census to the next — following the same person across ten years, then twenty, then a lifetime.

That linkage depends almost entirely on names. And the names, as they exist in the machine-readable census files, were transcribed by people working from handwritten enumerator sheets that are often faint, crowded, hurried or damaged. Where the transcriber guessed wrong, the link fails — silently. The person is simply absent from the linked sample, and no analysis built on that sample will ever know they were there.

What the project does

We retranscribe the names using machine learning, and show that doing so substantially increases the rate at which individuals can be matched across census rounds. The model performs best precisely where the existing transcriptions are weakest — where legibility of the original form is low, and human transcribers had the least to work with.

Why it is not only a matter of quantity

The more important finding is about who gets linked. Transcription failure is not evenly distributed. Names that are unfamiliar to the transcriber, spelled inconsistently, or recorded by an enumerator in a hurry are harder to read — and those characteristics are correlated with being an immigrant, being poor, being mobile, or belonging to a group the census historically recorded with less care.

The consequence is that a linked sample is not a random subset of the population. It systematically over-represents the settled, the native-born and the easily-spelled, and every estimate computed on it inherits that selection. Improving the transcriptions improves the linkage rate for the groups that were hardest to link — which changes not just how much data is available, but whom the data is about.

This is the same argument that runs through the unit's fourth research theme: machine-generated data carry systematic, not random, error, and the error is usually correlated with exactly the variables under study. Here the error was introduced by human transcribers rather than a model — but the logic, and the damage to downstream estimates, is identical.

Why this collaboration

The project brings together BDAD's handwriting-recognition work — the methods behind HANA and DARE, originally developed on Danish sources — with the record-linkage expertise of Sam Il Myoung Hwang and Munir Squires at the Vancouver School of Economics, whose work on the properties of automated linking methods and on measurement error in linked historical US census samples defines much of what is known about how these datasets behave.

It also demonstrates something the unit cares about: the transcription methods are not Denmark-specific. They were built on Danish police registers and parish records, and they transfer to nineteenth-century American enumerator sheets without being rebuilt from scratch.

 Scale. Where most of the unit's projects begin with records that have never been transcribed at all, this one begins with records that have been — by hand, at great expense, and imperfectly. It is a reminder that "already digitised" and "usable" are not the same claim, and that revisiting existing transcriptions can be as valuable as producing new ones.

 

Last Updated 17.09.2026