Evidence, data and collaboration

We build transparent ways to turn large, imperfect historical collections into evidence—combining bibliographic data, text, images, layout, scholarly interpretation and computational infrastructure.
How historical sources become evidence

A digitised source is not yet evidence. Historical collections arrive with catalogue records, OCR, page images and inherited classifications, all created for different purposes and with different kinds of uncertainty. Our work begins by asking what historical claim the material should support, and then constructing a transparent path from sources to evidence.

This changes how people work with large historical datasets. Instead of treating a corpus as a finished object, we connect bibliographic metadata, text, typography, layout and image evidence; test where those layers agree or fail; and preserve the decisions that turn them into research data. The result is interpretation at scale without surrendering scholarly judgement.

Reconstructing books

In Eighteenth Century Collections Online alone, more than 200,000 digitised volumes must be understood not merely as containers of OCR, but as made objects. We develop methods that break books into meaningful components: title pages and half-titles, frontispieces, tables of contents, dedications, advertisements, chapter openings, footnotes, marginal quotations, catchwords, page numbers, ornaments and blank leaves. New layout-detection and multimodal methods connect these components to page images and text.

This enables book history to operate at a scale that was previously impossible while keeping individual books visible. We can study how printers organised knowledge, how editions changed, how advertisements travelled, and how typographic and material choices shaped reading.

Critical editions that scale

Critical editions are core instruments of humanities research. We are developing computational tools that align clean or keyed texts, OCR witnesses and page images while preserving the identity of each witness and documenting editorial intervention. A modern reset, a scholarly edition, an early printed edition and an OCR transcription cannot simply be collapsed into one text.

Our work on David Hume provides a concrete instance: the close reasoning required to establish and annotate his texts becomes a model that can scale to large collections. The objective is not automatic editing, but computational support for comparison, provenance, uncertainty and responsible editorial decisions.

Adaptive Evidence Construction

We call this iterative way of working Adaptive Evidence Construction. A historical question guides the selection and linking of metadata, OCR, page images and layout features. Models then help identify patterns or candidate evidence; historians test the results against sources; and the data, code, uncertainty and decisions are revised and documented. Evidence is constructed adaptively because each stage can reveal that the question, source model or method needs to change.

Humanities researchers do not merely contribute domain knowledge after technical systems have been built. They design, govern and implement these systems in an integrated interdisciplinary environment. Historians, linguists, computer scientists, data scientists and research software specialists share responsibility for the research design and for the limits of its claims.

AI, annotation and human judgement

One current line of work is an annotator agent trained through a curriculum of cases reviewed by eighteenth-century historians. The agent learns to reason about difficult books through a versioned casebook, while independent validation keeps evaluation separate from instruction. The aim is an auditable research instrument that can explain why it proposes an annotation and where it remains uncertain.

Other work combines translation mining with meaning matching to study intellectual engagement across languages and editions. Rather than counting shared words alone, we investigate when translation, adaptation, disagreement or reuse changes the meaning of an argument.

HPC and durable research infrastructure

Large-scale historical interpretation requires serious infrastructure. We use CSC's Roihu environment for bounded CPU and GPU workloads and Allas object storage for durable research materials, with FIN-CLARIAH and DARIAH-FI supporting reusable national infrastructure. For example, experiments on Hume's typography begin with a controlled benchmark of 1,024 page images; image repair is applied only to the subset that needs it, and every derivative retains its provenance. In work on catchwords, validated page-to-page relations are released with manifests and checksums rather than left as temporary model output.

The distinction matters: fast scratch computation enables experimentation, but stable storage, versioned datasets and documented workflows make results inspectable and reusable. This is how computation becomes part of historical method.

Collaboration and open research

COMHIS develops open data, tools and analytical ecosystems with partners in Finland and internationally. Our work has connected and harmonised resources including Fennica, the English Short Title Catalogue, the Heritage of the Printed Book database and ECCO. We welcome collaboration on intellectual history, book history, critical editions, language technology, computer vision, research infrastructure and the responsible use of AI.

Research Groups