COMHIS took shape in 2013 when Mikko Tolonen, an intellectual historian, began working with a data scientist on a problem that conventional catalogue searches and close reading could not resolve: how could the movement of books, arguments and publishing structures be studied across whole collections?
The collaboration expanded because each result exposed a question that no single discipline could answer. Historians, computer scientists, linguists, librarians and research-software specialists joined through successive projects. Historical questions continued to lead, but data design, modelling, validation and interpretation became shared intellectual work.
The first phase concentrated on bibliographic data. National bibliographies and union catalogues made it possible to compare formats, page counts, publication places and languages across early modern print culture. This work established a principle that still guides COMHIS: works, editions, printings, surviving copies and catalogue records are different historical objects, and any explanation depends on knowing which one has been counted.
Computation extended Tolonen’s research on Mandeville, Hume, commercial society and Enlightenment publishing. Questions previously pursued through individual citation trails could now be tested across larger publishing systems without abandoning their intellectual-historical foundations.
The Research Council of Finland consortium Computational History and the Transformation of Public Discourse in Finland, 1640–1910 connected Helsinki researchers with the National Library of Finland and University of Turku. The programme widened from books and catalogues to newspapers, text reuse, language change and public discourse. Open data, reusable infrastructure and interdisciplinary workflows became part of the historical method.
RiCEP, NewsEye and HPC-HD developed research on book prices, publishing networks, register, OCR, canon formation and textual circulation. The Reception Reader connected hundreds of thousands of detected relationships to passages and page images, allowing researchers to move repeatedly between corpus-wide patterns and individual evidence.
Material features also moved to the foreground. Format, typography, layout and printers’ ornaments became evidence for production and circulation. Metadata, text and images were increasingly treated as interacting views of the same historical process.
The current phase moves from identifying similarity towards studying engagement and interpretation. Translation mining follows arguments across languages and transformations; meaning matching searches for less explicit conceptual uptake; multimodal models connect text to layout and image; and tools developed for the critical edition of Hume’s History of England bring collection-wide comparison into exacting editorial work.
CASCADE and MECANO added research on semantic change and canon formation. ReBeL and IMPRINT extend the programme towards implicit meaning, material evidence and intellectual networks. FIN-CLARIAH and CSC infrastructure provide the shared data, CPU and GPU computing, preservation and interfaces required to make the work cumulative.
AI changes the speed and range of this environment, but not its governing purpose. The annotator agent and Adaptive Evidence Construction are recent developments in a longer trajectory: intellectual history becoming a collaborative, inspectable practice capable of moving between close interpretation and hundreds of thousands of books.
For the wider argument, see Mikko Tolonen’s