Our Project Goals
Digitization, Curation, Provision and Research
Digitization
Curation
Provision
Research
Initiative
We are building the first comprehensive, machine-readable dataset on the development of German companies and financial markets since the late 19th century. This creates the foundation for systematically analyzing economic upheavals, crises, and long-term developments.
Challenge
Historical financial and corporate data exist in large quantities in Germany – but they are difficult to use. They are scattered across archives, yearbooks, and historical publications, often only as scanned documents or unstructured text. For research, this means: a substantial part of the work consists not of analysis, but of the laborious collection, digitization, and preparation of the data.
Relevance
Germany offers an exceptional economic-theoretical perspective. Extreme events such as the hyperinflation of 1923, the Great Depression, the reconstruction after World War II, and later phases of globalization make long-term analyses particularly valuable. Without systematically prepared data, however, many of these developments can only be captured in a fragmentary way.
Approach
In the GerHisFin project, we digitize, structure, and link historical financial and corporate data over long time periods.
Our goal is to build a consistent database that:
covers firm-level corporate data
integrates financial market data
is comparable across different time periods
This provides the basis for new forms of empirical research.
Research Questions
The data enable, among other things, analyses such as:
How did companies develop during the hyperinflation of 1923?
Which firms survived economic crises – and which did not?
How did capital markets change across political and economic regimes?
So far, such questions can only be investigated with considerable manual effort, or not at all.
Methods
A central component of the project is the use of modern data processing methods.
Historical sources are automatically curated and structured with the help of methods from the fields of machine learning and artificial intelligence. These include in particular:
Text recognition (OCR) / document understanding (DU)
Layout and table extraction
Entity identification and linking
For the first time, these technologies make it possible to process large volumes of historical data efficiently and reproducibly.
Results
In the long term, the project provides the following resources:
structured datasets on companies and financial markets
documented processing steps and metadata
interfaces for scientific use
accompanying research papers
The goal is to make the data sustainably available to the scientific community.
Target Audience
The project is aimed in particular at:
Researchers in economics and the social sciences
Economic historians
Data scientists with an interest in historical data
Institutions with an interest in long-term economic developments
Outlook
GerHisFin contributes to transferring historical data holdings into a modern research infrastructure.
In the long term, this will create a better understanding of economic developments – and lay the foundation for new research on crises, growth, and structural change.
Working Site University Library Mannheim
Digitization & Curation
The working site at the University Library Mannheim is responsible for the physical and digital curation of the source material as well as for building up the knowledge structure. Central tasks include the professional digitization of analog sources – including the Handbook of German Stock Corporations and patent registers –, metadata capture, automated full-text recognition (OCR/DU), and providing the digital copies via the digital library.
The working sites in Frankfurt (SAFE) and Mannheim (University of Mannheim) are responsible for the data-science processing, structuring, and provision of the research database. In particular, SAFE handles entity recognition, the extraction of relevant information from unstructured sources, and its conversion into machine-readable data formats. A methodological focus lies on the development of advanced linking procedures. The processed research data are made available to the research community together with metadata and persistent digital identifiers (DOI) via the repository of the SAFE Research Data Center.
Working Site Frankfurt (SAFE) and Mannheim (University)
Provision & Research
Scientific Advisory Board
The scientific advisory board provides academic guidance to the project and supports the further development of the research program. Its tasks include in particular scientific advice to the project partners, the assessment of new research projects, and the evaluation of project progress in the individual funding phases. The interdisciplinary composition of the board ensures that both the research agenda and the development of the data infrastructure are continuously reflected on and advanced from a scientific perspective.
Rui Esteves
Geneva Graduate Institute
Email
Mark Spoerer
University of Regensburg
Email
Latest News
We will keep you up to date
Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
- Written by Jan Kamlah
- Published on