Skip to main content

We make the
economic-historical
past future-ready

FAIR Datasets

Our Datasets

Our Project Goals

Digitization, Curation, Provision and Research

Financing Options icon

Digitization

Financing Solutions icon

Curation

Financial Modeling icon

Provision

Business Planning icon

Research

Initiative

We are building the first comprehensive, machine-readable dataset on the development of German companies and financial markets since the late 19th century. This creates the foundation for systematically analyzing economic upheavals, crises, and long-term developments.

Challenge

Historical financial and corporate data exist in large quantities in Germany – but they are difficult to use. They are scattered across archives, yearbooks, and historical publications, often only as scanned documents or unstructured text. For research, this means: a substantial part of the work consists not of analysis, but of the laborious collection, digitization, and preparation of the data.

Relevance

Germany offers an exceptional economic-theoretical perspective. Extreme events such as the hyperinflation of 1923, the Great Depression, the reconstruction after World War II, and later phases of globalization make long-term analyses particularly valuable. Without systematically prepared data, however, many of these developments can only be captured in a fragmentary way.

Approach

In the GerHisFin project, we digitize, structure, and link historical financial and corporate data over long time periods.

Our goal is to build a consistent database that:

  • covers firm-level corporate data

  • integrates financial market data

  • is comparable across different time periods

This provides the basis for new forms of empirical research.

Research Questions

The data enable, among other things, analyses such as:

  • How did companies develop during the hyperinflation of 1923?

  • Which firms survived economic crises – and which did not?

  • How did capital markets change across political and economic regimes?

So far, such questions can only be investigated with considerable manual effort, or not at all.

Methods

A central component of the project is the use of modern data processing methods.

Historical sources are automatically curated and structured with the help of methods from the fields of machine learning and artificial intelligence. These include in particular:

  • Text recognition (OCR) / document understanding (DU)

  • Layout and table extraction

  • Entity identification and linking

For the first time, these technologies make it possible to process large volumes of historical data efficiently and reproducibly.

Results

In the long term, the project provides the following resources:

  • structured datasets on companies and financial markets

  • documented processing steps and metadata

  • interfaces for scientific use

  • accompanying research papers

The goal is to make the data sustainably available to the scientific community.

Target Audience

The project is aimed in particular at:

  • Researchers in economics and the social sciences

  • Economic historians

  • Data scientists with an interest in historical data

  • Institutions with an interest in long-term economic developments

Outlook

GerHisFin contributes to transferring historical data holdings into a modern research infrastructure.

In the long term, this will create a better understanding of economic developments – and lay the foundation for new research on crises, growth, and structural change.

Working Site University Library Mannheim

Digitization & Curation


The working site at the University Library Mannheim is responsible for the physical and digital curation of the source material as well as for building up the knowledge structure. Central tasks include the professional digitization of analog sources – including the Handbook of German Stock Corporations and patent registers –, metadata capture, automated full-text recognition (OCR/DU), and providing the digital copies via the digital library.


The working sites in Frankfurt (SAFE) and Mannheim (University of Mannheim) are responsible for the data-science processing, structuring, and provision of the research database. In particular, SAFE handles entity recognition, the extraction of relevant information from unstructured sources, and its conversion into machine-readable data formats. A methodological focus lies on the development of advanced linking procedures. The processed research data are made available to the research community together with metadata and persistent digital identifiers (DOI) via the repository of the SAFE Research Data Center.

Working Site Frankfurt (SAFE) and Mannheim (University)

Provision & Research

Scientific Advisory Board

The scientific advisory board provides academic guidance to the project and supports the further development of the research program. Its tasks include in particular scientific advice to the project partners, the assessment of new research projects, and the evaluation of project progress in the individual funding phases. The interdisciplinary composition of the board ensures that both the research agenda and the development of the data infrastructure are continuously reflected on and advanced from a scientific perspective.

Rui Esteves

Source: Image archive of the Geneva Graduate Institute

Geneva Graduate Institute
Email

Profile

Sibylle Lehmann-Hasemeyer

Source: Sibylle Lehmann-Hasemeyer

University of Hohenheim
Email

Profile

Angelo Riva

Source: Image archive of INSEEC

Paris School of Economics
Email

Profile

Mark Spoerer

Source: Image archive of the University of Regensburg

University of Regensburg
Email

Profile

Latest News

We will keep you up to date

© GerHisFin