Hi, I'm Christoph 👋🏻

I'm a lead data scientist in the Language AI team at Bayer Pharma, where I work on agentic systems for pharma — language models that plan, use tools, and work through tasks across scientific and regulatory text.

Before moving to industry, I was a postdoc in the Machine Learning group at Humboldt University of Berlin, headed by Prof. Alan Akbik , and a member of the excellence cluster Science of Intelligence . Earlier, I was a researcher and doctoral student in the speech and language technology group at the German Research Center for AI (DFKI), supervised by Prof. Sebastian Möller and Prof. Hans Uszkoreit .

What I work on
  • Agentic systems for pharmaceutical research and development
  • Tool use, planning, and grounding in scientific and regulatory text
  • Language models in domains with little tolerance for error
  • Evaluating and analysing what models actually learn
  • I was promoted to Lead Data Scientist in the Language AI team at Bayer Pharma. 🎉

  • I joined Bayer Pharma as a Senior Data Scientist in the Language AI team.

  • Our DFG-funded project Modeling Neurogenesis for Continuous Learning has been approved as part of the Science of Intelligence cluster. 🎉

  • Our DFG-funded project Efficient Model Learning from Data with Partially Incorrect Labels has been approved as part of the Science of Intelligence cluster. 🎉

Selected publications
See all publications

Neural sequential transfer learning for relation extraction

Christoph Alt

Doctoral Thesis, Technische Universität Berlin, 2021

Show abstract

Relation extraction (RE) is concerned with developing methods and models that automatically detect and retrieve relational information from unstructured data. It is crucial to information extraction (IE) applications that aim to leverage the vast amount of knowledge contained in unstructured natural language text, for example, in web pages, online news, and social media; and simultaneously require the powerful and clean semantics of structured databases instead of searching, querying, and analyzing unstructured text directly. In practical applications, however, relation extraction is often characterized by limited availability of labeled data, due to the cost of annotation or scarcity of domain-specific resources. In such scenarios it is difficult to create models that perform well on the task. It therefore is desired to develop methods that learn more efficiently from limited labeled data and also exhibit better overall relation extraction performance, especially in domains with complex relational structure.

In this thesis, I propose to use transfer learning to address this problem, i.e., to reuse knowledge from related tasks to improve models, in particular, their performance and efficiency to learn from limited labeled data. I show how sequential transfer learning, specifically unsupervised language model pre-training, can improve performance and sample efficiency in supervised and distantly supervised relation extraction. In the light of improved modeling abilities, I observe that better understanding neural network-based relation extraction methods is crucial to gain insights that further improve their performance. I therefore present an approach to uncover the linguistic features of the input that neural RE models encode and use for relation prediction. I further complement this with a semi-automated analysis approach focused on model errors, datasets, and annotations. It effectively highlights controversial examples in the data for manual evaluation and allows to specify error hypotheses that can be verified automatically. Together, the researched approaches allow us to build better performing, more sample efficient relation extraction models, and advance our understanding despite their complexity. Further, it facilitates more comprehensive analyses of model errors and datasets in the future.

Probing Linguistic Features of Sentence-Level Representations in Neural Relation Extraction

Christoph Alt, Aleksandra Gabryszak, Leonhard Hennig

Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020

Show abstract

Despite the recent progress, little is known about the features captured by state-of-the-art neural relation extraction (RE) models. Common methods encode the source sentence, conditioned on the entity mentions, before classifying the relation. However, the complexity of the task makes it difficult to understand how encoder architecture and supporting linguistic knowledge affect the features learned by the encoder. We introduce 14 probing tasks targeting linguistic properties relevant to RE, and we use them to study representations learned by more than 40 different encoder architecture and linguistic feature combinations trained on two datasets, TACRED and SemEval 2010 Task 8. We find that the bias induced by the architecture and the inclusion of linguistic features are clearly expressed in the probing task performance. For example, adding contextualized word representations greatly increases performance on probing tasks with a focus on named entity and part-of-speech information, and yields better results in RE. In contrast, entity masking improves RE, but considerably lowers performance on entity type related probing tasks.

TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction Task

Christoph Alt, Aleksandra Gabryszak, Leonhard Hennig

Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020

Show abstract

TACRED is one of the largest, most widely used crowdsourced datasets in Relation Extraction (RE). But, even with recent advances in unsupervised pre-training and knowledge enhanced neural RE, models still show a high error rate. In this paper, we investigate the questions: Have we reached a performance ceiling or is there still room for improvement? And how do crowd annotations, dataset, and models contribute to this error rate? To answer these questions, we first validate the most challenging 5K examples in the development and test sets using trained annotators. We find that label errors account for 8% absolute F1 test error, and that more than 50% of the examples need to be relabeled. On the relabeled test set the average F1 score of a large baseline model set improves from 62.1 to 70.1. After validation, we analyze misclassifications on the challenging instances, categorize them into linguistically motivated error groups, and verify the resulting error hypotheses on three state-of-the-art RE models. We show that two groups of ambiguous relations are responsible for most of the remaining errors and that models may adopt shallow heuristics on the dataset when entities are not masked.

  • May 2026 - Current
    Bayer Pharma
    Lead Data Scientist, Language AI

    Leading work on agentic systems for pharma — language models that plan, call tools, and carry out tasks grounded in scientific and regulatory text.

  • Jun 2023 - Apr 2026
    Bayer Pharma
    Senior Data Scientist, Language AI

    Built language AI for pharmaceutical text in the Language AI team, taking models from prototype to production and moving into agentic systems.

Get in touch

If you want to talk about language AI, a paper, or anything else, email me or find me on these: