Colin Swaelens, Philology Meets Machine Learning: Digital Annotation and Semantic Modelling of Byzantine Greek Texts

Abstract

Diplomatic transcriptions of medieval texts pose particular challenges for digital annotation: their fragmentary preservation, non-standard orthography, and heterogeneous linguistic features complicate processing and limit the reuse of existing tools. To address these issues, we developed a full annotation pipeline for unedited medieval Greek, including part-of-speech tagging, morphological analysis, and lemmatisation. A manually created gold standard of 10,000 tokens was used to train and benchmark transformer- based language models, resulting in DBBErt, a system optimised for historical Greek. This pipeline not only improves the accuracy of linguistic annotation in low-resource settings, but also provides a foundation for more complex research questions. As a case in point, a semantic textual similarity benchmark was designed to investigate intertextual connections across Byzantine book epigrams, demonstrating how annotated data can enable new perspectives on meaning and reuse. Our work highlights both the potential and the limitations of existing annotation strategies: while transformer models significantly enhance processing quality, traditional embeddings remain competitive in some contexts, and software constraints require careful workarounds. More broadly, we will show how digital annotation fosters collaboration across philology, linguistics, and computational modelling, creating added value for the study of medieval texts and manuscript culture.

Practical information

This round table presentation will be given at Byzantium beyond Byzantium. 25th International Congress of Byzantine Studies, which takes places in Vienna on 24-29 August 2026.

Date & time: Saturday the 29th of August 2026, 08:30

Location: University of Vienna (Universitätsring 1, 1010 Vienna)