| Dataset abstract
This reproduction package includes the R code used to generate 3072 English-language sentence pairs featuring the word "girl" using an open-source, locally-run Large Language Model (LLM) and to replace all pronouns with placeholders. It also includes R code to translate these sentence pairs into German using the free-tier DeepL API, as well as to automatically retrieve the use of neuter and feminine pronouns referring to the hybrid noun "Mädchen". All instances where automatic retrieval failed were manually annotated. These data form the basis of our analyses and are included as CSV files. We further provide our analysis code for our descriptive statistics, Bayesian mixed-effects modelling, and data visualisation. Additionally, we provide a second replication dataset that was generated using the same procedure to test the consistency of the DeepL translations.
Note that the data and code are in German. (2026-03-03)
Article abstract - English
The grammatical gender of the German noun Mädchen (‘girl’) is neuter, even though, semanti-cally, it refers to a female individual. For this reason, German speakers frequently refer to Mäd-chen with feminine pronouns (Thurmair, 2006). We explore this grammatical vs. referential gender conflict by examining how the machine translation tool DeepL handles the pronominali-sation of this hybrid noun with both personal and possessive pronouns. Our analysis accounts for the following factors associated with gender- or sex-congruent pronominalisation: (1) type of pronoun, (2) linear distance in words between Mädchen and its pronominal resumption, (3) whether the NP is introduced with a definite or indefinite article (das/ein Mädchen), (4) age of the Mädchen, and (4) its associations with 96 stereotypical adjectives (mascu-line/feminine/positive/negative). We used an open-source Large Language Model (LLM) to systematically generate English items featuring all 3,072 combinations of these factors. The depronominalised items were translated into German using the DeepL API. We fitted a Bayesi-an logistic regression model to analyse the factors that led to higher probabilities of referential vs. grammatical gender congruency in the pronominalisation of Mädchen NPs. As expected, our results show that the probability of feminine pronominalisation increases significantly with increased linear distance and age. More surprisingly, it is lowest when the Mädchen is ascribed negative-feminine qualities. Furthermore, the introduction of a Mädchen NP with an indefinite article favours semantic congruency when taken up with a possessive pronoun. In contrast, the probability of grammatical congruency increases when there are fewer than six words between the introductory NP and a possessive pronoun rather than a personal pronoun is used. These results provides new insights into the role of prominence hierarchies in the pronominal refer-ence of hybrid nouns, whilst also supporting key findings from previous work on the repro-duction of gender biases in machine translation and language generation. (2026-07-18)
Article abstract - German
Das Genus des deutschen Substantivs Mädchen ist Neutrum, obwohl es semantisch auf eine weibliche Person referiert. Dieser Genus-Sexus-Konflikt führt dazu, dass Mädchen von deutschsprachigen Sprecher:innen häufig mit einem femininen Pronomen wiederaufgenommen wird (Thurmair, 2006). Die vorliegende Studie erforscht diesen Genus-Sexus-Konflikt und untersucht, wie das maschinelle Übersetzungstool DeepL das hybride Nomen Mädchen bei der Übersetzung aus dem Englischen durch Personal- und Possessivpronomen wiederaufnimmt. In unserer Analyse werden folgende Faktoren berücksichtigt, die mit einer genus- bzw. sexuskongruenten Pronominalisierung in Zusammenhang stehen: (1) die Pronomenart, (2) der Abstand in Wörtern zwischen Mädchen und seiner pronominalen Wiederaufnahme im Folgekontext, (3) die Definitheit der NP (das/ein Mädchen), (4) das Alter des Mädchens und (4) die Assoziationen von Mädchen mit 96 stereotypischen Adjektiven, die hinsichtlich Geschlechterstereotypizität (männlich/weiblich) und Konnotation (positiv/negativ) variieren. Mithilfe eines offenen Large Language Models (LLM) wurden erst englische Items generiert, die sämtliche 3.072 möglichen Kombinationen dieser Faktoren abdecken. Nachdem alle Vorkommen von Pronomen in den Items durch Platzhalter ersetzt wurden, wurden diese von der DeepL-API ins Deutsche übersetzt. Anschließend wenden wir ein bayessches logistisches Regressionsmodell an, das Aufschluss darüber gibt, welche Kombinationen mit einer sexuskongruenten gegenüber einer genuskongruenten Pronominalisierung assoziiert werden. Unsere Ergebnisse zeigen, dass die Wahrscheinlichkeit einer sexuskongruenten Pronominalisierung mit zunehmendem Alter des Mädchens steigt. Sie nimmt auch zu, wenn Mädchen mit dem Indefinitartikel eingeführt und im Folgesatz mit einem Possessivpronomen wiederaufgenommen wird. Am geringsten ist sie hingegen, wenn dem Mädchen ein weiblich-negativ konnotiertes Adjektiv zugeschrieben wird. Insgesamt verdeutlichen die Befunde, dass sowohl referentielle als auch kontextuelle Merkmale die Pronominalisierung hybrider Nomina beeinflussen. Zugleich zeigen sie, dass von DeepL generiertes Deutsch ähnliche Muster der Pronominalisierung aufweist wie die von Menschen produzierte Sprache. Unsere Studie liefert somit neue Erkenntnisse über die Rolle von Prominenzhierarchien bei der pronominalen Wiederaufnahme hybrider Nomen und stützt zugleich zentrale Ergebnisse vorangehender Arbeiten zur Reproduktion von Gender-Biases in maschineller Übersetzung und Sprachgenerierung. (2026-07-18) |