Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Humble Pleas from the Archives: Automatic Analysis and Information Extraction from Historical Petitions
Uppsala University, Disciplinary Domain of Humanities and Social Sciences, Faculty of Languages, Department of Linguistics and Philology.
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Description
Abstract [en]

Historical archives offer rich insights into past lived experiences, yet linguistic variation, non-standard orthography, and limited annotated resources challenge computational analysis. This dissertation investigates the application of NLP to historical text, focusing on 18th-century Swedish petitions. The work explores how different modelling paradigms can support both genre identification and extraction of historically meaningful information.

The research follows a stepwise methodology across four studies. First, petitions are classified alongside other historical text types using approaches ranging from traditional machine learning models to a Swedish BERT-based classifier, achieving strong results on in-domain data. Second, feature analysis is used to identify key linguistic markers of the petition genre, including thematic vocabulary and expressions of social hierarchy. Third, we explore automatic methods to identify rhetorical components, such as salutations and requests, using both low-resource techniques and LLMs, showing that while formulaic sections can be reliably detected, other parts remain inherently ambiguous. The inclusion of an English dataset further enables evaluation of cross-linguistic generalisation.

Finally, the dissertation addresses extraction and phrase normalisation of work-related expressions within the Gender and Work (GaW) framework. Experiments with large language models show promising results: although exact phrase matching is weak, string-level and semantic similarity indicate that models can locate relevant topical regions. Qualitative analysis further shows that models can detect plausible work-related expressions not present in the gold data, pointing towards hybrid human–machine workflows for improving coverage in historical research.

A key contribution of this work is the application of evaluation strategies that move beyond exact matching to incorporate string-level and semantic similarity, enabling a more nuanced assessment of model performance on noisy historical text. Overall, the findings highlight both the potential and limitations of current NLP methods for historical text, and demonstrate how computational approaches can support the analysis of complex archival material.

Place, publisher, year, edition, pages
Uppsala: Acta Universitatis Upsaliensis, 2026. , p. 58
Series
Studia Linguistica Upsaliensia, ISSN 1652-1366 ; 36
Keywords [en]
NLP for historical text, digital humanities, petitions, large language models, LLMs, text classification, feature analysis, rhetorical analysis, information extraction, text analysis, text normalisation, phrase normalisation, Swedish, prompting, NLP, language technology, computational linguistics
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
URN: urn:nbn:se:uu:diva-584819ISBN: 978-91-513-2864-5 (print)OAI: oai:DiVA.org:uu-584819DiVA, id: diva2:2055353
Public defence
2026-06-12, Humanistiska teatern, Thunbergsvägen 3H, Uppsala, 13:15 (English)
Opponent
Supervisors
Part of project
Speaking to One´s Superiors: Petitions as cultural heritage and sources of knowledge, Swedish Research CouncilAvailable from: 2026-05-22 Created: 2026-04-23 Last updated: 2026-05-22
List of papers
1. To the Most Gracious Highness, from Your Humble Servant: Analysing Swedish 18th Century Petitions Using Text Classification
Open this publication in new window or tab >>To the Most Gracious Highness, from Your Humble Servant: Analysing Swedish 18th Century Petitions Using Text Classification
2022 (English)In: Proceedings of the 6th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, Association for Computational Linguistics, 2022, p. 53-64Conference paper, Published paper (Refereed)
Abstract [en]

Petitions are a rich historical source, yet they have been relatively little used in historical research. In this paper, we aim to analyse Swedish texts from around the 18th century, and petitions in particular, using automatic means of text classification. We also test how text pre-processing and different feature representations affect the result, and we examine feature importance for our main class of interest – petitions. Our experiments show that the statistical algorithms NB, RF, SVM, and kNN are indeed very able to classify different genres of historical text. Further, we find that normalisation has a positive impact on classification, and that content words are particularly informative for the traditional models. A fine-tuned BERT model, fed with normalised data, outperforms all other classification experiments with a macro average F1 score at 98.8. However, using less computationally expensive methods, including feature representation with word2vec, fastText embeddings or even TF-IDF values, with a SVM classifier also show good results for both unnormalised and normalised data. In the feature importance analysis, where we obtain the features most decisive for the classification models, we find highly relevant characteristics of the petitions, namely words expressing signs of someone inferior addressing someone superior. 

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2022
Keywords
text classification, feature importance, petitions, Swedish, historical, 18th century, digital humanities, digital philology
National Category
Natural Language Processing
Identifiers
urn:nbn:se:uu:diva-491250 (URN)001698332000007 ()
Conference
The 6th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, October 2022, Gyeongju, Republic of Korea
Available from: 2022-12-20 Created: 2022-12-20 Last updated: 2026-05-29Bibliographically approved
2. Low-Resource Techniques for Analysing the Rhetorical Structure of Swedish Historical Petitions
Open this publication in new window or tab >>Low-Resource Techniques for Analysing the Rhetorical Structure of Swedish Historical Petitions
2023 (English)In: Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023) / [ed] Nikolai Ilinykh; Felix Morger; Dana Dannélls; Simon Dobnik; Beáta Megyesi; Joakim Nivre, Association for Computational Linguistics, 2023, p. 132-139Conference paper, Published paper (Refereed)
Abstract [en]

Natural language processing techniques can be valuable for improving and facilitating historical research. This is also true for the analysis of petitions, a source which has been relatively little used in historical research. However, limited data resources pose challenges for mainstream natural language processing approaches based on machine learning. In this paper, we explore methods for automatically segmenting petitions according to their rhetorical structure. We find that the use of rules, word embeddings, and especially keywords can give promising results for this task.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2023
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:uu:diva-518959 (URN)978-1-959429-73-9 (ISBN)
Conference
Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023), Tórshavn, the Faroe Islands, May 22, 2023
Available from: 2024-01-01 Created: 2024-01-01 Last updated: 2026-04-23Bibliographically approved
3. Finding the Plea: Evaluating the Ability of LLMs to Identify Rhetorical Structure in Swedish and English Historical Petitions
Open this publication in new window or tab >>Finding the Plea: Evaluating the Ability of LLMs to Identify Rhetorical Structure in Swedish and English Historical Petitions
2025 (English)In: Proceedings of the First Workshop on Natural Language Processing and Language Models for Digital Humanities / [ed] Isuri Nanomi Arachchige; Francesca Frontini; Ruslan Mitkov; Paul Rayson, Association for Computational Linguistics, 2025, p. 86-101Conference paper, Published paper (Refereed)
Abstract [en]

Large language models (LLMs) have shown impressive capabilities across many NLP tasks, but their effectiveness on fine-grained content annotation, especially for historical texts, remains underexplored. This study investigates how well GPT-4, Gemini, Mixtral, Mistral, and LLaMA can identify rhetorical sections (Salutatio, Petitio, and Conclusio) in 100 English and 100 Swedish petitions using few-shot prompting with varying levels of detail. Most models perform very well, achieving F1 scores in the high 90s for Salutatio, though Petitio and Conclusio prove more challenging, particularly for smaller models and Swedish data. Cross-lingual prompting yields mixed results, and models generally underestimate document difficulty. These findings demonstrate the strong potential of LLMs for assisting with nuanced historical annotation while highlighting areas for further investigation.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2025
Keywords
historical text, large language models, petitions, digital humanities, NLP, annotation, segmentation, rhetoric
National Category
Natural Language Processing
Identifiers
urn:nbn:se:uu:diva-584703 (URN)10.26615/978-954-452-106-6-008 (DOI)978-954-452-106-6 (ISBN)
Conference
The First Workshop on Natural Language Processing and Language Models for Digital Humanities, 11 September, 2025, Varna, Bulgaria
Funder
Swedish Research Council, 2018-06159
Available from: 2026-04-21 Created: 2026-04-21 Last updated: 2026-04-23Bibliographically approved
4. Uncovering Work from Words: LLM-Based Information Extraction from Historical Petitions
Open this publication in new window or tab >>Uncovering Work from Words: LLM-Based Information Extraction from Historical Petitions
2026 (English)Conference paper, Published paper (Refereed)
Abstract [en]

We investigate the extraction and normalisation of phrases describing work from 18th-century Swedish petitions using four LLMs: GPT-4o, Llama-3 70B/8B, and Mixtral-8x7B. Performance is evaluated across four configurations: isolated extraction, isolated normalisation, a staged pipeline, and a combined multitasking setup, using both full and filtered texts (with formal greetings and closing sections removed). While exact phrase matching remains low (F1 < .10), token-level and semantic similarity scores suggest that models consistently locate relevant topical regions. Semantic similarity scores must however be interpreted with caution, since they are often only marginally higher than an average baseline. Results reveal a “multitasking paradox”: combined extraction and normalisation improves phrase location for high-parameter models but degrades normalisation precision. Furthermore, normalisation benefits from the context of a staged pipeline compared to isolated tasks, while text filtering has only marginal effects. Despite a tendency towards over-prediction, qualitative analysis suggests that models can detect plausible work-related expressions missed by human annotators. These findings illustrate the challenges of historical extraction and suggest that hybrid human – machine workflows are a promising approach for enhancing coverage and interpretability in cultural heritage research.

Keywords
information extraction, petitions, large language models, historical Swedish, NLP, historical NLP, digital humanities, text normalisation
National Category
Natural Language Processing
Identifiers
urn:nbn:se:uu:diva-584788 (URN)
Conference
The Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026), Palma, Mallorca, 11 May, 2026
Funder
Swedish Research Council, 2018-06159
Available from: 2026-04-23 Created: 2026-04-23 Last updated: 2026-04-24

Open Access in DiVA

UUThesis_Lindqvist,E-2026(555 kB)146 downloads
File information
File name FULLTEXT01.pdfFile size 555 kBChecksum SHA-512
81c50c9d9629cb812c5ef0db0155b53d23bd0a3f9482da04420e493087cf9cea332d68d3ab5e0312173bfbf08a910163e33ff3f9ef858c9cdf3541196e971405
Type fulltextMimetype application/pdf

Authority records

Lindqvist, Ellinor

Search in DiVA

By author/editor
Lindqvist, Ellinor
By organisation
Department of Linguistics and Philology
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

isbn
urn-nbn

Altmetric score

isbn
urn-nbn
Total: 310 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf