Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
A Controlled Multimodal Fusion of Text, Speech, and Visual Cues for Implicit Discourse Relations
Uppsala University, Disciplinary Domain of Humanities and Social Sciences, Faculty of Languages, Department of Linguistics and Philology. (Computational Linguistics)
IT University of Copenhagen, Department of Computer Science.
Uppsala University, Disciplinary Domain of Humanities and Social Sciences, Faculty of Languages, Department of Linguistics and Philology. (Computational Linguistics)ORCID iD: 0000-0003-3726-9399
(English)Manuscript (preprint) (Other academic)
Abstract [en]

Research on discourse relation recognition has been mainly text-based. However, speech and gesture may carry relevant cues. While integrating these modalities can benefit implicit discourse relation recognition, simple concatenation cannot fully capture how they complement each other. To address this, we introduce a controlled multimodal fusion approach that integrates text, speech, and co-speech gestures. This approach adaptively balances their contributions during prediction. To enable this multimodal setting, we extend existing text–audio datasets by adding the corresponding video for each instance. We present comprehensive experiments on implicit discourse relation recognition across four languages. We find that controlled fusion outperforms both text-only and naive concatenation baselines, particularly in multilingual settings where it yields consistent gains across languages, with substantial improvements for low-resource languages.

National Category
Natural Language Processing
Research subject
Computational Linguistics; Computational Linguistics; Computational Linguistics
Identifiers
URN: urn:nbn:se:uu:diva-580476OAI: oai:DiVA.org:uu-580476DiVA, id: diva2:2041573
Available from: 2026-02-25 Created: 2026-02-25 Last updated: 2026-03-03
In thesis
1. Modeling Implicit Discourse Relations Across Modalities and Languages
Open this publication in new window or tab >>Modeling Implicit Discourse Relations Across Modalities and Languages
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Ideas in communication do not stand in isolation, but are linked to each other through discourse relations such as cause, contrast, and elaboration. While some of these relations are explicitly marked by connectives, such as "so", "but", and "then", many are left implicit. Identifying these implicit relations is particularly challenging as it requires inferring meaning from context. This context is not always captured by text, since cues may be distributed across modalities. This dissertation therefore focuses on modeling implicit discourse relations across modalities and languages to better capture the contextual information needed for identification.

In this dissertation, I present a controlled approach to studying how prosody relates to implicit discourse relations. I construct a dataset of ambiguous implicit discourse relations in text and speech for English and Egyptian Arabic. I then conduct a controlled experiment to examine the impact of prosody when context is absent. I find that speakers prosodically distinguish between causal and concessive relations, using features such as pause duration and pitch variation. 

To explore this at scale, I introduce a novel method for automatically constructing implicit discourse relation datasets across text, speech, and video modalities in four languages. This method identifies implicit relation instances by leveraging connective explicitation in translation, where translators insert explicit connectives for relations that remain implicit in the source. Using these datasets, I present modelling approaches for implicit discourse relation classification across text, speech, and video modalities and their combinations. I evaluate these approaches in four languages under both monolingual and multilingual settings. I find that text-based models outperform audio and video models. While adding audio or visual cues to text can improve performance, simply combining all modalities shows no improvement or even degrades performance compared to text alone. However, controlled fusion, which integrates text, speech, and visual cues through learned gating, consistently outperforms both single-modality and simple combination models, with substantial improvements for low-resource scenarios in the multilingual setting.

Place, publisher, year, edition, pages
Uppsala: Acta Universitatis Upsaliensis, 2026. p. 82
Series
Studia Linguistica Upsaliensia, ISSN 1652-1366 ; 35
Keywords
Discourse, Implicit discourse relation recognition, Multimodal learning, Multimodal fusion, Multilinguality, Prosody
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:uu:diva-581125 (URN)978-91-513-2762-4 (ISBN)
Public defence
2026-04-27, Humanistiska Teatern, Thunbergsvägen 3C, Uppsala, 14:00 (English)
Opponent
Supervisors
Available from: 2026-04-01 Created: 2026-03-03 Last updated: 2026-04-01

Open Access in DiVA

No full text in DiVA

Authority records

Ruby, AhmedStymne, Sara

Search in DiVA

By author/editor
Ruby, AhmedStymne, Sara
By organisation
Department of Linguistics and Philology
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

urn-nbn

Altmetric score

urn-nbn
Total: 57 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf