Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages
Cardiff Univ, Cardiff, Wales..
Imperial Coll London, London, England..
Univ Pretoria, Data Sci Social Impact Res Grp, Pretoria, South Africa..
Show others and affiliations
2024 (English)In: FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS: ACL 2024 / [ed] Martins, A Srikumar, V Ku, LW, Association for Computational Linguistics, 2024, p. 2512-2530Conference paper, Published paper (Refereed)
Abstract [en]

Exploring and quantifying semantic relatedness is central to representing language and holds significant implications across various NLP tasks. While earlier NLP research primarily focused on semantic similarity, often within the English language context, we instead investigate the broader phenomenon of semantic relatedness. In this paper, we present SemRel, a new semantic relatedness dataset collection annotated by native speakers across 13 languages: Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu. These languages originate from five distinct language families and are predominantly spoken in Africa and Asia - regions characterised by a relatively limited availability of NLP resources. Each instance in the SemRel datasets is a sentence pair associated with a score that represents the degree of semantic textual relatedness between the two sentences. The scores are obtained using a comparative annotation framework. We describe the data collection and annotation processes, challenges when building the datasets, baseline experiments, and their impact and utility in NLP.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2024. p. 2512-2530
National Category
Natural Language Processing Comparative Language Studies and Linguistics Studies of Specific Languages
Identifiers
URN: urn:nbn:se:uu:diva-558305DOI: 10.18653/v1/2024.findings-acl.147ISI: 001356731802036Scopus ID: 2-s2.0-85205313385ISBN: 979-8-89176-099-8 (electronic)OAI: oai:DiVA.org:uu-558305DiVA, id: diva2:1965965
Conference
62nd Annual Meeting of the Association-for-Computational-Linguistics (ACL) / Student Research Workshop (SRW), AUG 11-16, 2024, Bangkok, THAILAND
Available from: 2025-06-09 Created: 2025-06-09 Last updated: 2025-10-07Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Beloucif, Meriem

Search in DiVA

By author/editor
Beloucif, Meriem
By organisation
Department of Linguistics and Philology
Natural Language ProcessingComparative Language Studies and LinguisticsStudies of Specific Languages

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 60 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf