Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Social Media Sentiment Analysis on Multilingual Dataset
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology.
2024 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 HE creditsStudent thesis
Abstract [en]

This thesis presents using unlabeled data and soft targets from a LLM to train a pre-trained model for sentiment analysis. Inspired by the study “What do LLMs Know about Financial Markets? A Case Study on Reddit Market Sentiment Analysis” by Deng et al. (2023), this thesis focuses on enhancing sentiment analysis in social media, particularly for YouTube comments that often lack actual labels. Throughout the thesis, three different approaches were tested: (i) We evaluated the performance of sentiment analysis using LLM. (ii) We use the LLM's results as hints for a fine- tuned model to make predictions and then evaluate those results. (iii) We use the LLM's results as soft targets to fine-tune a pre-trained model for improved sentiment analysis.

Our final implementation takes the best approach, and its results show that leveraging the knowledge from LLMs to fine-tune a pre-trained model, even with half the size of the dataset, yields similar sentiment analysis results to model using deep learning methods on the entire dataset or the LLM alone. Additionally, we establish an evaluation method to assessment of our approach.

Place, publisher, year, edition, pages
2024. , p. 45
Series
IT ; mDA 24 011
Keywords [en]
Sentiment Analysis, Unsupervised learning, Social media, Nature Language Processing, Chain-of-Thought
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:uu:diva-536121OAI: oai:DiVA.org:uu-536121DiVA, id: diva2:1888601
External cooperation
CreatorDB
Subject / course
Language Technology
Educational program
Master's Programme in Data Science
Presentation
2024-06-07, Ångströmlaboratoriet, Uppsala University, Lägerhyddsvägen 1, 752 37 Uppsala, Sweden, Sweden, 10:00 (English)
Supervisors
Examiners
Available from: 2024-08-16 Created: 2024-08-13 Last updated: 2025-02-07Bibliographically approved

Open Access in DiVA

fulltext(5080 kB)1597 downloads
File information
File name FULLTEXT01.pdfFile size 5080 kBChecksum SHA-512
caf75f27ad96c99856b5e9c833d2b6f321d61684a3b0a2256281c9acff1387bd9de82167aea621b8661450a7dde20416cb0981800bdb21e54a87d388f40b9e94
Type fulltextMimetype application/pdf

By organisation
Department of Information Technology
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar
Total: 1598 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 1534 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf