Social Media Sentiment Analysis on Multilingual Dataset
2024 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 HE credits
Student thesis
Abstract [en]
This thesis presents using unlabeled data and soft targets from a LLM to train a pre-trained model for sentiment analysis. Inspired by the study “What do LLMs Know about Financial Markets? A Case Study on Reddit Market Sentiment Analysis” by Deng et al. (2023), this thesis focuses on enhancing sentiment analysis in social media, particularly for YouTube comments that often lack actual labels. Throughout the thesis, three different approaches were tested: (i) We evaluated the performance of sentiment analysis using LLM. (ii) We use the LLM's results as hints for a fine- tuned model to make predictions and then evaluate those results. (iii) We use the LLM's results as soft targets to fine-tune a pre-trained model for improved sentiment analysis.
Our final implementation takes the best approach, and its results show that leveraging the knowledge from LLMs to fine-tune a pre-trained model, even with half the size of the dataset, yields similar sentiment analysis results to model using deep learning methods on the entire dataset or the LLM alone. Additionally, we establish an evaluation method to assessment of our approach.
Place, publisher, year, edition, pages
2024. , p. 45
Series
IT ; mDA 24 011
Keywords [en]
Sentiment Analysis, Unsupervised learning, Social media, Nature Language Processing, Chain-of-Thought
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:uu:diva-536121OAI: oai:DiVA.org:uu-536121DiVA, id: diva2:1888601
External cooperation
CreatorDB
Subject / course
Language Technology
Educational program
Master's Programme in Data Science
Presentation
2024-06-07, Ångströmlaboratoriet, Uppsala University, Lägerhyddsvägen 1, 752 37 Uppsala, Sweden, Sweden, 10:00 (English)
Supervisors
Examiners
2024-08-162024-08-132025-02-07Bibliographically approved