Social Media Sentiment Analysis on Multilingual Dataset
2024 (engelsk)Independent thesis Advanced level (degree of Master (Two Years)), 20 hp
Oppgave
Abstract [en]
This thesis presents using unlabeled data and soft targets from a LLM to train a pre-trained model for sentiment analysis. Inspired by the study “What do LLMs Know about Financial Markets? A Case Study on Reddit Market Sentiment Analysis” by Deng et al. (2023), this thesis focuses on enhancing sentiment analysis in social media, particularly for YouTube comments that often lack actual labels. Throughout the thesis, three different approaches were tested: (i) We evaluated the performance of sentiment analysis using LLM. (ii) We use the LLM's results as hints for a fine- tuned model to make predictions and then evaluate those results. (iii) We use the LLM's results as soft targets to fine-tune a pre-trained model for improved sentiment analysis.
Our final implementation takes the best approach, and its results show that leveraging the knowledge from LLMs to fine-tune a pre-trained model, even with half the size of the dataset, yields similar sentiment analysis results to model using deep learning methods on the entire dataset or the LLM alone. Additionally, we establish an evaluation method to assessment of our approach.
sted, utgiver, år, opplag, sider
2024. , s. 45
Serie
IT ; mDA 24 011
Emneord [en]
Sentiment Analysis, Unsupervised learning, Social media, Nature Language Processing, Chain-of-Thought
HSV kategori
Identifikatorer
URN: urn:nbn:se:uu:diva-536121OAI: oai:DiVA.org:uu-536121DiVA, id: diva2:1888601
Eksternt samarbeid
CreatorDB
Fag / kurs
Language Technology
Utdanningsprogram
Master's Programme in Data Science
Presentation
2024-06-07, Ångströmlaboratoriet, Uppsala University, Lägerhyddsvägen 1, 752 37 Uppsala, Sweden, Sweden, 10:00 (engelsk)
Veileder
Examiner
2024-08-162024-08-132025-02-07bibliografisk kontrollert