Social Media Sentiment Analysis on Multilingual Dataset
2024 (Engelska)Självständigt arbete på avancerad nivå (masterexamen), 20 hp
Studentuppsats (Examensarbete)
Abstract [en]
This thesis presents using unlabeled data and soft targets from a LLM to train a pre-trained model for sentiment analysis. Inspired by the study “What do LLMs Know about Financial Markets? A Case Study on Reddit Market Sentiment Analysis” by Deng et al. (2023), this thesis focuses on enhancing sentiment analysis in social media, particularly for YouTube comments that often lack actual labels. Throughout the thesis, three different approaches were tested: (i) We evaluated the performance of sentiment analysis using LLM. (ii) We use the LLM's results as hints for a fine- tuned model to make predictions and then evaluate those results. (iii) We use the LLM's results as soft targets to fine-tune a pre-trained model for improved sentiment analysis.
Our final implementation takes the best approach, and its results show that leveraging the knowledge from LLMs to fine-tune a pre-trained model, even with half the size of the dataset, yields similar sentiment analysis results to model using deep learning methods on the entire dataset or the LLM alone. Additionally, we establish an evaluation method to assessment of our approach.
Ort, förlag, år, upplaga, sidor
2024. , s. 45
Serie
IT ; mDA 24 011
Nyckelord [en]
Sentiment Analysis, Unsupervised learning, Social media, Nature Language Processing, Chain-of-Thought
Nationell ämneskategori
Språkbehandling och datorlingvistik
Identifikatorer
URN: urn:nbn:se:uu:diva-536121OAI: oai:DiVA.org:uu-536121DiVA, id: diva2:1888601
Externt samarbete
CreatorDB
Ämne / kurs
Språkteknologi
Utbildningsprogram
Masterprogram i dataanalys
Presentation
2024-06-07, Ångströmlaboratoriet, Uppsala University, Lägerhyddsvägen 1, 752 37 Uppsala, Sweden, Sweden, 10:00 (Engelska)
Handledare
Examinatorer
2024-08-162024-08-132025-02-07Bibliografiskt granskad