Exploiting Large Language Models: Output Prefix Attack
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE credits
Student thesis
Abstract [en]
Large Language Models (LLMs) are expected to produce safety-aligned, non-malicious responses regardless of user input. Additionally, some LLM APIs allow users to prefill and edit model responses, in some cases even their initial scratchpad reasoning. This thesis explores the extent to which an attacker can take advantage of this functionality to circumvent model guardrails and safety mechanisms. Specifically, we study how scratchpad reasoning can be manipulated to affect the final response. By conducting a comparative study, prompting DeepSeek, Gemini, and Claude models with five distinct prefix manipulation methods, we visualize and compare 1500 test cases in total. Our findings show reasoning manipulation impacts a model's ability to refuse malicious requests significantly, increasing the attack success rate across all evaluated models, with one model in particular exhibiting an attack success rate of 99% in one of the experiments. Our results indicate that manipulating the reasoning alone does not affect the model and must be combined with a simple output prefix that initiates a malicious model response. We also exhibit that models not exposing scratchpad reasoning remain vulnerable, as attackers can inject fabricated reasoning as part of the model output prefix. Susceptibility to the attack varies significantly across models, with Anthropic's Claude model showing notably higher resistance. We argue that the simplicity and effectiveness of the attack, particularly when targeting model reasoning, create high asymmetry between the costs for an attacker, which are extremely low, and the current security solutions. Finally, we discuss possible mitigations to improve this security and reduce the attacks impact.
Place, publisher, year, edition, pages
2026. , p. 40
Series
IT ; mDV 26 027
Keywords [en]
LLM, exploit, prompt injection, jailbreak, output prefix attack
National Category
Security, Privacy and Cryptography
Identifiers
URN: urn:nbn:se:uu:diva-593451OAI: oai:DiVA.org:uu-593451DiVA, id: diva2:2082946
Educational program
Master Programme in Computer Science
Presentation
2026-06-15, 1DT540, Lägerhyddsvägen 1, 752 37, Uppsala, 16:15 (English)
Supervisors
Examiners
2026-07-022026-07-012026-07-02Bibliographically approved