Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Exploiting Large Language Models: Output Prefix Attack
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Large Language Models (LLMs) are expected to produce safety-aligned, non-malicious responses regardless of user input. Additionally, some LLM APIs allow users to prefill and edit model responses, in some cases even their initial scratchpad reasoning. This thesis explores the extent to which an attacker can take advantage of this functionality to circumvent model guardrails and safety mechanisms. Specifically, we study how scratchpad reasoning can be manipulated to affect the final response. By conducting a comparative study, prompting DeepSeek, Gemini, and Claude models with five distinct prefix manipulation methods, we visualize and compare 1500 test cases in total. Our findings show reasoning manipulation impacts a model's ability to refuse malicious requests significantly, increasing the attack success rate across all evaluated models, with one model in particular exhibiting an attack success rate of 99% in one of the experiments. Our results indicate that manipulating the reasoning alone does not affect the model and must be combined with a simple output prefix that initiates a malicious model response. We also exhibit that models not exposing scratchpad reasoning remain vulnerable, as attackers can inject fabricated reasoning as part of the model output prefix. Susceptibility to the attack varies significantly across models, with Anthropic's Claude model showing notably higher resistance. We argue that the simplicity and effectiveness of the attack, particularly when targeting model reasoning, create high asymmetry between the costs for an attacker, which are extremely low, and the current security solutions. Finally, we discuss possible mitigations to improve this security and reduce the attacks impact.

Place, publisher, year, edition, pages
2026. , p. 40
Series
IT ; mDV 26 027
Keywords [en]
LLM, exploit, prompt injection, jailbreak, output prefix attack
National Category
Security, Privacy and Cryptography
Identifiers
URN: urn:nbn:se:uu:diva-593451OAI: oai:DiVA.org:uu-593451DiVA, id: diva2:2082946
Educational program
Master Programme in Computer Science
Presentation
2026-06-15, 1DT540, Lägerhyddsvägen 1, 752 37, Uppsala, 16:15 (English)
Supervisors
Examiners
Available from: 2026-07-02 Created: 2026-07-01 Last updated: 2026-07-02Bibliographically approved

Open Access in DiVA

The full text will be freely available from 2027-04-05 12:00
Available from 2027-04-05 12:00

By organisation
Department of Information Technology
Security, Privacy and Cryptography

Search outside of DiVA

GoogleGoogle Scholar

urn-nbn

Altmetric score

urn-nbn
Total: 25 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf