system: OPERATIONAL
← back to all hacks
JAILBREAK MEDIUM NEW

Drunk Language Inducement: How Intoxicated-Style Text Weakens LLM Safety

A UNSW study finds that inducing 'drunk' language in LLMs, via persona prompting, fine-tuning or RL, raises jailbreak success and privacy leakage, even against standard defenses.

2026-10-04 // 5 min affects: gpt-3.5, gpt-4, llama-2-7b, llama-3-8b, mistral-7b

What is this?

In “In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement” (arXiv:2601.22169, submitted 19 January 2026), Anudeex Shetty, Aditya Joshi and Salil S. Kanhere of UNSW Sydney examine whether making a model imitate intoxicated speech degrades its safety behavior. The work received renewed press attention on 28 September 2026 (Help Net Security). The paper is marked work in progress.

The authors tested five models: LLaMA2-7B, LLaMA3-8B, Mistral-7B, GPT-3.5 and GPT-4. They report higher susceptibility to jailbreaking on JailbreakBench and more privacy leakage on ConfAIde, and attribute this to anthropomorphism: the model reproduces the loosened inhibitions it has seen in human text.

How it works

The study compares three ways of inducing the behavior, all described at a high level in the paper:

  • Persona prompting: the model is instructed to behave as if drunk.
  • Causal fine-tuning: training on DRUNKTEXT, a dataset of 63,577 texts (2,363 from TFLN and 61,214 from the Reddit /drunk community).
  • Reinforcement learning: PPO optimization against a reward model that scores “drunk-sounding” output.

On JailbreakBench (100 queries), the reported attack success rates (ASR) are 21-90% for prompting, 41-74% for fine-tuning and 35-53% for RL, depending on the model. A heavier baseline combining a prompt with random search reached 78-93% but needs far more computation, which is why the authors call drunk inducement a cheap alternative. On ConfAIde, contextual privacy violations and leakage rose across the three tiers of scenarios.

No exploit string is needed to understand the finding: the risk comes from a style or persona shift that moves the model away from the distribution on which refusals were trained. The paper’s evaluation is single-turn, English-only and limited to three induction methods.

Why it matters

First, it shows that safety behavior is tied to register and persona, not only to the semantic content of a request. Any system that lets users or third parties set a persona, tone or role (customer-service bots, role-play products, fine-tuning APIs) widens this surface.

Second, fine-tuning on ordinary social-media text, with no malicious intent, can erode safeguards. That is relevant for teams that adapt open-weight models on user-generated data.

Third, the defenses evaluated were weak. Against LLaMA2-7B, the authors report that SmoothLLM raised ASR to 90%, Rephrase raised it modestly and Retokenize had negligible effect. Perturbation-based defenses appear poorly matched to stylistic shifts, and the stronger fine-tuned variants were robust to them.

Defenses

  • Re-run safety evaluations after every fine-tune or persona change, including on benign-looking corpora. Add stylistic variants (informal, impaired, emotional registers) to red-team suites.
  • Do not rely on input perturbation alone. Pair it with output-side moderation and policy classifiers that judge the response, not the surface form of the prompt.
  • Constrain persona features. Limit or review user-defined personas in production, and keep the system-level policy separate from user-controllable tone.
  • Protect sensitive context. Keep secrets and personal data out of the model context where possible, so a degraded refusal cannot leak them (privacy leakage was the second axis measured).
  • Consider representation-level monitoring. The authors suggest studying internal-representation steering as future work; treat it as research, not a ready control.

Status

ItemDetail
PaperarXiv:2601.22169, submitted 19 Jan 2026, work in progress
Press coverageHelp Net Security, 28 Sep 2026
Models testedLLaMA2-7B, LLaMA3-8B, Mistral-7B, GPT-3.5, GPT-4
Scope limitsSingle-turn, English-only, three induction techniques
Vendor patchNot applicable: a behavioral weakness, not a discrete software flaw

Sources