K2

K² · Artificial intelligence

AI models can't tell when a hacker put words in their mouth

The paper examines how reliably an LLM can recognize that its own prior response was elicited by an adversarial prefill attack, extending prior work on LLM introspection from benign tasks to safety contexts.

QN
UA

Quang Minh Nguyen, Uzair Ahmed, Taegyoon Kim

3 authors · cs.CL

arXiv preprintArtificial intelligenceJun 2026 · ~70s read

The 30-second scan

Trick an AI into saying something harmful, then ask it: "Did you mean to say that?" Most of the time, it says yes.

  1. Across ten open-weight instruction-tuned LLMs ranging from 3B to 70B parameters and four safety benchmarks, no model reliably recognizes its own compromised outputs.
  2. Models claim intent on prefilled responses at an average rate of 27.3%.
  3. The introspective signal stems largely from safety- and refusal-related reasoning.
Ten open-weight instruction-tuned LLMs ranging from 3B to 70B parametersFour safety benchmarksModels claim intent on prefilled responses at an average rate of 27.3%