K² · Artificial intelligence
AI models can't tell when a hacker put words in their mouth
The paper examines how reliably an LLM can recognize that its own prior response was elicited by an adversarial prefill attack, extending prior work on LLM introspection from benign tasks to safety contexts.
QN
UA
Quang Minh Nguyen, Uzair Ahmed, Taegyoon Kim
3 authors · cs.CL
The 30-second scan
Trick an AI into saying something harmful, then ask it: "Did you mean to say that?" Most of the time, it says yes.
- Across ten open-weight instruction-tuned LLMs ranging from 3B to 70B parameters and four safety benchmarks, no model reliably recognizes its own compromised outputs.
- Models claim intent on prefilled responses at an average rate of 27.3%.
- The introspective signal stems largely from safety- and refusal-related reasoning.
Ten open-weight instruction-tuned LLMs ranging from 3B to 70B parametersFour safety benchmarksModels claim intent on prefilled responses at an average rate of 27.3%