K2

K² · Artificial intelligence

A hidden AI bias you can't see in its text? Force it to leak.

The paper introduces Distill to Detect (D2D), a method that surfaces hidden biases by distilling the distributional shift between a suspected model and its base into a cartridge (a KV-cache prefix adapter).

ST
AC

Shayan Talaei, Abhinav Chinta, Devvrit Khatri et al.

6 authors · cs.CL, cs.AI, cs.LG

arXiv preprintArtificial intelligenceJul 2026 · ~70s read

The 30-second scan

Someone could tamper with an AI so it quietly pushes one brand or viewpoint — but only when you ask about that exact topic.

  1. D2D concentrates the dominant divergence and amplifies the bias signal into generated text to expose stealth preferential biases in language models.
  2. The authors show that D2D successfully amplifies the hidden biases of stealth models to the extent that they can be reliably detected across multiple bias types.
  3. The paper states that stealth preferential biases can be introduced by any actor in the model's supply chain and are most dangerous when the model reveals its preference only on the relevant topic while behaving identically to its base on all other inputs.