K2

K² · Artificial intelligence

Can an AI guess your strategy just by watching you play — and testing you?

The authors introduce RevengeBench, a benchmark for reconstructing an agent's hidden decision program as executable code from only its behavioral traces in a game environment.

BR
SD

Babak Rahmani, Sebastian Dziadzio, Joschka Strüber et al.

5 authors · cs.LG

arXiv preprintArtificial intelligenceJun 2026 · ~60s read

The 30-second scan

Watch someone play a game long enough and you start to guess their strategy.

  1. RevengeBench consists of 75 LLM-generated, Elo-calibrated policies across five game environments, drawn from CodeClash tournament trajectories.
  2. The learner observes a hidden target policy playing against sampled opponents and designs behavioral probes in the form of custom opponent policies that elicit informative behavior.
  3. The learner submits an executable hypothesis that is evaluated using continuous action-distance metrics.
75 LLM-generated, Elo-calibrated policies in the benchmarkfive game environmentstwelve frontier LLMs evaluated