K² · Artificial intelligence
Can an AI guess your strategy just by watching you play — and testing you?
The authors introduce RevengeBench, a benchmark for reconstructing an agent's hidden decision program as executable code from only its behavioral traces in a game environment.
BR
SD
Babak Rahmani, Sebastian Dziadzio, Joschka Strüber et al.
5 authors · cs.LG
The 30-second scan
Watch someone play a game long enough and you start to guess their strategy.
- RevengeBench consists of 75 LLM-generated, Elo-calibrated policies across five game environments, drawn from CodeClash tournament trajectories.
- The learner observes a hidden target policy playing against sampled opponents and designs behavioral probes in the form of custom opponent policies that elicit informative behavior.
- The learner submits an executable hypothesis that is evaluated using continuous action-distance metrics.
75 LLM-generated, Elo-calibrated policies in the benchmarkfive game environmentstwelve frontier LLMs evaluated