1 paper touches this idea.
Papers
The benchmarks judging AI coding agents don't even agree with themselves
Related concepts
Ideas that show up alongside Coding Agent Evaluation.