1 paper touches this idea.
Papers
AI models can't tell when a hacker put words in their mouth
Related concepts
Ideas that show up alongside Model Interpretability.