K2

The Morning PaperSaturday, September 26

AI models can't tell when a hacker put words in their mouth

AI models can't reliably tell when hackers inject fake instructions into their prompts—and training them to recognize it often backfires.

~70s readRead today’s paper. Start a streak.

This week’s episodeThe strongest hint of alien life yet

Trending now

Translate any paper

A new paper lands every morning.