LLMs respond differently to harmful prompts when AI watermarking is used


SynthID can cause models to follow harmful instructions they would otherwise refuse.
Quelle: Originalartikel öffnen


SynthID can cause models to follow harmful instructions they would otherwise refuse.
Quelle: Originalartikel öffnen