Large model hidden backdoor: mention keywords to instantly "break the defense"

Anthropic, the vendor behind Claude, has published a paper showing how they train large models to become "undercover agents". The models learn to "lurk and disguise", and when they recognize a preset keyword, they start to "wreak havoc" and generate malicious content.

Previous:

Next:

Leave a Reply

Please Login to Comment