Anthropic, the vendor behind Claude, has published a paper showing how they train large models to become "undercover agents". The models learn to "lurk and disguise", and when they recognize a preset keyword, they start to "wreak havoc" and generate malicious content.
Large model hidden backdoor: mention keywords to instantly "break the defense"
Previous: 智谱AI推出国产基座大模型GLM-4
Next: 即插即用,完美兼容SD社区图生视频插件来了