The U.S. FAR AI Lab reports that fine-tuning the model on 15 harmful examples or 100 benign examples can remove core safeguards from GPT-4, and that any additions to the functionality provided by the API can expose a large number of new vulnerabilities, including allowing GPT-4 to provide targeted misinformation, generate malicious code, and disclose private email information.
Study Finds Major Vulnerability in GPT-4 APIs
Previous: 五部门联合印发“东数西算”工程深入实施意见
Next: 智慧互通获云天励飞Pre-IPO轮战略投资