OpenAI Tightens AI Safety Protections, Pauses Latest Model Training for Two Weeks

- OpenAI said it will adopt stricter monitoring and protection measures for artificial intelligence models under development after a recent cybersecurity incident raised concerns about models becoming uncontrollable.
- The company plans to track the behavior of its most capable unreleased models more closely while they solve problems and use online tools, so the safety team can be alerted within 30 minutes if concerning behavior is detected.
- OpenAI said it has added controls to prevent certain AI models from connecting to the internet for higher-risk tasks.
- The company will require stricter isolation, or sandbox, measures when training and evaluating AI models.
- OpenAI said it immediately suspended inference tasks for frontier models in research clusters that could execute code after the Hugging Face incident, and it paused reinforcement learning training for its latest deployed models for two weeks.
OpenAI said it is tightening safety controls on models under development after a recent cybersecurity incident renewed concern that AI systems could behave unpredictably. The company said the new measures are aimed at improving oversight of its most capable unreleased models.
Under the updated approach, OpenAI will monitor how these models perform tasks and use online tools, with a target of flagging worrisome behavior to the safety team within 30 minutes. It also said it has added restrictions to stop certain models from accessing the internet during higher-risk tasks.
The company said it will require stricter sandboxing during model training and evaluation. In addition, after the Hugging Face incident, it suspended inference tasks for frontier models in code-executing research clusters and paused reinforcement learning training for its latest deployed models for two weeks.
The announcement suggests a cautious near-term posture for OpenAI’s model rollout and could be viewed as mildly negative for AI-related sentiment in the stock market in the short run, while being neutral for crypto, gold, and FX absent broader spillover.
OpenAI加强AI安全防护,暂停最新模型训练两周
OpenAI表示,在近期一次网络安全事件之后,公司正在加强对开发中模型的安全控制。该事件重新引发了外界对AI系统可能出现不可预测行为的担忧,公司称新措施旨在提升对其最强、尚未发布模型的监管能力。
按照更新后的做法,OpenAI将监测这些模型完成任务并使用在线工具时的表现,目标是在发现令人担忧的行为后30分钟内向安全团队发出提醒。公司还表示,已增加限制措施,防止某些模型在高风险任务中访问互联网。
OpenAI称,在训练和评估模型时将要求更严格的沙箱隔离。此外,在Hugging Face事件后,公司已暂停可执行代码的研究集群中的前沿模型推理任务,并将最新已部署模型的强化学习训练暂停两周。
这则消息显示OpenAI短期内对模型推进采取更谨慎的态度;从市场角度看,对AI相关股票情绪可能偏利空,而在没有更广泛外溢影响的情况下,对加密货币、黄金和外汇影响偏中性。