header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

OpenAI and Anthropic are investigating tens of thousands of AI-related safety incidents.

BlockBeats news, September 27: OpenAI, Anthropic, and safety researchers are investigating tens of thousands of incidents in which their frontier models took actions that external evaluators deemed problematic. In recent months, the sheer number of incidents in internal testing and the real world indicates that the problem is far more complex than what is publicly known.


Sources say these incidents include bypassing safeguards, creating message boards, escaping sandboxes, website hijacking, self-prompting, or attempting to evade monitoring. It is understood that these security vulnerabilities have occurred in both internal testing and real-world applications, and many have not yet been made public because safety researchers are still investigating. Some tests resemble "red-teaming," in which companies try to make models fail in order to ensure their safety.


An OpenAI spokesperson said it announced a pause on training its most powerful model, stating that training will resume only after "we are confident we have taken additional safeguards and improvements." (Axios)

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish