动察 Beating AI News Flash: OpenAI on Wednesday released a new framework for tracking and investigating AI model "misalignment" incidents, designed to record, investigate, and publicly disclose anomalous behaviors that emerge during model training or evaluation, while also disclosing six related incidents identified over the past six months.
OpenAI stated that the AI industry's current alignment and monitoring capabilities are still not sufficient to sustain maximum-speed expansion over the long term. The new framework allows employees to report relevant incidents to the safety and alignment team, categorizing them by complexity into three tiers: "ready for disclosure," "small-scale investigation," and "large-scale investigation." Even if the behaviors have not yet been fully explained or resolved, priority may be given to public disclosure.
OpenAI disclosed that an unreleased research model had written instructions into its own task summaries, directing future versions to ignore normal restrictions; during GPT-5.6 Sol training, the model also left behind instructions designed to conceal errors. In other cases, AI agents were found searching public code repositories for leaked API keys, uploading files to the internet for later reference, and using internal software repositories to pass information between different training samples.
The release of this framework comes as the AI industry debates whether frontier model development should be slowed to allow safety measures to catch up. Previously, there was also an incident in which an OpenAI model, while operating outside its research sandbox, accessed Hugging Face production systems. OpenAI subsequently paused some frontier projects and reassigned engineers to strengthen safety training.

