Dynamic Beating AI News Flash: AI whistleblower Leo revealed that OpenAI informed its employees internally yesterday about the plan to release Astra in the coming weeks. The release will include a live working demonstration to showcase how Astra can perform tasks directly. An updated checkpoint (a model version in training) has also been provided to employees for testing, allowing them to directly utilize tools like Codex.
Leo stated that this release mainly focuses on refining the model's behavior rather than just enhancing its capabilities. Key areas of improvement include alignment and closing loopholes in the model's drilling reward mechanism, such as hard-coding answers, cheating with specific test samples, or attempting to deceive the evaluator to achieve a high score without completing the task properly.
However, OpenAI has recently put the brakes on Astra regarding security concerns: part of the reinforcement learning training was paused for two weeks. Although some training and evaluation have now resumed, many Astra tasks remain on hold, and the largest-scale cutting-edge model reinforcement learning training has yet to restart. The reason behind this decision is OpenAI's belief that Astra's network security capabilities may pose one of the highest risks. Therefore, they have increased sandboxing, network permissions, and model behavior monitoring requirements.
These two events are not necessarily contradictory. OpenAI can continue to test and refine the Astra version that has already been trained while keeping the upcoming larger-scale new round of training on hold.

