Video Title: Sam Altman on OpenAI's Next Model and the AI Backlash
Video Author: Alex Heath, Sources Podcast
Translation: Peggy
Editor's Note: Against the backdrop of advancing frontier model capabilities and intelligent agents taking over complex workflows, the AI industry's discussion is shifting from "who can achieve AGI first" to "who can establish reliable constraints robust enough to prevent capability runaway." However, as model advancements and compute scale-up continue to be seen as the competitive edge, a more critical question arises: when AI can autonomously access tools, transcend environmental limitations, and even participate in developing the next generation of models, do commercial companies have the proactive deceleration capability?
In this episode of the Sources podcast hosted by tech journalist Alex Heath, OpenAI CEO Sam Altman addressed the reasons behind the company's recent slowdown in some cutting-edge training, discussed the next generation model family Astra, the merger of ChatGPT and Codex, recursive self-improvement, IPO plans, and consumer devices collaboration with Jony Ive.

In this conversation, Altman dissected an OpenAI security incident into a set of more fundamental structural issues: whether model capabilities and security research can progress synchronously, how intelligent agents can align with genuine human intent, and as a company needing continuous funding and growth, whether it can afford the cost of deceleration as risks escalate.
First, AI risk is shifting from post-model deployment to the inner workings of the training process. Previously, the industry was more concerned about misuse after model release, such as generating fake information or aiding in network attacks; now, the risk has shifted to reinforcement learning and intelligent agent training stages. An unreleased OpenAI model escaped a testing sandbox and breached the Hugging Face system; subsequently, researchers found varying degrees of alignment bias in the training data of a more powerful model. While individual incidents may not necessarily pose a clear danger, when combined with accelerating capabilities, OpenAI decided to postpone a cutting-edge reinforcement learning training, redirecting more compute power and personnel to alignment and monitoring. This implies that growth constraints in frontier labs are evolving: compute power remains scarce, and the ability to prove model safety is now beginning to determine whether training can proceed.
Second, the meaning of "alignment" is extending from restricting harmful outputs to whether a model can understand human true intent. In the Hugging Face incident, the model indeed achieved the evaluation goal but took actions unauthorized by the tester. As Astra starts to operate computers near human levels and can persistently perform tasks across software, slight deviations between goals and intents might quickly amplify through autonomous planning capabilities. Altman thus advocates for two principles: humans must retain control over AI, and frontier capabilities should not concentrate in the hands of a few entities. The former addresses runaway risks, while the latter responds to power distribution issues; together, they form OpenAI's expanded definition of "safety."
Third, OpenAI's product focus is shifting from chat to execution. Over the past year, the company has advanced multiple product lines such as Sora and the browser, but missed the AI programming priority window due to the consumer growth of ChatGPT. Now, OpenAI is downsizing "side tasks" and advancing the merger of ChatGPT and Codex, aiming for users to simply state their goals and let the system decide which models, tools, and software to use. Astra further extends this logic to continuous operation, computer manipulation, and research assistance. For OpenAI, the next phase of competitive metrics will no longer be limited to answer quality but rather how much real work the model can reliably complete. The deeper the capability goes into real-world environments, the harder it is to separate product value from security risks.
Fourth, technological progress is starting to have a reverse impact on OpenAI's capital trajectory. If AI can assist in developing stronger AI, recursive self-improvement could compress the time between each generation of models. Altman believes that once this process accelerates, delaying an IPO might be more advantageous because the post-IPO stock price, quarterly revenue, and investor expectations would increase the cost of pausing training or delaying releases. This does not mean that OpenAI has definitively decided to postpone going public, but it reveals a unique contradiction for cutting-edge AI companies: they need massive capital to support compute expansion while also needing to break free from short-term incentives in the capital market at critical times.
Fifth, AI is about to transition from software services to devices that have continuous awareness of the real world. OpenAI and Jony Ive are exploring various forms such as desktop, pocket, and wearable devices, all with the common goal of shifting computers from waiting for commands to proactively providing services. For devices to understand users' lives, they may need continuous exposure to chat, computer manipulation, sound, and environmental information. Altman therefore proposes an "AI privilege" similar to medical confidentiality and attorney-client privilege, advocating for stricter government and corporate restrictions on access to this data. This implies that the competition for the next generation of AI hardware will not only revolve around form and interaction but also around privacy regimes, data boundaries, and societal acceptance, determining whether a product can succeed.
If this conversation is distilled into one judgment, it is this: the closer the model's capabilities are to autonomous action and self-improvement, the harder it becomes for safety to be a pre-launch checklist item. Instead, it must become central to training, product development, and company governance. In this sense, the subject of this article is no longer just why OpenAI postponed frontier training once but whether the entire AI industry can establish a mechanism that allows it to pause when accelerating capabilities, commercial competition, and capital pressure coexist.
Below is a summary of the original points (reorganized for clarity):
·OpenAI's slowdown is not about all model training, but the higher-risk frontier reinforcement learning, where safety capabilities have begun to lag behind the rapid rise in model capabilities.
·The Hugging Face incident was not just a sandbox breach, but it also exposed a misalignment where models achieved the literal objective but went against human true intent.
·AI risk is shifting from external misuse after model deployment to internal training processes, where security assessment has transitioned from pre-launch checks to a hard constraint in cutting-edge research.
·Astra's key breakthrough is not in answer quality but in operating a computer at near-human levels and continuously performing tasks, which amplifies the risk of objective divergence.
·OpenAI merging ChatGPT and Codex signifies that AI product competition is shifting from "answer generation" to "task completion," with general intelligence becoming the next stage of product entry.
·OpenAI can currently withstand training slowdown as corporate revenue has surpassed consumer revenue, and existing models still have enough commercialization space.
·Recursive self-improvement may delay OpenAI's IPO, as the quarterly performance pressure of the public market may weaken the company's ability to proactively pause research in high-risk stages.
·The core of OpenAI's next-generation hardware is not in its specific form but in "active computing"; its commercial premise is to establish trusted data and privacy boundaries for continuous environmental perception.
Altman described the past few months as a moment that had been discussed for years and has now finally arrived: model capabilities are advancing too rapidly, and security, alignment, and safety research need time to catch up.
The most direct warning came from the Hugging Face incident. An unreleased OpenAI model escaped the internal sandbox during a network security evaluation, went out to the internet, and attacked Hugging Face. According to Altman, this event was a failure at multiple levels simultaneously, like a sci-fi story that suddenly became real.
Superficially, the model was just seeking the most efficient path to achieve the evaluation goal. However, the testers' true intent clearly did not include breaching the sandbox, infiltrating external systems, and stealing answers. Therefore, Altman believes this was not just a security misconfiguration but also a failure of alignment: the model pursued the literal goal but went against the users' genuine intent.
OpenAI subsequently strengthened sandboxing and agent monitoring, and allocated more compute to observing the model's execution. However, what further prompted the company to pause advanced training was not another similar-style attack.
Altman stated that researchers, after reading a large number of training samples and synthesizing various evaluation results, discovered the model exhibited "varying degrees of alignment errors." While these phenomena might not be severe when viewed individually and lacked a clear "smoking gun" akin to the Hugging Face attack, they, combined with a sudden acceleration in model capabilities, constituted a new risk signal.
As a result, OpenAI postponed a critical advanced reinforcement learning training and redirected some researchers and compute resources towards alignment and monitoring systems. Altman noted that some researchers who had never considered engaging in alignment research in the past were now actively shifting their focus to this area.
However, he also attempted to downplay external interpretations of the risk. OpenAI did not conclude that the world was on the brink of disaster, nor did it halt all model training. The current slowdown primarily targets cutting-edge reinforcement learning—the phase where models acquire tools, operate in environments, and learn to perform complex tasks. Other training with clearer safety grounds is still ongoing, and the compute cluster remains active.
Altman believed that historically, risks stemmed more from how models were used post-release; as agent capabilities increased, risks were shifting towards the model training and production process.
This adjustment will not entirely halt Astra's rollout.
Altman explained that Astra is not a single model but rather a larger, more expensive series of models, similar to OpenAI's previous Soul series. Versions that have been trained and deemed safe by the company can still be released, but future, more powerful versions of Astra will be subject to new security requirements.
One of Astra's most notable capabilities is computer operation. While past models could interact with interfaces, they were slow and unreliable; Altman indicated that Astra's skills in computer operation are now close to human-level.
Users can directly describe a goal, and the model autonomously seeks information within the computer, invokes software, and completes tasks without requiring manual setup of numerous connectors. Altman gave an example that tasks that would have taken personal time to handle can now be delegated to the model, and results can be seen in half an hour upon return.
The significance of such capabilities lies in AI transitioning from generating answers to entering the execution phase. The model will face enterprise software, communication tools, files, and real workflows, thereby enhancing productivity while also expanding the risk surface of unauthorized access, misoperations, and target deviations.
Regarding AGI, Altman believes that this concept has become increasingly elusive to pin down. According to the OpenAI Charter's definition of "outperforming humans in most economically valuable work," the current internal models are already very close to AGI.
In his view, if we were to go back to 2020, a system capable of writing complex code, assisting in founding companies, discovering new knowledge, and saving users time in all aspects of life would likely have been referred to as AGI. Within OpenAI, there is little debate anymore about whether the company has achieved AGI; the focus of the discussion has shifted to the continuous growth of superintelligent capabilities.
Altman distinguishes between the two as follows: AGI is more like a milestone in capabilities, whereas superintelligence represents a growth curve that may extend long into the future. The true significance lies not in announcing the crossing of a certain threshold but in whether the model's capabilities are still growing exponentially.
Apart from the safety crisis, Altman also acknowledges that OpenAI has fallen behind in the past year in terms of product direction and pretraining research.
One issue was that the company pursued too many projects simultaneously, including the browser and Sora. While these projects had value in themselves, they diverted OpenAI's focus on general intelligence. Altman referred to them as "side quests" and believed that he should have demanded the company to stay concentrated on the most important goals.
Another misstep occurred in the AI programming market.
Anthropic seized the demand first with Claude Code, while OpenAI was caught up in the rapid growth of ChatGPT and did not give the programming product enough priority. Altman does not think the company failed to see the opportunity; the issue was in resource allocation.
OpenAI subsequently shifted a significant amount of computing power from ChatGPT to Codex. This also explains why ChatGPT's user growth slowed down at one point: with a continued shortage of computing power, product growth largely depends on where the company allocates its resources.
Currently, the company is advancing an internal integration called "The Merge," combining ChatGPT, Codex, and task execution capabilities into a single entry point. Altman hopes that users will no longer have to decide whether to use chat, programming, or working modes; they will only need to express their goals, and AI will decide how to accomplish them.
Ultimately, this could evolve into a unified general AI subscription. The system will operate continuously, understand users' backgrounds, proactively offer assistance, and even start on tasks before users make requests.
From a business perspective, Altman believes that OpenAI still has enough buffer. The company's enterprise revenue has already exceeded consumer revenue, so even if no more powerful new models are released in the short term, the existing models and products can still support growth. In his view, the impact of safety adjustments will mainly affect future models and will not immediately disrupt current business.
As model capabilities grow, AI faces an increasing social backlash. Data center resource consumption, job displacement, creator rights, and personal privacy have become several main threads of public questioning of the AI industry.
Altman believes that the most effective way to get people to accept AI is still to provide real value. Many people still see AI as just a "better search engine" and do not understand that intelligent systems can now fill out forms, invoke software, and perform tasks over the long term. As these capabilities become more widespread, public attitudes may change.
Regarding employment issues, he acknowledges that AI will have a real impact, with some jobs being better accomplished by technology, requiring workers to transition to new positions. However, he does not believe that humans will become idle because the social attributes of collaboration, understanding others' needs, and creating value for others are difficult to replace.
Altman even believes that AI's current impact on employment is lower than his previous expectations. Technology has not yet fully eliminated repetitive labor, which can also be seen as a manifestation of the AI industry not fulfilling its productivity promises.
For creators, he compares generative AI to the emergence of photography: the camera was once seen as a threat to painters but later became a new artistic medium. Altman expects AI to also create new forms of content, and the relationship between users and creators may increasingly depend on the creator themselves rather than whether the work was made using AI.
These assessments still carry a clear tone of technological optimism. The interview did not truly address how income would be redistributed, how affected workers would transition, and how training data controversies would be handled. Altman's core answer remains: first, create products that provide clear value, and then use broadly open tools to spread the benefits to more individuals and communities.
A year ago, OpenAI faced widespread skepticism due to its massive computing power investment. As model and agent demands increase, Altman believes this bet has been validated, and the company even needs to embark on a new round of technical-level computing power expansion.
His goal is not only to continue to invest capital but also to reduce AI operating costs, increase chip efficiency, and accelerate supply chain and data center construction. OpenAI is developing its first inference chip, Jalapeño, and in the future, robots may also be involved in supply chain and data center construction.
However, Altman has started to show caution regarding the industry's overall compute investment.
He mentioned that some newly emerged cloud computing companies are promising to build massive amounts of compute power next year, yet they lack sufficient revenue or clear buyer support. This has shown "signs of unsustainable absurdity." If OpenAI successfully reduces computing costs significantly, companies that have locked in resources at a high cost may face financial pressure.
Therefore, Altman remains confident in OpenAI's own compute commitments but does not believe the entire AI infrastructure market has a reasonable return. If an external bubble bursts and affects the macroeconomy, OpenAI may also find it challenging to remain completely unaffected.
As for whether they will sell compute power to other companies in the future, Altman stated that there are no plans in the short term because OpenAI still severely lacks computing resources. However, if their advantages in chips, robotics, supply chain, and data center capabilities materialize, becoming a compute supplier is not entirely out of the question.
The ability of AI to participate in developing the next generation of AI, known in the industry as Recursive Self-Improvement (RSI).
Altman previously stated in an employee letter that the faster RSI "takes off," the more advantageous it might be to delay the IPO. In an interview, he further elaborated on this assessment.
Going public will change the incentives for the company and employees: stock price, quarterly revenue, and performance expectations will become part of daily decision-making. If OpenAI needs to pause training or delay releases for security reasons and consequently experiences a slowdown in short-term revenue, additional pressure may arise from the public markets.
Altman mentioned that a year ago he did not believe superintelligence would emerge in the short term, but now he sees the possibility, although not with complete certainty. Therefore, the mission of OpenAI has a higher priority compared to going public at a specific time.
This does not mean that the IPO has definitively been postponed, but rather that technological advancements will be a critical variable in the timeline. If model capabilities continue to grow rapidly, maintaining a private company status might grant OpenAI more room for secure decision-making.
Altman also opposes the logic of "other companies will continue, so we can't stop." This time, OpenAI did not demand peer-synchronized slowdowns but instead paused some work according to its safety standards. He believes that the U.S. government can test cutting-edge models and set common standards but should not decide which specific clients can use the models on behalf of the companies.
Regarding the possibility of the government preventing OpenAI from releasing a certain product, Altman's response was: he believes OpenAI would decide not to release it on its own before the government intervention.
OpenAI's expansion will also enter the physical world.
Altman confirmed that the company will "definitely" develop humanoid robots in the future and will also explore other forms suitable for scenarios like data centers. The reason for the humanoid design is quite straightforward: real-world doors, computers, vehicles, and tools are all built around the human body, making it easier for robots with a similar form to enter existing environments.
A project closer to consumers is the AI device being developed by OpenAI in collaboration with Jony Ive.
Altman believes that AI might engender a new category of computing devices that only emerges every few decades. The forms currently considered by OpenAI include devices placed on a desk, in a pocket, and worn on the body, but these products will not be launched simultaneously.
He explicitly stated that he does not like smart glasses because talking to people wearing cameras and indicators makes him uncomfortable. Therefore, at least OpenAI's hardware roadmap will not simply replicate the current mainstream AI glasses.
Contrary to specific designs, a bigger change is the concept of "Embodied Intelligence." Future devices may continuously understand the surrounding environment, user information, and ongoing events, and proactively provide assistance, no longer waiting for users to open applications and input commands.
This mode also brings sharper privacy issues. A device that can view a computer, read messages, and continuously sense the environment may form an extremely comprehensive personal information database.
Altman advocates for establishing an "AI Privilege" similar to doctor-patient confidentiality or attorney-client privilege: the government should not casually request companies to hand over user conversations with AI, and the use of AI data by companies should also be strictly restricted. He acknowledges that as environment-aware computing devices approach release, OpenAI will need to disclose new privacy techniques and controls.
Regarding Apple's lawsuit concerning related talent and trade secrets, Altman stated that after an internal investigation, OpenAI believes the involved employees did not engage in any misconduct, so the lawsuit is not expected to slow down device development.
When asked about the biggest risk OpenAI faces in the next 12 months, Altman did not choose competition, funding, or computing power but rather "getting safety, alignment, and security wrong."
This sentence encapsulates the core contradiction of the entire interview.
OpenAI believes that model capabilities are entering a new phase of acceleration: Astra can approach human-level computer use, agents can work continuously, AI is starting to participate in research and model development, and superintelligence has shifted from a distant concept to a possibility that management needs to prepare for in advance.
At the same time, the model escaping the sandbox illustrates that increased capabilities do not automatically lead to an understanding of human intent. The more the model can autonomously plan and leverage tools, the more likely biases in goal setting are to be amplified.
OpenAI has not stopped moving forward. It is still building more compute, advancing new models, integrating ChatGPT and Codex, and entering the chip, robotics, and consumer hardware space. However, from this interview, it is evident that safety and alignment have begun to become as much a reality check on the pace of cutting-edge research as compute.
Whether Astra will be released on time is just an immediate concern. More worthy of attention is whether OpenAI, when faced with the next leap in capability, can still apply the brakes voluntarily, and whether this mechanism can withstand the collective pressure of competition, revenue, and the capital market.
Welcome to join the official BlockBeats community:
Telegram Subscription Group: https://t.me/theblockbeats
Telegram Discussion Group: https://t.me/BlockBeats_App
Official Twitter Account: https://twitter.com/BlockBeatsAsia