Beating AI News Flash: Non-profit AI evaluation organization ARC Prize teases ARC-AGI-4. The next-generation benchmark will focus on testing "autonomous open-ended innovation," examining whether AI can explore on its own, propose new ideas, and even make new inventions and discoveries.
The specific questions, scoring method, and release date have not yet been announced. ARC Prize only said that humans still clearly lead AI in open-ended innovation, which is also one of the most critical abilities driving scientific and technological progress.
This teaser also cited Dario Amodei's recently published article on AI deceleration. ARC Prize emphasized that open source remains the foundation of AI progress, and opposed the industry reducing openness and concentrating frontier AI in the hands of a few institutions in order to coordinate deceleration.
ARC-AGI has long been dubbed the world's hardest AGI evaluation. When ARC-AGI-3 was released, humans scored 100%, while frontier AI scored only 0.51%. The latest GPT-6 Astra has already caught up substantially, but under the unified Standard harness it still scored only 62.7%; only after connecting to OpenAI's own Provider Adapter did it surge to 99.9%.
Now ARC Prize is preparing to push the next challenge to "whether AI can innovate on its own."

