Dynamic Beating AI News: Scale AI has launched RSI Bench, specifically designed to test whether AI can autonomously conduct AI research like a researcher.
The AI will directly access papers, code, and computational resources, then conduct its own experiments, training runs, method modifications, and assess whether it can improve upon the original AI.
The initial testing involved Claude Opus 5 and GPT-5.6 Sol. Both models were able to conduct experiments, tune parameters, iterate on solutions autonomously, and some tasks did surpass the given baselines.
However, current AIs are not yet proficient at "inventing." When challenged with the latest research, most attempts still rely on existing methods and parameter tuning, with few proposing truly novel ideas.

