header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

OpenScience self-tests at 75.7%, surpassing Codex: Research Agents begin running experiments in their own loop.

Beating AI News Flash: Y Combinator Winter 2026 batch (YC W26) company Synthetic Sciences has officially released the open-source research Agent OpenScience. It can search papers, process data, write code, and call scientific research databases. The newly added Autoresearch also allows AI to run experiments continuously on its own.


For example, to train a model, researchers only need to tell it to "bring down the validation set loss." OpenScience will first run a baseline, then propose modifications on its own, and execute experiments round after round. If the results improve, it keeps them; if they worsen, it rolls back, and then decides what to try next based on previous results. Users can also set limits in advance on the number of runs, total duration, and stopping conditions, or specify "only modify the optimizer next."


OpenScience automates the experimental iteration that originally required researchers to repeatedly watch over: proposing plans, executing experiments, comparing metrics, eliminating failed plans, and then continuing to the next round. Experiments can run locally or be handed off to remote GPUs, SSH servers, and Slurm/PBS clusters. The code, results, and judgments of each round of experiments are recorded for later review.


It also supports direct login with a ChatGPT/Codex subscription, and can also connect to your own API Key, local models, or use the official pay-as-you-go Ace.


In official self-testing, OpenScience completed 53 of 70 Terminal-Bench Science tasks, scoring 75.7%. By comparison, the public score for Codex + GPT-6 Astra is 68.1%.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish