Beating AI News Flash: Anthropic has added an automated tuning workflow to Claude Code's official claude-api Skill. build-eval first creates a test set and scoring method from real conversations, bugs, or manual cases, and then hillclimb has Claude Code repeatedly modify the system prompt, tool descriptions, model, and reasoning effort. After each change, the tests are rerun; if performance improves, the change is kept, and if it worsens, it is rolled back.
To prevent Claude from optimizing only for the test questions, it does not tune on all samples together. Some cases are used to identify problems and adjust configurations, while another batch is reserved to check whether the changes can generalize to unseen problems. If the earlier score rises but the held-out test results do not improve, the change may be judged as overfitting and withdrawn.
Anthropic demonstrated this process with 44 customer support tickets. Claude Code successively cleaned up the prompt, added business rules, tried different models, and reduced reasoning effort. In the end, it found a Sonnet 5 low-effort configuration. On 14 tickets that were not involved in tuning, accuracy rose from 78.6% with the initial configuration to 90.5%, and cost dropped to about one-fifth of the original.

