Article

Claude Code Self-directed AI application tuning: modify Prompt, switch models, and roll back when performance worsens

Beating AI breaking news: Anthropic has added Claude Code's official claude-api Skill an automated tuning workflow.build-eval It first builds a test set and evaluation method from real conversations, Bugs, or human-curated cases,hillclimb then has Claude Code repeatedly modify the system Prompt, tool descriptions, models, and reasoning intensity. After each change, it reruns the tests; improvements are kept, while regressions are rolled back.

To prevent Claude from optimizing only for the test questions, it does not tune against all samples at once. Some cases are used to identify problems and modify the configuration, while another batch is reserved to check whether the changes work on unseen problems. If the score on the first set rises but the held-out test results do not improve, the change may be deemed overfitting and reverted.

Anthropic demonstrated the workflow using 44 customer support tickets.Claude Code first cleaned up Prompt, added business rules, tried different models, and reduced reasoning intensity. It ultimately found a Sonnet 5 low-effort configuration. On 14 customer support tickets that were not used for tuning, accuracy rose from 78.6%under the initial configuration to 90.5%, while costs fell to approximately one-fifth of the original.

Original link https://m.theblockbeats.info/flash/369545