BlockBeats calls Anthropic's Claude Code official claude-api Skill's addition of an automated tuning workflow
BIBIBI
AT A GLANCE
The report says that Anthropic added an automated tuning workflow to Claude Code's claude-api Skill, allowing changes to be tested and degraded configurations to be rolled back.
Article
Claude Code Self-directed AI application tuning: modify Prompt, switch models, and roll back when performance worsens
Beating AI breaking news: Anthropic has added Claude Code's official claude-api Skill an automated tuning workflow.build-eval It first builds a test set and evaluation method from real conversations, Bugs, or human-curated cases,hillclimb then has Claude Code repeatedly modify the system Prompt, tool descriptions, models, and reasoning intensity. After each change, it reruns the tests; improvements are kept, while regressions are rolled back.
To prevent Claude from optimizing only for the test questions, it does not tune against all samples at once. Some cases are used to identify problems and modify the configuration, while another batch is reserved to check whether the changes work on unseen problems. If the score on the first set rises but the held-out test results do not improve, the change may be deemed overfitting and reverted.
Anthropic demonstrated the workflow using 44 customer support tickets.Claude Code first cleaned up Prompt, added business rules, tried different models, and reduced reasoning intensity. It ultimately found a Sonnet 5 low-effort configuration. On 14 customer support tickets that were not used for tuning, accuracy rose from 78.6%under the initial configuration to 90.5%, while costs fell to approximately one-fifth of the original.
Original link https://m.theblockbeats.info/flash/369545
Key points
01
The report claims that Anthropic added an automated tuning workflow to Claude Code's official claude-api Skill.
02
build-eval It creates a test set and evaluation method from real conversations, Bugs, or human-curated cases.
03
hillclimb It repeatedly modifies the system Prompt, tool descriptions, models, and reasoning intensity, and retests after each change.
04
The report describes the rule as follows: keep changes when performance improves and roll them back when it worsens.
05
Some cases are used to identify problems and modify the configuration, while another batch checks performance on unseen problems.
06
The report says that Anthropic used 44 customer support tickets to demonstrate the process.
07
The report claims that it ultimately found a Sonnet 5 low-effort configuration and, on 14 customer support tickets not used for tuning, raised accuracy from 78.6%to 90.5%, while reducing costs to approximately one-fifth of the original.
AI-assisted interpretation
The following is analysis, separate from reported facts. Verify important claims independently.
This workflow is essentially like having Claude Code automatically experiment with different AI application configurations: first prepare the questions and evaluation criteria, then modify the prompts, tool descriptions, models, or reasoning intensity before taking the test again. If the score improves, keep the new configuration; if it worsens, restore the old configuration. It also sets aside some cases from tuning to check whether the new configuration has merely memorized the test questions.
Why it matters to readers
If the report is accurate, this workflow can turn prompt engineering, model selection, and cost adjustment into a repeatably tested process, using rollbacks and held-out cases to reduce the risk of configurations deteriorating with each change or adapting only to specific test questions.
Beginners can think of it as an automated parameter-tuning tool for AI applications, but representative cases and a clear evaluation method must still be prepared first; the material does not explain how ordinary users can obtain, install, or reproduce the workflow.
Risks and unknowns
The material is a secondhand account from a news channel and provides no Anthropic official announcement or Skill source text for verification.
44 The contents of the customer support tickets, evaluation criteria, and baseline configuration were not disclosed.
How costs are calculated, and whether “approximately one-fifth of the original” generalizes to other tasks, cannot be confirmed from the material.
Related Developments
Loading event timeline…
The material does not specify the workflow’s scope of availability, release date, installation method, or applicable limitations.
Sonnet 5 The specific parameters and complete reproduction results for the low-effort configuration were not provided.
Related concepts
System Prompt
This term is not in the glossary yet. Browse related concepts in the glossary.