Cua says it has open-sourced two S1 Computer Use decision-making models and Cua-Bench-S1
Cua says it has open-sourced Nano and 4B, two S1 models for selecting the next action in Computer Use, and released the Cua-Bench-S1 evaluation set.
Loading…
Connect scattered updates into a story you can understand.
Cua says it has open-sourced Nano and 4B, two S1 models for selecting the next action in Computer Use, and released the Cua-Bench-S1 evaluation set.
Artificial Analysis reports that Claude Opus 5.5 max ranked first with 58 points; GPT-6 Sol and Luna performed similarly to the previous generation, but their per-task costs fell significantly.
OpenAI will expand third-party participation in safety testing during the early stages of model training, evaluation, and deployment, and is in discussions with organizations including METR and Redwood Research.
According to the FT, OpenAI recommended that the US promote unified international assessments of AI models’ capabilities and risks, while establishing cross-border security communication and threat-sharing mechanisms.