❯ Open-source agent OpenScience claims to outscore Codex on a science benchmark, thanks to workflow design
53 of 70 tasksStartup Synthetic Sciences states on its GitHub page that its open-source research agent OpenScience solved 53 of 70 tasks on the Terminal-Bench-Science benchmark, scoring 75.7%. According to posts on social media, the top entry on a September 23 leaderboard mirror, Codex with GPT-6 Astra, scored 68.1%. The result is self-reported; full run traces are public, and no third-party replication has appeared yet.
From literature to write-upOpenScience is free under the Apache 2.0 license, runs as a desktop app, in a browser or from the command line, and works with models from different providers. Users describe a research task in plain language, and it searches literature, forms hypotheses, writes and runs experiment code, analyzes data and writes up results, with every step visible. Terminal-Bench-Science, hosted by Stanford University and the Laude Institute, uses tasks written by experts in the life, physical and earth sciences to test agents on real research workflows.
Same model, different workflowOpenScience itself runs on GPT-6 Astra and Sol, so its lead of more than seven points over Codex comes from tools and steps designed around research work. If third parties reproduce the result, general-purpose coding agents from OpenAI and Anthropic will face direct competition from vertical tools like this in specialist fields.
▪ SIGNALWhen the underlying model is the same, the win comes down to workflow design, an uncomfortable reminder for companies selling only general-purpose assistants.