importantSYS.SOURCE: Terminal-Bench-Science• 2026-08-28T00:06:51Z
Evaluating AI Agents on Scientific Research Workflows via Terminal-Bench-Science
Terminal-Bench-Science 0.1 introduces a benchmark evaluating AI agents on real scientific research workflows across multiple disciplines, with Claude Opus 5 achieving the highest resolution rate of 30%. The project emphasizes collaboration with scientists to create evolving, expert-curated tasks that reflect actual research challenges.
*** END OF TRANSMISSION ***