Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

55 points | by matt_d 4 hours ago

13 comments