ResearchSimon Willison

Claude Fable 5.1 made me a really nice animated pelican

#ai#benchmark#animation#pelican#coding

English

Claude Fable 5.1 has shown significant improvements in coding and problem-solving tasks, achieving a 52.6% score on the new Terminal-Bench-Science benchmark. The author tested its capabilities by generating an animated pelican, highlighting varying reasoning levels and the resulting output quality.

中文

Claude Fable 5.1 在编码和问题解决任务中显示出显著的改进,在新的 Terminal-Bench-Science 基准测试中获得了 52.6% 的分数。作者通过生成动画鹈鹕测试其能力,突出了不同推理水平和生成输出的质量。