Hackernews posts about Benchmarks
- My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw” (frogs.vaguespac.es)
- GLM-5.3 Artificial Analysis Benchmarks (artificialanalysis.ai)
- Goodhart's Law Comes for Every Benchmark You Trust (cacm.acm.org)
- Show HN: Artificial Analysis tool to create custom benchmarks for any use case (artificialanalysis.ai)
- Modern-Fs-Benchmark (github.com)
- Show HN: Do Codex skills save tokens? A six-run task-size benchmark (codex-howto-benchmark.nguyenvantamdk2.chatgpt.site)
- Agents on Rails: the first benchmark report (rubyonrails.org)
- Goodhart's Law Comes for Every Benchmark You Trust (cacm.acm.org)
- Musk assembled a full-stack AI coding play while everyone watched benchmarks (pub.towardsai.net)
- A Lightweight Open-Source LLM Benchmark Tool – Compare Any Model on OpenRouter (cheikhhseck.medium.com)
- PostgreSQL Autovacuum Internals and Benchmark (percona.community)
- DeepSeek just dropped official V4 Pro GA benchmarks (twitter.com)
- Handwritten-edit benchmark: Fable 5 is #1, Opus 4.8 regresses 55% on miscounting (dorrit.pairsys.ai)
- 1Password's new benchmark teaches AI agents how not to get scammed (1password.github.io)