Hackernews posts about Benchmarks
- GLM 5.2 beats Claude in our benchmarks (semgrep.dev)
- Kimi K3, and what we can still learn from the pelican benchmark (simonwillison.net)
- Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers (senior-swe-bench.snorkel.ai)
- Lies, Damn Lies and Database Benchmarks (questdb.com)
- Claude Sonnet 5 – benchmark results (artificialanalysis.ai)
- FrontierFinance: The largest open benchmark for investor workflows (research.samaya.ai)
- Show HN: Benchmark your eng team's AI agent maturity in 5 minutes (agent-benchmarks.com)
- Benchmarks Are Dead (For Us) (poetiq.ai)
- Benchmark object storage in objects/s, not GB/s (fractalbits.com)
- Perfectly Hitting the Wrong Target: The Story of an AI Code Review Benchmark (shrsv.hexmos.com)
- Ollama vs. Llama.cpp – Quick Benchmark (blawg.pages.dev)
- Every AI Memory Benchmark Has an Asterisk (tenureai.dev)
- DOJ Requires Egg Producers to End Coordinated Benchmark Manipulation (www.justice.gov)
- Political Neutrality Benchmark of popular AI models (neutralityproject.org)
- GPT-5.5-Cyber Tops Mythos 5 on Cybersecurity Benchmark (twitter.com)
- Show HN: Is grep enough? A transparent benchmark for agentic code navigation (entelligentsia.github.io)