Hackernews posts about Ben
- GLM 5.2 beats Claude in our benchmarks (semgrep.dev)
- Kimi K3, and what we can still learn from the pelican benchmark (simonwillison.net)
- Lost city discovered beneath Egypt's desert with ancient church (www.dailymail.com)
- Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers (senior-swe-bench.snorkel.ai)
- What Emily Bender meant by "stochastic parrots" (spectrum.ieee.org)
- Benchmarking coding agents on Databricks' multi-million line codebase (www.databricks.com)
- Benchmarking 15 “E-Waste” GPUs with Modern Workloads (esologic.com)
- Ben Bernanke Joins Anthropic Oversight Trust (www.anthropic.com)
- An interactive explorer for Benford's Law across real datasets (vatsalbakshi.com)
- Claude Sonnet 5 – benchmark results (artificialanalysis.ai)
- Discord banned 8k users for posting benign grid images (www.theverge.com)
- Ben Thompson is wrong: US frontier labs are right to be panicking (larrysalibra.com)
- The Richest Country Is Pretty Mid Now [Benn Jordan][video] (www.youtube.com)
- The Seed Beneath the Snow (eli.li)
- FrontierFinance: The largest open benchmark for investor workflows (research.samaya.ai)
- Show HN: Benchmark your eng team's AI agent maturity in 5 minutes (agent-benchmarks.com)
- Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 (www.gilesthomas.com)
- Giving a domain a hill to climb: benchmarking as data activation (sparsethought.com)