Hackernews posts about Evals
- Astra and Fable still hack on simple variants of alignment evals from 2025 (www.lesswrong.com)
- Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x (frontierharness.org)
- My LLM eval cried wolf. Here's what I measured (digline.dev)
- Show HN: Coder Eval – A Framework for Evals (github.com)
- I turned my game into an eval (crux.lakin.dev)
- Deep dive on evals for agents [video] (www.youtube.com)
- Evals Aren't Dead, but They're Low Res (hackbot.dad)
- Astra and Fable still hack on simple variants of alignment evals from 2025 (www.lesswrong.com)
- The Hugging Face Controversy, Evals for Product Teams-Food 4 Agile Thought 560 (age-of-product.com)
- Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety... (storage.googleapis.com)
- The index was lying and the eval knew (gil-neto.com)
- Run SDK: secure eval for your agents (vercel.com)
- Secure-Eval-Worker: Least-Privilege JavaScript Execution in Node.js (blog.platformatic.dev)
- Grug-Brained Evals (2025) (softwaredoug.com)
- How to Build Effective Evals for AI Agents (www.kdnuggets.com)
- Ollama 0.33.3 changed what prompt_eval_duration measures (lognebudo.github.io)
- Bad evals, my own: five exercises from two LLM judges (digline.dev)