Hackernews posts about VLLM
- Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025) (www.aleksagordic.com)
- Kimi K3 on vLLM: Up to 370 Tokens/sec (vllm.ai)
- vLLM Recipes (recipes.vllm.ai)
- C++ Version of vLLM (github.com)
- Inferact vLLM creators are hiring (twitter.com)
- vLLM Serving Experiments on H100s – config beats the baseline on p95 TTFT,ITL (efficientagent.substack.com)
- Show HN: ExANS – Lossless KV cache compression at 622 GB/s on H100 (www.theopenlake.com)
- Show HN: Gainz.fast – Local Inference, Faster (gainz.fast)
- The Inference Engine Guide for K3 Deployment (twitter.com)
- Fluent – Voice computer use agent for Windows (VLM and accessibility tree) (fluentforall.com)
- The Best Way to Make Your (Small) Vision Language Model Smarter (loganbolton.github.io)