Hackernews posts about Llama 405B
- Running Llama 3.1 405B (a11ce.com)
- We fine-tuned Llama 405B on AMD GPUs (publish.obsidian.md)
- Llama 405B 506 tokens/second on an H200 (developer.nvidia.com)
- Show HN: Fast and Cheap Llama-405B (centml.ai)
- Llama 405B up to 142 tok/s on Nvidia H200 SXM (old.reddit.com)
- (hot take) Llama 405B is big enough for AGI (twitter.com)
- The Future of AI: Synthetic Data Gen with Llama 3.1 405B and Raft (techcommunity.microsoft.com)
- Cerebras achieves 2,500T/s on Llama 4 Maverick (400B) (www.cerebras.ai)