Hackernews posts about MoE
- Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 (www.gilesthomas.com)
- AntGroup releases Ling 3.0 Flash MoE model (twitter.com)
- Ornith 1.5 35B-A3B: SOTA 35B MoE model (huggingface.co)
- DeepSeek V4 Pro goes GA: 1.6T MoE flagship exits preview (ainexusdaily.vercel.app)
- Mixture-of-Kittens: An MoE training megakernel for NVL72 (twitter.com)
- Instella-Moe: An Open Mixture-of-Experts Language Model (rocm.blogs.amd.com)
- Imprint – Fine-tune MoE LLMs bigger than your RAM (github.com)
- MoE routing is just branch prediction (ssenthilnathan3.github.io)
- Train-infer mismatch for Open-weight MoE RL in Open-source code (kiddyboots216.github.io)
- Mixture of Experts (Moe): How Transformers Scale Without Activating Everything (chizkidd.github.io)
- The Voyage 4 model family: shared embedding space with MoE architecture (blog.voyageai.com)
- Generate LoRA adapters from Skill.MDs *MoE-models now too (terradev.cloud)
- Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 (www.gilesthomas.com)
- Compute-optimal is not cluster-optimal (szha.ai)
- Log is non-monotonic in PHP and Lua (purplesyringa.moe)
- Beyond "Clean Code": Why Your Comments Matter (blog.moertel.com)