Smart
News
news
interests
topics
domains
about
library
Topics
/
Llama 3 405B
Hackernews posts about Llama 3 405B
RSS feed for Llama 3 405B
01
new
Llama
405B
506 tokens/second on an H200
developer.nvidia.com
▲ 21
moondistance
722d
💬 5
save
saved
02
new
Llama
3
.1
405B
now runs at 969 tokens/s on Cerebras Inference
cerebras.ai
▲ 427
benchmarkist
686d
💬 156
save
saved
03
new
HuggingFace - Tencent launches Hunyuan Large which outperforms
Llama
3
.1
405B
huggingface.co
▲ 23
janik-io
700d
💬 1
save
saved
04
new
Benchmarking
Llama
3
.1
405B
on 8x AMD MI300X GPUs
dstack.ai
▲ 11
latchkey
727d
💬 3
save
saved
05
new
Running
Llama
3
.1
405B
a11ce.com
▲ 4
a11ce
39d
discuss
save
saved
06
new
How to Run Meta
Llama
3
.1
405B
with Nebius AI Studio API
nebius.com
▲ 1
dsaed
708d
discuss
save
saved
07
new
Show HN: How to guide on training
Llama
-
405B
using PyTorch distributed APIs
github.com
▲ 3
lambda-research
721d
💬 4
save
saved
08
new
Ask HN: AI Cloud Computing, is it cheaper than the OpenAI API in the end?
▲ 2
calipsow
712d
💬 4
save
saved
09
new
Vakgpt: Open-Source Chat Wrapper for Rapid AI Prototyping
▲ 1
krishna-vakx
496d
discuss
save
saved
10
new
AllenAI Tulu
3
405B
available for chat and download
▲ 12
soundworlds
609d
discuss
save
saved
11
new
Show HN: Slash your LLM Inference Costs with Overnight Processing
▲ 5
Blue_Cosma
727d
💬 1
save
saved
12
new
Show HN: Most Efficient Batch API for Open-Source and Custom Models
withexxa.com
▲ 1
Blue_Cosma
718d
discuss
save
saved
13
new
Show HN: OS Megakernel that match M5 Max Tok/w at 2x the Throughput on RTX 3090
github.com
▲ 6
GreenGames
181d
💬 1
save
saved
14
new
Show HN: ZSE – Open-source LLM inference engine with
3
.9s cold starts
github.com
▲ 58
zyoralabs
222d
💬 9
save
saved
·
edit
Keyboard shortcuts
j / k
next / previous story
o or Enter
open the story
c
open the comments
s
save for later
h
history & saved stories
?
show this help
Got it