llama3.1:8b on A6000, A100, H100 — 2026-04-18

2026-04-18T02:28:26.000Z → 2026-04-18T02:29:20.701Z

On 2026-04-18, llama3.1:8b ran across A6000, A100, H100. H100 finished first at 193.8 tok/s, 24× faster than A6000. The per-million-token cost: $3.57 (H100) vs $12.00 (A6000) — the surprise: the headline GPU H100 was also the cheapest per-million tokens, coming in 3.4× less than A6000.

Podium

H100

1st

thunder-h100

peak tok/s: 193.8
avg tok/s: 193.8
$ / 1M tok: $3.57

A100

2nd

thunder-a100

peak tok/s: 14.9
avg tok/s: 14.9
$ / 1M tok: $14.54

A6000

3rd

thunder-a6000

peak tok/s: 8.1
avg tok/s: 8.1
$ / 1M tok: $12.00