llama3.2:1b on A6000, A100, H100 — 2026-04-17

2026-04-17T21:53:43.000Z → 2026-04-17T21:54:19.177Z

On 2026-04-17, llama3.2:1b ran across A6000, A100, H100. H100 finished first at 590.0 tok/s, 22× faster than A6000. The per-million-token cost: $1.17 (H100) vs $3.60 (A6000) — the surprise: the headline GPU H100 was also the cheapest per-million tokens, coming in 3.1× less than A6000.

Podium

H100

1st

thunder-h100

peak tok/s: 590.0
avg tok/s: 590.0
$ / 1M tok: $1.17

A100

2nd

thunder-a100

peak tok/s: 44.2
avg tok/s: 44.2
$ / 1M tok: $4.90

A6000

3rd

thunder-a6000

peak tok/s: 27.0
avg tok/s: 27.0
$ / 1M tok: $3.60