the bench

where the gates land · rows marked predicted or computed are roofline-derived · cpu-interpret and simulated rows are verified off-chip · hardware rows replace both

Every number, with provenance.

Nothing quantitative appears on this site outside an instrument, and every value states what it is. Each record carries chip, dtype, shapes, date, and the lab that produced it; gate verdicts on the chapter pages resolve to rows here.

records
10

The record

id kernel metric value baseline chip dtype shapes date
P-001 matmul NxNxN predicted floor (roofline) 697.7 µs chip constants v5e (predicted) bf16 4096³ 2026-07-26
P-002 matmul skinny 8xNxN predicted floor (roofline) 41.1 µs chip constants v5e (predicted) bf16 8x4096x4096 2026-07-26
P-003 elementwise add predicted floor (roofline) 122.8 µs chip constants v5e (predicted) bf16 4096² 2026-07-26
P-004 softmax rows predicted floor (roofline) 81.8 µs chip constants v5e (predicted) bf16 4096² 2026-07-26
P-005 naive attention spill computed HBM round-trip 268 MB score matrix size any (computed) bf16 seq 8192 2026-07-26
C-001 flash forward (lab 3.2) max err vs reference ≤ 1e-4 jax.nn.softmax @ v cpu interpret f32 512x1024x64 2026-07-26
C-002 attention custom_vjp grads (lab 3.3) max err vs jax.grad ≤ 1e-5 autodiff of reference cpu interpret f32 128x256x64 2026-07-26
C-003 causal flash, blocks skipped (lab 3.4) max err vs masked reference ≤ 1e-4 dense masked attention cpu interpret f32 512x512x64 2026-07-26
C-004 ring all-gather (lab 4.1) match vs lax.all_gather exact collective 8 simulated devices f32 8 shards x 4x128 2026-07-26
C-005 ring attention (lab 4.2) max err vs full attention ≤ 1e-5 single-device reference 8 simulated devices f32 128x256x64 sharded 2026-07-26
compare

Your chip against the record

Ran a lab? Paste the results blob its final cell printed. Parsed in your browser; nothing you paste is sent anywhere.

waiting for a results blob ·

Parsed in your browser. Nothing you paste here is sent anywhere.