AttentionSpan
The inner loop is back_
A language model running entirely in your browser on hand-written WGSL kernels. Qwen2.5-0.5B, int4-quantised, no ML framework.
Watch the attention, live in 3D →Detecting WebGPU…
Loading weights…
WebGPU isn't available here
This demo runs the model on your GPU via WebGPU. Your browser or driver doesn't expose it. Try a recent Chrome/Edge on a desktop, or update your graphics driver. A recorded run will live here.
—tokens/sec
—ms to first token
0tokens
Hand-written WGSL
—tok/s
TTFT — ms
transformers.js · ONNX Runtime
—tok/s
TTFT — ms