Picchio caught 28 Qwen3.8-27B layers on the CPU and a 14.3× decode slowdown.
RTX 5090 · 32 GB66/66 GPU · 81.5 tok/s
RTX 4070 SUPER · 12 GB38/66 GPU · 5.7 tok/s
Same Q4_K_M GGUF · llama.cpp · ctx 4096 · 1 request · decode
python picchio.pyzPaste your result
Parsed locally in your browser. Zero data collected.
Real runs
| Model | Machine | Engine | Placement | Prefill | Decode | Wall | J/tok | Verdict |
|---|---|---|---|---|---|---|---|---|
| Loading results… | ||||||||