PICCHIO Star on GitHub ↗

Picchio caught 28 Qwen3.8-27B layers on the CPU and a 14.3× decode slowdown.

RTX 5090 · 32 GB66/66 GPU · 81.5 tok/s
RTX 4070 SUPER · 12 GB38/66 GPU · 5.7 tok/s

Same Q4_K_M GGUF · llama.cpp · ctx 4096 · 1 request · decode

python picchio.pyz

Paste your result

Parsed locally in your browser. Zero data collected.

Real runs

ModelMachineEnginePlacementPrefillDecodeWallJ/tokVerdict
Loading results…