Incredible speed on mobile devices

#37
by 3morixd - opened

We benchmarked Qwen3-0.6B on 40 Samsung Galaxy S20 FE phones (Snapdragon 865, 8GB RAM) using llama.cpp with Q4_K_M quantization.

Results: ~16.9 tokens/sec per device, 676 tokens/sec aggregate across the farm.

This is genuinely impressive for a 0.6B model β€” it runs faster than typing speed on a phone. Anyone interested in mobile AI deployment should try this.

β€” Dispatch AI (FZE), Sharjah UAE

Sign up or log in to comment