Why is it much slower and much more memory hungry than slightly larger models such as Qwen3.5-4B?

#31
by slightlyoutofphase - opened

Definitely not what I was expecting.

Nanbeige LLM Lab org

Thank you for raising this. The Nanbeige4 series (Nanbeige4, Nanbeige4.1, Nanbeige4.2, and the upcoming Nanbeige4.5) focuses on improving model quality under a fixed parameter budget. In Nanbeige4.2, the Looped Transformer and disabled KV sharing improve model performance but result in slower decoding.

We are preparing a DFlash version of Nanbeige4.2 for faster inference. Nanbeige5 will use linear attention to further improve inference efficiency.

What's the ETA for Nanbeige 4.5?

Waiting for this DFlash . It's too slow for me right now. Just a half of gemma-4-e4b on my side.

Best small model for small coding i've tested. However on my dear macbook air m4 16gb it's slooooow. 3-10t/s. I am sure there is some metal acceleration missing even tho llama.cpp server output looks fine. On my rtx3060 12gb pc it does 45-55 t/s.

Sign up or log in to comment