On a Motorola Edge 70, llama.cpp's NPU backend produced nonsense. The chip's matrix unit has no FP16, and the software assumed it did. Instead of waiting for a fix, we asked what the unit can do: integer math. Three days later the NPU is the fastest way to read a prompt on that phone, with the same answers as the CPU.
I worked with an AI coding agent. I set the direction and the checkpoints; the agent ran the probes, wrote the kernels and measured. Every claim had to hold on the real phone before it counted. The dead ends stay in the picture: they are where most of the learning was.
Motorola Edge 70, Snapdragon 7 Gen 4. The goal: local models that are actually fast.
The speed was fake: the matrix unit wrote zeros and infinities. Attention and the Qwen3.5 layers failed their tests too. Filed upstream as #29473.
"That could be our winning case."
Same Hexagon generation (v73), different result. The difference had to be in the chip, not the code.
Qualcomm's runtime on this chip: fp16=false. Five of 21 Snapdragons have no FP16 in the matrix unit. llama.cpp picks FP16 by generation, so the two v73 parts among them are exposed.
"If we teach the matrix unit to count in integers, that's the breakthrough."
And a condition before investing: show what is actually new.
Then the same inputs on the phone: 12 of 12 output hashes match. An emulator is a hypothesis until hardware agrees.
Weights to int8, activations as two 8-bit planes, the exact 32-bit sum stitched back on the vector unit. 118 t/s, already 2.3× the CPU.
Correct first, then fast, then clean enough for upstream. Publish code and numbers; keep the SDK lab private (its licence).
Matrix unit computes tile i while the vector unit finishes i−1 and prepares i+1. 118 → 126 → 147 t/s.
Checking 8 draft tokens costs the NPU more than the CPU. Measured, dropped.
All unit tests passed. Only the full quality run caught it: a decay product underflowed to 0/0 on fast-forgetting heads. Moved to log space.
3.3× the CPU, 2.3× the GPU, byte-identical greedy answers. The chip is detected at start-up; no settings needed.
The matrix unit got faster, the vector unit became the wall. Quality cost +0.3–1.6%. Not worth it.
Both share one memory bus. Together: +15% over the best alone. Generation is limited by memory, not compute.
8192 rows of 12 bytes each, one memory wait per row. Now: one transfer in, one vector gather, one out. Decode 8.31 → 8.99 t/s.
Side result: a CPU kernel for PrismML's ternary models, 1.9× faster prompts on the same phone, opened as PR #290.
Current version v1.4. Motorola Edge 70, 12 GB. llama-bench, 512-token prompt. CPU is the best of 2/4/6/8 threads; GPU is the stock OpenCL backend.
| Model | NPU | GPU | CPU | vs CPU |
|---|---|---|---|---|
| Qwen3.5-4B Q4_0 | 171.9 | 73.9 | 51.7 | 3.33× |
| Llama-3.2-1B Q4_0 | 550.6 | 307.2 | 226.8 | 2.43× |
| Qwen3.5-4B Q4_0, mixed quants | 105.9 | 76.2 | 47.1 | 2.25× |
| Quality, wikitext-2, 20 × 512 | Perplexity | Time |
|---|---|---|
| CPU | 10.8533 | 4:02 |
| NPU | 10.8355 | 1:08 |
The matrix unit (HMX) on this chip multiplies 8-bit integers. Everything else is arranged so the answer comes back exact.
Re-quantized per column against the largest block scale in each K chunk. That scale rides in the unit's output conversion.
Fed through the uh:2x1 mode; a second bias word cancels the unsigned offset exactly.
Two 16-bit stores per chunk, joined on the vector unit (HVX) into the exact 32-bit sum, then rescaled.
HMX computes tile i while HVX reduces i−1 and converts weights for i+1. Attention runs the same way.
Qualcomm AI Hub, QNN 2.50, one probe model per chipset. On the two v73 parts without FP16, stock llama.cpp takes the FP16 path anyway. 7 Gen 4 is confirmed on a real phone; QCM6690 is untested.