Case · On-device AI · llama.cpp · Snapdragon 7 Gen 4

The phone's NPU returned garbage. Now it reads prompts 3.3× faster than the CPU.

On a Motorola Edge 70, llama.cpp's NPU backend produced nonsense. The chip's matrix unit has no FP16, and the software assumed it did. Instead of waiting for a fix, we asked what the unit can do: integer math. Three days later the NPU is the fastest way to read a prompt on that phone, with the same answers as the CPU.

0×
prompt speed vs the best CPU setting (Qwen3.5-4B, 512 tokens)
0×
vs the phone's GPU (Adreno 722, OpenCL)
10.84
perplexity on the NPU, 10.85 on the CPU; greedy answers byte-identical
0
Snapdragon chips checked for FP16; 5 lack it
How we got here

The reasoning, step by step.

I worked with an AI coding agent. I set the direction and the checkpoints; the agent ran the probes, wrote the kernels and measured. Every claim had to hold on the real phone before it counted. The dead ends stay in the picture: they are where most of the learning was.

Me · direction
Agent · probes, code, numbers
26.09 Question

The phone has an NPU. Can llama.cpp use it?

Motorola Edge 70, Snapdragon 7 Gen 4. The goal: local models that are actually fast.

26.09 Test

437 tokens/s on the prompt. The answer is gibberish.

The speed was fake: the matrix unit wrote zeros and infinities. Attention and the Qwen3.5 layers failed their tests too. Filed upstream as #29473.

26.09 Hypothesis

A hunch: maybe nobody has solved this yet.

"That could be our winning case."

26.09 Evidence

The maintainer: same models run clean on IQ-9075.

Same Hexagon generation (v73), different result. The difference had to be in the chip, not the code.

27.09 Test Evidence

A probe model on 21 chipsets through Qualcomm AI Hub.

Qualcomm's runtime on this chip: fp16=false. Five of 21 Snapdragons have no FP16 in the matrix unit. llama.cpp picks FP16 by generation, so the two v73 parts among them are exposed.

27.09 Decision

Turn the error into a resource.

"If we teach the matrix unit to count in integers, that's the breakthrough."

And a condition before investing: show what is actually new.

27.09 Test

Map the undocumented integer mode on the SDK emulator.

Then the same inputs on the phone: 12 of 12 output hashes match. An emulator is a hypothesis until hardware agrees.

27.09 Build

First integer matmul. First correct answer from Qwen on this NPU.

Weights to int8, activations as two 8-bit planes, the exact 32-bit sum stitched back on the vector unit. 118 t/s, already 2.3× the CPU.

27.09 Decision

Take it end to end, not just a bug report.

Correct first, then fast, then clean enough for upstream. Publish code and numbers; keep the SDK lab private (its licence).

27.09 Measure

Pipeline the units; attention goes integer too.

Matrix unit computes tile i while the vector unit finishes i−1 and prepares i+1. 118 → 126 → 147 t/s.

27.09 Dead end

Speculative decoding with the NPU as verifier

Checking 8 draft tokens costs the NPU more than the CPU. Measured, dropped.

27.09 Build Bug

Qwen3.5's recurrent layers, eight tokens per pass. Perplexity: 69.5.

All unit tests passed. Only the full quality run caught it: a decay product underflowed to 0/0 on fast-forgetting heads. Moved to log space.

27.09 Measure

172 tokens/s. Perplexity 10.84 vs 10.85 on the CPU.

3.3× the CPU, 2.3× the GPU, byte-identical greedy answers. The chip is detected at start-up; no settings needed.

27.09 Dead end

8-bit activations for 2× matrix speed

The matrix unit got faster, the vector unit became the wall. Quality cost +0.3–1.6%. Not worth it.

28.09 Dead end

Decode on CPU and NPU at once

Both share one memory bus. Together: +15% over the best alone. Generation is limited by memory, not compute.

28.09 Build

Two slow copies found in decode. Fixed.

8192 rows of 12 bytes each, one memory wait per row. Now: one transfer in, one vector gather, one out. Decode 8.31 → 8.99 t/s.

28.09 Ship

Public repo, upstream discussion, PR prepared.

Side result: a CPU kernel for PrismML's ternary models, 1.9× faster prompts on the same phone, opened as PR #290.

Results

Faster than the CPU and the GPU, on the same phone.

Current version v1.4. Motorola Edge 70, 12 GB. llama-bench, 512-token prompt. CPU is the best of 2/4/6/8 threads; GPU is the stock OpenCL backend.

Prompt processing, tokens per second

Higher is better
NPU, integer HMXGPUCPU
Table view
ModelNPUGPUCPUvs CPU
Qwen3.5-4B Q4_0171.973.951.73.33×
Llama-3.2-1B Q4_0550.6307.2226.82.43×
Qwen3.5-4B Q4_0, mixed quants105.976.247.12.25×

Qwen3.5-4B prompt speed by version

Tokens per second; the dashed line is the CPU. v1.4 kept the prompt speed and made generation 8% faster (8.31 → 8.99 t/s).
Quality, wikitext-2, 20 × 512PerplexityTime
CPU10.85334:02
NPU10.83551:08
How it works

Integer math, exact sums.

The matrix unit (HMX) on this chip multiplies 8-bit integers. Everything else is arranged so the answer comes back exact.

01 · weights

q4_0 → int8

Re-quantized per column against the largest block scale in each K chunk. That scale rides in the unit's output conversion.

02 · activations

16-bit as two 8-bit planes

Fed through the uh:2x1 mode; a second bias word cancels the unsigned offset exactly.

03 · accumulate

Coarse + fine, stitched

Two 16-bit stores per chunk, joined on the vector unit (HVX) into the exact 32-bit sum, then rescaled.

04 · pipeline

Both units busy

HMX computes tile i while HVX reduces i−1 and converts weights for i+1. Attention runs the same way.

Who needs this

Five of 21 Snapdragons have no FP16 matrix math.

Qualcomm AI Hub, QNN 2.50, one probe model per chipset. On the two v73 parts without FP16, stock llama.cpp takes the FP16 path anyway. 7 Gen 4 is confirmed on a real phone; QCM6690 is untested.

no FP16: needs the integer pathFP16 available
Working with an AI agent

Who did what.

Me

  • Picked the goal and the bet: a phone NPU as a real LLM engine
  • Reframed the bug as the opportunity: integers, not a workaround
  • Asked for the novelty case before investing time
  • Set the bar: correct on hardware, same quality as the CPU
  • Chose what to publish, and what to keep private for licence reasons

Agent (Claude Code)

  • Probed the chip: AI Hub runs, emulator, hash checks on the phone
  • Wrote the integer kernels, the pipeline and the auto-detection
  • Measured every version; kept the logs and the dead ends
  • Drafted the write-ups; I edit and publish them

Honest limits

  • The matrix unit speeds up prompts. Generation is memory-bound; the CPU is still slightly faster (Qwen3.5-4B: 9.7 vs 9.0 tokens/s).
  • Only q4_0 weights run on the matrix unit, so mixed-quant models gain less (2.25×).
  • Tested on one phone. Not upstream yet: discussion in #29473.