SHARD
End-to-end VLM inference on the Rockchip RK3588 NPU
Reverse-engineered undocumented NPU hardware limits (32KB SRAM, 4KB page) to deploy a vision-language model that failed with the standard vendor tools, using novel “Nano-Tiling” and “Graph Surgery” techniques.
- <2.0s latency, a 15x speedup over the CPU baseline.
- FP32-equivalent fidelity (>0.999 cosine similarity).
- Published as SHARD: A Compatibility Framework for Deploying Transformer Models on Edge NPUs, EuroMLSys’26.
Writeup: Reverse-Engineering the RK3588 NPU