
Prepare the model
Start with the desired LLM
Export the model with the qai_hub_models Python package
qai_hub_models Python package to export the model.This process:- Downloads the model weights.
- Uploads them to AI Hub for compilation.
- Generates QNN binaries split into multiple parts for NPU execution.
- Creates a deployable folder (
genie_bundle) with all required assets (context binaries, configs, tokenizer).
- AI Hub supports quantization (typically 4-bit internally, though weights may be stored as 8-bit for compatibility).
- Export scripts handle splitting large models into prompt processors and token generator components.
Deploy the model
Install the Qualcomm AI Runtime (QAIRT) SDK on the target device
Copy the compiled binaries and configuration files to the device
Use Genie CLI tools or the Genie dialog API for inference
genie-t2t-run) or Genie dialog API for inference.Ensure the target device meets the requirements
- Hexagon architecture: v73 or newer
- Required RAM:
- 16 GB for 7B models
- ~12 GB for 3B models
Run the model on-device using Genie APIs
- Genie manages multiple binaries and execution orders for optimal NPU utilization.
Important notes
-
AI Hub advantages
- automatically handles model compilation, quantization, and splitting.
- Provides pre-optimized models and bring your own model (BYOM) support.
-
Genie
- Simplifies inference by abstracting complex execution steps.
- Offers APIs for text-to-text and dialogue-based interactions.
-
Customization
- Export flow defaults to 4-bit quantization for runtime efficiency.
- No direct option to store weights as 4-bit; they remain 8-bit but load as 4-bit during execution.

