10.2 TFLM Runtime & Model Embedding
This phase covers two things: integrating the TensorFlow Lite Micro (TFLM) runtime into the Q2390 / IQ2390 MCU firmware, and embedding the pre-built int8 model (produced in Phase 1) directly into the firmware image. It includes the non-obvious C++ runtime setup required on SDLLVM/RISC-V.10.2.1 Overview
The integration has five distinct parts:- Fetch TFLM source — TFLM source is not vendored; it must be fetched manually
- Inspect model & verify op support — extract op types from the model and confirm TFLM supports them
- Configure LLVM libc++ — SDLLVM doesn’t configure libc++ by default; 5 specific fixes are required
- Create the inference module — build the module with embedded model C arrays, inference harness (
tflite_infer.cc), Kconfig header, and CMakeLists.txt - Integrate into firmware build — call
tflite_infer_run()frommain.cand register the module in CMake
10.2.2. Part 1 — Fetch TFLite Micro Source
Path convention: Throughout this phase,The TFLM module stubs (<wasp_proc>refers to the root directory of your MCU firmware workspace — the directory containingzephyr/,modules/,config/, etc. Replace it with your actual path wherever it appears (e.g./path/to/your_workspace/wasp_proc).
CMakeLists.txt, Kconfig) exist at wasp_proc/zephyr/zephyr/modules/tflite-micro/ but the actual C++ source tree is absent — it must be cloned manually.
The tflite-micro project is defined in the west manifest (zephyr/zephyr/submanifests/optional.yaml) in the optional group, which Zephyr’s own west.yml explicitly disables with group-filter: [-optional]. This means west update tflite-micro will not work — use git clone instead.
10.2.2.1. Clone the source
Inside the Docker build container:10.2.2.2. Verify
10.2.3. Part 2 — Inspect Model & Verify Op Support
Before writing any firmware code, we need two things from the model:- Which ops the model uses — the TFLM runtime requires every operator used by the model to be explicitly registered before inference can run. If any operator is missing, the firmware will fail at startup with an unresolved op error. This step identifies the exact set of operators your model needs, so you know what to register when integrating the inference harness in Part 5.
- Input/output quantization parameters (scale + zero_point) — if the model is int8-quantized, float inputs must be pre-quantized to int8 before passing to the inference harness.
10.2.3.1. Option A — flatc (full topology decode — recommended)
Decodes the model binary into human-readable JSON using the TFLM schema. Gives you: op types, tensor names, shapes, and quantization params for all tensors.
Step 1 — Run flatc to decode the model:
cnn1d_minimal_int8.json in the current directory.
Step 2 — Extract the op list from the JSON:
10.2.3.2. Option B — tf.lite.Interpreter (quant params only, no flatc needed)
Use when flatc is not available. Gives you input/output quantization params (scale, zero_point). Does not expose op types — use Option A for the op list.
tf.lite.Interpreterdoes not expose op types — use Option A (flatc) to get the op list.
10.2.3.3. All ops supported by this TFLM build
The authoritative source for which ops TFLM supports ismicro_mutable_op_resolver.h — every Add*() method declared in that file is a supported op for the exact revision downloaded.
zephyr_20240627 — 114 ops:
The resolver method for each op isAdd+ op name — e.g.Conv2D→AddConv2D(),FullyConnected→AddFullyConnected().
10.2.3.4. Reference — ops confirmed for the example model
All ops incnn1d_minimal_int8.tflite, their TFLM kernel files, and the resolver method:
If an op is not registered, you will see one of these errors:
Compile-time (wrong or missing method name):
AllocateTensors() (op not registered):
10.2.4. Part 3 — Configure LLVM libc++ (5-Step Fix)
10.2.4.1. Why the default approach fails
To use TFLM (written in C++), Zephyr needs a C++ standard library. The natural starting point is enablingCONFIG_LIBCXX_LIBCPP=y. On SDLLVM, however, this triggers a Kconfig dependency chain that breaks the build. The failure appears as an error in a completely unrelated kernel file:
10.2.4.2. The 4-step fix
10.2.4.2.1. Step 1 — Remove select REQUIRES_FULL_LIBCPP from TFLM Kconfig
File: zephyr/zephyr/modules/tflite-micro/Kconfig
Remove the select REQUIRES_FULL_LIBCPP line from inside config TENSORFLOW_LITE_MICRO. This is the root of the dependency cascade — removing it prevents Zephyr from triggering the broken REQUIRES_FULL_LIBC path on SDLLVM.
10.2.4.2.2. Step 2 — Use EXTERNAL_MODULE_LIBCPP, not EXTERNAL_LIBCPP
No action here — this setting is included in the new file created in section 2.4.4.
10.2.4.2.3. Step 3 — Inject libc++ includes via -I, not -isystem
No action here — this call is part of the complete block added in section 2.4.3.
-I directories are searched before any -isystem directory. By injecting libc++ headers as -I BEFORE, clang finds libc++‘s own <stdio.h>/<string.h> wrappers first, which chain to musl via #include_next — the correct resolution order for C++ translation units.
10.2.4.2.4. Step 4 — Compile C++ TUs with -DNDEBUG
No action here — this flag is part of the complete block added in section 2.4.3.
musl’s assert() macro expands to __assert_fail(), which picolibc does not provide. -DNDEBUG compiles all assert calls out in C++ translation units. Scoped to $<$<COMPILE_LANGUAGE:CXX>:-DNDEBUG> so C code is untouched.
10.2.4.3. Complete CMake configuration block
File:modules/hal/qcom/core/config/CMakeLists.txt (pre-existing file — add this complete block at the end). Gated on CONFIG_EXTERNAL_MODULE_LIBCPP so it only activates when the TFLM Kconfig fragment is merged in.
10.2.4.4. shikra_lpaicp_tflite.conf
10.2.4.4.1. Step 1 — Create shikra_lpaicp_tflite.conf
Create a new file at zephyr/kernel/config/shikra_lpaicp_tflite.conf:
10.2.4.4.2. Step 2 — Register the file in config.yml
File: zephyr/kernel/config/config.yml (pre-existing file — add a new entry):
Why a separate file instead of adding toshikra_lpaicp.conf? The existing entry uses regex(^|shikra_)lpaicp_.*, which matches both the hardware target (SHIKRA_LPAICP_TEST) and the QEMU simulation variant that uses Zephyr-SDK GCC. Adding TFLM’s SDLLVM-specific libc++ configuration toshikra_lpaicp.confwould break the GCC-based QEMU build. The narrower regexshikra_lpaicp_.*matches only the hardware target.
10.2.5. Part 4 — Create the Inference Module
10.2.5.1. Step 1 — Generate src/model_data.cpp and src/input_data.cpp
Create the module directories first, then convert the model and input files produced in Phase 1 into C byte arrays. Run the commands from the directory containing the Phase 1 source files.<wasp_proc> with the path to your wasp_proc MCU code base.
The generated files look like this:
alignas(8)on the model array: TFLM’s FlatBuffers parser requires the model byte array to be at least 4-byte aligned.alignas(8)guarantees this regardless of where the linker places the symbol.
10.2.5.2. Step 2 — Create inc/tflite_infer.h
File: modules/hal/qcom/core/tflite_infer/inc/tflite_infer.h (new file).
10.2.5.2.1. Tuning TFLI_ARENA_KB
TFLI_ARENA_KB sets the size of the tensor arena — the contiguous block of memory TFLM uses for input/output tensors, intermediate activation buffers between layers, and its internal allocator metadata. It is allocated from the system heap via k_malloc at the start of tflite_infer_run().
What goes wrong if the value is wrong:
Both failures cause
tflite_infer_run() to return early — g_tfli_batches remains 0 and all latency globals stay at their initial values.
How to find the right value — measure with g_tfli_arena_used:
- Set
TFLI_ARENA_KBto a generous starting value (e.g.32) — large enough to guaranteeAllocateTensors()succeeds. - Build, flash, boot, and run
read_tflite_infer.cmmin T32. - Confirm
g_tfli_status == 0andg_tfli_batches >= 1. - Read
g_tfli_arena_usedfrom the T32 Var.View window — this is the exact byte count TFLM required. - Compute the minimum safe value and update the header:
- Rebuild and reflash. Verify
g_tfli_arena_usedis still belowg_tfli_arena_size.
Configured value for the reference model:RAM impact ofTFLI_ARENA_KB = 4(4,096 bytes). This is the value used in the reference codebase forcnn1d_minimal_int8.tflite. Measureg_tfli_arena_usedon-target after a successful run to confirm utilisation for your model.
TFLI_ARENA_KB:
The arena is runtime-allocated from the system heap — it does not contribute directly to static image size. However, the heap backing buffer (kheap_buf__system_heap) is a statically linked section sized by CONFIG_HEAP_MEM_POOL_SIZE (set to 32768 in shikra_lpaicp.conf). Reducing TFLI_ARENA_KB alone does not reduce static RAM. To reclaim static RAM, also lower CONFIG_HEAP_MEM_POOL_SIZE to match — ensure the new size covers TFLI_ARENA_KB * 1024 + 16 plus any other k_malloc callers.
10.2.5.3. Step 3 — Create CMakeLists.txt
File: modules/hal/qcom/core/tflite_infer/CMakeLists.txt (new file). Gated on CONFIG_TENSORFLOW_LITE_MICRO.
10.2.5.4. Step 4 — Create src/tflite_infer.cc
File: modules/hal/qcom/core/tflite_infer/src/tflite_infer.cc (new file). Runs TFLI_ITERS timed Invoke() calls per batch (after one warm-up), then publishes per-batch min/avg/max latency in microseconds, QTMR ticks, and CPU cycles into the g_tfli_* volatile globals.
Note on__attribute__((retain)):g_tfli_arena_sizeandg_tfli_itersare only written at initialisation and never read by firmware code. Without__attribute__((retain)), the linker’s--gc-sectionspass silently drops their ELF sections even though they arevolatile—volatileprevents the compiler from optimising them away but does not protect against linker GC. Theretainattribute marks the section as always-live, ensuring the symbols remain visible to T32.
10.2.6. Part 5 — Integrate into Firmware Build
10.2.6.1. Step 1 — Integrate tflite_infer_run() into main.c
File: modules/hal/qcom/core/main/src/main.c
Placement:tflite_infer_run()goes after theCONFIG_QC_CLK_READYblock. If that block is not present, placing it directly after the initialLOG_INF/printklines is equivalent.
10.2.6.2. Step 2 — Register the module in CMake
File:modules/hal/qcom/core/config/CMakeLists.txt (pre-existing file — add one line at the end of the add_subdirectory_ifdef block):
add_subdirectory_ifdef line:
10.2.7. RAM Note
The Shikra LPAICP firmware has approximately 40 KB of free RAM before TFLM is added. TFLM’s static code footprint (~25 KB for the 3-op subset), model data (~2 KB), and heap allocation can bring the image close to the 3 MB limit. If the build fails with:log.conf in zephyr/kernel/config/config.yml:
log.conf enables Zephyr’s deferred logging subsystem: a 16 KB ring buffer and a dedicated background thread with a 4 KB stack, consuming ~20 KB of static RAM. Since this firmware uses the T32 RAM console for output (no UART), the deferred logging thread has no delivery path. Commenting it out reclaims ~20 KB while CONFIG_LOG_MODE_MINIMAL=y from shikra_lpaicp.conf remains active.
10.2.8. Kconfig Flag Reference
All Kconfig symbols added or modified for TFLM integration:
Inference harness parameters (
TFLI_ARENA_KB, TFLI_ITERS, TFLI_BATCH_SLEEP_MS) are #define constants in inc/tflite_infer.h. Edit that file directly to tune without recompiling any other file. See section 2.5.2.1 for the arena sizing procedure.
