Skip to main content

10.2 TFLM Runtime & Model Embedding

This phase covers two things: integrating the TensorFlow Lite Micro (TFLM) runtime into the Q2390 / IQ2390 MCU firmware, and embedding the pre-built int8 model (produced in Phase 1) directly into the firmware image. It includes the non-obvious C++ runtime setup required on SDLLVM/RISC-V.

10.2.1 Overview

The integration has five distinct parts:
  1. Fetch TFLM source — TFLM source is not vendored; it must be fetched manually
  2. Inspect model & verify op support — extract op types from the model and confirm TFLM supports them
  3. Configure LLVM libc++ — SDLLVM doesn’t configure libc++ by default; 5 specific fixes are required
  4. Create the inference module — build the module with embedded model C arrays, inference harness (tflite_infer.cc), Kconfig header, and CMakeLists.txt
  5. Integrate into firmware build — call tflite_infer_run() from main.c and register the module in CMake

10.2.2. Part 1 — Fetch TFLite Micro Source

Path convention: Throughout this phase, <wasp_proc> refers to the root directory of your MCU firmware workspace — the directory containing zephyr/, modules/, config/, etc. Replace it with your actual path wherever it appears (e.g. /path/to/your_workspace/wasp_proc).
The TFLM module stubs (CMakeLists.txt, Kconfig) exist at wasp_proc/zephyr/zephyr/modules/tflite-micro/ but the actual C++ source tree is absent — it must be cloned manually. The tflite-micro project is defined in the west manifest (zephyr/zephyr/submanifests/optional.yaml) in the optional group, which Zephyr’s own west.yml explicitly disables with group-filter: [-optional]. This means west update tflite-micro will not work — use git clone instead.

10.2.2.1. Clone the source

Inside the Docker build container:
The revision is read directly from the manifest so this works for any LPAICP release without modification. The resulting detached HEAD state is expected and correct.

10.2.2.2. Verify


10.2.3. Part 2 — Inspect Model & Verify Op Support

Before writing any firmware code, we need two things from the model:
  1. Which ops the model uses — the TFLM runtime requires every operator used by the model to be explicitly registered before inference can run. If any operator is missing, the firmware will fail at startup with an unresolved op error. This step identifies the exact set of operators your model needs, so you know what to register when integrating the inference harness in Part 5.
  2. Input/output quantization parameters (scale + zero_point) — if the model is int8-quantized, float inputs must be pre-quantized to int8 before passing to the inference harness.
Decodes the model binary into human-readable JSON using the TFLM schema. Gives you: op types, tensor names, shapes, and quantization params for all tensors. Step 1 — Run flatc to decode the model:
This produces cnn1d_minimal_int8.json in the current directory. Step 2 — Extract the op list from the JSON:
Step 3 — Cross-check each op against the supported op list in 1.3.3.

10.2.3.2. Option B — tf.lite.Interpreter (quant params only, no flatc needed)

Use when flatc is not available. Gives you input/output quantization params (scale, zero_point). Does not expose op types — use Option A for the op list.
tf.lite.Interpreter does not expose op types — use Option A (flatc) to get the op list.

10.2.3.3. All ops supported by this TFLM build

The authoritative source for which ops TFLM supports is micro_mutable_op_resolver.h — every Add*() method declared in that file is a supported op for the exact revision downloaded.
Complete supported op list for release tag zephyr_20240627 — 114 ops:
The resolver method for each op is Add + op name — e.g. Conv2D → AddConv2D(), FullyConnected → AddFullyConnected().

10.2.3.4. Reference — ops confirmed for the example model

All ops in cnn1d_minimal_int8.tflite, their TFLM kernel files, and the resolver method: If an op is not registered, you will see one of these errors: Compile-time (wrong or missing method name):
Runtime at AllocateTensors() (op not registered):

10.2.4. Part 3 — Configure LLVM libc++ (5-Step Fix)

10.2.4.1. Why the default approach fails

To use TFLM (written in C++), Zephyr needs a C++ standard library. The natural starting point is enabling CONFIG_LIBCXX_LIBCPP=y. On SDLLVM, however, this triggers a Kconfig dependency chain that breaks the build. The failure appears as an error in a completely unrelated kernel file:

10.2.4.2. The 4-step fix

10.2.4.2.1. Step 1 — Remove select REQUIRES_FULL_LIBCPP from TFLM Kconfig

File: zephyr/zephyr/modules/tflite-micro/Kconfig Remove the select REQUIRES_FULL_LIBCPP line from inside config TENSORFLOW_LITE_MICRO. This is the root of the dependency cascade — removing it prevents Zephyr from triggering the broken REQUIRES_FULL_LIBC path on SDLLVM.

10.2.4.2.2. Step 2 — Use EXTERNAL_MODULE_LIBCPP, not EXTERNAL_LIBCPP

No action here — this setting is included in the new file created in section 2.4.4.

10.2.4.2.3. Step 3 — Inject libc++ includes via -I, not -isystem

No action here — this call is part of the complete block added in section 2.4.3. -I directories are searched before any -isystem directory. By injecting libc++ headers as -I BEFORE, clang finds libc++‘s own <stdio.h>/<string.h> wrappers first, which chain to musl via #include_next — the correct resolution order for C++ translation units.

10.2.4.2.4. Step 4 — Compile C++ TUs with -DNDEBUG

No action here — this flag is part of the complete block added in section 2.4.3. musl’s assert() macro expands to __assert_fail(), which picolibc does not provide. -DNDEBUG compiles all assert calls out in C++ translation units. Scoped to $<$<COMPILE_LANGUAGE:CXX>:-DNDEBUG> so C code is untouched.

10.2.4.3. Complete CMake configuration block

File: modules/hal/qcom/core/config/CMakeLists.txt (pre-existing file — add this complete block at the end). Gated on CONFIG_EXTERNAL_MODULE_LIBCPP so it only activates when the TFLM Kconfig fragment is merged in.

10.2.4.4. shikra_lpaicp_tflite.conf

10.2.4.4.1. Step 1 — Create shikra_lpaicp_tflite.conf

Create a new file at zephyr/kernel/config/shikra_lpaicp_tflite.conf:

10.2.4.4.2. Step 2 — Register the file in config.yml

File: zephyr/kernel/config/config.yml (pre-existing file — add a new entry):
Why a separate file instead of adding to shikra_lpaicp.conf? The existing entry uses regex (^|shikra_)lpaicp_.*, which matches both the hardware target (SHIKRA_LPAICP_TEST) and the QEMU simulation variant that uses Zephyr-SDK GCC. Adding TFLM’s SDLLVM-specific libc++ configuration to shikra_lpaicp.conf would break the GCC-based QEMU build. The narrower regex shikra_lpaicp_.* matches only the hardware target.

10.2.5. Part 4 — Create the Inference Module

10.2.5.1. Step 1 — Generate src/model_data.cpp and src/input_data.cpp

Create the module directories first, then convert the model and input files produced in Phase 1 into C byte arrays. Run the commands from the directory containing the Phase 1 source files.
Replace <wasp_proc> with the path to your wasp_proc MCU code base. The generated files look like this:
alignas(8) on the model array: TFLM’s FlatBuffers parser requires the model byte array to be at least 4-byte aligned. alignas(8) guarantees this regardless of where the linker places the symbol.

10.2.5.2. Step 2 — Create inc/tflite_infer.h

File: modules/hal/qcom/core/tflite_infer/inc/tflite_infer.h (new file).

10.2.5.2.1. Tuning TFLI_ARENA_KB

TFLI_ARENA_KB sets the size of the tensor arena — the contiguous block of memory TFLM uses for input/output tensors, intermediate activation buffers between layers, and its internal allocator metadata. It is allocated from the system heap via k_malloc at the start of tflite_infer_run(). What goes wrong if the value is wrong: Both failures cause tflite_infer_run() to return early — g_tfli_batches remains 0 and all latency globals stay at their initial values. How to find the right value — measure with g_tfli_arena_used:
  1. Set TFLI_ARENA_KB to a generous starting value (e.g. 32) — large enough to guarantee AllocateTensors() succeeds.
  2. Build, flash, boot, and run read_tflite_infer.cmm in T32.
  3. Confirm g_tfli_status == 0 and g_tfli_batches >= 1.
  4. Read g_tfli_arena_used from the T32 Var.View window — this is the exact byte count TFLM required.
  5. Compute the minimum safe value and update the header:
  1. Rebuild and reflash. Verify g_tfli_arena_used is still below g_tfli_arena_size.
Configured value for the reference model: TFLI_ARENA_KB = 4 (4,096 bytes). This is the value used in the reference codebase for cnn1d_minimal_int8.tflite. Measure g_tfli_arena_used on-target after a successful run to confirm utilisation for your model.
RAM impact of TFLI_ARENA_KB: The arena is runtime-allocated from the system heap — it does not contribute directly to static image size. However, the heap backing buffer (kheap_buf__system_heap) is a statically linked section sized by CONFIG_HEAP_MEM_POOL_SIZE (set to 32768 in shikra_lpaicp.conf). Reducing TFLI_ARENA_KB alone does not reduce static RAM. To reclaim static RAM, also lower CONFIG_HEAP_MEM_POOL_SIZE to match — ensure the new size covers TFLI_ARENA_KB * 1024 + 16 plus any other k_malloc callers.

10.2.5.3. Step 3 — Create CMakeLists.txt

File: modules/hal/qcom/core/tflite_infer/CMakeLists.txt (new file). Gated on CONFIG_TENSORFLOW_LITE_MICRO.

10.2.5.4. Step 4 — Create src/tflite_infer.cc

File: modules/hal/qcom/core/tflite_infer/src/tflite_infer.cc (new file). Runs TFLI_ITERS timed Invoke() calls per batch (after one warm-up), then publishes per-batch min/avg/max latency in microseconds, QTMR ticks, and CPU cycles into the g_tfli_* volatile globals.
Note on __attribute__((retain)): g_tfli_arena_size and g_tfli_iters are only written at initialisation and never read by firmware code. Without __attribute__((retain)), the linker’s --gc-sections pass silently drops their ELF sections even though they are volatile — volatile prevents the compiler from optimising them away but does not protect against linker GC. The retain attribute marks the section as always-live, ensuring the symbols remain visible to T32.

10.2.6. Part 5 — Integrate into Firmware Build

10.2.6.1. Step 1 — Integrate tflite_infer_run() into main.c

File: modules/hal/qcom/core/main/src/main.c
Placement: tflite_infer_run() goes after the CONFIG_QC_CLK_READY block. If that block is not present, placing it directly after the initial LOG_INF / printk lines is equivalent.

10.2.6.2. Step 2 — Register the module in CMake

File: modules/hal/qcom/core/config/CMakeLists.txt (pre-existing file — add one line at the end of the add_subdirectory_ifdef block):
Place it after the last existing add_subdirectory_ifdef line:

10.2.7. RAM Note

The Shikra LPAICP firmware has approximately 40 KB of free RAM before TFLM is added. TFLM’s static code footprint (~25 KB for the 3-op subset), model data (~2 KB), and heap allocation can bring the image close to the 3 MB limit. If the build fails with:
Comment out log.conf in zephyr/kernel/config/config.yml:
log.conf enables Zephyr’s deferred logging subsystem: a 16 KB ring buffer and a dedicated background thread with a 4 KB stack, consuming ~20 KB of static RAM. Since this firmware uses the T32 RAM console for output (no UART), the deferred logging thread has no delivery path. Commenting it out reclaims ~20 KB while CONFIG_LOG_MODE_MINIMAL=y from shikra_lpaicp.conf remains active.

10.2.8. Kconfig Flag Reference

All Kconfig symbols added or modified for TFLM integration: Inference harness parameters (TFLI_ARENA_KB, TFLI_ITERS, TFLI_BATCH_SLEEP_MS) are #define constants in inc/tflite_infer.h. Edit that file directly to tune without recompiling any other file. See section 2.5.2.1 for the arena sizing procedure.