ONNX-MLIR to TFLite is an experimental compiler pipeline that converts static ONNX models directly to TensorFlow Lite FlatBuffers through MLIR. It is built as a downstream extension of ONNX-MLIR and reuses its ONNX importer, ONNX Dialect, shape inference, and canonicalization infrastructure.
The project does not construct a TensorFlow graph, Keras model, or SavedModel. Unsupported configurations fail conversion instead of falling back to Select TF, Flex, or custom operators.
The default compilation path is:
ONNX protobuf
|
| onnx-mlir --EmitONNXIR
v
ONNX Dialect MLIR
|
| onnx-mlir-opt --convert-onnx-to-tfl
v
TFL Dialect MLIR (unoptimized)
|
| TensorFlow/LiteRT litert-opt
| optional; enabled by default
v
TFL Dialect MLIR (optimized)
|
| TensorFlow flatbuffer_translate
v
TFLite FlatBuffer
Artifacts and intermediate IR appear as nodes; the executable responsible for
each transformation appears on the connecting arrow. With
--no-optimize-tfl, the litert-opt stage is bypassed and the unoptimized TFL
Dialect MLIR is passed directly to flatbuffer_translate.
The onnx-to-tflite driver orchestrates these stages, verifies intermediate
results, and publishes the output only after FlatBuffer validation succeeds.
The ONNX-to-TFL stage is implemented with MLIR Dialect Conversion and requires the source ONNX operations to be fully legalized. TFL optimization remains in the TFL Dialect and applies TensorFlow/LiteRT canonicalization, fusion, and cleanup passes before FlatBuffer export.
ONNX-MLIR and TensorFlow are pinned to different LLVM/MLIR revisions. To avoid mixing incompatible MLIR C++ ABIs in one process, the pipeline uses textual TFL MLIR as the checked interface between the ONNX-MLIR tools and TensorFlow tools. TensorFlow reparses and verifies that IR using its authoritative TFL Dialect implementation.
- Static, ranked tensor compilation with FP32 as the primary activation type and selected integer and boolean tensor support.
- Layout-aware lowering: the stable default uses NHWC for rank-4 activations in TFLite, while other graph ranks retain their ONNX axis order.
- Compile-time conversion of convolution filters and broadcast parameters to the layouts expected by TFLite.
- Direct lowering to builtin TFLite operations without TensorFlow graph conversion or runtime fallback operators.
- Optional TensorFlow/LiteRT optimization, enabled by default and disabled with
--no-optimize-tflfor debugging or A/B comparison. - Automatic buffer-offset export for large models using common adjacent ONNX external-data layouts.
- Per-pass verification, FlatBuffer identifier checks, TensorFlow round-trip parsing, and atomic output publication.
See the operator support matrix for the exact supported types, ranks, attributes, and opset revisions. Only configurations listed there are claimed to be supported.
Prebuilt Linux x86_64 runtime packages are available from GitHub Releases. This is the recommended way to try the compiler without building LLVM/MLIR, ONNX-MLIR, and TensorFlow/LiteRT from source.
Download both the .tar.gz archive and its accompanying .sha256 file from
the latest release, then verify and extract the package:
sha256sum -c onnx-mlir-to-tflite-runtime-*.tar.gz.sha256
mkdir onnx-to-tflite-runtime
tar -xzf onnx-mlir-to-tflite-runtime-*.tar.gz \
--strip-components=1 \
-C onnx-to-tflite-runtime
cd onnx-to-tflite-runtimeThe packaged driver is located at ./bin/onnx-to-tflite. See the
model conversion guide for usage and command-line
options.
The prebuilt package is intended for Linux x86_64 and requires the glibc and
libstdc++ ABI versions listed in its README.md and BUILD_INFO.txt. Ubuntu
24.04 or newer is recommended. Build from source when the package is not
compatible with the target system.
The complete pipeline requires building both the pinned LLVM/MLIR toolchain
and two TensorFlow/LiteRT tools: litert-opt for optional TFL optimization and
flatbuffer_translate for required FlatBuffer export. TensorFlow is invoked
as a separate toolchain and is not linked into the ONNX-MLIR executables.
For a complete bootstrap build from an activated Python environment:
python -m pip install -r requirements.txt onnxruntime tensorflow
PYTHON_BIN="$(command -v python)" \
BUILD_JOBS=8 \
./scripts/bootstrap_and_build.shThe script clones pinned LLVM and TensorFlow revisions inside this repository, builds all required tools, and runs an end-to-end smoke test. Generated source and build trees are ignored by Git. See the build guide for prerequisites, repository layout, build products, configuration options, and instructions for reusing existing TensorFlow tools.
./build/bin/onnx-to-tflite \
/path/to/model.onnx \
-o /path/to/model.tfliteSee Convert an ONNX model for runtime-package usage, layout selection, diagnostics, large-model export, and the complete option reference.
ONNX commonly represents rank-4 activations as NCHW, while TFLite uses NHWC.
The default legacy mode exposes rank-4 TFLite graph inputs and outputs as
NHWC and keeps convolution and pooling regions in NHWC. It does not insert
runtime transposes solely to preserve an NCHW external ABI. Callers must
therefore adapt rank-4 boundary tensors when comparing or integrating the
generated model.
Rank-1, rank-2, rank-3, and rank-5-or-higher graph boundaries retain their ONNX axis order. Local layout conversions required by individual operations, including Conv3D fallback, are represented explicitly inside the graph.
See the model conversion guide for layout-mode selection and the layout policy for convolution filter layouts, broadcast rules, rank-sensitive behavior, and Conv3D reduction details.
For models whose constants would exceed ordinary FlatBuffer metadata limits,
the driver can use TensorFlow's buffer-offset format. Common adjacent external
data files such as model.onnx.data are detected automatically once the model
size reaches the configured threshold. The resulting .tflite remains a
single self-contained file, but its consumer must support TFLite buffer
offsets.
Use --use-buffer-offset to enable this mode explicitly for other external
data layouts.
Successful conversion includes all of the following checks:
- MLIR verification after each conversion stage by default.
- Rejection of unlegalized ONNX operations.
- TFLite
TFL3file-identifier validation. - FlatBuffer import back into MLIR through TensorFlow's parser.
- Atomic publication only after validation succeeds.
Run the ONNX-to-TFL MLIR regression suite with:
llvm-project/build/bin/llvm-lit \
-sv build/test/mlir/conversion/onnx_to_tflThe end-to-end similarity utility handles the rank-4 boundary layout contract, compares ONNX Runtime and TFLite results, and reports cosine similarity, Euclidean distance, relative Euclidean distance, RMSE, and maximum absolute error:
python test/e2e/run_similarity.py \
--onnx test/models/mlp.onnx \
--tflite build/test/models/mlp.tfliteThe project focuses on statically shaped inference graphs. Dynamic shapes, general control flow, quantization, sparse tensors, and arbitrary custom or Flex operators are outside the current scope. Some complex operations are supported only for constrained static configurations or through ONNX-MLIR importer decomposition.
Review the known limitations before relying on a configuration that is not explicitly covered by the operator support matrix.
| Document | Description |
|---|---|
| Model conversion | Driver usage, layout selection, diagnostics, and command-line options |
| Build guide | Complete LLVM/MLIR, ONNX-MLIR, and TensorFlow/LiteRT build procedure |
| Compiler optimizations | ONNX preprocessing, layout propagation, lowering-time rank reduction, and TFL structural rewrites |
| Operator support | Supported ONNX operations, opsets, types, ranks, attributes, and test coverage |
| Architecture | Cross-version MLIR boundary, conversion design, optimization, and export details |
| Layout policy | Rank-sensitive layout contract and convolution lowering rules |
| Known limitations | Unsupported configurations and implementation constraints |
The original ONNX-MLIR documentation remains available under docs/ and at
onnx.ai/onnx-mlir.
This repository retains the history and compiler infrastructure of upstream ONNX-MLIR and adds the ONNX-to-TFL conversion and TFLite export pipeline.
The upstream base is ONNX-MLIR commit
7cea64f8aa2fbc56f859df768211a95dc2c72ad6.
All project-specific compiler, tooling, test, and documentation changes are
maintained independently from that revision.
See the repository LICENSE for licensing information.
This is an independent project and is not an official ONNX or ONNX-MLIR release.