Skip to content

About

A general-purpose ONNX → TFLite compiler built on MLIR, producing native TensorFlow Lite FlatBuffers.

Topics

Resources

Code of conduct

Contributing

Stars

11 stars

Watchers

0 watching

Forks

Latest commit

 

History

2,661 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ONNX-MLIR to TFLite compiler pipeline

ONNX-MLIR to TFLite

ONNX-MLIR to TFLite is an experimental compiler pipeline that converts static ONNX models directly to TensorFlow Lite FlatBuffers through MLIR. It is built as a downstream extension of ONNX-MLIR and reuses its ONNX importer, ONNX Dialect, shape inference, and canonicalization infrastructure.

The project does not construct a TensorFlow graph, Keras model, or SavedModel. Unsupported configurations fail conversion instead of falling back to Select TF, Flex, or custom operators.

Compiler pipeline

The default compilation path is:

ONNX protobuf
    |
    |  onnx-mlir --EmitONNXIR
    v
ONNX Dialect MLIR
    |
    |  onnx-mlir-opt --convert-onnx-to-tfl
    v
TFL Dialect MLIR (unoptimized)
    |
    |  TensorFlow/LiteRT litert-opt
    |  optional; enabled by default
    v
TFL Dialect MLIR (optimized)
    |
    |  TensorFlow flatbuffer_translate
    v
TFLite FlatBuffer

Artifacts and intermediate IR appear as nodes; the executable responsible for each transformation appears on the connecting arrow. With --no-optimize-tfl, the litert-opt stage is bypassed and the unoptimized TFL Dialect MLIR is passed directly to flatbuffer_translate.

The onnx-to-tflite driver orchestrates these stages, verifies intermediate results, and publishes the output only after FlatBuffer validation succeeds.

The ONNX-to-TFL stage is implemented with MLIR Dialect Conversion and requires the source ONNX operations to be fully legalized. TFL optimization remains in the TFL Dialect and applies TensorFlow/LiteRT canonicalization, fusion, and cleanup passes before FlatBuffer export.

ONNX-MLIR and TensorFlow are pinned to different LLVM/MLIR revisions. To avoid mixing incompatible MLIR C++ ABIs in one process, the pipeline uses textual TFL MLIR as the checked interface between the ONNX-MLIR tools and TensorFlow tools. TensorFlow reparses and verifies that IR using its authoritative TFL Dialect implementation.

Key properties

  • Static, ranked tensor compilation with FP32 as the primary activation type and selected integer and boolean tensor support.
  • Layout-aware lowering: the stable default uses NHWC for rank-4 activations in TFLite, while other graph ranks retain their ONNX axis order.
  • Compile-time conversion of convolution filters and broadcast parameters to the layouts expected by TFLite.
  • Direct lowering to builtin TFLite operations without TensorFlow graph conversion or runtime fallback operators.
  • Optional TensorFlow/LiteRT optimization, enabled by default and disabled with --no-optimize-tfl for debugging or A/B comparison.
  • Automatic buffer-offset export for large models using common adjacent ONNX external-data layouts.
  • Per-pass verification, FlatBuffer identifier checks, TensorFlow round-trip parsing, and atomic output publication.

See the operator support matrix for the exact supported types, ranks, attributes, and opset revisions. Only configurations listed there are claimed to be supported.

Quick Start

Prebuilt Linux x86_64 runtime packages are available from GitHub Releases. This is the recommended way to try the compiler without building LLVM/MLIR, ONNX-MLIR, and TensorFlow/LiteRT from source.

Download both the .tar.gz archive and its accompanying .sha256 file from the latest release, then verify and extract the package:

sha256sum -c onnx-mlir-to-tflite-runtime-*.tar.gz.sha256
mkdir onnx-to-tflite-runtime
tar -xzf onnx-mlir-to-tflite-runtime-*.tar.gz \
  --strip-components=1 \
  -C onnx-to-tflite-runtime
cd onnx-to-tflite-runtime

The packaged driver is located at ./bin/onnx-to-tflite. See the model conversion guide for usage and command-line options.

The prebuilt package is intended for Linux x86_64 and requires the glibc and libstdc++ ABI versions listed in its README.md and BUILD_INFO.txt. Ubuntu 24.04 or newer is recommended. Build from source when the package is not compatible with the target system.

Build from source

The complete pipeline requires building both the pinned LLVM/MLIR toolchain and two TensorFlow/LiteRT tools: litert-opt for optional TFL optimization and flatbuffer_translate for required FlatBuffer export. TensorFlow is invoked as a separate toolchain and is not linked into the ONNX-MLIR executables.

For a complete bootstrap build from an activated Python environment:

python -m pip install -r requirements.txt onnxruntime tensorflow

PYTHON_BIN="$(command -v python)" \
BUILD_JOBS=8 \
./scripts/bootstrap_and_build.sh

The script clones pinned LLVM and TensorFlow revisions inside this repository, builds all required tools, and runs an end-to-end smoke test. Generated source and build trees are ignored by Git. See the build guide for prerequisites, repository layout, build products, configuration options, and instructions for reusing existing TensorFlow tools.

Convert a model

./build/bin/onnx-to-tflite \
  /path/to/model.onnx \
  -o /path/to/model.tflite

See Convert an ONNX model for runtime-package usage, layout selection, diagnostics, large-model export, and the complete option reference.

Layout contract

ONNX commonly represents rank-4 activations as NCHW, while TFLite uses NHWC. The default legacy mode exposes rank-4 TFLite graph inputs and outputs as NHWC and keeps convolution and pooling regions in NHWC. It does not insert runtime transposes solely to preserve an NCHW external ABI. Callers must therefore adapt rank-4 boundary tensors when comparing or integrating the generated model.

Rank-1, rank-2, rank-3, and rank-5-or-higher graph boundaries retain their ONNX axis order. Local layout conversions required by individual operations, including Conv3D fallback, are represented explicitly inside the graph.

See the model conversion guide for layout-mode selection and the layout policy for convolution filter layouts, broadcast rules, rank-sensitive behavior, and Conv3D reduction details.

Large models

For models whose constants would exceed ordinary FlatBuffer metadata limits, the driver can use TensorFlow's buffer-offset format. Common adjacent external data files such as model.onnx.data are detected automatically once the model size reaches the configured threshold. The resulting .tflite remains a single self-contained file, but its consumer must support TFLite buffer offsets.

Use --use-buffer-offset to enable this mode explicitly for other external data layouts.

Verification and tests

Successful conversion includes all of the following checks:

  1. MLIR verification after each conversion stage by default.
  2. Rejection of unlegalized ONNX operations.
  3. TFLite TFL3 file-identifier validation.
  4. FlatBuffer import back into MLIR through TensorFlow's parser.
  5. Atomic publication only after validation succeeds.

Run the ONNX-to-TFL MLIR regression suite with:

llvm-project/build/bin/llvm-lit \
  -sv build/test/mlir/conversion/onnx_to_tfl

The end-to-end similarity utility handles the rank-4 boundary layout contract, compares ONNX Runtime and TFLite results, and reports cosine similarity, Euclidean distance, relative Euclidean distance, RMSE, and maximum absolute error:

python test/e2e/run_similarity.py \
  --onnx test/models/mlp.onnx \
  --tflite build/test/models/mlp.tflite

Current scope

The project focuses on statically shaped inference graphs. Dynamic shapes, general control flow, quantization, sparse tensors, and arbitrary custom or Flex operators are outside the current scope. Some complex operations are supported only for constrained static configurations or through ONNX-MLIR importer decomposition.

Review the known limitations before relying on a configuration that is not explicitly covered by the operator support matrix.

Documentation

Document Description
Model conversion Driver usage, layout selection, diagnostics, and command-line options
Build guide Complete LLVM/MLIR, ONNX-MLIR, and TensorFlow/LiteRT build procedure
Compiler optimizations ONNX preprocessing, layout propagation, lowering-time rank reduction, and TFL structural rewrites
Operator support Supported ONNX operations, opsets, types, ranks, attributes, and test coverage
Architecture Cross-version MLIR boundary, conversion design, optimization, and export details
Layout policy Rank-sensitive layout contract and convolution lowering rules
Known limitations Unsupported configurations and implementation constraints

The original ONNX-MLIR documentation remains available under docs/ and at onnx.ai/onnx-mlir.

Project lineage and license

This repository retains the history and compiler infrastructure of upstream ONNX-MLIR and adds the ONNX-to-TFL conversion and TFLite export pipeline.

The upstream base is ONNX-MLIR commit 7cea64f8aa2fbc56f859df768211a95dc2c72ad6. All project-specific compiler, tooling, test, and documentation changes are maintained independently from that revision.

See the repository LICENSE for licensing information.

This is an independent project and is not an official ONNX or ONNX-MLIR release.

About

A general-purpose ONNX → TFLite compiler built on MLIR, producing native TensorFlow Lite FlatBuffers.

Topics

Resources

Code of conduct

Contributing

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages