- Install
uv- Unix:
curl -LsSf https://astral.sh/uv/install.sh | sh - Windows:
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
- Unix:
- Build ty:
cargo build --bin ty --release cdinto the benchmark directory:cd scripts/ty_benchmark- Install npm 11.10.0 or newer, which supports the dependency cooldown in
.npmrc - Install Pyright:
npm ci --ignore-scripts - Run benchmarks:
uv run benchmark
Requires hyperfine 1.20 or newer.
Run with:
uv run --python 3.14 benchmarkMeasures how long it takes to type check a project without a pre-existing cache.
You can run the benchmark with --single-threaded to measure the check time when using a single thread only.
Run with:
uv run --python 3.14 benchmark --warmMeasures how long it takes to recheck a project if there were no changes.
Note: Of the benchmarked type checkers, only mypy supports caching.
Measures how long it takes for a newly started LSP to return the diagnostics for the files open in the editor.
Run with:
uv run --python 3.14 pytest src/benchmark/test_lsp_diagnostics.py::test_fetch_diagnosticsNote: Use -v -s to see the set of diagnostics returned by each type checker.
Measure how long it takes to recheck all open files after making a single change in a file.
Run with:
uv run --python 3.14 pytest src/benchmark/test_lsp_diagnostics.py::test_incremental_editNote: This benchmark uses pull diagnostics for type checkers that support this operation (ty), and falls back to publish diagnostics otherwise (Pyright, Pyrefly).
The tested type checkers implement Python's type system to varying degrees and some projects only successfully pass type checking using a specific type checker. We benchmark against the latest version of each type checker, but some projects may use another version in practice, leading to different diagnostic results between benchmarking and reality.
The benchmark script supports snapshotting the results when running with --snapshot and --accept.
The goal of those snapshots is to catch accidental regressions. For example, if a project adds
new dependencies that we fail to install. They are not intended as a testing tool. E.g. the snapshot runner doesn't account for platform differences so that
you might see differences when running the snapshots on your machine.