Discussed in #13378
Originally posted by ollz272 June 16, 2026
Add orm/_loading_cy.py with _InstancesBatch, a whole-chunk row processor
built by _instance_processor() for the non-refresh, non-polymorphic load path.
It replaces the per-row [proc(row) for row in fetch] comprehension in
instances(), hoisting all per-query-constant conditions out of the row loop:
- identity-map lookups go against
WeakInstanceDict._dict directly;
InstanceState construction is inlined for the default
ClassManager / InstanceState combination;
- for tuple rows, the primary key and quick column populators are applied by
position, with itemgetter indexes recovered by probing the getters against a
range object.
Rows for instances already in the identity map delegate to _instance(), which
keeps the partial-population, populate_existing and version-check paths. Runs as
plain Python when the extension is not compiled. Most of the speedup is the
batch restructuring itself (it wins uncompiled too); -O3 compilation adds a
smaller further gain (~2–6%), to weigh against the cost of a new compiled module.
Benchmarks
Python 3.14, in-memory SQLite, 50 repeats, extension compiled at -O3.
| case |
before |
after |
Δ |
| subquery_o2m |
15.725 |
11.066 |
−29.6% |
| selectin_o2m |
14.760 |
11.920 |
−19.2% |
| selectin_o2m_few_big |
25.880 |
20.967 |
−19.0% |
| selectin_nested |
14.908 |
12.148 |
−18.5% |
| selectin_m2o |
13.849 |
11.354 |
−18.0% |
| plain_wide |
6.794 |
5.884 |
−13.4% |
| joined_m2o |
12.551 |
11.109 |
−11.5% |
| plain_small |
2.130 |
1.929 |
−9.4% |
| selectin_m2m |
9.681 |
9.538 |
−1.5% |
| joined_o2m |
15.860 |
16.721 |
+5.4% |
Big improvements across the board, though a slight regression on joined loads in the o2m case. I believe this will be due to some double counting going on so could potentially gate this somehow, but also o2m is not recommended for joinedloads anyway so could be a regression worth taking when the m2o case improves.
Discussed in #13378
Originally posted by ollz272 June 16, 2026
Add
orm/_loading_cy.pywith_InstancesBatch, a whole-chunk row processorbuilt by
_instance_processor()for the non-refresh, non-polymorphic load path.It replaces the per-row
[proc(row) for row in fetch]comprehension ininstances(), hoisting all per-query-constant conditions out of the row loop:WeakInstanceDict._dictdirectly;InstanceStateconstruction is inlined for the defaultClassManager/InstanceStatecombination;position, with itemgetter indexes recovered by probing the getters against a
rangeobject.Rows for instances already in the identity map delegate to
_instance(), whichkeeps the partial-population, populate_existing and version-check paths. Runs as
plain Python when the extension is not compiled. Most of the speedup is the
batch restructuring itself (it wins uncompiled too);
-O3compilation adds asmaller further gain (~2–6%), to weigh against the cost of a new compiled module.
Benchmarks
Python 3.14, in-memory SQLite, 50 repeats, extension compiled at
-O3.Big improvements across the board, though a slight regression on joined loads in the o2m case. I believe this will be due to some double counting going on so could potentially gate this somehow, but also o2m is not recommended for joinedloads anyway so could be a regression worth taking when the m2o case improves.