The problem
SQLAlchemy generates Python source at runtime and exec()s it in several
places. Because those exec() calls compile against the default "<string>"
filename and register nothing with the linecache module, the resulting
functions have no retrievable source. Any traceback that passes through
generated code shows an opaque frame with no source line, pdb cannot list or
step through it meaningfully, and inspect.getsource() fails.
The most user-visible case is orm/instrumentation.py::_generate_init(), which
builds the __init__ wrapper installed on every mapped class. It is therefore
on the stack for every mapped object construction. A user who passes a bad
argument to a mapped class currently gets:
Traceback (most recent call last):
File "test.py", line 17, in <module>
User(name=42)
~~~~^^^^^^^^^
File "<string>", line 4, in __init__
File "/.../sqlalchemy/orm/state.py", line 597, in _initialize_instance
with util.safe_reraise():
~~~~~~~~~~~~~~~~~^^
File "/.../sqlalchemy/util/langhelpers.py", line 164, in __exit__
raise exc_value.with_traceback(exc_tb)
File "/.../sqlalchemy/orm/state.py", line 595, in _initialize_instance
manager.original_init(*mixed[1:], **kwargs)
~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^
File "<stdin>", line 12, in __init__
AttributeError: 'int' object has no attribute 'upper'
The bare File "<string>", line 4, in __init__ frame in the middle is ours,
and it is unattributed and unreadable.
The solution
Compile generated source against a unique synthetic filename and register that
filename in linecache.cache. The traceback machinery does not read source out
of the frame; the frame carries only a filename and a line number, and
linecache is what resolves that pair to source text. Registering an entry is
sufficient to make generated frames render fully:
fname = "<sqlalchemy generated ...>"
linecache.cache[fname] = (len(code), None, code.splitlines(True), fname)
exec(compile(code, fname, "exec"), env)
Prototyping this against _generate_init() turns the frame above into:
File "<sqlalchemy generated __init__ for class 'User'>", line 4, in __init__
return new_state._initialize_instance(self, name)
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^
inspect.getsource() on the generated __init__ also starts working, where
today it raises:
>>> inspect.getsource(User.__init__)
def __init__(self, name):
new_state = class_manager._new_state_if_none(self)
if new_state:
return new_state._initialize_instance(self, name)
else:
return original_init(self, name)
pdb can likewise list and step through generated frames. The technique also
composes with PEP 657 fine-grained error locations — the carets above require
the source line to be resolvable.
The None in the second slot of the tuple is the mtime. linecache.checkcache()
explicitly skips entries whose mtime is None, so a synthetic entry is never
invalidated by a cache sweep and nothing ever attempts to os.stat() a
filename that does not exist.
Is this a supported practice, or a trick?
It is supported, and the non-file source case is explicitly designed for.
- CPython's own standard library does exactly this.
timeit.Timer compiles
its generated template against dummy_src_name = "<timeit-src>" and, in
print_exc(), registers linecache.cache[dummy_src_name] = (len(self.src), None, self.src.split("\n"), dummy_src_name). Its docstring states: "The
advantage over the standard traceback is that source lines in the compiled
template will be displayed." idlelib uses the technique as well.
linecache.lazycache(filename, module_globals) is public API added in
Python 3.5, documented as: "Capture enough detail about a non-file-based
module to permit getting its lines later via getline()." Its existence
establishes that source without a backing file is a case the module is meant
to serve.
- PEP 302.
linecache.getline() is documented to fall back to a PEP 302
__loader__.get_source() in module_globals when the file is not found on
disk — the general protocol for "source that isn't a file."
- The
mtime is None branch in checkcache() carries the comment # no-op for files loaded via a __loader__, i.e. the exemption exists precisely for
sources with no file behind them.
attrs ships a production implementation in attr/_make.py::_linecache_and_compile(),
including filename-collision handling, used for its generated __init__ /
__eq__ / etc. (Note: pydantic does not use this technique.)
Two implementation routes exist. The direct linecache.cache[...] = (...)
assignment above works with any filename, including the conventional
<angle-bracket> form. lazycache() is lazier but requires a real
__loader__ object exposing get_source(), and it explicitly rejects
filenames of the form <...>, so adopting it would mean abandoning the
angle-bracket convention that timeit and attrs both use. The direct route
is recommended.
Where it should be applied
util/langhelpers.py::_exec_code_in_env() is already a shared chokepoint, so
fixing it there covers several call sites at once; the remaining exec() sites
can then be routed through it.
Apply:
| location |
generates |
util/langhelpers.py:350 _exec_code_in_env() |
central helper; backs @decorator and HasTraversalDispatch._generate_dispatcher() (traversal + cache key dispatchers) |
orm/instrumentation.py:735 _generate_init() |
the ORM __init__ wrapper — highest user visibility |
util/langhelpers.py:1040 monkeypatch_proxied_specials() |
proxy methods |
util/langhelpers.py:2039 attrsetter() |
def set(obj, value) |
Do not apply — these gain nothing:
| location |
why not |
sql/lambdas.py:1264 _rewrite_code_obj() |
the generated make_cells() is called once and discarded; it exists only to manufacture a closure tuple, which is then grafted onto the user's own lambda, and that retains its real filename. No persistent frame to name. |
orm/clsregistry.py:592 |
one-shot eval() of a relationship string; errors are already caught and re-raised with the name embedded via _raise_for_name() |
util/typing.py:265,267 eval_expression() |
one-shot eval() of an annotation; already wrapped as NameError(f"Could not de-stringify annotation {expression!r}") |
testing/plugin/pytestplugin.py:645 |
test infrastructure, not shipped behavior (could be done for consistency) |
Implementation notes
- Filenames must be unique per generated function, since
linecache.cache
is keyed on them. attrsetter() generates one function per attribute name,
_generate_init() one per mapped class, _generate_dispatcher() one per
class; if they all claim the same key they will clobber each other and
display the wrong source. The discriminator needs to be in the filename, e.g.
<sqlalchemy generated __init__ for class 'User'>. attrs handles the
residual collision case with cache.setdefault() plus a disambiguating
counter.
- Because
mtime=None makes these entries immune to checkcache(), they are
never evicted. That is acceptable here — all four target sites generate per
class or per attribute name, so the entry count is bounded by the
application's class/attribute count. It would be a leak if anything ever
generated per instance, which none of these do.
- The cost is retaining the generated source string in memory for the lifetime
of the process, per generated function. Dispatcher generation is lazy today
(only 9 of 333 HasCacheKey subclasses generate a dispatcher across two
representative statements), so in practice this is small.
- Worth deciding whether to gate this behind a flag.
attrs and timeit both
do it unconditionally.
Related
This came up while evaluating options for reducing cache key generation
overhead (#12514), where one candidate is expanding the existing per-class code
generation in sql/cache_key.py to emit the whole traversal body rather than a
list of dispatch tuples. That would put substantially more logic into generated
code, which makes generated-code debuggability a prerequisite rather than a
nicety. The linecache change stands on its own, though, and is worth doing
independently of whatever happens with that.
The problem
SQLAlchemy generates Python source at runtime and
exec()s it in severalplaces. Because those
exec()calls compile against the default"<string>"filename and register nothing with the
linecachemodule, the resultingfunctions have no retrievable source. Any traceback that passes through
generated code shows an opaque frame with no source line,
pdbcannot list orstep through it meaningfully, and
inspect.getsource()fails.The most user-visible case is
orm/instrumentation.py::_generate_init(), whichbuilds the
__init__wrapper installed on every mapped class. It is thereforeon the stack for every mapped object construction. A user who passes a bad
argument to a mapped class currently gets:
The bare
File "<string>", line 4, in __init__frame in the middle is ours,and it is unattributed and unreadable.
The solution
Compile generated source against a unique synthetic filename and register that
filename in
linecache.cache. The traceback machinery does not read source outof the frame; the frame carries only a filename and a line number, and
linecacheis what resolves that pair to source text. Registering an entry issufficient to make generated frames render fully:
Prototyping this against
_generate_init()turns the frame above into:inspect.getsource()on the generated__init__also starts working, wheretoday it raises:
pdbcan likewise list and step through generated frames. The technique alsocomposes with PEP 657 fine-grained error locations — the carets above require
the source line to be resolvable.
The
Nonein the second slot of the tuple is the mtime.linecache.checkcache()explicitly skips entries whose mtime is
None, so a synthetic entry is neverinvalidated by a cache sweep and nothing ever attempts to
os.stat()afilename that does not exist.
Is this a supported practice, or a trick?
It is supported, and the non-file source case is explicitly designed for.
timeit.Timercompilesits generated template against
dummy_src_name = "<timeit-src>"and, inprint_exc(), registerslinecache.cache[dummy_src_name] = (len(self.src), None, self.src.split("\n"), dummy_src_name). Its docstring states: "Theadvantage over the standard traceback is that source lines in the compiled
template will be displayed."
idlelibuses the technique as well.linecache.lazycache(filename, module_globals)is public API added inPython 3.5, documented as: "Capture enough detail about a non-file-based
module to permit getting its lines later via
getline()." Its existenceestablishes that source without a backing file is a case the module is meant
to serve.
linecache.getline()is documented to fall back to a PEP 302__loader__.get_source()inmodule_globalswhen the file is not found ondisk — the general protocol for "source that isn't a file."
mtime is Nonebranch incheckcache()carries the comment# no-op for files loaded via a __loader__, i.e. the exemption exists precisely forsources with no file behind them.
attrsships a production implementation inattr/_make.py::_linecache_and_compile(),including filename-collision handling, used for its generated
__init__/__eq__/ etc. (Note: pydantic does not use this technique.)Two implementation routes exist. The direct
linecache.cache[...] = (...)assignment above works with any filename, including the conventional
<angle-bracket>form.lazycache()is lazier but requires a real__loader__object exposingget_source(), and it explicitly rejectsfilenames of the form
<...>, so adopting it would mean abandoning theangle-bracket convention that
timeitandattrsboth use. The direct routeis recommended.
Where it should be applied
util/langhelpers.py::_exec_code_in_env()is already a shared chokepoint, sofixing it there covers several call sites at once; the remaining
exec()sitescan then be routed through it.
Apply:
util/langhelpers.py:350_exec_code_in_env()@decoratorandHasTraversalDispatch._generate_dispatcher()(traversal + cache key dispatchers)orm/instrumentation.py:735_generate_init()__init__wrapper — highest user visibilityutil/langhelpers.py:1040monkeypatch_proxied_specials()util/langhelpers.py:2039attrsetter()def set(obj, value)Do not apply — these gain nothing:
sql/lambdas.py:1264_rewrite_code_obj()make_cells()is called once and discarded; it exists only to manufacture a closure tuple, which is then grafted onto the user's own lambda, and that retains its real filename. No persistent frame to name.orm/clsregistry.py:592eval()of a relationship string; errors are already caught and re-raised with the name embedded via_raise_for_name()util/typing.py:265,267eval_expression()eval()of an annotation; already wrapped asNameError(f"Could not de-stringify annotation {expression!r}")testing/plugin/pytestplugin.py:645Implementation notes
linecache.cacheis keyed on them.
attrsetter()generates one function per attribute name,_generate_init()one per mapped class,_generate_dispatcher()one perclass; if they all claim the same key they will clobber each other and
display the wrong source. The discriminator needs to be in the filename, e.g.
<sqlalchemy generated __init__ for class 'User'>.attrshandles theresidual collision case with
cache.setdefault()plus a disambiguatingcounter.
mtime=Nonemakes these entries immune tocheckcache(), they arenever evicted. That is acceptable here — all four target sites generate per
class or per attribute name, so the entry count is bounded by the
application's class/attribute count. It would be a leak if anything ever
generated per instance, which none of these do.
of the process, per generated function. Dispatcher generation is lazy today
(only 9 of 333
HasCacheKeysubclasses generate a dispatcher across tworepresentative statements), so in practice this is small.
attrsandtimeitbothdo it unconditionally.
Related
This came up while evaluating options for reducing cache key generation
overhead (#12514), where one candidate is expanding the existing per-class code
generation in
sql/cache_key.pyto emit the whole traversal body rather than alist of dispatch tuples. That would put substantially more logic into generated
code, which makes generated-code debuggability a prerequisite rather than a
nicety. The
linecachechange stands on its own, though, and is worth doingindependently of whatever happens with that.