Skip to content

register runtime-generated code with linecache so tracebacks, pdb and inspect.getsource() work #13505

Description

@zzzeek

The problem

SQLAlchemy generates Python source at runtime and exec()s it in several
places. Because those exec() calls compile against the default "<string>"
filename and register nothing with the linecache module, the resulting
functions have no retrievable source. Any traceback that passes through
generated code shows an opaque frame with no source line, pdb cannot list or
step through it meaningfully, and inspect.getsource() fails.

The most user-visible case is orm/instrumentation.py::_generate_init(), which
builds the __init__ wrapper installed on every mapped class. It is therefore
on the stack for every mapped object construction. A user who passes a bad
argument to a mapped class currently gets:

Traceback (most recent call last):
  File "test.py", line 17, in <module>
    User(name=42)
    ~~~~^^^^^^^^^
  File "<string>", line 4, in __init__
  File "/.../sqlalchemy/orm/state.py", line 597, in _initialize_instance
    with util.safe_reraise():
         ~~~~~~~~~~~~~~~~~^^
  File "/.../sqlalchemy/util/langhelpers.py", line 164, in __exit__
    raise exc_value.with_traceback(exc_tb)
  File "/.../sqlalchemy/orm/state.py", line 595, in _initialize_instance
    manager.original_init(*mixed[1:], **kwargs)
    ~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^
  File "<stdin>", line 12, in __init__
AttributeError: 'int' object has no attribute 'upper'

The bare File "<string>", line 4, in __init__ frame in the middle is ours,
and it is unattributed and unreadable.

The solution

Compile generated source against a unique synthetic filename and register that
filename in linecache.cache. The traceback machinery does not read source out
of the frame; the frame carries only a filename and a line number, and
linecache is what resolves that pair to source text. Registering an entry is
sufficient to make generated frames render fully:

fname = "<sqlalchemy generated ...>"
linecache.cache[fname] = (len(code), None, code.splitlines(True), fname)
exec(compile(code, fname, "exec"), env)

Prototyping this against _generate_init() turns the frame above into:

  File "<sqlalchemy generated __init__ for class 'User'>", line 4, in __init__
    return new_state._initialize_instance(self, name)
           ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^

inspect.getsource() on the generated __init__ also starts working, where
today it raises:

>>> inspect.getsource(User.__init__)
def __init__(self, name):
    new_state = class_manager._new_state_if_none(self)
    if new_state:
        return new_state._initialize_instance(self, name)
    else:
        return original_init(self, name)

pdb can likewise list and step through generated frames. The technique also
composes with PEP 657 fine-grained error locations — the carets above require
the source line to be resolvable.

The None in the second slot of the tuple is the mtime. linecache.checkcache()
explicitly skips entries whose mtime is None, so a synthetic entry is never
invalidated by a cache sweep and nothing ever attempts to os.stat() a
filename that does not exist.

Is this a supported practice, or a trick?

It is supported, and the non-file source case is explicitly designed for.

  • CPython's own standard library does exactly this. timeit.Timer compiles
    its generated template against dummy_src_name = "<timeit-src>" and, in
    print_exc(), registers linecache.cache[dummy_src_name] = (len(self.src), None, self.src.split("\n"), dummy_src_name). Its docstring states: "The
    advantage over the standard traceback is that source lines in the compiled
    template will be displayed."
    idlelib uses the technique as well.
  • linecache.lazycache(filename, module_globals) is public API added in
    Python 3.5, documented as: "Capture enough detail about a non-file-based
    module to permit getting its lines later via getline()."
    Its existence
    establishes that source without a backing file is a case the module is meant
    to serve.
  • PEP 302. linecache.getline() is documented to fall back to a PEP 302
    __loader__.get_source() in module_globals when the file is not found on
    disk — the general protocol for "source that isn't a file."
  • The mtime is None branch in checkcache() carries the comment # no-op for files loaded via a __loader__, i.e. the exemption exists precisely for
    sources with no file behind them.
  • attrs ships a production implementation in attr/_make.py::_linecache_and_compile(),
    including filename-collision handling, used for its generated __init__ /
    __eq__ / etc. (Note: pydantic does not use this technique.)

Two implementation routes exist. The direct linecache.cache[...] = (...)
assignment above works with any filename, including the conventional
<angle-bracket> form. lazycache() is lazier but requires a real
__loader__ object exposing get_source(), and it explicitly rejects
filenames of the form <...>, so adopting it would mean abandoning the
angle-bracket convention that timeit and attrs both use. The direct route
is recommended.

Where it should be applied

util/langhelpers.py::_exec_code_in_env() is already a shared chokepoint, so
fixing it there covers several call sites at once; the remaining exec() sites
can then be routed through it.

Apply:

location generates
util/langhelpers.py:350 _exec_code_in_env() central helper; backs @decorator and HasTraversalDispatch._generate_dispatcher() (traversal + cache key dispatchers)
orm/instrumentation.py:735 _generate_init() the ORM __init__ wrapper — highest user visibility
util/langhelpers.py:1040 monkeypatch_proxied_specials() proxy methods
util/langhelpers.py:2039 attrsetter() def set(obj, value)

Do not apply — these gain nothing:

location why not
sql/lambdas.py:1264 _rewrite_code_obj() the generated make_cells() is called once and discarded; it exists only to manufacture a closure tuple, which is then grafted onto the user's own lambda, and that retains its real filename. No persistent frame to name.
orm/clsregistry.py:592 one-shot eval() of a relationship string; errors are already caught and re-raised with the name embedded via _raise_for_name()
util/typing.py:265,267 eval_expression() one-shot eval() of an annotation; already wrapped as NameError(f"Could not de-stringify annotation {expression!r}")
testing/plugin/pytestplugin.py:645 test infrastructure, not shipped behavior (could be done for consistency)

Implementation notes

  • Filenames must be unique per generated function, since linecache.cache
    is keyed on them. attrsetter() generates one function per attribute name,
    _generate_init() one per mapped class, _generate_dispatcher() one per
    class; if they all claim the same key they will clobber each other and
    display the wrong source. The discriminator needs to be in the filename, e.g.
    <sqlalchemy generated __init__ for class 'User'>. attrs handles the
    residual collision case with cache.setdefault() plus a disambiguating
    counter.
  • Because mtime=None makes these entries immune to checkcache(), they are
    never evicted. That is acceptable here — all four target sites generate per
    class or per attribute name, so the entry count is bounded by the
    application's class/attribute count. It would be a leak if anything ever
    generated per instance, which none of these do.
  • The cost is retaining the generated source string in memory for the lifetime
    of the process, per generated function. Dispatcher generation is lazy today
    (only 9 of 333 HasCacheKey subclasses generate a dispatcher across two
    representative statements), so in practice this is small.
  • Worth deciding whether to gate this behind a flag. attrs and timeit both
    do it unconditionally.

Related

This came up while evaluating options for reducing cache key generation
overhead (#12514), where one candidate is expanding the existing per-class code
generation in sql/cache_key.py to emit the whole traversal body rather than a
list of dispatch tuples. That would put substantially more logic into generated
code, which makes generated-code debuggability a prerequisite rather than a
nicety. The linecache change stands on its own, though, and is worth doing
independently of whatever happens with that.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ormrobots involved 🤖issue/discussion/PR that is located using machine processes such as LLMs / fuzzers etcuse casenot really a feature or a bug; can be support for new DB features or user use cases not anticipated

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions