Details
-
Umbrella
-
Status: Open
-
Major
-
Resolution: Unresolved
-
5.0.0
-
None
-
None
Description
Following the serializer and eval-type refactor (SPARK-55388, SPARK-55384, SPARK-55724), the per-eval-type execution logic now lives in read_udfs in python/pyspark/worker.py. While that consolidation was intentional, read_udfs has grown into a large central dispatcher: every eval type is a branch in a single if/elif chain that selects a serializer, parses offsets, defines an execution closure, and returns the same runner contract.
This structure makes the shared execution lifecycle implicit and couples all eval types into one function. Adding or maintaining an eval type means extending the central chain rather than working within a self-contained unit, which raises the cost and risk of every change and makes the common contract hard to see and enforce.
This umbrella tracks the work to give Python UDF eval-type execution an extensible, self-describing structure so that each eval type can be reasoned about, tested, and evolved independently, with no change to the user-facing UDF API or the on-the-wire format.