Uploaded image for project: 'Spark'
  1. Spark
  2. SPARK-59415

Introduce an extensible execution model for Python UDF eval types in the PySpark worker

    XMLWordPrintableJSON

Details

    • Umbrella
    • Status: Open
    • Major
    • Resolution: Unresolved
    • 5.0.0
    • None
    • PySpark
    • None

    Description

      Following the serializer and eval-type refactor (SPARK-55388, SPARK-55384, SPARK-55724), the per-eval-type execution logic now lives in read_udfs in python/pyspark/worker.py. While that consolidation was intentional, read_udfs has grown into a large central dispatcher: every eval type is a branch in a single if/elif chain that selects a serializer, parses offsets, defines an execution closure, and returns the same runner contract.

      This structure makes the shared execution lifecycle implicit and couples all eval types into one function. Adding or maintaining an eval type means extending the central chain rather than working within a self-contained unit, which raises the cost and risk of every change and makes the common contract hard to see and enforce.

      This umbrella tracks the work to give Python UDF eval-type execution an extensible, self-describing structure so that each eval type can be reasoned about, tested, and evolved independently, with no change to the user-facing UDF API or the on-the-wire format.

      Attachments

        Activity

          People

            Unassigned Unassigned
            yiconghuang Yicong Huang
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

            Dates

              Created:
              Updated: