Skip to content

Proposal: speed up parser.parse() (ISO-8601 fast-path + tokenizer hot-path) — gauging appetite before a PR #1519

Description

@tfoutrein

Hi — before opening a large PR I'd like to check your appetite and preferred shape. This is strictly zero behavioral change perf work.

For context: I'm doing this as part of a meetup talk on self-improvement harnesses — the thesis being that you can point one at a real, widely-used library and get genuine, conformance-preserving speedups. dateutil is the case study (a hot transitive dependency, and the per-element fallback in pandas.to_datetime), and I'd genuinely like the improvements to land, not just live on a slide.

I have a fork (branched from current master) with a few commits touching only src/dateutil/parser/_parser.py:

  • a fast-path for ISO-8601 inputs at the top of parse() — today parse() never dispatches to isoparse() (that's a separate module since Implement dedicated isoparser #489/Add isoparse() function. #424), so well-formed ISO strings pay the full general-parser cost. A precompiled regex recognises complete ISO forms and builds the datetime directly; anything non-ISO or out-of-range falls back to the existing slow path unchanged;
  • a bulk-scan tokenizer in _timelex.split (replacing the per-char read(1) loop);
  • unrolled _result.__init__ / __len__, an exception-free per-token float() guard, and a leaner numeric path.

No := and no f-strings were introduced (kept py2.7-friendly on purpose).

Correctness (the priority, given how widely the lib is used)

I ran the full test suite on both master and the fork with identical tests: the failure sets are identical (43 failures, all in test_tz.py / test_tz_prop.py — environmental, local tzdata I/O), 1988 passed, 0 failures in parser/isoparser, 0 introduced regressions. I also ran a differential harness over 2546 inputs (parse(s) datetime equality + exception parity) and the Hypothesis property tests (green).

Performance (workload-dependent — both ends, honestly)

Workload Result Note
Deliberately diversified corpus ×1.64 (−39%) lower bound — over-represents rare manual formats that need the general parser
Representative ISO-dominant corpus (ETL / logs / pandas per-row) ×6.42 (−84%), 534 → 83 ms fast-path bypasses ~90% of the ISO inputs in that corpus

The big number is conditional on ISO-heavy traffic; the diversified mix is the realistic floor. I'm aware reproducibility is the hard part of perf claims (cf. #580) — happy to follow whatever benchmark methodology you'd want, with an explicit workload.

Overlap with existing work

The ISO fast-path looks net-new (I didn't find a change that adds ISO dispatch to parse() — point me at one if it exists). The tokenizer micro-opts overlap with the open #1451 (string construction in get_token) and #1384 (inline methods) — I'd rather reference/coordinate with those than duplicate, or consolidate. The perf appetite in #505 / #732 is what prompted me to ask first.

For the PR

I'd ensure py2.7 compatibility (the fork runs on 3.11; one concrete item I've already spotted is that I bind str.isalpha, which I'd switch to text_type.isalpha so unicode tokens work on py2.7 — on 3.x the two coincide), run the existing pre-commit (darker --isort, trailing-whitespace, debug-statements), follow PEP 8 with dateutil's class-naming exception, add a changelog.d/<PR#>.misc.rst fragment, add targeted tests (ISO forms incl. Z / +HH:MM / out-of-range fallback + a bytes-input tokenizer test), and add myself to AUTHORS.md. Happy to contribute under the existing Apache-2.0 + BSD-3 dual license (no CLA, per CONTRIBUTING).

Questions

  1. Is there appetite for this?
  2. Preferred shape — ISO fast-path alone first, then the tokenizer opts as a separate PR (coordinated with Perf: Optimize string construction in get_token #1451 / Optimize get_token by with inline methods #1384)? One PR per optimization?
  3. Any specific benchmark methodology you'd want for the perf claims?

Thanks!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions