You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Hi — before opening a large PR I'd like to check your appetite and preferred shape. This is strictly zero behavioral change perf work.
For context: I'm doing this as part of a meetup talk on self-improvement harnesses — the thesis being that you can point one at a real, widely-used library and get genuine, conformance-preserving speedups. dateutil is the case study (a hot transitive dependency, and the per-element fallback in pandas.to_datetime), and I'd genuinely like the improvements to land, not just live on a slide.
I have a fork (branched from current master) with a few commits touching onlysrc/dateutil/parser/_parser.py:
a fast-path for ISO-8601 inputs at the top of parse() — today parse() never dispatches to isoparse() (that's a separate module since Implement dedicated isoparser #489/Add isoparse() function. #424), so well-formed ISO strings pay the full general-parser cost. A precompiled regex recognises complete ISO forms and builds the datetime directly; anything non-ISO or out-of-range falls back to the existing slow path unchanged;
a bulk-scan tokenizer in _timelex.split (replacing the per-char read(1) loop);
unrolled _result.__init__ / __len__, an exception-free per-token float() guard, and a leaner numeric path.
No := and no f-strings were introduced (kept py2.7-friendly on purpose).
Correctness (the priority, given how widely the lib is used)
I ran the full test suite on both master and the fork with identical tests: the failure sets are identical (43 failures, all in test_tz.py / test_tz_prop.py — environmental, local tzdata I/O), 1988 passed, 0 failures in parser/isoparser, 0 introduced regressions. I also ran a differential harness over 2546 inputs (parse(s) datetime equality + exception parity) and the Hypothesis property tests (green).
Performance (workload-dependent — both ends, honestly)
Workload
Result
Note
Deliberately diversified corpus
×1.64 (−39%)
lower bound — over-represents rare manual formats that need the general parser
Representative ISO-dominant corpus (ETL / logs / pandas per-row)
×6.42 (−84%), 534 → 83 ms
fast-path bypasses ~90% of the ISO inputs in that corpus
The big number is conditional on ISO-heavy traffic; the diversified mix is the realistic floor. I'm aware reproducibility is the hard part of perf claims (cf. #580) — happy to follow whatever benchmark methodology you'd want, with an explicit workload.
Overlap with existing work
The ISO fast-path looks net-new (I didn't find a change that adds ISO dispatch to parse() — point me at one if it exists). The tokenizer micro-opts overlap with the open #1451 (string construction in get_token) and #1384 (inline methods) — I'd rather reference/coordinate with those than duplicate, or consolidate. The perf appetite in #505 / #732 is what prompted me to ask first.
For the PR
I'd ensure py2.7 compatibility (the fork runs on 3.11; one concrete item I've already spotted is that I bind str.isalpha, which I'd switch to text_type.isalpha so unicode tokens work on py2.7 — on 3.x the two coincide), run the existing pre-commit (darker --isort, trailing-whitespace, debug-statements), follow PEP 8 with dateutil's class-naming exception, add a changelog.d/<PR#>.misc.rst fragment, add targeted tests (ISO forms incl. Z / +HH:MM / out-of-range fallback + a bytes-input tokenizer test), and add myself to AUTHORS.md. Happy to contribute under the existing Apache-2.0 + BSD-3 dual license (no CLA, per CONTRIBUTING).
Hi — before opening a large PR I'd like to check your appetite and preferred shape. This is strictly zero behavioral change perf work.
For context: I'm doing this as part of a meetup talk on self-improvement harnesses — the thesis being that you can point one at a real, widely-used library and get genuine, conformance-preserving speedups.
dateutilis the case study (a hot transitive dependency, and the per-element fallback inpandas.to_datetime), and I'd genuinely like the improvements to land, not just live on a slide.I have a fork (branched from current
master) with a few commits touching onlysrc/dateutil/parser/_parser.py:parse()— todayparse()never dispatches toisoparse()(that's a separate module since Implement dedicated isoparser #489/Add isoparse() function. #424), so well-formed ISO strings pay the full general-parser cost. A precompiled regex recognises complete ISO forms and builds thedatetimedirectly; anything non-ISO or out-of-range falls back to the existing slow path unchanged;_timelex.split(replacing the per-charread(1)loop);_result.__init__/__len__, an exception-free per-tokenfloat()guard, and a leaner numeric path.No
:=and no f-strings were introduced (kept py2.7-friendly on purpose).Correctness (the priority, given how widely the lib is used)
I ran the full test suite on both
masterand the fork with identical tests: the failure sets are identical (43 failures, all intest_tz.py/test_tz_prop.py— environmental, local tzdata I/O), 1988 passed, 0 failures inparser/isoparser, 0 introduced regressions. I also ran a differential harness over 2546 inputs (parse(s)datetime equality + exception parity) and the Hypothesis property tests (green).Performance (workload-dependent — both ends, honestly)
The big number is conditional on ISO-heavy traffic; the diversified mix is the realistic floor. I'm aware reproducibility is the hard part of perf claims (cf. #580) — happy to follow whatever benchmark methodology you'd want, with an explicit workload.
Overlap with existing work
The ISO fast-path looks net-new (I didn't find a change that adds ISO dispatch to
parse()— point me at one if it exists). The tokenizer micro-opts overlap with the open #1451 (string construction inget_token) and #1384 (inline methods) — I'd rather reference/coordinate with those than duplicate, or consolidate. The perf appetite in #505 / #732 is what prompted me to ask first.For the PR
I'd ensure py2.7 compatibility (the fork runs on 3.11; one concrete item I've already spotted is that I bind
str.isalpha, which I'd switch totext_type.isalphaso unicode tokens work on py2.7 — on 3.x the two coincide), run the existing pre-commit (darker --isort, trailing-whitespace, debug-statements), follow PEP 8 with dateutil's class-naming exception, add achangelog.d/<PR#>.misc.rstfragment, add targeted tests (ISO forms incl.Z/+HH:MM/ out-of-range fallback + a bytes-input tokenizer test), and add myself toAUTHORS.md. Happy to contribute under the existing Apache-2.0 + BSD-3 dual license (no CLA, per CONTRIBUTING).Questions
Thanks!