Repository navigation
substr() with replacement slowdown on large strings #23967
Description
Activity
This might have always been the case, or might be a problem with the
OP_SUBSTR_LEFToptimization, which was introduced in the 5.41.x dev cycle. (If the latter, it might boil down to the performance of the OOK mechanism.)I'll take a look.
Reacted by Bartosz Jarzynav5.40.1 shows similar behaviour, so
OP_SUBSTR_LEFTdoesn't seem to have introduced a regression. Will look some more later.Possibly what's happening is a 2x string copy in the
substr_replacecase, courtesy ofPerl_sv_backoff. This double copy is plausibly why the longer the string, the more pronounced the effect.- After the first iteration of
my $cpy = $string;,$cpyis a COW copy of$string. [FLAGS = (POK,IsCOW,pPOK)] - After the first substr,
$cpyis a full-buffer copy of$string, with the OOK hack applied. [FLAGS = (POK,OOK,pPOK)] - At the subsequent iterations of
my $cpy = $string, the OOK hack is undone, incurring a full copy of the full buffer less the offset 6 chars. The contents of$string's buffer are then copied into$cpys buffer again. [FLAGS = (POK,pPOK)] - Rinse & repeat steps 2 & 3.
The
SvOK_off(dsv)at https://github.com/Perl/perl5/blob/blead/sv.c#L4651C15-L4651C23 triggers the de-OOKing.That macro is defined as:
#define SvOK_off(sv) (assert_not_ROK(sv) assert_not_glob(sv) \ SvFLAGS(sv) &= ~(SVf_OK| \ SVf_IVisUV|SVf_UTF8), \ SvOOK_off(sv))With
SvOOK_off(sv)being:#define SvOOK_off(sv) ((void)(SvOOK(sv) && (sv_backoff(sv),0)))Maybe it's possible for
Perl_sv_backoffto only copy the existing buffer contents ifSVf_OKis true. I'll give it a go to find out.- After the first iteration of
PS This expectation can't be achieved without major work on the interpreter:
Expected behavior I would expect substr_replace case above to keep running at more or less the speed of substr regardless of the input size.As currently written, the
substr_replacecase will do a complete copy of$stringon each iteration, whereas thesubstrcase only copies 6 chars on each iteration. The best that we can do currently is try to ensure thatsubstr_replaceis only doing one copy, not two!If the interpreter supported multiple
viewsinto a COW string (probably replacing the OOK hack completely) the expected behaviour could be supported, but that's not something we currently have. (The expectation that all strings have a terminating null byte seems a considerable complicating factor to that idea.)Reacted by Bartosz JarzynaMaybe it's possible for
Perl_sv_backoffto only copy the existing buffer contents ifSVf_OKis true.This might be a good idea in general for the likes of
Perl_sv_setsv_flags, but doesn't immediately help this ticket 'cosPerl_sv_backoffis being called here byPerl_leave_scope. Will look at that callsite next.Thanks. I think at the very minimum,
substr_replaceshould never be slower thanregexp- is this true if we fix the double copying bug?Thanks. I think at the very minimum,
substr_replaceshould never be slower thanregexp- is this true if we fix the double copying bug?Should be. With longer strings, the (1x) copy will dominate both cases and they should end up with similar throughput.
- added 2 commits that reference this issue
on Dec 3, 2025 The two commits above were merged 7 weeks ago and haven't resulted in any BBC reports thus far. Unless there are any objections, I suggest we close this issue.
- addedClosable?We might be able to close this ticket, but we need to check with the reporterWe might be able to close this ticket, but we need to check with the reporter
on Jan 21, 2026 The two commits above were merged 7 weeks ago and haven't resulted in any BBC reports thus far. Unless there are any objections, I suggest we close this issue.
No objection heard to closing this ticket; hence, closing now.
- removedClosable?We might be able to close this ticket, but we need to check with the reporterWe might be able to close this ticket, but we need to check with the reporter
on Mar 11, 2026
Description
substrwith fourth argument (REPLACEMENT) is slower than alternative ways of doing the same thing. The difference is more pronounced as the size of the string grows.Test program:
When ran with a small number argument (like
5), bothsubstrandsubstr_replaceare way faster thanregexp(around +200%),substr_replacebeing marginally slower (not unexpected).Weird things start happening around number
10000(100 kB string).substr_replacestarts to fall off, dropping below the speed ofregexp. With argument1000000(10 MB string)substr_replaceis running at less than half the speed ofregexp.Steps to Reproduce
(see script above)
Expected behavior
I would expect
substr_replacecase above to keep running at more or less the speed ofsubstrregardless of the input size.Perl configuration