Describe the bug
HierarchicalDocumentSplitter.split(Document) mishandles a segment buffer that holds only whitespace.
SegmentBuilder.isNotEmpty() checks the raw (untrimmed) length, but SegmentBuilder.toString() trims the result. So once a part won't fit and the buffer is flushed, if the buffer contains only whitespace, segmentText comes back as "". The flush-guard !segmentText.equals(overlap) is then usually true (since overlap starts as null), so the code tries to add an empty TextSegment — or, if subSplitter is null, falls straight into the "doesn't fit" branch with a misleading message, since the whitespace was never actually too big for the max size.
To Reproduce
new DocumentByCharacterSplitter(1, 0).split(Document.from("a b"));
// RuntimeException: The text "b..." (1 characters long) doesn't fit into the maximum
// segment size (1 characters), and there is no subSplitter defined to split it further.
new DocumentByCharacterSplitter(1, 0).split(Document.from(" a"));
// IllegalArgumentException: text cannot be null or blank
new DocumentByCharacterSplitter(2, 0).split(Document.from("ab cd"));
// RuntimeException: The text "c..." (1 characters long) doesn't fit into the maximum
// segment size (2 characters), and there is no subSplitter defined to split it further.
In every case the splitter should just skip the whitespace and continue — nothing here is actually too big for maxSegmentSize.
Expected behavior
split("a b") with maxSegmentSize=1 should return ["a", "b"] (or similar), not throw.
Found via: fuzz-testing all six built-in splitters (DocumentByCharacterSplitter, DocumentByWordSplitter, DocumentBySentenceSplitter, DocumentByLineSplitter, DocumentByParagraphSplitter, DocumentSplitters.recursive) against ~20k randomly generated texts with small maxSegmentSize values, checking for exceptions and oversized/lost content.
Describe the bug
HierarchicalDocumentSplitter.split(Document)mishandles a segment buffer that holds only whitespace.SegmentBuilder.isNotEmpty()checks the raw (untrimmed) length, butSegmentBuilder.toString()trims the result. So once a part won't fit and the buffer is flushed, if the buffer contains only whitespace,segmentTextcomes back as"". The flush-guard!segmentText.equals(overlap)is then usually true (sinceoverlapstarts asnull), so the code tries to add an emptyTextSegment— or, ifsubSplitterisnull, falls straight into the "doesn't fit" branch with a misleading message, since the whitespace was never actually too big for the max size.To Reproduce
In every case the splitter should just skip the whitespace and continue — nothing here is actually too big for
maxSegmentSize.Expected behavior
split("a b")withmaxSegmentSize=1should return["a", "b"](or similar), not throw.Found via: fuzz-testing all six built-in splitters (
DocumentByCharacterSplitter,DocumentByWordSplitter,DocumentBySentenceSplitter,DocumentByLineSplitter,DocumentByParagraphSplitter,DocumentSplitters.recursive) against ~20k randomly generated texts with smallmaxSegmentSizevalues, checking for exceptions and oversized/lost content.