Describe the bug
CSVToDocument(conversion_mode="row") raises _csv.Error for CSV input whose records are separated by carriage returns (\r). The same input with LF or CRLF separators produces the expected documents. This was found with a local reproduction on current main, not a production incident.
To reproduce
from haystack.components.converters import CSVToDocument
from haystack.dataclasses import ByteStream
source = ByteStream(data=b"text,author\rfirst,Ada\rsecond,Bob\r")
result = CSVToDocument(conversion_mode="row").run(
sources=[source], content_column="text"
)
print([doc.content for doc in result["documents"]])
Actual behavior
_csv.Error: new-line character seen in unquoted field - do you need to open the file with newline=''?
The error occurs while reading reader.fieldnames.
Expected behavior
The result should contain ['first', 'second'], with the corresponding author metadata and row numbers 0 and 1. Quoted multiline fields should retain their original newline characters.
Root cause and proposed fix
The converter passes io.StringIO(data) to csv.DictReader. Its default newline setting does not split CR-separated records. Passing io.StringIO(data, newline="") lets the CSV reader handle the line endings without translating quoted field content, consistent with Python's CSV input guidance.
A small local patch and regression test are ready. Before the fix, the parameterized LF/CRLF/CR test gives 1 failed / 2 passed; after the fix, all 20 tests in test_csv_to_document.py pass. The test also checks quoted multiline content and row metadata.
System
- macOS arm64, Python 3.12.14
- Haystack 3.3.0-rc0,
main at bb5b39f42edb5cd9f69e5d9025d48d389c08598f
- Official Hatch test environment; no model, GPU, or external API required
AI assistance: Codex generated this report, the reproduction, and the proposed patch, and ran the local checks.
Describe the bug
CSVToDocument(conversion_mode="row")raises_csv.Errorfor CSV input whose records are separated by carriage returns (\r). The same input with LF or CRLF separators produces the expected documents. This was found with a local reproduction on currentmain, not a production incident.To reproduce
Actual behavior
The error occurs while reading
reader.fieldnames.Expected behavior
The result should contain
['first', 'second'], with the corresponding author metadata and row numbers 0 and 1. Quoted multiline fields should retain their original newline characters.Root cause and proposed fix
The converter passes
io.StringIO(data)tocsv.DictReader. Its default newline setting does not split CR-separated records. Passingio.StringIO(data, newline="")lets the CSV reader handle the line endings without translating quoted field content, consistent with Python's CSV input guidance.A small local patch and regression test are ready. Before the fix, the parameterized LF/CRLF/CR test gives 1 failed / 2 passed; after the fix, all 20 tests in
test_csv_to_document.pypass. The test also checks quoted multiline content and row metadata.System
mainatbb5b39f42edb5cd9f69e5d9025d48d389c08598fAI assistance: Codex generated this report, the reproduction, and the proposed patch, and ran the local checks.