feat(node): add blobMode to return blob bytes from queries - #4353
Open
BilalAtique wants to merge 4 commits into
Open
BilalAtique wants to merge 4 commits into
BilalAtique wants to merge 4 commits into
Conversation
The bytes come from a pinned snapshot. The query reruns (up to three times) when the table version changes while it runs.
There was a problem hiding this comment.
The earlier mixed-version race and documentation gap are addressed. Blob bytes are fetched from a pinned table snapshot, and a changed query version causes a full retry; the blobMode example now appears in the generated API reference.
If the table changes during all three attempts, this mode returns an error. That visible, documented limit is acceptable here, but callers on heavily updated tables may need to retry the read later.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #4187.
Node queries can now return blob bytes instead of descriptors.
QueryExecutionOptionsgetsblobMode?: "descriptions" | "bytes", used bytoArrow()andtoArray():"descriptions"stays the default, so nothing changes for existing callers.How it works
With
"bytes", the query also reads_rowid, then fills each top-level blob v2 column with the bytes fromfetchBlobs(asLargeBinary)._rowidstays out of the result unless the caller asked for it withwithRowId(). Renamed columns work too, for exampleselect(new Map([["picture", "image"]])). Blobs nested in a struct or a list keep their descriptors.This is the same approach Python uses for blob v2.
to_pandas(blob_mode="bytes")reads descriptors and_rowid, then swaps in the bytes viafetch_blobs(from #3578), and the plan agreed in #4186 does the same forto_arrowandto_list. I also looked at passing Lance'sBlobHandling::AllBinarythrough the Rust query instead. I didn't go that way becauseQueryRequestand the remote query protocol have no field for it, so it would only work on local tables. Going throughfetchBlobsworks on remote tables as well, and it works for vector, full-text and take queries, not just plain scans.Two details:
withRowId()on the caller's query would leak_rowidinto its later executions. Instead, nativeexecutetakes an optionalwithRowIdflag that adds_rowidto a clone for that one execution.To get at
fetchBlobs, the query classes now carry the table they read (Query,VectorQuery,TakeQuery, and theAutoQuerybehindsearch()). The TypeScript API is unchanged apart from the new option and theBlobModetype.Tests
New tests in
table.test.tscovertoArrowandtoArraywithblobMode: "bytes"(column type, no blob marker, nulls),_rowidonly when requested, a renamed blob column, a take query, a second execution of the same query still returning descriptors and no_rowid, a selection with no blob column, an empty result, and an invalidblobMode. The whole Node suite passes. The Node docs underdocs/src/jsare regenerated.A
"lazy"mode that returnsBlobFilehandles could only work intoArray(), since Arrow can't hold JS objects, so I left it for later.