Repository navigation
Cursor Read doesn't work in cluster #239
Description
Activity
hi @tjohnson-gala , thank you for using NRedisStack and letting us know your case,,
i see its been a long time since you report the issue,, thank you for your patience. Is this still a valid issue on your side ?
If yes, it would help speeding up if you can provide your code snippet to reproduce the issue.
Upon your feedback, i ll plan to work on this next week, let me know.- addedwaiting-for-feedbackWe need additional information before we can continueWe need additional information before we can continue
on Feb 3, 2025 I believe this is a routing error caused by atypical routing of the
FT.CURSOR READcommand, specifically whereby thecursor_idis only valid on the node that created it; normally, commands are either routed by a key, or are keyless and can be issued anywhere. This is an unusual special-case.To resolve this, the first thing we'd need is an API in SE.Redis to get a pre-pinned connection, for example (new API):
IDatabase GetPinnedDatabase([RedisKey key], [int db], [CommandFlags flags]) // naming is hard
- if a key is provided on a clustered setup, we return a single-endpoint instance that matches the provided key (and flags, for primary/replica etc)
- otherwise, a single-endpoint instance (matching the flags, for primary/replica etc) is selected arbitrarily
That would give us the starting point for this, however: it would be awkward to use this with the existing
SearchCommands[Async]API, as we'd need to carry this state between calls, which does not currently exist. The caller could do this externally, but that then creates a burden on them.Thinking aloud, I wonder if the real answer here is a new primitive that wraps the cursor, i.e.
AggregationRequest r = ... // existing API var cursor = ft.AggregateCursor(index, r); // new method class FTCursor // question: should we : IDisposable, IAsyncDisposable and wrap FT.CURSOR DEL, or just let it timeout? { IDatabase pinnedDb; long cursorId; // todo: some methods to fetch results }
The "fetch" API is left intentionally vague because I wonder if we should expose that as
I[Async]Enumerable<Row>like how SE.Redis hides the implementation details ofSCAN/HSCANetc. For example:AggregationRequest r = ... // existing API await foreach (var row in ft.AggregateCursorAsync(index, r)) { // tada }
Note: SE.Redis does have
IServerwhich is perhaps similar; the slight wrinkle, though, is thatIServerintentionally does not carry theint databaseinformation, and we could perhaps do with a betterGetServerAPI. I'll make a proposal over in SE.Redis.SE.Redis API implementation here: StackExchange/StackExchange.Redis#2936
PR is in the review queue. By necessity, there is an API change required here (although I've made it non-breaking - existing non-cluster code works, after all).
You have two options with the new PR:
- use the new
Cursor{Read|Del}[Async]API that takes anAggregationResultas the parameter (rather than the cursor), and keep passing in the previous result - switch to the new
Aggregate[Async]EnumerableAPI which works comparably, but which returns anI[Async]Enumerable<Row>that deals with all the messy bits internally
- use the new
NRedisStack Version:
v0.8.0, also tried upgrading to v0.11.0
Redis Stack Version:
Not using a Redis Stack image. We've buiilt our own image for cluster, using only the modules we need.
Redis cluster v7.2.1 with RediSearch v2.8.6 and RedisJSON v2.6.6
Description:
Aggregate cursor read fails, saying "cursor not found"
Rarely, it will work and return the next set of results, only to fail on the next iteration.
From the cli, I can use aggregate and cursor read to successfully iterate over all the results.
I have tried some tests where I will run the aggregate from NRedisStack and cursor read from cli, and vice-versa. It usually fails with "cursor not found". In the case where the cli cursor read works with NRedisStack's aggregate, I am able to keep using the cursor and iterate over all the results.
From these details, it appears that NRedisStack's aggregate and cursor commands are served by random nodes in the cluster, but only the node that did the aggregate command can do the subsequent cursor commands.