You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Is your feature request related to a problem? Please describe.
The Python get_online_features path against the Redis online store is slow for wide reads, and the cost is on the client, not the server. Our shape is 4 FeatureViews, ~90 fields, 350 to 500 entities per call, served from a FastAPI process that does other work at the same time.
Two separate problems show up:
HMGET on listpack-encoded hashes scans once per requested field. Redis and Valkey keep a hash as a listpack up to hash-max-listpack-entries (default 128 on Redis 7 and Valkey 8). Our hashes have ~120 fields, so every HMGET of ~90 fields does ~90 linear scans. Measured on Valkey 8 with INFO commandstats, same 500 keys:
command
server time per command
server time per 500-entity read
HMGET, 94 named fields
25.8 us
12.9 ms
HGETALL (123 fields)
5.6 us
2.8 ms
HGETALL returns about a third more bytes and is still 4.6x cheaper for the server. This is the optimization HGETALL optimization for Redis retrieval in Python #3337 asked for in 2022 (the Java server switches to HGETALL above 50 features); that issue was closed by the stale bot without a change.
The Python client work holds the GIL in thousands of short bursts per read. redis-py packs every command, hiredis parses every reply, and each socket receive releases and reacquires the GIL. Then the SDK turns every ValueProto blob into a Python object. In a process with other threads running Python, every one of those reacquires waits out the switch interval, so the read time depends on how busy the rest of the process is.
same field plan, one HGETALL per entity through valkey-glide (Rust core), decode in Python
43 ms
160 ms
The 6x gap under contention is the part that matters in production. It does not show up in a single-threaded benchmark.
Describe the solution you'd like
Two changes to RedisOnlineStore, independent of each other:
In _read_features_per_fv / the batched get_online_features, issue HGETALL instead of HMGET when the number of requested hash fields crosses a threshold (the Java server uses 50), and pick the requested fields out of the reply. Everything else stays the same: same _redis_key, same _mmh3 field names, same _ts:<view> presence check.
An opt-in client for the Redis online store that runs the pipeline off the GIL. valkey-glide (valkey-glide-sync on PyPI) is an official Valkey client with a Rust core; a non-atomic Batch is one FFI call, so the whole fetch runs with the GIL released and returns once. It speaks RESP to Redis and Valkey and supports TLS and cluster mode. A connection_string-compatible client: glide option on RedisOnlineStoreConfig would keep feature_store.yaml as the single place the store is configured.
Describe alternatives you've considered
The Go feature server. It needs the file registry (we use the SQL registry) and a separate Python transformation service for on-demand views, and the Python wheel does not ship the Go library.
get_online_features_async. Same Python parsing on the same GIL; measured within a few percent of the sync path.
Lowering hash-max-listpack-entries on the server so hashes become hashtables. That does remove most of the HMGET cost (7.4 ms to 1.1 ms server time in our test) but costs ~40% more memory and is an operator-side change; HGETALL gets most of the same win from the client.
Reading OnlineResponse.proto directly instead of to_dict(). Worth about 20%, but the redis-py fetch is still the part that stalls under contention.
We have both changes running in a fork of the read path and are happy to contribute either as a PR if there is interest in the direction.
Additional context
Related: #3337 (HGETALL request, closed stale), #4711 and #6337 (single pipeline across feature views, merged; the numbers above are with that in place), #3649 (registry from_proto overhead, a separate cost).
Is your feature request related to a problem? Please describe.
The Python
get_online_featurespath against the Redis online store is slow for wide reads, and the cost is on the client, not the server. Our shape is 4 FeatureViews, ~90 fields, 350 to 500 entities per call, served from a FastAPI process that does other work at the same time.Two separate problems show up:
HMGET on listpack-encoded hashes scans once per requested field. Redis and Valkey keep a hash as a listpack up to
hash-max-listpack-entries(default 128 on Redis 7 and Valkey 8). Our hashes have ~120 fields, so every HMGET of ~90 fields does ~90 linear scans. Measured on Valkey 8 withINFO commandstats, same 500 keys:HGETALL returns about a third more bytes and is still 4.6x cheaper for the server. This is the optimization HGETALL optimization for Redis retrieval in Python #3337 asked for in 2022 (the Java server switches to HGETALL above 50 features); that issue was closed by the stale bot without a change.
The Python client work holds the GIL in thousands of short bursts per read. redis-py packs every command, hiredis parses every reply, and each socket receive releases and reacquires the GIL. Then the SDK turns every
ValueProtoblob into a Python object. In a process with other threads running Python, every one of those reacquires waits out the switch interval, so the read time depends on how busy the rest of the process is.Same machine, same data, Feast 0.66 (which already has the single-pipeline read from feat: Addresses performance issues in the Redis online store #6337), 500 entities, ~90 fields across 4 views, p50 of 40 runs:
get_online_features(redis-py + hiredis)The 6x gap under contention is the part that matters in production. It does not show up in a single-threaded benchmark.
Describe the solution you'd like
Two changes to
RedisOnlineStore, independent of each other:In
_read_features_per_fv/ the batchedget_online_features, issue HGETALL instead of HMGET when the number of requested hash fields crosses a threshold (the Java server uses 50), and pick the requested fields out of the reply. Everything else stays the same: same_redis_key, same_mmh3field names, same_ts:<view>presence check.An opt-in client for the Redis online store that runs the pipeline off the GIL. valkey-glide (
valkey-glide-syncon PyPI) is an official Valkey client with a Rust core; a non-atomicBatchis one FFI call, so the whole fetch runs with the GIL released and returns once. It speaks RESP to Redis and Valkey and supports TLS and cluster mode. Aconnection_string-compatibleclient: glideoption onRedisOnlineStoreConfigwould keepfeature_store.yamlas the single place the store is configured.Describe alternatives you've considered
get_online_features_async. Same Python parsing on the same GIL; measured within a few percent of the sync path.hash-max-listpack-entrieson the server so hashes become hashtables. That does remove most of the HMGET cost (7.4 ms to 1.1 ms server time in our test) but costs ~40% more memory and is an operator-side change; HGETALL gets most of the same win from the client.OnlineResponse.protodirectly instead ofto_dict(). Worth about 20%, but the redis-py fetch is still the part that stalls under contention.We have both changes running in a fork of the read path and are happy to contribute either as a PR if there is interest in the direction.
Additional context
Related: #3337 (HGETALL request, closed stale), #4711 and #6337 (single pipeline across feature views, merged; the numbers above are with that in place), #3649 (registry
from_protooverhead, a separate cost).