You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Feast has no Lance support today — no issue or PR references it. This proposes LanceSource, but the more useful framing is that Lance is a second implementation of the abstraction already agreed in #6499, and it is the motivating use case that makes #5782, #5652 and #5330 concrete rather than theoretical.
Filing it as one issue because those three are only separable on paper: an embedding source is useless without a version pin, a declared vector schema, and a choice of execution engine.
Format-specific sources have a poor record here — #1494 (Hudi) was closed wontfix, #1533 (Delta) has been open since 2021. So this is deliberately not "add another format".
In #6499, @ntkathole proposed moving catalog connection config onto the DataSource so the offline store stays generic and other engines (Trino, DuckDB) can serve the same tables; @falloficaruss agreed. That abstraction is currently specified against exactly one format and one catalog API. Lance stresses it in two ways Iceberg does not:
A non-Iceberg format inside a REST catalog. Apache Polaris serves Lance through its generic-tables API, on a different path from its Iceberg API. If the source only models Iceberg tables, catalog-backed formats need a parallel code path — cheaper to settle before IcebergRestCatalogSource lands than after.
A non-JVM read path. Lance has a native Python/Rust reader, so this source can be served with no Spark at all. That is a real test of whether the DataSource is engine-agnostic, rather than Spark-agnostic in name only.
Proposal
LanceSource
LanceSource(
catalog="my_catalog", # catalog-backednamespace="my_namespace",
table="item_embeddings",
# or path-based:# uri="s3://bucket/item_embeddings.lance",version=3, # or tag="candidate" — #5782timestamp_field="generated_at",
)
A naming question worth settling early: if #6499 lands as IcebergRestCatalogSource, then a sibling called LanceSource mixes two axes — one named for a catalog, one for a format. Either the catalog connection is a shared, reusable piece that both formats reference, or the names should agree on an axis. I don't have a strong preference, but it is easier to decide now than to rename later.
Vector schema is declared on the FeatureView, using fields Field already has:
The read-path half is already solved: HybridOfflineStore (#5541, shipped in 0.51.0) routes by data source type, so both stores can coexist in one project.
The remaining gap is #5330 — batch_engine / batch_configs on FeatureView are still placeholders, so a project cannot choose the engine per FeatureView. An embedding pipeline needs exactly that: Spark for generation, native for interactive retrieval, one project. #6359 looks like the structural prerequisite, since the offline-store retrieval path currently bypasses the compute engine.
Precedent for one source across engines: FileSource is served by both the Dask and DuckDB offline stores.
@jfw-ppi asked whether the response schema would stay FeatureView-defined rather than varying by store or snapshot. Proposed:
The FeatureView schema is the contract. A pinned version selects data, never shape.
At resolve time, validate the pinned version's actual schema against the declared schema.
Incompatibility fails explicitly. A vector-length change is a hard error, not a silently different response.
Related, and an argument for validating rather than trusting declarations: in Field.__eq__ the vector_index and vector_search_metric comparisons are currently commented out, so only vector_length participates in equality:
orself.vector_length!=other.vector_length# or self.vector_index != other.vector_index# or self.vector_search_metric != other.vector_search_metric
A metric change from COSINE to L2 is therefore invisible to feast apply / plan. Happy to split that out into its own bug if preferred.
This assumes #5652's direction: vector config belongs to the feature definition, not global online_store config. The Field attributes already exist, but online stores still read dimension and metric from global config, so two feature views with different dimensions cannot coexist in one project.
Scope
Not proposing Lance as an online store. Lance is a format plus indexes, not a low-latency KV service, and its write path is columnar while online_write_batch is row-oriented proto. Offline / retrieval axis only.
Summary
Feast has no Lance support today — no issue or PR references it. This proposes
LanceSource, but the more useful framing is that Lance is a second implementation of the abstraction already agreed in #6499, and it is the motivating use case that makes #5782, #5652 and #5330 concrete rather than theoretical.Filing it as one issue because those three are only separable on paper: an embedding source is useless without a version pin, a declared vector schema, and a choice of execution engine.
Why this is not #1494 / #1533
Format-specific sources have a poor record here — #1494 (Hudi) was closed wontfix, #1533 (Delta) has been open since 2021. So this is deliberately not "add another format".
In #6499, @ntkathole proposed moving catalog connection config onto the
DataSourceso the offline store stays generic and other engines (Trino, DuckDB) can serve the same tables; @falloficaruss agreed. That abstraction is currently specified against exactly one format and one catalog API. Lance stresses it in two ways Iceberg does not:IcebergRestCatalogSourcelands than after.DataSourceis engine-agnostic, rather than Spark-agnostic in name only.Proposal
LanceSourceA naming question worth settling early: if #6499 lands as
IcebergRestCatalogSource, then a sibling calledLanceSourcemixes two axes — one named for a catalog, one for a format. Either the catalog connection is a shared, reusable piece that both formats reference, or the names should agree on an axis. I don't have a strong preference, but it is easier to decide now than to rename later.Vector schema is declared on the FeatureView, using fields
Fieldalready has:Two engines, one source — #5330
pylance+ DataFusion/DuckDBlance-sparkThe read-path half is already solved:
HybridOfflineStore(#5541, shipped in 0.51.0) routes by data source type, so both stores can coexist in one project.The remaining gap is #5330 —
batch_engine/batch_configsonFeatureVieware still placeholders, so a project cannot choose the engine per FeatureView. An embedding pipeline needs exactly that: Spark for generation, native for interactive retrieval, one project. #6359 looks like the structural prerequisite, since the offline-store retrieval path currently bypasses the compute engine.Precedent for one source across engines:
FileSourceis served by both the Dask and DuckDB offline stores.Pin semantics — answering @jfw-ppi on #5782
@jfw-ppi asked whether the response schema would stay FeatureView-defined rather than varying by store or snapshot. Proposed:
FeatureViewschema is the contract. A pinned version selects data, never shape.Related, and an argument for validating rather than trusting declarations: in
Field.__eq__thevector_indexandvector_search_metriccomparisons are currently commented out, so onlyvector_lengthparticipates in equality:A metric change from
COSINEtoL2is therefore invisible tofeast apply/ plan. Happy to split that out into its own bug if preferred.Vector config placement — #5652
This assumes #5652's direction: vector config belongs to the feature definition, not global
online_storeconfig. TheFieldattributes already exist, but online stores still read dimension and metric from global config, so two feature views with different dimensions cannot coexist in one project.Scope
Not proposing Lance as an online store. Lance is a format plus indexes, not a low-latency KV service, and its write path is columnar while
online_write_batchis row-oriented proto. Offline / retrieval axis only.Suggested sequencing
LanceSource+ native (pylance) offline store, path-based only — landable without waiting on feat: Extend Feast's DataSource to natively support Iceberg REST Catalog-backed tables #6499version/tagpinning + schema validation (Supporting snapshot based time travel for Feast #5782)LanceSourceRelated: #6499, #5782, #5652, #5330, #6359, #5541, #2406