Skip to content

feat: Lance as an engine-agnostic, catalog-backed, version-pinned data source #6899

Description

@haoxu0

Summary

Feast has no Lance support today — no issue or PR references it. This proposes LanceSource, but the more useful framing is that Lance is a second implementation of the abstraction already agreed in #6499, and it is the motivating use case that makes #5782, #5652 and #5330 concrete rather than theoretical.

Filing it as one issue because those three are only separable on paper: an embedding source is useless without a version pin, a declared vector schema, and a choice of execution engine.

Why this is not #1494 / #1533

Format-specific sources have a poor record here — #1494 (Hudi) was closed wontfix, #1533 (Delta) has been open since 2021. So this is deliberately not "add another format".

In #6499, @ntkathole proposed moving catalog connection config onto the DataSource so the offline store stays generic and other engines (Trino, DuckDB) can serve the same tables; @falloficaruss agreed. That abstraction is currently specified against exactly one format and one catalog API. Lance stresses it in two ways Iceberg does not:

  1. A non-Iceberg format inside a REST catalog. Apache Polaris serves Lance through its generic-tables API, on a different path from its Iceberg API. If the source only models Iceberg tables, catalog-backed formats need a parallel code path — cheaper to settle before IcebergRestCatalogSource lands than after.
  2. A non-JVM read path. Lance has a native Python/Rust reader, so this source can be served with no Spark at all. That is a real test of whether the DataSource is engine-agnostic, rather than Spark-agnostic in name only.

Proposal

LanceSource

LanceSource(
    catalog="my_catalog",        # catalog-backed
    namespace="my_namespace",
    table="item_embeddings",
    # or path-based:
    # uri="s3://bucket/item_embeddings.lance",

    version=3,                   # or tag="candidate"  — #5782
    timestamp_field="generated_at",
)

A naming question worth settling early: if #6499 lands as IcebergRestCatalogSource, then a sibling called LanceSource mixes two axes — one named for a catalog, one for a format. Either the catalog connection is a shared, reusable piece that both formats reference, or the names should agree on an axis. I don't have a strong preference, but it is easier to decide now than to rename later.

Vector schema is declared on the FeatureView, using fields Field already has:

Field(name="embedding", dtype=Array(Float32),
      vector_index=True, vector_length=160, vector_search_metric="COSINE")

Two engines, one source — #5330

Path Engine Use
Native pylance + DataFusion/DuckDB local dev, CI, notebooks, eval-set construction
Spark lance-spark existing batch/GPU clusters

The read-path half is already solved: HybridOfflineStore (#5541, shipped in 0.51.0) routes by data source type, so both stores can coexist in one project.

The remaining gap is #5330 — batch_engine / batch_configs on FeatureView are still placeholders, so a project cannot choose the engine per FeatureView. An embedding pipeline needs exactly that: Spark for generation, native for interactive retrieval, one project. #6359 looks like the structural prerequisite, since the offline-store retrieval path currently bypasses the compute engine.

Precedent for one source across engines: FileSource is served by both the Dask and DuckDB offline stores.

Pin semantics — answering @jfw-ppi on #5782

@jfw-ppi asked whether the response schema would stay FeatureView-defined rather than varying by store or snapshot. Proposed:

  • The FeatureView schema is the contract. A pinned version selects data, never shape.
  • At resolve time, validate the pinned version's actual schema against the declared schema.
  • Incompatibility fails explicitly. A vector-length change is a hard error, not a silently different response.

Related, and an argument for validating rather than trusting declarations: in Field.__eq__ the vector_index and vector_search_metric comparisons are currently commented out, so only vector_length participates in equality:

or self.vector_length != other.vector_length
# or self.vector_index != other.vector_index
# or self.vector_search_metric != other.vector_search_metric

A metric change from COSINE to L2 is therefore invisible to feast apply / plan. Happy to split that out into its own bug if preferred.

Vector config placement — #5652

This assumes #5652's direction: vector config belongs to the feature definition, not global online_store config. The Field attributes already exist, but online stores still read dimension and metric from global config, so two feature views with different dimensions cannot coexist in one project.

Scope

Not proposing Lance as an online store. Lance is a format plus indexes, not a low-latency KV service, and its write path is columnar while online_write_batch is row-oriented proto. Offline / retrieval axis only.

Suggested sequencing

  1. LanceSource + native (pylance) offline store, path-based only — landable without waiting on feat: Extend Feast's DataSource to natively support Iceberg REST Catalog-backed tables #6499
  2. Catalog-backed addressing, coordinated with feat: Extend Feast's DataSource to natively support Iceberg REST Catalog-backed tables #6499
  3. version / tag pinning + schema validation (Supporting snapshot based time travel for Feast #5782)
  4. Spark offline store support for LanceSource
  5. Per-FeatureView engine selection (Support compute engine configs into FeatureView definition #5330, after Wire ComputeEngine.get_historical_features() into the standard retrieval path to replace per-store BFV transformation duplication #6359)

Related: #6499, #5782, #5652, #5330, #6359, #5541, #2406

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions