Skip to content

Vector length validation is unreachable from the batch path, skipped unless the vector is the first field, and off by default #6907

Description

@haoxu0

Summary

vector_length is declared on Field and there is a validator for it, but the validator is reachable from only one code path and has three conditions that silently disable it. The net effect is that a feature view can declare vector_length=160, be materialized with 200-dimensional vectors, and Feast will not complain.

All line numbers are against master as of filing.

1. Validation only runs on the online-write path

_validate_vector_features is defined at feature_store.py:3618 and has exactly one call site, feature_store.py:3665, inside _get_feature_view_and_df_for_online_write.

So:

entry point validated?
write_to_online_store (L3677) yes
materialize (L2835) no
materialize_incremental (L2592) no
get_historical_features (L1988) no

Embedding workloads are overwhelmingly batch. The path that actually carries vectors at volume is the unvalidated one.

2. The check assumes the vector is the first feature

feature_store.py:3629:

if feature_view.features and feature_view.features[0].vector_index:

If the vector field is not at index 0, features[0].vector_index is False and the whole check is skipped. Declaring Field(name="id", ...) before Field(name="embedding", vector_index=True, vector_length=160) is enough to disable validation, with no warning.

There is already a helper that does this correctly — utils.py:1959 _get_feature_view_vector_field_metadata(), which scans feature_view.schema for vector_index and raises if there is more than one (L1965). Several online stores already use it. The validator should too.

3. vector_length=0 is the default, and 0 means "skip"

feature_store.py:3631:

if feature_view.features[0].vector_length != 0:

and the default is 0 (field.py:60). So Field(name="embedding", dtype=Array(Float32), vector_index=True) — a perfectly natural declaration — validates nothing. Opt-in-by-accident is the wrong default for a correctness check.

Suggestion: when vector_index=True, require vector_length, or infer it once from the first batch and then enforce it consistently. Either is better than silently doing nothing.

4. df.iterrows() does not scale

feature_store.py:3632 validates with a per-row Python loop. For embedding tables this is the dominant cost and effectively forces users to turn validation off. A vectorized length check over the column (or an Arrow-level check) is equivalent and ~free.

5. Field.__eq__ ignores two of the three vector attributes

field.py:91-93:

or self.vector_length != other.vector_length
# or self.vector_index != other.vector_index
# or self.vector_search_metric != other.vector_search_metric

Only vector_length participates in equality. Changing vector_search_metric from COSINE to L2, or flipping vector_index, is therefore invisible to feast apply / feast plan — a change that alters retrieval semantics produces no diff.

If these are commented out deliberately (e.g. to avoid churn on registries written before the fields existed), that is worth a comment saying so, because as written it reads like an oversight.

Why this matters

Reviewing an unrelated production embedding pipeline recently, I found the same class of bug with a worse outcome: instead of raising, it silently truncated vectors to the declared width (vector.take(dimension)), and the declared width itself fell back to "length of the first row" when absent from config. Nothing downstream verified it, and the resulting recall loss was invisible because no quality metric was attached to the artifact.

Feast gets the hard part right — it raises rather than truncates. The problem is only that the raise is unreachable from the batch path, skipped when the vector is not the first field, and off by default. Those are cheap to fix and they are what makes the guarantee real.

Proposed changes

  1. Call _validate_vector_features from the materialize / historical-retrieval paths, not only online write.
  2. Resolve the vector field via _get_feature_view_vector_field_metadata() instead of features[0].
  3. Require vector_length when vector_index=True, or enforce a consistent inferred value.
  4. Replace iterrows() with a vectorized length check.
  5. Decide Field.__eq__ intent for vector_index / vector_search_metric — include them, or document why not.

Happy to send this as a PR; filing first since (3) and (5) are behaviour decisions rather than clear-cut fixes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions