Configure cluster settings for ES|QL Data Federation

The data sources feature adds the following cluster settings. For general guidance on how to apply cluster settings across deployment types, refer to configure Elasticsearch.

The object-count limits and authentication gates are operator-managed and take effect without a restart. Settings marked Dynamic can also be updated at runtime. The remaining settings require a node restart.

Note

Some of these settings were renamed. Where a row lists two names, the badge on each shows which versions accept it. Unless a note says otherwise, the older name is no longer registered, and a node that still sets it in elasticsearch.yml fails to start.

These settings cap the number of data sources and datasets a cluster can hold.

Setting Default Description
esql.data_sources.max_count 100 Maximum number of data sources that can be defined. Range 0–1000.
esql.datasets.max_count 1000 Maximum number of datasets that can be defined. Range 0–10,000.

These settings control how many concurrent requests each node sends to external storage and how long it retries throttled requests.

Setting Default Description
esql.external.max_concurrent_requests allocated processors * 3, minimum 4 and maximum 100, further limited so concurrent 8 MiB reads stay within a quarter of heap (or half of indices.breaker.request.limit when tighter) Maximum concurrent cloud API requests per storage scheme, per node. A positive value is still capped by that memory term, so an old explicit 16 cannot skip the budget. 0 removes the permit limit. Range 0–500. On tiny heaps the parse floor of 4 can exceed that memory term. Node-scoped: changing this setting, or tightening the request breaker, takes effect after a restart.
esql.external.throttle_max_retry_duration 30 Maximum total time, in seconds, spent retrying throttled cloud API requests before failing the query. 0 removes the budget. Range 0–300 seconds.
esql.external.max_concurrent_segmenters
esql.external.max_concurrent_segmentators
0 Maximum number of file segmentation tasks that run concurrently. 0 derives the value automatically. Range 0–4096.
esql.external.admission.rescue.enabled true When a byte-budget FIFO head has waited without a grant, admit it over the cap as a plain hold so the queue can progress. At most one rescued hold is live. Disable to observe a hang for diagnosis. Dynamic.

These settings limit how many files and objects glob patterns can discover. Not all caps behave the same way when exceeded — some abort the query, while others fall back to a slower discovery path.

Setting Default Description
esql.external.max_listed_objects 1,000,000 Hard cap on objects visited while listing a glob, including keys that do not match the pattern and keys dropped by exclusion. Applied independently to each glob listing. A comma-separated resource of N globs therefore does N listings; a rewrite-empty fallback can list the same glob again. The kept-files cap (esql.external.max_discovered_files) is shared across that list. Range 1–10,000,000. Dynamic.
esql.external.max_discovered_files 25,000 Hard cap on files kept after listing filters (_file.*) during glob expansion before the query aborts. Protects against degenerate globs. Range 1–1,000,000. Dynamic.
esql.external.max_glob_expansion 100 Cap on concrete paths generated by brace expansion before falling back to listing. Range 1–10,000. Dynamic.

These settings limit how much CPU and read-thread time a single compressed object can consume. The first 1 MB of decompressed data is always allowed; the ratio limit applies after that.

Setting Default Description
esql.external.max_decompression_ratio 200 Maximum ratio of decompressed to compressed bytes for gzip and other stream-only codecs. A read fails with a 400 error if the object expands beyond this multiple of its compressed size. 0 disables the check. Dynamic.
esql.external.max_decompression_ratio.zstd 2000 Maximum decompression ratio for zstd-compressed objects. Overrides esql.external.max_decompression_ratio for zstd, which can legitimately reach higher ratios than gzip. 0 disables the check for zstd. Dynamic.
esql.external.schema_max_fields 1000 Default maximum number of fields or columns a file's schema can have. For NDJSON, it counts objects as well as leaf fields. For CSV, TSV, and Parquet, it counts columns. If a file's schema exceeds this limit, the query fails with an HTTP 400 error. The same limit applies to the number of columns a dataset declares with mappings, and to the width of the schema merged from all files under schema_resolution: union_by_name. With dynamic: false, a declared schema is held to the limit by its number of declared columns rather than by the width of the files it reads, since only the declared columns are read. Parquet files wider than 100,000 columns can still be refused, because the planner reads a footer at that limit to check the declared types. With dynamic: true, each file's inferred schema is held to the limit as well. A dataset overrides it with its schema_max_fields setting. Range 1–100,000. Applied at node startup only.

These settings control which authentication modes data sources can use.

Setting Default Description
esql.external.managed_identity.enabled
esql.datasource.managed_identity.enabled
false Enables auth: "managed_identity" (the node's own cloud identity through the instance metadata service (IMDS)). Operator-only. Intended for single-cloud, single-tenant deployments. Never enable in serverless or multi-tenant clusters. Refer to the managed identity row in authentication models for guidance.
esql.external.federated_identity.enabled
esql.datasource.federated_identity.enabled
false Enables auth: "federated_identity" (OIDC-to-STS token exchange). Operator-only. Available on Elastic Cloud Hosted and Serverless; not available on self-managed, Elastic Cloud Enterprise, or Elastic Cloud on Kubernetes. For setup details, refer to connect with federated identity.

These settings control the external-source cache, which stores inferred schemas, file listings, and the footers of columnar files.

Setting Default Description
esql.external.cache.enabled
esql.source.cache.enabled
true Enables the external-source cache (inferred schemas and file listings). Dynamic.
esql.external.cache.size
esql.source.cache.size
0.5% of heap
0.4% of heap
Memory budget for the cache. Applied at node startup only.
esql.source.cache.schema.ttl — Deprecated and ignored.
Use esql.external.cache.schema.ttl instead. It doesn't inherit this key's value.
esql.external.cache.listing.ttl
esql.source.cache.listing.ttl
5m How long a file listing stays fresh after it is written. A file added or removed becomes visible on the next query once this elapses. A file overwritten in place is noticed then too, though a columnar file rewritten at the same byte length can still be read from a cached footer until esql.external.cache.footer.ttl lapses. A memoized dataset total expires on it too, but for text formats it is re-folded from measurements that never expire, so expiry recomputes the total without re-measuring. Lower it when changes must be visible sooner. Applied at node startup only.
esql.external.cache.schema.ttl 20m How long a schema inferred from a file is reused before it's inferred again. Statistics measured by reading a text file are not covered: they are reused while their entry is held. Set to 0 to remove the limit on schemas; the memoized dataset total stays on esql.external.cache.listing.ttl. Applied at node startup only.
esql.external.cache.footer.size 0.5% of heap Memory budget for cached raw footer bytes (for example, Parquet footers), which are reused across the resolution, split discovery, and execution phases of a query and across back-to-back queries. The budget applies per columnar format reader. Accepts a percentage of heap or an absolute size, and must be greater than zero. Applied at node startup only.
esql.external.cache.footer.parsed.size 1% of heap Memory budget for cached deserialized footers, which avoid re-parsing a footer in every query phase. A parsed footer costs several times its serialized form and grows with column count rather than file size, so raise this when querying wide schemas across large file sets. Applies per columnar format reader, like esql.external.cache.footer.size. Applied at node startup only.
esql.external.cache.footer.ttl 5m How long a cached footer survives after it is stored. Shared by the raw and parsed footer caches. Footer entries are keyed by path and file length rather than modification time, so a file overwritten in place at the same length can be served from the cache until its entry expires. Lower this if your data files are mutated in place. Applied at node startup only.
esql.external.cache.footer.coalesce true When true, concurrent queries that load the same Parquet footer share one object-store GET. Applied at node startup only.
Note

The esql.source.cache.* keys are accepted as deprecated fallbacks and emit a deprecation warning. You can set them in elasticsearch.yml, but you can't update them through the cluster settings API. Use the esql.external.cache.* keys for new configuration.

These settings supply the default for an ES|QL query setting when a query does not specify one.

Setting Default Description
esql.query.settings.wildcards_match_datasets false Whether a wildcard in FROM also matches registered datasets. When false, a wildcard resolves to indices, data streams, aliases, and views, and a dataset is reached by its exact name. A query overrides this in the _query request body or with SET wildcards_match_datasets. Refer to query across datasets and indices. Dynamic.