Loaded content tier data nodes
Some content-tier data nodes show materially higher CPU utilization than their peers on the same tier. Content nodes that absorb disproportionate indexing or search traffic can become the cluster bottleneck.
For a complete list of insights, refer to AutoOps insights.
| Field | Value |
|---|---|
| Component | Elasticsearch |
| Severity | Medium |
| Scope | Node |
| Domains | performance, resource-utilization |
You can customize these settings to adjust when AutoOps detects this event and presents the insight. Refer to AutoOps event settings for details.
The default customization settings are:
| Setting | Type | Default |
|---|---|---|
| CPU percent above peer baseline for content-tier data nodes | Integer | 180 |
| Consecutive samples meeting the imbalance conditions on content-tier data nodes | Integer | 3 |
| Minimum CPU utilization to consider a content-tier data node loaded | Integer | 1 |
| Search queue threshold | Integer | 5 |
| Write queue threshold | Integer | 1 |
Raising these thresholds reduces noise but delays detection. Lowering them triggers the insight sooner but can increase alerts during minor blips.
The following is an example of what you might see when this insight is triggered. Real insights use live data and links from your deployment or cluster.
CPU utilization on es-data-01 and es-data-02 is sustained above the peer baseline for content-tier data nodes.
Indices with high search activity:
logs-prod-000045Indices with high indexing activity:
logs-prod-000045
AutoOps shows different recommendations depending on how their conditions match your deployment or cluster.
Rebalance index shards
Condition: Shown when index shards are not well spread over data nodes.
Node es-data-01 holds 1000 shards while es-data-02 holds 12. Move shard 0 of index logs-prod-000045 from es-data-01 to es-data-02 using the action below.
POST _cluster/reroute
{
"commands": [
{
"move": {
"index": "logs-prod-000045",
"shard": 0,
"from_node": "es-data-01",
"to_node": "es-data-02"
}
}
]
}
Requires the manage cluster privilege. Requires Elasticsearch 8.0.0 or later. This action changes cluster or index configuration.
Rollover indices
Condition: Shown when high-indexing activity detected and if index has less primary shards than available data nodes.
If logs-prod-000045 are time-based, roll them over and set 1 primary shards on the new write index so you can avoid a heavy segment merge on the current indices.
Rollover indices by shard size
Condition: Shown when high-indexing activity detected and if primary shard size > 30GB.
If logs-prod-000045 are time-based, roll them over when average shard size reaches your target.
Review index templates and mappings
Condition: Shown when high-indexing activity detected and for indexes with highest indexing latency.
Review templates and mappings for logs-prod-000045; they strongly affect indexing performance. Fix oversized mappings, unnecessary fields, and inefficient index settings.
Review indexing slow logs
Condition: Always shown for this insight.
Enable indexing slow logs with the action below on logs-prod-000045, then review the logs to find slow index operations. On Elasticsearch 8.14+, set index.indexing.slowlog.include.user to true to record which user triggered a slow operation.
PUT logs-prod-000045/_settings
{
"index.indexing.slowlog.threshold.index.warn": "10s",
"index.indexing.slowlog.threshold.index.info": "5s",
"index.indexing.slowlog.threshold.index.debug": "2s",
"index.indexing.slowlog.threshold.index.trace": "500ms",
"index.indexing.slowlog.source": "1000"
}
Requires the manage index privilege. Requires Elasticsearch 8.0.0 or later. This action changes cluster or index configuration.
Add data node
Condition: Shown when high-indexing activity detected and if index has more primary shards as available data nodes.
Add a data node to increase capacity and reduce pressure on the existing nodes.
Increase replica count
Condition: Shown when high-searching activity detected and if index has no replica.
Set number_of_replicas to 1 on logs-prod-000045 (currently 2) using the action below.
PUT logs-prod-000045/_settings
{
"index": {
"number_of_replicas": 1
}
}
Requires the manage index privilege. Requires Elasticsearch 8.0.0 or later. This action changes cluster or index configuration.
Enable and review search slow logs
Condition: Always shown for this insight.
Enable search slow logs with the action below, then review the slow log to find expensive queries. See Slow logs for configuration details.
PUT logs-prod-000045/_settings
{
"index.search.slowlog.threshold.query.warn": "10s",
"index.search.slowlog.threshold.query.info": "5s",
"index.search.slowlog.threshold.query.debug": "2s",
"index.search.slowlog.threshold.query.trace": "500ms",
"index.search.slowlog.threshold.fetch.warn": "1s",
"index.search.slowlog.threshold.fetch.info": "800ms",
"index.search.slowlog.threshold.fetch.debug": "500ms",
"index.search.slowlog.threshold.fetch.trace": "200ms"
}
Requires the manage index privilege. Requires Elasticsearch 8.0.0 or later. This action changes cluster or index configuration.
Limit merge thread count
Condition: Shown when high-merging activity is detected.
On spinning disks, set index.merge.scheduler.max_thread_count to 1 for affected indices to reduce I/O contention. Use the action below.
PUT _settings
{
"index.merge.scheduler.max_thread_count": 1
}
Requires the manage index privilege. Requires Elasticsearch 8.0.0 or later. This action changes cluster or index configuration.
When load is uneven across content-tier data nodes, busier nodes become bottlenecks for search and indexing while quieter peers sit underused. Common causes include clients that do not spread requests across the tier, and hot indices whose shards sit on only a few content nodes.
If the imbalance persists, latency and queue pressure tend to concentrate on the overloaded content nodes first.