Loaded content tier data nodes

Some content-tier data nodes show materially higher CPU utilization than their peers on the same tier. Content nodes that absorb disproportionate indexing or search traffic can become the cluster bottleneck.

Note

For a complete list of insights, refer to AutoOps insights.

Field Value
Component Elasticsearch
Severity Medium
Scope Node
Domains performance, resource-utilization

You can customize these settings to adjust when AutoOps detects this event and presents the insight. Refer to AutoOps event settings for details.

The default customization settings are:

Setting Type Default
CPU percent above peer baseline for content-tier data nodes Integer 180
Consecutive samples meeting the imbalance conditions on content-tier data nodes Integer 3
Minimum CPU utilization to consider a content-tier data node loaded Integer 1
Search queue threshold Integer 5
Write queue threshold Integer 1
Tip

Raising these thresholds reduces noise but delays detection. Lowering them triggers the insight sooner but can increase alerts during minor blips.

The following is an example of what you might see when this insight is triggered. Real insights use live data and links from your deployment or cluster.

CPU utilization on es-data-01 and es-data-02 is sustained above the peer baseline for content-tier data nodes.

  • Indices with high search activity: logs-prod-000045

  • Indices with high indexing activity: logs-prod-000045

Note

AutoOps shows different recommendations depending on how their conditions match your deployment or cluster.

When load is uneven across content-tier data nodes, busier nodes become bottlenecks for search and indexing while quieter peers sit underused. Common causes include clients that do not spread requests across the tier, and hot indices whose shards sit on only a few content nodes.

If the imbalance persists, latency and queue pressure tend to concentrate on the overloaded content nodes first.