Skip to content

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

☁️ Azure Smart Detector Alert Rule Terraform Module

Runs one of Azure's six built-in smart detectors against Application Insights telemetry and routes its detections to action groups. Targets hashicorp/azurerm ~> 4.0.

Terraform azurerm Module Type Resources Posture


🧩 Overview

  • ⚙️ Creates one azurerm_monitor_smart_detector_alert_rule, named this, with its required action_group block.
  • 🔴 One of the six detectors already exists. Azure creates a Failure Anomalies alert rule automatically for every Application Insights component. Declaring that detector here builds a second rule, and nothing in the plan says so.
  • 🔀 The other five sit on the far side of a migration. They are configurable as alert rules only after a component has been migrated to alert-based smart detection; before that, the same detectors are configured through a different resource type that a sibling module owns.
  • 🧮 Converts Sev0–Sev4 into the bare integer form a sibling alert resource uses, so severities can be compared across the family without re-deriving anything.
  • 🛡️ Rejects the four shapes this resource accepts and then quietly fails at: an empty action-group ID set, a subscription- or resource-group-level scope, and IDs that differ only by letter case in either ID collection.
  • 📣 Reports what it cannot check — inert scopes, an ineffective throttle, and durations too compound to parse.

💡 Why it matters: every failure mode on this resource is silent. A rule with no action group IDs, a rule scoped at a subscription, a duplicate-by-case scope, a second Failure Anomalies rule — each applies cleanly, shows no drift, and either never fires or tells nobody when it does. This module turns those into plan-time errors and named outputs.


❤️ Support this project

If this module saves you time:


🗺️ Where this fits in the family

flowchart TB
  rg["terraform-azurerm-resource-group"]
  appi["terraform-azurerm-application-insights"]
  ag["terraform-azurerm-monitor-action-group"]
  this["terraform-azurerm-monitor-smart-detector-alert-rule"]
  classic["terraform-azurerm-application-insights-smart-detection-rule"]
  sqr["terraform-azurerm-monitor-scheduled-query-rules-alert-v2"]
  supp["terraform-azurerm-monitor-alert-processing-rule-suppression"]

  rg -->|"name"| this
  appi -->|"id as a scope"| this
  ag -->|"id"| this
  this -->|"raises alerts"| supp
  classic -.->|"same detectors before migration"| this
  sqr -.->|"sibling alert type, integer severity"| this

  style this fill:#0078D4,stroke:#004578,color:#ffffff
  style appi fill:#004578,stroke:#004578,color:#ffffff
  style rg fill:#F3F6F9,stroke:#8A8886,color:#201F1E
  style ag fill:#F3F6F9,stroke:#8A8886,color:#201F1E
  style classic fill:#F3F6F9,stroke:#8A8886,color:#201F1E
  style sqr fill:#F3F6F9,stroke:#8A8886,color:#201F1E
  style supp fill:#F3F6F9,stroke:#8A8886,color:#201F1E
Loading

The dashed edges are the two relationships worth understanding before you use this module. The classic Application Insights smart detection rule module configures the same detectors on a component that has not been migrated. The scheduled query rule v2 module is the neighbouring alert type, and it takes severity as a bare integer where this one takes a string.


🧬 What this module builds

flowchart TB
  subgraph inputs["Inputs"]
    i1["detector_type, one of six"]
    i2["scope_resource_ids"]
    i3["severity Sev0 to Sev4"]
    i4["frequency, ISO 8601"]
    i5["action_group with ids"]
    i6["throttling_duration, optional"]
  end

  this["azurerm_monitor_smart_detector_alert_rule.this"]

  subgraph outputs["Outputs"]
    o1["id, name"]
    o2["severity_ordinal, the integer form"]
    o3["targets_the_auto_created_failure_anomalies_rule"]
    o4["throttle_shorter_than_frequency, reported"]
    o5["scopes_that_are_not_application_insights_components"]
  end

  i1 --> this
  i2 --> this
  i3 --> this
  i4 --> this
  i5 --> this
  i6 --> this
  this --> o1
  this --> o2
  this --> o3
  this --> o4
  this --> o5

  style this fill:#0078D4,stroke:#004578,color:#ffffff
  style inputs fill:#F3F6F9,stroke:#8A8886,color:#201F1E
  style outputs fill:#F3F6F9,stroke:#8A8886,color:#201F1E
Loading

Resource inventory

Resource Count Notes
azurerm_monitor_smart_detector_alert_rule.this 1 The keystone
action_group block exactly 1 Required by the provider, MaxItems: 1
timeouts block 0–1 All four operations supported

✅ Provider / Versions

Item Value
Terraform >= 1.12.0
hashicorp/azurerm ~> 4.0
Provider block None in this module — the caller configures the provider, its auth, and its features {} block
ARM resource provider Microsoft.AlertsManagement — not Microsoft.Insights

Schema notes that bite

  • 🔴 The provider enforces no relationship between any two fields. Its source declares no CustomizeDiff, ConflictsWith, RequiredWith or ExactlyOneOf, and the documentation carries no note blocks at all.
  • 🔴 action_group.ids is required but has no minimum length. An empty set satisfies the provider.
  • 🔴 Both ID collections are hashed case-insensitively by the provider while Terraform's set(string) is case-sensitive — two IDs differing only in case never converge.
  • 🔴 scope_resource_ids accepts any Resource ID, including a subscription or resource group, which is valid for other alert types and inert for this one.
  • ⚠️ severity is a string here and an integer on the scheduled query rule v2. Both scales are inverted.
  • ⚠️ frequency and throttling_duration have no published value sets — only the ISO 8601 format.
  • ⚠️ timeouts takes Go durations (30m) while frequency takes ISO 8601 (PT30M).
  • ⚠️ No location argument exists. Smart detection is a global service.
  • Only name and resource_group_name are force-new.

🔑 Required Azure RBAC Roles / Permissions

Principal Permission Scope Why
The Terraform identity Microsoft.AlertsManagement/smartDetectorAlertRules/write The rule's resource group Create and update
The Terraform identity Microsoft.AlertsManagement/smartDetectorAlertRules/read The rule Refresh and plan
The Terraform identity Microsoft.AlertsManagement/smartDetectorAlertRules/delete The rule Destroy
The Terraform identity Microsoft.Insights/components/read Each scoped component Azure validates the scope on create
The Terraform identity Microsoft.Insights/actionGroups/read Each action group To reference it

Monitoring Contributor covers all of the above and more. A custom role carrying the three smartDetectorAlertRules operations plus the two reads is the least-privilege form — and note the resource provider, because a role scoped only to Microsoft.Insights/* cannot manage this record.

Plan access is read-only in the useful sense. id is the only computed attribute, so a refresh retrieves nothing sensitive.

⚠️ A notification consequence no Azure role governs. The action group Azure attaches to the auto-created Failure Anomalies rule notifies whoever holds Monitoring Reader or Monitoring Contributor at subscription scope, by role rather than by name. Granting read access can therefore subscribe someone to production alert mail as a side effect.


Azure Prerequisites

  • Microsoft.AlertsManagement registered on the subscription.
  • An existing resource group. Its region does not affect the rule — this resource has no location.
  • At least one existing Application Insights component to scope at.
  • At least one existing action group, since the provider requires the block.
  • For the five non-Failure-Anomalies detectors: the component must already be migrated to alert-based smart detection. Not checkable from Terraform.
  • For Failure Anomalies: nothing to create. Import the rule Azure already made.

📁 Module Structure

terraform-azurerm-monitor-smart-detector-alert-rule/
├── providers.tf    # required_version, pinned azurerm; no provider block
├── variables.tf    # 12 deeply-typed inputs, 23 validations
├── main.tf         # the keystone plus derived facts the plan cannot show
├── outputs.tf      # id first, then identity, then the reported rules
├── README.md       # this file
├── SCOPE.md        # the cross-module contract
├── LICENSE         # MIT
└── .gitignore

⚙️ Quick Start

provider "azurerm" {
  features {}
}

module "trace_severity_alert" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-shop-trace-severity"
  resource_group_name = var.resource_group_name

  detector_type      = "TraceSeverityDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev2"
  frequency          = "PT15M"

  action_group = {
    ids = [var.action_group_id]
  }
}

The caller configures the provider, its authentication, and the mandatory features {} block. This module declares none of them.


🔌 Cross-Module Contract

Consumes

Input Type Source
resource_group_name Resource group name terraform-azurerm-resource-group → name
scope_resource_ids Application Insights component IDs terraform-azurerm-application-insights → id
action_group.ids Action Group IDs terraform-azurerm-monitor-action-group → id

Emits

Output Description
id Rule Resource ID, emitted first
name / resource_group_name Identity
detector_type The detector in use
targets_the_auto_created_failure_anomalies_rule The one to read before applying
requires_component_migrated_to_alert_based_smart_detection Which side of the migration this assumes
severity_ordinal The integer form a sibling resource uses
throttle_shorter_than_frequency A computed rule, reported not enforced
scopes_that_are_not_application_insights_components Probably-inert scopes

📚 Example Library

1 · The minimum call
module "minimum" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-minimum"
  resource_group_name = var.resource_group_name

  detector_type      = "RequestPerformanceDegradationDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev3"
  frequency          = "PT1H"

  action_group = {
    ids = [var.action_group_id]
  }
}

ℹ️ Six of the twelve inputs are required, because the provider requires them. There is no smaller call.

2 · Failure Anomalies, and why this one is different
module "failure_anomalies" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-shop-failures"
  resource_group_name = var.resource_group_name

  detector_type      = "FailureAnomaliesDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev1"
  frequency          = "PT1M"

  action_group = {
    ids = [var.action_group_id]
  }
}

⚠️ Azure already created this rule. A Failure Anomalies alert rule is provisioned automatically with every Application Insights component. Applying the above creates a second rule against the same component, with its own action groups — so a detection notifies two sets of people through two rules, and only one of them is in Terraform state.

The output targets_the_auto_created_failure_anomalies_rule is true here. Prefer importing:

terraform import 'module.failure_anomalies.azurerm_monitor_smart_detector_alert_rule.this' \
  "/subscriptions/$SUB/resourceGroups/rg-obs/providers/Microsoft.AlertsManagement/smartDetectorAlertRules/failure anomalies - appi-shop"
3 · Guarding against the second Failure Anomalies rule
module "failure_anomalies" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-shop-failures"
  resource_group_name = var.resource_group_name

  detector_type      = "FailureAnomaliesDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev1"
  frequency          = "PT5M"

  action_group = {
    ids = [var.action_group_id]
  }
}

check "failure_anomalies_was_imported_not_created" {
  assert {
    condition = !module.failure_anomalies.targets_the_auto_created_failure_anomalies_rule
    error_message = join(" ", [
      "This configuration declares FailureAnomaliesDetector, which Azure creates automatically for every",
      "Application Insights component. Import the existing rule instead of creating a duplicate, or",
      "acknowledge the duplicate deliberately by removing this check.",
    ])
  }
}

💡 A check block warns without blocking the plan, which is the right severity for a fact that is sometimes deliberate. A caller who genuinely wants a second rule removes the check and has a record of the decision.

4 · The empty action-group set the provider accepts
# REJECTED at plan time by this module.
module "notifies_nobody" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-silent"
  resource_group_name = var.resource_group_name

  detector_type      = "MemoryLeakDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev2"
  frequency          = "PT30M"

  action_group = {
    ids = []
  }
}

🔒 The provider requires the action_group block and requires its ids attribute — but sets no minimum on the set's contents, so ids = [] is accepted and produces a rule that detects and tells nobody. No error, no drift, nothing in the plan. This module refuses it:

action_group.ids must contain at least one Action Group Resource ID. The provider accepts an empty
set, which produces a rule that fires and notifies nobody.
5 · A scope that can never fire
# REJECTED at plan time by this module.
module "subscription_scoped" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-whole-subscription"
  resource_group_name = var.resource_group_name

  detector_type      = "ExceptionVolumeChangedDetector"
  scope_resource_ids = ["/subscriptions/00000000-0000-0000-0000-000000000000"]
  severity           = "Sev2"
  frequency          = "PT15M"

  action_group = {
    ids = [var.action_group_id]
  }
}

⚠️ Metric alert rules and activity log alert rules accept a subscription or resource group as their scope. Smart detector alert rules do not — but the provider validates only that the string parses as some Resource ID, so this applies cleanly and never fires. This module rejects the two scope shapes that are recognisably wrong while only reporting resource-level scopes it cannot prove are wrong.

6 · A scope this module reports rather than rejects
module "mixed_scopes" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-mixed"
  resource_group_name = var.resource_group_name

  detector_type = "TraceSeverityDetector"
  scope_resource_ids = [
    var.application_insights_id,
    var.log_analytics_workspace_id,
  ]
  severity  = "Sev3"
  frequency = "PT15M"

  action_group = {
    ids = [var.action_group_id]
  }
}

output "inert_scopes" {
  value = module.mixed_scopes.scopes_that_are_not_application_insights_components
}

ℹ️ Every built-in detector reads Application Insights telemetry, so a workspace scope is very likely inert — but the platform documentation names Application Insights without saying nothing else can ever work. This suite does not reject what its sources do not clearly forbid, so the entry is named in an output instead.

7 · The duplicate that never converges
# REJECTED at plan time by this module.
module "case_variant_scopes" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-dupes"
  resource_group_name = var.resource_group_name

  detector_type = "DependencyPerformanceDegradationDetector"
  scope_resource_ids = [
    "/subscriptions/$SUB/resourceGroups/rg-obs/providers/Microsoft.Insights/components/appi-shop",
    "/subscriptions/$SUB/resourceGroups/rg-obs/providers/Microsoft.Insights/components/APPI-SHOP",
  ]
  severity  = "Sev2"
  frequency = "PT15M"

  action_group = {
    ids = [var.action_group_id]
  }
}

🔒 Terraform's set(string) is case-sensitive, so it keeps both entries. The provider hashes this set with HashStringIgnoreCase, so it keeps one. The result is a plan that never reaches a steady state, and neither tool explains why. The same check guards action_group.ids.

8 · Throttling repeat notifications
module "throttled" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-throttled"
  resource_group_name = var.resource_group_name

  detector_type      = "RequestPerformanceDegradationDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev2"
  frequency          = "PT15M"

  throttling_duration = "PT6H"

  action_group = {
    ids = [var.action_group_id]
  }
}

💡 Unset means every detection notifies, which is this module's default because hearing less about an ongoing problem should be typed deliberately rather than inherited.

9 · A throttle that does nothing, reported not rejected
module "ineffective_throttle" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-pointless-throttle"
  resource_group_name = var.resource_group_name

  detector_type      = "MemoryLeakDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev3"
  frequency          = "PT1H"

  throttling_duration = "PT5M"

  action_group = {
    ids = [var.action_group_id]
  }
}

check "throttle_has_an_effect" {
  assert {
    condition     = module.ineffective_throttle.throttle_shorter_than_frequency != true
    error_message = "The throttle is shorter than the evaluation frequency, so it expires before the next evaluation and suppresses nothing."
  }
}

ℹ️ The provider documents no rule about this pairing, so the module computes the answer and leaves the refusal to the composition. throttle_shorter_than_frequency is null when either duration is unset or too compound to parse — hence != true rather than == false.

10 · Severity across two resource types
module "smart_detector" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-severity-demo"
  resource_group_name = var.resource_group_name

  detector_type      = "TraceSeverityDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev1"
  frequency          = "PT15M"

  action_group = {
    ids = [var.action_group_id]
  }
}

output "severity_comparison" {
  value = {
    as_this_resource_expresses_it = module.smart_detector.severity
    as_an_integer                 = module.smart_detector.severity_ordinal
    with_its_meaning              = module.smart_detector.severity_label
  }
}

⚠️ This resource takes "Sev1"; azurerm_monitor_scheduled_query_rules_alert_v2 takes 1 for the same concept. Both scales are inverted — 0 is the most severe. severity_ordinal exists so a composition comparing the two does not have to strip the prefix itself.

11 · A custom subject and webhook payload
module "custom_notification" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-custom-notify"
  resource_group_name = var.resource_group_name

  detector_type      = "ExceptionVolumeChangedDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev2"
  frequency          = "PT15M"

  action_group = {
    ids           = [var.action_group_id]
    email_subject = "Exception volume anomaly on the shop front end"
    webhook_payload = jsonencode({
      service = "shop-frontend"
      team    = "payments"
      runbook = "https://runbooks.example.com/exception-volume"
    })
  }
}

🔒 Neither field takes effect unless a referenced action group actually has an email receiver or a webhook receiver — which lives in the action group resource, not here, so this module reports rather than validates it. Do not put a bearer token in the payload: it would sit in Terraform state in plaintext. Configure the credential on the action group's webhook receiver instead.

12 · Compound durations, and the gap this module admits
module "compound_duration" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-compound"
  resource_group_name = var.resource_group_name

  detector_type      = "MemoryLeakDetector"
  scope_resource_ids = [var.application_insights_id]
  severity           = "Sev3"
  frequency          = "P1DT2H"

  action_group = {
    ids = [var.action_group_id]
  }
}

output "what_was_not_parsed" {
  value = module.compound_duration.durations_this_module_could_not_parse
}

ℹ️ P1DT2H is legal and is passed through untouched. Because this resource publishes no set of legal durations — unlike the scheduled query rule v2, whose duration arguments each have one — no exact lookup table is possible, so only the single-unit forms are converted to minutes. frequency_minutes is null here and durations_this_module_could_not_parse names the field, rather than the module guessing.

13 · Many detectors over one component with `for_each`
locals {
  detectors = {
    request_perf    = { detector_type = "RequestPerformanceDegradationDetector", severity = "Sev2" }
    dependency_perf = { detector_type = "DependencyPerformanceDegradationDetector", severity = "Sev2" }
    exceptions      = { detector_type = "ExceptionVolumeChangedDetector", severity = "Sev1" }
    trace_severity  = { detector_type = "TraceSeverityDetector", severity = "Sev3" }
    memory_leak     = { detector_type = "MemoryLeakDetector", severity = "Sev3" }
  }
}

module "detectors" {
  source   = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"
  for_each = local.detectors

  name                = "sdar-shop-${each.key}"
  resource_group_name = var.resource_group_name

  detector_type      = each.value.detector_type
  scope_resource_ids = [var.application_insights_id]
  severity           = each.value.severity
  frequency          = "PT15M"

  throttling_duration = "PT6H"

  action_group = {
    ids = [var.action_group_id]
  }

  tags = { workload = "shop", managed_by = "terraform" }
}

output "all_require_migration" {
  value = {
    for k, m in module.detectors : k => m.requires_component_migrated_to_alert_based_smart_detection
  }
}

💡 These are the five detectors that Failure Anomalies is not, so every one reports requires_component_migrated_to_alert_based_smart_detection = true. All five need the component migrated to alert-based smart detection first — and before migration they are configured through the classic resource type instead.

14 · 🏗️ End-to-end composition
provider "azurerm" {
  features {}
}

module "observability_rg" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-observability-eastus2"
  location = "eastus2"
  tags     = { workload = "shop", managed_by = "terraform" }
}

module "workspace" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-log-analytics-workspace.git?ref=v1.0.0"

  name                = "law-shop-eastus2"
  location            = module.observability_rg.location
  resource_group_name = module.observability_rg.name
  tags                = { workload = "shop", managed_by = "terraform" }
}

module "app_insights" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-application-insights.git?ref=v1.0.0"

  name                = "appi-shop-eastus2"
  location            = module.observability_rg.location
  resource_group_name = module.observability_rg.name
  application_type    = "web"

  # The ARM Resource ID of the workspace, not its workspace_id GUID.
  workspace_id = module.workspace.id

  tags = { workload = "shop", managed_by = "terraform" }
}

module "oncall" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-action-group.git?ref=v1.0.0"

  name                = "ag-shop-oncall"
  resource_group_name = module.observability_rg.name
  short_name          = "shoponcall"
  tags                = { workload = "shop", managed_by = "terraform" }
}

module "exception_anomalies" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-smart-detector-alert-rule.git?ref=v1.0.0"

  name                = "sdar-shop-exception-anomalies"
  resource_group_name = module.observability_rg.name

  detector_type      = "ExceptionVolumeChangedDetector"
  scope_resource_ids = [module.app_insights.id]
  severity           = "Sev1"
  frequency          = "PT15M"

  throttling_duration = "PT6H"
  description         = "Exception volume anomaly on the shop front end. Runbook: exception-volume."

  action_group = {
    ids           = [module.oncall.id]
    email_subject = "Exception volume anomaly - shop front end"
  }

  tags = { workload = "shop", managed_by = "terraform" }
}

module "quiet_hours" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-monitor-alert-processing-rule-suppression.git?ref=v1.0.0"

  name                = "apr-shop-quiet-hours"
  resource_group_name = module.observability_rg.name
  scopes              = [module.observability_rg.id]
  tags                = { workload = "shop", managed_by = "terraform" }
}

check "the_alert_can_actually_reach_someone" {
  assert {
    condition     = module.exception_anomalies.action_group_count > 0
    error_message = "The smart detector rule has no action groups, so a detection would notify nobody."
  }
}

check "every_scope_is_application_insights" {
  assert {
    condition     = length(module.exception_anomalies.scopes_that_are_not_application_insights_components) == 0
    error_message = "A scope is not an Application Insights component, so this detector will very likely never fire against it."
  }
}

check "the_throttle_is_not_pointless" {
  assert {
    condition     = module.exception_anomalies.throttle_shorter_than_frequency != true
    error_message = "The throttle is shorter than the evaluation frequency and therefore suppresses nothing."
  }
}

output "detection_posture" {
  value = {
    rule_id                = module.exception_anomalies.id
    severity               = module.exception_anomalies.severity_label
    notifies               = module.exception_anomalies.action_group_count
    duplicate_of_azures    = module.exception_anomalies.targets_the_auto_created_failure_anomalies_rule
    needs_migrated_component = module.exception_anomalies.requires_component_migrated_to_alert_based_smart_detection
    silent_on_resolution   = module.exception_anomalies.no_notification_is_sent_when_a_detection_auto_resolves
  }
}

🏗️ Note the two facts the composition surfaces that no individual plan would. module.workspace.id feeds Application Insights because that resource wants the workspace's ARM Resource ID, not the workspace_id GUID the same module also emits. And silent_on_resolution is true for every configuration — anything downstream that waits for a recovery signal from this rule waits forever.


📥 Inputs

Required (7)

Name Type Description
name string Rule name. Force-new.
resource_group_name string Resource group name. Force-new.
detector_type string One of six detectors, case-sensitive.
scope_resource_ids set(string) Application Insights component IDs.
severity string Sev0–Sev4, inverted scale.
frequency string ISO 8601 duration.
action_group object(...) Required by the provider; at least one ID required here.

Optional (5)

Name Type Default Description
description string null What the rule is for.
enabled bool true Whether it evaluates.
throttling_duration string null Re-notification throttle, ISO 8601.
tags map(string) {} Tags; max 50.
timeouts object(...) null Go durations; all four operations.
Full input schemas
variable "detector_type" {
  type = string
  # Case-sensitive. One of:
  #   FailureAnomaliesDetector                  (Azure creates this one automatically)
  #   RequestPerformanceDegradationDetector
  #   DependencyPerformanceDegradationDetector
  #   ExceptionVolumeChangedDetector
  #   TraceSeverityDetector
  #   MemoryLeakDetector
}

variable "scope_resource_ids" {
  type = set(string)
  # Each entry must be a resource-level Resource ID. Subscription and resource-group scopes are rejected.
  # Entries differing only by letter case are rejected.
}

variable "severity" {
  type = string
  # "Sev0" through "Sev4", case-sensitive. Sev0 is the most severe.
  # Note: the scheduled query rule v2 resource takes the bare integer 0-4 for the same concept.
}

variable "frequency" {
  type = string
  # ISO 8601 duration, for example PT15M. No value set is published for this field.
  # A Go duration such as 15m is rejected with that explanation.
}

variable "action_group" {
  type = object({
    ids             = set(string)          # at least one Action Group Resource ID
    email_subject   = optional(string)     # only effective with an email receiver
    webhook_payload = optional(string)     # valid JSON; only effective with a webhook receiver
  })
}

variable "throttling_duration" {
  type    = string
  default = null
  # ISO 8601. Unset means every detection notifies.
}

variable "timeouts" {
  type = object({
    create = optional(string)
    read   = optional(string)
    update = optional(string)
    delete = optional(string)
  })
  default = null
  # Go durations here: 30m, not PT30M.
}

This resource supports tags but has no location argument, so this module carries the tags half of the universal tail and omits any region input.


🧾 Outputs

Output Description Notes
id Rule Resource ID Emitted first
name Rule name
resource_group_name Containing resource group Governs RBAC, not placement
detector_type Detector in use
targets_the_auto_created_failure_anomalies_rule Duplicate-of-Azure warning Read this one
requires_component_migrated_to_alert_based_smart_detection Migration assumption
detection_window_is_fixed_regardless_of_frequency Failure Anomalies only Constant per detector
severity / severity_ordinal / severity_label Three shapes of one value Cross-family comparison
enabled Whether it evaluates
frequency / frequency_minutes As supplied / in minutes null if compound
durations_this_module_could_not_parse The admitted gap Usually empty
throttling_duration / throttling_minutes As supplied / in minutes null if unset
throttles_renotification Whether a throttle is set
throttle_shorter_than_frequency Computed rule null when unknowable
scope_resource_ids / scope_count The scopes Sorted
scopes_that_are_not_application_insights_components Probably-inert scopes Reported, not rejected
scope_subscription_ids / scopes_span_multiple_subscriptions Where telemetry lives
action_group_ids / action_group_count Notification targets
notification_is_required_but_an_empty_id_set_is_not_rejected_by_the_provider Constant true Why a check exists
no_notification_is_sent_when_a_detection_auto_resolves Constant true Automation design
has_email_subject / has_webhook_payload Presence only Payload never emitted
email_subject_and_webhook_payload_depend_on_receivers_this_module_cannot_see Constant true
rule_is_global_and_has_no_location_argument Constant true
tags Applied tags
accepts_no_credential Constant true With one caveat

No secret is emitted. Nothing here is marked sensitive, because nothing here is a secret.


🧠 Architecture Notes

The duplicate nobody sees. Azure provisions a Failure Anomalies alert rule with every Application Insights component. Terraform has no way to notice: the create succeeds, the state is consistent, and the component now has two rules whose action groups differ. The auto-created rule also survives deletion of the component it watched. targets_the_auto_created_failure_anomalies_rule exists so this is a value you can assert on rather than a fact you have to remember.

Two resource types, one detector. The five non-Failure-Anomalies detectors exist in two forms: the classic Application Insights smart detection rule, and — after a component is migrated to alert-based smart detection — this alert rule. A sibling module owns the classic form. Managing the same detector through both at once is a collision Terraform cannot detect, because the two resource types share no identity to conflict on. This module cannot see whether a component has been migrated, so it reports which side of the line the configuration assumes and leaves the choice visible.

Why several checks exist that duplicate nothing. The provider validates each field alone and relates none of them. action_group.ids is required with no minimum, so an empty set is legal. scope_resource_ids is checked only for ARM-ID syntax, so a subscription scope is legal. Neither produces an error, drift, or any plan output — the rule simply never fires, or fires and tells nobody. Those are the two shapes this module refuses outright.

Case sensitivity in two directions at once. Terraform's set(string) is case-sensitive; the provider hashes both of this resource's ID sets with HashStringIgnoreCase. Two IDs differing only in case are therefore two values to Terraform and one to Azure, and the plan never converges. The module rejects that input because neither tool explains it.

Durations, twice over. frequency and throttling_duration take ISO 8601 (PT30M); timeouts takes Go durations (30m). Neither accepts the other's spelling, and confusing them is common enough that both directions carry a pointed error message. Unlike the scheduled query rule v2 in this family, this resource publishes no set of legal durations, so the module parses the single-unit forms and names anything else in durations_this_module_could_not_parse instead of approximating.

Silence on recovery. Smart detection auto-resolves a fired alert after 8–24 hours without invoking any action group. There is no resolution notification to wire automation to.

No location, no zones. Smart detection is a global service and the resource has no location. The resource group's region determines nothing about the rule.


🧱 Design Principles

Concern This module's default (empty call) Opt-out
Evaluation enabled = true — a disabled detection rule is a silent gap set false
Re-notification no throttle — every detection notifies set throttling_duration
Notification reachability at least one action group ID required none; the module will not accept an empty set
Scope validity subscription and resource-group scopes rejected none; they cannot work
Duplicate IDs case-only duplicates rejected none; they never converge
Secrets none accepted, none emitted none

⚠️ This resource has no empty call. Seven of the twelve inputs are required because the provider requires six of them, so the suite's usual "the empty call produces the safe resource" rule cannot apply. Where a required argument is security-relevant — severity and detector_type — the module instead validates the value set, names the legal values in the error message, and emits the choice as an output.

Three provider defaults were examined and none was flipped. enabled was already the detection-preserving choice, an unset throttle was already the loudest setting, and tags defaults to empty. Both meaningful defaults are stated in the types rather than left implicit.


🚀 Runbook

terraform init -backend=false
terraform validate
terraform fmt -check

Pin the module with ?ref=v1.0.0 — never a branch. This module is plan-only in CI; a human applies.


🧪 Testing

terraform validate proves the configuration parses and every type resolves. terraform fmt -check proves the formatting. Neither runs a variable validation block on this module as a called module — but terraform console does, when this directory is the root:

echo 'null' | terraform console -var-file=bad.tfvars

All 23 validations were proven to fire this way, identified by line number rather than message text.

Four checks necessarily co-fire with a companion, and that is structural rather than accidental: a wrong-case detector_type or severity is by definition also outside the case-sensitive value set, and a Go-duration frequency or throttling_duration also fails the leading-P form check. In each pair the broad check catches the value and the narrow one supplies the explanation, so isolating the narrow check alone is impossible by construction.

What only a real plan or apply can exercise: whether the scoped component exists, whether it has been migrated to alert-based smart detection, whether a Failure Anomalies rule already exists for it, and whether the referenced action groups have the receivers that email_subject and webhook_payload depend on.


💬 Example Output

id                                              = "/subscriptions/.../resourceGroups/rg-observability-eastus2/providers/Microsoft.AlertsManagement/smartDetectorAlertRules/sdar-shop-exception-anomalies"
name                                            = "sdar-shop-exception-anomalies"
detector_type                                   = "ExceptionVolumeChangedDetector"
targets_the_auto_created_failure_anomalies_rule = false
requires_component_migrated_to_alert_based_smart_detection = true
severity                                        = "Sev1"
severity_ordinal                                = 1
severity_label                                  = "Sev1 - error"
frequency                                       = "PT15M"
frequency_minutes                               = 15
throttling_duration                             = "PT6H"
throttling_minutes                              = 360
throttle_shorter_than_frequency                 = false
durations_this_module_could_not_parse            = []
scope_count                                     = 1
scopes_that_are_not_application_insights_components = []
scopes_span_multiple_subscriptions               = false
action_group_count                              = 1
no_notification_is_sent_when_a_detection_auto_resolves = true
rule_is_global_and_has_no_location_argument      = true
accepts_no_credential                            = true

🔍 Troubleshooting

Symptom Cause Fix
Two alert emails for one failure spike FailureAnomaliesDetector was created here alongside the rule Azure made automatically Import the existing rule; check targets_the_auto_created_failure_anomalies_rule
The rule exists and has never fired Scope is not an Application Insights component Check scopes_that_are_not_application_insights_components; scope at the component
The rule exists and has never fired, scope is correct Component not migrated to alert-based smart detection Migrate the component, or use the classic smart detection rule resource instead
Every plan shows a change to scope_resource_ids Two IDs differ only by letter case The module now rejects this; supply each ID once
A detection notified nobody action_group.ids was empty The module now rejects this
Automation never sees a recovery Smart detection sends nothing on auto-resolution Do not wait for a resolution signal; see no_notification_is_sent_when_a_detection_auto_resolves
Custom email subject ignored The referenced action group has no email receiver Add an email receiver to the action group
Webhook payload ignored The referenced action group has no webhook receiver Add a webhook receiver to the action group
frequency rejected as a Go duration 15m was supplied where PT15M was wanted Use ISO 8601 here; Go durations belong in timeouts
A timeout appears to have no effect Misspelled timeouts key, silently discarded by object type conversion Check the spelling against the four declared keys
Raising frequency did not widen the detection window Failure Anomalies uses a fixed 20-minute window Expected; see detection_window_is_fixed_regardless_of_frequency
Custom role cannot manage the rule Role scoped to Microsoft.Insights/* only This record is under Microsoft.AlertsManagement

🔗 Related Docs


💙 "Infrastructure as Code should be standardized, consistent, and secure."