Skip to content

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

☁️ Azure Subscription Policy Remediation Terraform Module

Runs a bounded Azure Policy remediation task at subscription scope (azurerm_subscription_policy_remediation), with the narrower discovery mode by default and every pace control emitted. Targets hashicorp/azurerm ~> 4.0.

Terraform Provider Module Type Resources


🧩 Overview

  • 🔧 Creates one remediation task at subscription scope as a keystone resource named this.
  • ⚡ This is the family's one resource that takes action — it modifies or deploys to resources across every resource group in the subscription.
  • 🔍 Exposes resource_discovery_mode, which the management-group-scoped remediation resource does not — defaulted to the narrower ExistingNonCompliant, with ReEvaluateCompliance an explicit opt-in.
  • 🛑 failure_percentage is a circuit breaker, validated as a 0.0–1.0 fraction so the common "10 means 10%" mistake fails at plan.
  • 📏 Bounds a run with resource_count and slows it with parallel_deployments.
  • 🌎 Stages a rollout region by region with location_filters.
  • 📤 Emits the discovery mode and every pace control, so a state or plan review can see whether the job was bounded.

💡 Why it matters: A remediation at subscription scope is a change wave, not a report. The service's own defaults are deliberately fast, so these knobs are what stand between one policy and a subscription-wide modification. ReEvaluateCompliance compounds that: it re-scans before acting, so the set of resources it changes is not knowable from the current compliance report.

❤️ Support this project

If this module saves you time, please consider supporting its continued development:


🗺️ Where this fits in the family

flowchart LR
  sub["terraform-azurerm-subscription"]
  psd["terraform-azurerm-policy-set-definition"]
  uai["terraform-azurerm-user-assigned-identity"]
  pa["terraform-azurerm-subscription-policy-assignment"]
  exm["terraform-azurerm-subscription-policy-exemption"]
  this["terraform-azurerm-subscription-policy-remediation"]
  rem["azurerm_subscription_policy_remediation"]

  sub -->|"subscription_id"| this
  psd -->|"policy_definition_reference_id"| this
  uai -->|"identity attached to the assignment"| pa
  pa -->|"policy_assignment_id"| this
  exm -->|"removes resources from reach"| this
  this -->|"creates"| rem
  rem -->|"modifies or deploys to resources"| sub

  classDef me fill:#0078D4,stroke:#004578,color:#fff;
  classDef keystone fill:#004578,stroke:#001f3f,color:#fff;
  classDef sib fill:#eef2f7,stroke:#b8c4d0,color:#1b1b1b;
  class this me;
  class rem keystone;
  class sub,psd,uai,pa,exm sib;
Loading

🧬 What this module builds

flowchart TB
  ident["name: force-new, so each run is its own task"]
  target["policy_assignment_id: DeployIfNotExists or Modify"]
  ref["policy_definition_reference_id for an initiative"]
  disc["resource_discovery_mode defaults ExistingNonCompliant"]
  reeval["ReEvaluateCompliance re-scans first, wider reach"]
  pace["pace limits: resource_count, parallel_deployments"]
  brake["failure_percentage: 0.0 to 1.0 circuit breaker"]
  loc["location_filters: stage region by region"]
  this["terraform-azurerm-subscription-policy-remediation"]
  rem["azurerm_subscription_policy_remediation.this"]
  msi["assignment managed identity performs the change"]
  res["non-compliant resources across the subscription"]
  out["outputs: id, resource_discovery_mode, resource_count, failure_percentage"]

  ident -->|"identity"| this
  target -->|"what to remediate"| this
  ref -->|"which policy inside the initiative"| this
  disc -->|"narrower default"| this
  reeval -->|"explicit opt-in"| disc
  pace -->|"bounds the batch"| this
  brake -->|"aborts on systemic failure"| this
  loc -->|"restricts the sweep"| this
  this -->|"creates"| rem
  rem -->|"runs as"| msi
  msi -->|"acts on"| res
  rem -->|"emits"| out

  classDef me fill:#0078D4,stroke:#004578,color:#fff;
  classDef keystone fill:#004578,stroke:#001f3f,color:#fff;
  classDef sib fill:#eef2f7,stroke:#b8c4d0,color:#1b1b1b;
  class this me;
  class rem keystone;
  class ident,target,ref,disc,reeval,pace,brake,loc,msi,res,out sib;
Loading

Resource inventory

Resource Count Role
azurerm_subscription_policy_remediation.this 1 The keystone remediation task, with its timeouts block.

✅ Provider / Versions

Requirement Value
Terraform >= 1.12.0
hashicorp/azurerm ~> 4.0
Provider block None in this module — the caller configures provider "azurerm" { features {} }, auth, and subscription.

Schema notes that bite (verified against the live provider schema):

  • subscription_id is a Resource ID (/subscriptions/<guid>), not a bare GUID.
  • This variant exposes resource_discovery_mode, which the management-group-scoped remediation resource does not. Values are ExistingNonCompliant (the service default) and ReEvaluateCompliance. The second triggers a fresh compliance evaluation before remediating, so the set of resources it changes is not knowable from the current compliance report — it is both slower and wider-reaching. This is the main behavioural difference between the two scopes.
  • name and subscription_id are force-new. name being force-new is the right behaviour: each run should be its own auditable task rather than a reused record.
  • failure_percentage is a fraction between 0.0 and 1.0, not a percentage out of 100. Passing 10 intending "10%" is out of range at the API; this module rejects it at plan.
  • A remediation is a task, not a desired-state record. Deleting the Terraform resource cancels or removes the task; it does not revert changes the task already made.
  • The task runs asynchronously after apply. Terraform reports success once the job is accepted, not once the resources are compliant.
  • Only DeployIfNotExists and Modify effects are remediable. Pointing a task at an Audit or Deny assignment produces a job that finds no work and reports success.
  • Omitting policy_definition_reference_id against an assignment that assigns an initiative fails at apply, not at plan — the provider cannot tell which kind of definition the assignment references.
  • The remediation borrows the assignment's managed identity; it has none of its own. Missing roles on that identity, not on the caller, is the usual reason a job changes nothing.
  • This resource type has no tags surface.

🔑 Required Azure RBAC Roles / Permissions

  • Resource Policy Contributor at the target subscription, which covers Microsoft.Authorization/remediations/*.
  • The policy assignment's own managed identity — not the caller's — needs whatever rights the policy's roleDefinitionIds demand on every resource it will modify. A remediation that starts but changes nothing is usually this identity missing a role assignment at the target scope.
  • The deploying identity needs Microsoft.Authorization/policyAssignments/read on the assignment, including when that assignment lives at a management group above this subscription.

Azure Prerequisites

  • An existing subscription, and its Resource ID (/subscriptions/<guid>) — not a bare GUID.
  • An existing policy assignment at or above this subscription with a remediable effect — DeployIfNotExists or Modify. An Audit, AuditIfNotExists, or Deny assignment has nothing to remediate.
  • The assignment must carry a managed identity, and that identity must already hold the roles the policy requires. Policy identity role propagation is asynchronous; a remediation created immediately after the assignment can find no permissions.
  • With the default resource_discovery_mode = "ExistingNonCompliant", an initial compliance evaluation must already have run, because the job acts only on resources already recorded as non-compliant. ReEvaluateCompliance removes that prerequisite at the cost of a wider, slower run.
  • Where the assignment assigns an initiative, the reference ID of the specific policy to remediate.
  • The caller configures the provider "azurerm" { features {} } block, auth, and subscription.

📁 Module Structure

terraform-azurerm-subscription-policy-remediation/
├── providers.tf   # required_version >= 1.12.0; azurerm ~> 4.0; no provider block
├── variables.tf   # deeply-typed schemas, validated discovery mode + pace controls, timeouts tail
├── main.tf        # keystone azurerm_subscription_policy_remediation.this; dynamic timeouts
├── outputs.tf     # id, name, scope, assignment, reference, discovery mode, every pace control
├── README.md      # this document
├── SCOPE.md       # cross-module contract
├── LICENSE        # MIT
└── .gitignore     # canonical library ignore set

⚙️ Quick Start

provider "azurerm" {
  features {}
}

module "remediation" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-subscription-policy-remediation.git?ref=v1.0.0"

  name                 = "remediate-storage-https"
  subscription_id      = "/subscriptions/00000000-0000-0000-0000-000000000000"
  policy_assignment_id = var.baseline_assignment_id

  # Bound the first run: the narrower discovery mode, one region, a small batch,
  # low concurrency, and an early abort.
  resource_discovery_mode = "ExistingNonCompliant"
  location_filters        = ["eastus2"]
  resource_count          = 50
  parallel_deployments    = 5
  failure_percentage      = 0.1

  timeouts = { create = "2h" }
}

ℹ️ The caller owns the provider, its authentication, and the mandatory features {} block. This module never declares them.


🔌 Cross-Module Contract

Consumes

Input Type Source module
subscription_id string terraform-azurerm-subscription (subscription_resource_id) or caller (a /subscriptions/<guid> Resource ID)
policy_assignment_id string terraform-azurerm-subscription-policy-assignment (id) or terraform-azurerm-management-group-policy-assignment (id)
policy_definition_reference_id string an initiative module's policy_definition_reference_ids output
location_filters list(string) caller

Emits

Output Description Consumed by
id Remediation Resource ID (first) audit inventories, run tracking
name Remediation task name operational review
subscription_id The scope the job runs at composition wiring
policy_assignment_id The assignment being remediated audit trace back to the control
policy_definition_reference_id The initiative reference remediated audit scoping
resource_discovery_mode How resources were discovered change-control review — ReEvaluateCompliance means the acted-on set is not knowable from the prior report
resource_count The ceiling on resources acted on change-control review — null means unbounded
parallel_deployments Concurrency change-control review
failure_percentage The abort threshold change-control review — null means no explicit breaker
location_filters Regions the job is restricted to change-control review

📚 Example Library

1 · Minimal remediation against a standalone definition
module "remediation" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-subscription-policy-remediation.git?ref=v1.0.0"

  name                 = "remediate-diagnostics"
  subscription_id      = "/subscriptions/00000000-0000-0000-0000-000000000000"
  policy_assignment_id = var.assignment_id
}

⚠️ With no pace controls this job runs at the service defaults — effectively unbounded across the subscription. It does at least use the narrower ExistingNonCompliant discovery mode by default. Read example 3 before shipping this.

2 · The `subscription_id` form that catches people out
subscription_id = "/subscriptions/00000000-0000-0000-0000-000000000000" # ✅ a Resource ID
# subscription_id = "00000000-0000-0000-0000-000000000000"              # ❌ a bare GUID

⚠️ Wire it from terraform-azurerm-subscription's subscription_resource_id output — not its id, which is the alias Resource ID — or from data.azurerm_subscription.current.id. Do not compose the string by hand.

3 · A bounded first run
resource_discovery_mode = "ExistingNonCompliant" # already the default; stated for auditability
resource_count          = 25                     # act on at most 25 resources
parallel_deployments    = 2                      # two at a time, so a failure is noticeable
failure_percentage      = 0.05                   # abort once 5% of deployments fail
location_filters        = ["eastus2"]

🔒 This is the shape to use the first time a policy remediates anything. A bounded batch you can inspect beats a change wave you have to explain.

4 · The two discovery modes
resource_discovery_mode = "ExistingNonCompliant" # default: act only on already-known non-compliance
resource_discovery_mode = "ReEvaluateCompliance" # re-scan the subscription first, then remediate
resource_count          = 50                     # bound it — the acted-on set is not knowable in advance
failure_percentage      = 0.1
timeouts                = { create = "4h" }      # the evaluation pass runs before any work starts

⚠️ ReEvaluateCompliance triggers a fresh compliance evaluation before acting, so the set of resources it changes cannot be predicted from the current compliance report. Use it when compliance data is known to be stale, and always pair it with the pace controls. This module defaults to the narrower mode and emits the chosen one as an output.

ℹ️ This field does not exist at all on the management-group-scoped remediation resource — which is why that scope can only ever work from already-known non-compliance.

5 · Remediating one policy inside an initiative
policy_definition_reference_id = module.baseline_initiative.policy_definition_reference_ids["require-https-storage"]

⚠️ Required whenever the assignment assigns an initiative. Omitting it fails at apply rather than at plan, because the provider cannot tell from the ID which kind of definition the assignment references.

6 · The circuit breaker, in the right units
failure_percentage = 0.1 # ✅ 10% of deployments may fail before the job aborts
failure_percentage = 10 # ❌ fails at plan
Error: Invalid value for variable

  failure_percentage must be between 0.0 and 1.0 — it is a fraction, not a percentage out of 100.

💡 Without a breaker, a remediation failing for a systemic reason keeps trying across the whole subscription. With a low one it stops early and leaves the rest untouched.

7 · Staging a rollout region by region
# Run 1 — a single region.
location_filters = ["eastus2"]

# Run 2, after confirming the result — widen.
# location_filters = ["eastus2", "eastus", "centralus"]

# Run 3 — everywhere.
# location_filters = null

💡 Region filtering is the simplest way to stage a remediation on a subscription with workloads in several regions. Each widening is a reviewable diff — and since name is force-new, each run is its own task.

8 · Slowing a large job down deliberately
parallel_deployments = 1 # one resource at a time
resource_count       = 200
timeouts             = { create = "4h" }

ℹ️ A modification applied to hundreds of resources at once leaves no room to notice it going wrong. Serial execution costs wall-clock time and buys observability — raise the create timeout to match.

9 · Validating pace controls at plan time
resource_count       = 0  # ❌ "resource_count must be greater than 0 when set."
parallel_deployments = -1 # ❌ "parallel_deployments must be greater than 0 when set."

💡 Both are rejected before any resource is touched. Leave a control null to take the service default rather than passing zero.

10 · A deliberately unbounded sweep, once the policy is trusted
module "full_sweep" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-subscription-policy-remediation.git?ref=v1.0.0"

  name                 = "remediate-https-all-regions"
  subscription_id      = data.azurerm_subscription.current.id
  policy_assignment_id = var.assignment_id

  # Deliberately unset: this policy has already been staged through three bounded runs.
  resource_count       = null
  parallel_deployments = 10
  failure_percentage   = 0.1 # keep the breaker even on a trusted policy
  location_filters     = null

  timeouts = { create = "6h" }
}

🔒 An unbounded run is a reasonable final step, not a first one. Keep failure_percentage set even here — it costs nothing and stops a systemic failure from running to completion.

11 · Remediating an assignment inherited from a management group
policy_assignment_id = var.central_baseline_assignment_id # assigned above this subscription
subscription_id      = data.azurerm_subscription.current.id  # but remediated only here

💡 This is a useful pattern: the control is enforced centrally, but the remediation is scoped to one subscription so its blast radius stays that subscription's. The assignment's identity still needs the policy's roles at the resources being changed.

12 · Working with an exemption in place
module "remediation" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-subscription-policy-remediation.git?ref=v1.0.0"

  name                           = "remediate-remaining-https"
  subscription_id                = data.azurerm_subscription.current.id
  policy_assignment_id           = module.baseline_assignment.id
  policy_definition_reference_id = module.baseline_initiative.policy_definition_reference_ids["require-https-storage"]

  depends_on = [azurerm_subscription_policy_exemption.legacy_waiver] # ensure the exemption exists before the sweep runs
}

ℹ️ The depends_on here is deliberate: without it, Terraform may create the remediation before the exemption, and the job would act on resources you intended to exclude. This is one of the few places in this library where an explicit dependency is not expressible as an attribute reference.

13 · Name-stamping each run
name = "remediate-https-2026-08" # a date- or ticket-stamped name per run

⚠️ Changing name is force-new, which is exactly right here: each run is its own auditable task. Reusing one name across runs hides the history — and a remediation is a task, not a desired-state record.

14 · Several remediations from a keyed map
locals {
  sweeps = {
    https-storage = { reference = "require-https-storage", batch = 50 }
    diagnostics   = { reference = "require-diagnostics", batch = 25 }
  }
}

module "sweeps" {
  source   = "git::https://github.com/microsoftexpert/terraform-azurerm-subscription-policy-remediation.git?ref=v1.0.0"
  for_each = local.sweeps

  name                           = "remediate-${each.key}"
  subscription_id                = data.azurerm_subscription.current.id
  policy_assignment_id           = module.baseline_assignment.id
  policy_definition_reference_id = module.baseline_initiative.policy_definition_reference_ids[each.value.reference]

  resource_count       = each.value.batch
  parallel_deployments = 5
  failure_percentage   = 0.1
}

💡 One task per policy keeps each job's blast radius, pace, and outcome separately reviewable.

15 · Auditing what a run was allowed to do
output "remediation_review" {
  description = "The blast-radius envelope of each remediation. A ReEvaluateCompliance mode or a null resource_count is the row to question."
  value = {
    for k, m in module.sweeps : k => {
      id                   = m.id
      assignment           = m.policy_assignment_id
      reference            = m.policy_definition_reference_id
      discovery_mode       = m.resource_discovery_mode
      resource_count       = m.resource_count
      parallel_deployments = m.parallel_deployments
      failure_percentage   = m.failure_percentage
      location_filters     = m.location_filters
    }
  }
}

💡 The discovery mode and every pace control are outputs precisely so this table can be built from state, without querying Azure. That is the point of emitting them.

16 · 🏗️ End-to-end composition
provider "azurerm" {
  features {}
}

data "azurerm_subscription" "current" {}

# 1 · The identity the assignment will use to perform remediation.
module "policy_identity" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-user-assigned-identity.git?ref=v1.0.0"

  name                = "id-policy-remediation"
  resource_group_name = var.governance_resource_group_name
  location            = "eastus2"
}

# 2 · The initiative bundling the remediable policies.
module "baseline_initiative" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-policy-set-definition.git?ref=v1.0.0"

  name         = "platform-baseline"
  display_name = "Platform Baseline Controls"

  policy_definition_references = {
    require-https-storage = { policy_definition_id = var.https_only_builtin_definition_id }
    require-diagnostics   = { policy_definition_id = var.diagnostics_definition_id }
  }
}

# 3 · The assignment — carrying the identity, which is what actually performs changes.
module "baseline_assignment" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-subscription-policy-assignment.git?ref=v1.0.0"

  name                 = "platform-baseline"
  subscription_id      = data.azurerm_subscription.current.id
  policy_definition_id = module.baseline_initiative.id
  location             = "eastus2"

  identity = {
    type         = "UserAssigned"
    identity_ids = [module.policy_identity.id]
  }
}

# 4 · Grant that identity the rights the policy's roleDefinitionIds demand.
module "policy_identity_roles" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-role-assignments.git?ref=v1.0.0"

  scope = data.azurerm_subscription.current.id

  role_assignments = {
    contributor = {
      role_definition_name = "Contributor"
      principal_id         = module.baseline_assignment.identity_principal_id
    }
  }
}

# 5 · The bounded remediation — this module.
module "https_remediation" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-subscription-policy-remediation.git?ref=v1.0.0"

  name                           = "remediate-https-2026-08"
  subscription_id                = data.azurerm_subscription.current.id
  policy_assignment_id           = module.baseline_assignment.id
  policy_definition_reference_id = module.baseline_initiative.policy_definition_reference_ids["require-https-storage"]

  resource_discovery_mode = "ExistingNonCompliant" # the narrower default
  location_filters        = ["eastus2"]
  resource_count          = 50
  parallel_deployments    = 5
  failure_percentage      = 0.1

  timeouts = { create = "2h" }

  # The identity's role assignment must land before the job runs, or it finds no permissions.
  depends_on = [module.policy_identity_roles]
}

💡 This wiring shows the intended dependency order: identity → initiative → assignment → role grant → remediation. Steps 1 and 4 matter more than they look: the remediation borrows the assignment's identity, and policy identity role propagation is asynchronous — the depends_on is what keeps a job from starting before its permissions exist. Output names on sibling modules are illustrative; match them to the versions you pin.


📥 Inputs

Required: name, subscription_id (a Resource ID), policy_assignment_id.

Targeting: policy_definition_reference_id (required when the assignment assigns an initiative).

Discovery (secure default): resource_discovery_mode ("ExistingNonCompliant", validated).

Pace and blast radius (all validated, all emitted): resource_count, parallel_deployments, failure_percentage, location_filters.

Universal tail: timeouts. This resource type does not support tags.

Full object() schemas
variable "name"                 { type = string } # force-new — each run is its own task
variable "subscription_id"      { type = string } # "/subscriptions/<guid>" Resource ID; force-new
variable "policy_assignment_id" { type = string } # DeployIfNotExists or Modify effects only

variable "policy_definition_reference_id" {
  type    = string # required by the service when the assignment assigns an initiative
  default = null
}

variable "resource_discovery_mode" {
  type    = string # ExistingNonCompliant (default, narrower) | ReEvaluateCompliance (re-scans first)
  default = "ExistingNonCompliant"
  # NOTE: this field does not exist on the management-group-scoped remediation resource.
}

variable "location_filters" {
  type    = list(string) # null = every region
  default = null
}

variable "resource_count" {
  type    = number # null = service default (unbounded); validated > 0 when set
  default = null
}

variable "parallel_deployments" {
  type    = number # null = service default; validated > 0 when set
  default = null
}

variable "failure_percentage" {
  type    = number # a FRACTION 0.0–1.0, not a percentage out of 100; validated
  default = null
}

variable "timeouts" {
  type    = object({ create = optional(string), read = optional(string), update = optional(string), delete = optional(string) })
  default = null
}

🧾 Outputs

Output Description Kind
id The Azure Resource ID of the policy remediation task Passthrough
name The name of the remediation task Passthrough
subscription_id The subscription the remediation runs at Passthrough
policy_assignment_id The policy assignment being remediated Passthrough
policy_definition_reference_id The policy definition reference within an initiative being remediated, or null for a standalone definition Passthrough
resource_discovery_mode How resources to remediate were discovered: ExistingNonCompliant (narrower) or ReEvaluateCompliance (re-scans first) Passthrough
resource_count The ceiling on resources this job will act on, or null for the service default (unbounded) Passthrough
parallel_deployments How many resources are remediated concurrently, or null for the service default Passthrough
failure_percentage The failure fraction that aborts the job, or null for the service default (no explicit circuit breaker) Passthrough
location_filters The regions the remediation is restricted to, or null for every region Passthrough
failure_percentage_is_a_fraction Always true, and the name invites exactly the wrong reading Constant
re_evaluates_compliance_before_remediating True when the remediation re-scans compliance before acting, rather than trusting the last evaluation Derived
remediation_is_a_one_shot_job_not_a_standing_policy Always true, and it is the most common misconception about this resource Constant
scope_is_capped_but_not_targeted True when resource_count or parallel_deployments is set Derived

No secret is accepted or emitted. A remediation carries scope, discovery mode, and pace configuration only.

🧠 Architecture Notes

  • This resource acts; the rest of the family reports. A policy definition, initiative, assignment, or exemption changes what Azure says about compliance. A remediation changes the resources themselves, across every resource group in the subscription. Treat a plan that adds one as a change-control event, not a configuration edit.
  • resource_discovery_mode is the scope's defining difference. The management-group-scoped remediation resource has no such field and can only work from already-known non-compliance. Here you can choose, and the choice is a blast-radius decision: ReEvaluateCompliance re-scans first, so the set of resources changed is not predictable from the current compliance report. The module defaults to the narrower ExistingNonCompliant, validates the pair, and emits the chosen mode.
  • Pace controls are blast-radius limits, not tuning. The module deliberately leaves all four at the service default rather than inventing ceilings — a number picked here would be wrong for most callers and would silently truncate real remediations. Instead they are documented as limits and emitted as outputs, so a review can see whether a job was bounded. A null resource_count in state means unbounded; that is the finding, and it is visible without querying Azure.
  • The identity is not this module's. A remediation borrows the assignment's managed identity. The module accepts no identity input on purpose, because modelling one here would imply a permission boundary that does not exist. When a job runs and changes nothing, check that identity's role assignments first.
  • Task semantics, not desired state. Deleting the Terraform resource cancels or removes the task; it does not revert what the task already did. Each run should be its own name-stamped resource so the history is auditable, and name being force-new supports exactly that.
  • Asynchronous completion. Apply succeeds when the job is accepted. Resource compliance follows later, and the create timeout does not cover it — with ReEvaluateCompliance there is an evaluation pass before any work even starts.
  • Ordering against an exemption is not expressible as a reference. An exemption narrows what a remediation reaches, but there is no attribute to reference, so an explicit depends_on is the correct tool in that one case.
  • Remediating a centrally-assigned control is a useful pattern. The policy_assignment_id may live at a management group above this subscription while the remediation is scoped here — which keeps the job's blast radius to one subscription even though the control is enforced tenant-wide.
  • features {} dependence. The module carries no provider {} block. If it appears not to initialize in isolation, the cause is a missing caller-side provider "azurerm" { features {} }.
  • No tags. This resource type has no taggable surface, so the universal tail carries timeouts only.

🧱 Design Principles

Concern Secure default (empty call) Opt-out (caller must type it)
Discovery breadth resource_discovery_mode = "ExistingNonCompliant" set "ReEvaluateCompliance" to re-scan first
Unit correctness on the breaker failure_percentage enforced as a 0.0–1.0 fraction — (no opt-out)
Positive pace values resource_count / parallel_deployments validated > 0 leave null for the service default
Blast-radius visibility discovery mode and every pace control emitted as outputs — (no opt-out; that is the point)
Identity boundary no identity input — the assignment's identity is used — (attach the identity to the assignment)
Run history name force-new, so each run is its own task — (reusing a name hides the sequence)
Long-running jobs timeouts exposed with guidance to raise create accept the defaults

🚀 Runbook

terraform init -backend=false
terraform validate
terraform fmt -check
  • Pin the module with ?ref=v1.0.0 — never a branch.
  • This library is plan-only during authoring; a human runs terraform plan / apply from CI against real credentials. For this module in particular, that human review is the control — the resource makes changes to real infrastructure.

🧪 Testing

  • terraform validate proves the configuration is internally consistent and type-correct against the pinned provider schema, and exercises all four validation {} blocks — the resource_discovery_mode enum, the 0.0–1.0 failure_percentage fraction, and the positive resource_count / parallel_deployments.
  • terraform fmt -check enforces canonical formatting.
  • Neither command calls Azure. Only terraform plan (run by a human, from CI) exercises the ARM API — the module ships without any cloud apply. Whether the assignment's effect is remediable, whether its identity holds the required roles, whether policy_definition_reference_id is required for that assignment, and whether subscription_id is a well-formed Resource ID are all apply-time facts.

💬 Example Output

Apply complete! Resources: 1 added, 0 changed, 0 destroyed.

Outputs:

id                             = "/subscriptions/00000000-0000-0000-0000-000000000000/providers/Microsoft.PolicyInsights/remediations/remediate-https-2026-08"
name                           = "remediate-https-2026-08"
subscription_id                = "/subscriptions/00000000-0000-0000-0000-000000000000"
policy_assignment_id           = "/subscriptions/00000000-0000-0000-0000-000000000000/providers/Microsoft.Authorization/policyAssignments/platform-baseline"
policy_definition_reference_id = "require-https-storage"
resource_discovery_mode        = "ExistingNonCompliant"
resource_count                 = 50
parallel_deployments           = 5
failure_percentage             = 0.1
location_filters               = [
  "eastus2",
]

🔍 Troubleshooting

Symptom Cause Fix
Provider configuration not present / features error No caller-side provider "azurerm" { features {} }. Add the provider block with features {} in the root module.
Apply fails: invalid scope subscription_id was passed as a bare GUID. Use the /subscriptions/<guid> Resource ID form.
Plan error: resource_discovery_mode must be one of A value outside the pair. Use ExistingNonCompliant or ReEvaluateCompliance.
Plan error: failure_percentage must be between 0.0 and 1.0 A value like 10 intending "10%". Pass the fraction — 0.1.
Plan error: resource_count must be greater than 0 A zero or negative ceiling. Leave it null for the service default, or pass a positive number.
Apply succeeds but nothing is remediated The assignment's effect is Audit / Deny, or compliance has not been evaluated yet under ExistingNonCompliant. Confirm the policy uses DeployIfNotExists or Modify; let the initial evaluation run, or use ReEvaluateCompliance with pace controls set.
Job changed far more than expected ReEvaluateCompliance re-scanned and found resources the prior compliance report did not list. Bound the run with resource_count and location_filters; check the resource_discovery_mode output on past runs.
Job starts but every deployment fails The assignment's managed identity lacks the roles the policy's roleDefinitionIds demand. Grant those roles at the target scope, wait for propagation, then create a new task.
Apply fails: policy_definition_reference_id required The assignment assigns an initiative and no reference was given. Set it from the initiative's policy_definition_reference_ids output.
Terraform reports success but resources are still non-compliant The task runs asynchronously; apply confirms acceptance, not completion. Check the remediation's progress in Azure Policy; raise timeouts only if acceptance itself is slow.
Destroying the resource did not revert the changes A remediation is a task, not a desired-state record. Revert with a compensating change; deleting the task only removes the job.
The job acted on resources an exemption was meant to exclude The remediation was created before the exemption existed. Add depends_on for the exemption module — there is no attribute reference to express this ordering.

🔗 Related Docs


💙 "Infrastructure as Code should be standardized, consistent, and secure."