Skip to content

Latest commit

Β 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

☁️ Azure Site Recovery Replication Recovery Plan Terraform Module

Orders a multi-machine failover into sequenced boot groups, with automation and manual gates around each step, targeting hashicorp/azurerm ~> 4.0.

Terraform azurerm Module Type Resources Posture


🧩 Overview

  • 🎬 Manages azurerm_site_recovery_replication_recovery_plan β€” the record that decides in what order an application comes back, and what runs before and after each stage.
  • πŸ”΄ Boot group ORDER is the failover sequence. Group 2 starts only after every machine in group 1 has failed over and started, so the list index is the plan's entire behaviour β€” which is why this module models the groups as a list, not the keyed map it uses for repeated blocks elsewhere.
  • πŸ”΄ The provider documents four conditional requirements and enforces none of them β€” and enforces one prohibition it does not document. This module covers all five at plan time.
  • πŸ”΄ Three documented Azure limits are absent from the schema: at most 7 groups, at most 100 protected items per plan, and no machine in two groups. All three are validated here.
  • ⚠️ A manual action stops the failover until a human acknowledges it, so a plan containing one cannot run unattended.
  • ⚠️ The shutdown step is skipped during a test failover, so a drill never exercises that phase or its actions.
  • ✏️ The groups are editable in place; the identity is not. Refining a sequence after a drill neither replaces the plan nor disturbs the protected items it references.

πŸ’‘ Why it matters: a recovery plan's value is entirely in its sequence, and a sequence is not readable from a plan diff. Terraform will happily show three groups changing without saying that the change alters the order an application recovers in. So this module derives the shape β€” items per group in order, which actions are manual, which never run in a drill, which never run on failback β€” and validates the four documented requirements and three documented limits that the provider leaves to Azure.


❀️ Support this project

If this module saved you time:


πŸ—ΊοΈ Where this fits in the family

This module is terraform-azurerm-site-recovery-replication-recovery-plan, one of the fourteen blue nodes β€” all fourteen authored site_recovery_* modules share this diagram. The three dark nodes are not modules: each is a single ARM resource type that several of the blue modules write, two of them by three resource types each.

flowchart TB
  RG["terraform-azurerm-resource-group"]
  VAULT["terraform-azurerm-recovery-services-vault"]
  FABRIC["terraform-azurerm-site-recovery-fabric"]
  HVS["terraform-azurerm-site-recovery-services-vault-hyperv-site"]
  PC["terraform-azurerm-site-recovery-protection-container"]
  POLICY["terraform-azurerm-site-recovery-replication-policy"]
  HVPOL["terraform-azurerm-site-recovery-hyperv-replication-policy"]
  VMPOL["terraform-azurerm-site-recovery-vmware-replication-policy"]
  PCM["terraform-azurerm-site-recovery-protection-container-mapping"]
  HVASSOC["terraform-azurerm-site-recovery-hyperv-replication-policy-association"]
  VMASSOC["terraform-azurerm-site-recovery-vmware-replication-policy-association"]
  NM["terraform-azurerm-site-recovery-network-mapping"]
  HVNM["terraform-azurerm-site-recovery-hyperv-network-mapping"]
  RVM["terraform-azurerm-site-recovery-replicated-vm"]
  RRP["terraform-azurerm-site-recovery-replication-recovery-plan"]
  ARMF["ONE ARM type, TWO resource types: vaults replicationFabrics. Same name plus same vault means one object, silently"]
  ARMP["ONE ARM type, THREE resource types: vaults replicationPolicies. A2A, Hyper-V and VMware policies are interchangeable at the API"]
  ARMM["ONE ARM type, THREE resource types: replicationProtectionContainerMappings. The hardest group to assert on"]
  VM["terraform-azurerm-linux-virtual-machine or windows-virtual-machine: wire its id, NOT its virtual_machine_id"]
  VNET["terraform-azurerm-virtual-network: the failover network, the test network and the source network"]
  RUNBOOK["an Azure Automation runbook: a plan action names it by id and it runs with the automation account's own permissions. NOT bounded by vault RBAC"]
  VMWVM["terraform-azurerm-site-recovery-vmware-replicated-vm"]
  HOSTS["the Hyper-V hosts themselves, registered with a five-day vault key. NO azurerm resource does this"]
  APPL["the VMware replication appliance, deployed on-premises and registered against the vault. NO azurerm resource does this either"]
  VMM["the System Center VMM server, registered by installing the Site Recovery Provider on it. NO azurerm resource does this either"]
  VCENTER["the vCenter machines and the credentials stored on the appliance. source_vm_name and physical_server_credential_name are FRIENDLY names matched at apply time, and that credential is root or admin ON the source machine"]

  RG -->|"resource_group_name and location"| VAULT
  VAULT -->|"name, NOT id, so there is no dependency edge from a literal"| FABRIC
  VAULT -->|"id, NOT name: this sibling takes the vault ID instead"| HVS
  VAULT -->|"name, and NO fabric: a policy belongs to the vault"| POLICY
  VAULT -->|"id, a second convention on the same parent"| HVPOL
  VAULT -->|"id, and the association takes the SAME vault argument"| VMPOL
  VAULT -->|"id: BOTH the fabric and the container are discovered under it at apply time"| VMASSOC
  VAULT -->|"name plus resource group, a third convention again"| RVM
  VAULT -->|"id: the ONLY Resource ID this module resolves anything from"| HVNM
  VAULT -->|"id, and BOTH fabrics as ids: a fourth convention"| RRP
  FABRIC -->|"hardcoded instance type: an Azure fabric"| ARMF
  HVS -->|"hardcoded instance type: HyperVSite, absent from Swagger"| ARMF
  POLICY -->|"hardcoded A2A payload"| ARMP
  HVPOL -->|"hardcoded Hyper-V to Azure payload"| ARMP
  VMPOL -->|"hardcoded InMageRcm payload: the MODERNIZED VMware provider"| ARMP
  PCM -->|"container taken by name and by id"| ARMM
  HVASSOC -->|"container discovered under the fabric at apply time"| ARMM
  VMASSOC -->|"fabric AND container both discovered under the vault"| ARMM
  FABRIC -->|"name, so a Hyper-V site's name is equally legal here"| PC
  FABRIC -->|"name, for BOTH sides: ARM names, not friendly names"| NM
  FABRIC -->|"id, for BOTH sides, and neither is compared to the vault"| RRP
  PC -->|"name for the source side, id for the target side"| PCM
  POLICY -->|"id, where the container beside it is taken by name"| PCM
  HVPOL -->|"id, the only consumer of a Hyper-V policy"| HVASSOC
  HVS -->|"id, and the protection container is discovered from it"| HVASSOC
  VMPOL -->|"id, the only consumer of a VMware policy"| VMASSOC
  VNET -->|"id, twice. Only the SOURCE id is parsed by the provider"| NM
  VNET -->|"id, the target only. The SOURCE network is a VMM friendly name"| HVNM
  FABRIC -->|"source fabric by NAME, target fabric by ID: one resource, two conventions"| RVM
  PC -->|"source container by NAME, target container by ID"| RVM
  POLICY -->|"id, and a Hyper-V or VMware policy id would also be accepted"| RVM
  VM -->|"the machine being protected"| RVM
  VNET -->|"the failover network and the test network, both Optional AND Computed"| RVM
  PCM -->|"makes a container pair eligible before any item can replicate"| RVM
  NM -->|"decides where an Azure-to-Azure failover lands"| RVM
  HVNM -->|"decides where a VMM Hyper-V failover lands, for every machine on the source network"| VMM
  RVM -->|"ids, into ORDERED boot groups: group 2 starts only after group 1 has finished"| RRP
  VMWVM -->|"ids, into the same ORDERED boot groups as the Azure-to-Azure items"| RRP
  VAULT -->|"id ALONE, a fifth convention: the fabric AND the container are discovered from it, and there must be exactly ONE container"| VMWVM
  VMPOL -->|"id, its second consumer: the policy this VMware item replicates under"| VMWVM
  VNET -->|"the failover network and the test network. test_network_id is read back and NEVER updated"| VMWVM
  VMASSOC -->|"makes the discovered container eligible before any VMware machine can replicate"| VMWVM
  RUNBOOK -->|"id, from a plan pre-action or post-action"| RRP
  APPL -->|"appliance_name is its FRIENDLY name, matched against the fabric process servers. Microsoft: it cannot be changed once set"| VMWVM
  VCENTER -->|"source_vm_name and physical_server_credential_name, both matched by name against records the appliance discovered"| VMWVM
  HOSTS -->|"register into the site out of band, invisible to Terraform"| HVS
  HOSTS -->|"registration is what creates the container this record needs"| HVASSOC
  APPL -->|"required before any VMware machine can be enrolled"| VMASSOC
  VMM -->|"the fabric AND the source network are found under it by FRIENDLY NAME at apply time"| HVNM

  classDef me fill:#0078D4,stroke:#004578,color:#ffffff
  classDef keystone fill:#004578,stroke:#002438,color:#ffffff
  classDef sibling fill:#F3F6F9,stroke:#8A9BA8,color:#1B1F23
  class FABRIC,HVS,PC,POLICY,HVPOL,VMPOL,PCM,HVASSOC,VMASSOC,NM,HVNM,RVM,RRP,VMWVM me
  class ARMF,ARMP,ARMM keystone
  class RG,VAULT,VM,VNET,RUNBOOK,VCENTER,HOSTS,APPL,VMM sibling
Loading

This module sits at the end of the Azure-to-Azure chain: it consumes the replicated items terraform-azurerm-site-recovery-replicated-vm creates and orders them. Note the grey RUNBOOK node β€” a plan action can name an Automation runbook, and what that runbook does is bounded by the Automation account's permissions rather than by anything in this plan.


🧬 What this module builds

flowchart TB
  IN_NAME["name: force-new, and the provider's own validator is UNANCHORED at the start"]
  IN_VAULT["recovery_vault_id: split into three values. Anchored, because a fabric ID extends it"]
  IN_FABS["source_recovery_fabric_id and target_recovery_fabric_id: neither is compared to the vault, or to each other"]
  IN_BOOT["boot_recovery_groups: an ORDERED LIST. Up to 7 groups, up to 100 items, no item in two groups"]
  IN_FG["failover_recovery_group: exactly one, actions only, brackets the whole plan's failover"]
  IN_SG["shutdown_recovery_group: exactly one, actions only. The shutdown step is SKIPPED in a test failover"]
  IN_A2A["azure_to_azure_settings: optional, all fields force-new. Both pairings enforced by the provider"]
  THIS["azurerm_site_recovery_replication_recovery_plan.this"]
  RULES["FOUR documented conditional requirements the provider enforces NOWHERE, plus ONE prohibition it enforces and does not document"]
  ORDER["group order IS the failover sequence: group 2 starts only after every machine in group 1 has started"]
  OUT_PATHS["four comparable paths, including this record's own FULL path"]
  OUT_SHAPE["the plan's shape: group count, items per group, action names by type, what runs in a drill, what runs on failback"]
  OUT_CONST["nine constants: a manual action pauses the job, three unenforced Azure limits, a destroy that removes only the ordering"]

  IN_NAME --> THIS
  IN_VAULT --> THIS
  IN_FABS --> THIS
  IN_BOOT --> THIS
  IN_FG --> THIS
  IN_SG --> THIS
  IN_A2A --> THIS
  IN_BOOT -->|"list index, not a map key"| ORDER
  IN_BOOT --> RULES
  IN_FG -->|"same seven action rules, repeated because a validation sees only its own variable"| RULES
  IN_SG --> RULES
  IN_VAULT --> OUT_PATHS
  IN_FABS --> OUT_PATHS
  IN_NAME --> OUT_PATHS
  IN_BOOT --> OUT_SHAPE
  IN_FG --> OUT_SHAPE
  IN_SG --> OUT_SHAPE
  THIS --> OUT_CONST

  classDef me fill:#0078D4,stroke:#004578,color:#ffffff
  classDef keystone fill:#004578,stroke:#002438,color:#ffffff
  classDef sibling fill:#F3F6F9,stroke:#8A9BA8,color:#1B1F23
  class THIS keystone
  class RULES,ORDER me
  class IN_NAME,IN_VAULT,IN_FABS,IN_BOOT,IN_FG,IN_SG,IN_A2A,OUT_PATHS,OUT_SHAPE,OUT_CONST sibling
Loading
Resource Cardinality Note
azurerm_site_recovery_replication_recovery_plan single (this) The keystone. Five nested block families, three of them action-bearing.

A standalone module: one keystone resource, no for_each children. Four top-level arguments plus five blocks.


βœ… Provider / Versions

Item Value
Terraform >= 1.12.0
Provider hashicorp/azurerm ~> 4.0 (validated against 4.81.0)
Provider block None here. The caller configures provider "azurerm" { features {} }, auth and subscription.
ARM API version Microsoft.RecoveryServices 2024-04-01
tags Not supported β€” this resource has no tags and no location. Tag the vault instead.

Schema notes that bite

  • πŸ”΄ boot_recovery_group is a TypeList with MinItems: 1 and NO maximum. The order is the failover sequence and Azure's documented cap is seven.
  • πŸ”΄ Four documented -> notes are enforced nowhere: fabric_location required for runbook and script actions, runbook_id for runbook actions, manual_action_instruction for manual actions, script_path for script actions. All four fields are plain Optional, and an omitted value is transmitted as an empty string.
  • πŸ”΄ One prohibition IS enforced and is not documented: fabric_location must not be set for a manual action, raised as a Go error inside the expand β€” at apply time.
  • πŸ”΄ The name validator is missing its leading anchor. StringMatch(regexp.MustCompile("[a-zA-Z][a-zA-Z0-9-]{1,63}[a-zA-Z0-9]$"), …) matches a suffix, while the message it carries insists the name "should start with a letter".
  • ⚠️ replicated_protected_items and runbook_id are validated by azure.ValidateResourceID β€” a shape check that accepts any Resource ID. This module's anchored rules are the only type checks on either.
  • ⚠️ Force-new: name, recovery_vault_id, both fabric IDs, and every field inside azure_to_azure_settings. The three groups and all their actions are not, and timeouts has four keys.
  • ⚠️ Neither fabric argument is compared to the vault, or to the other. A plan spanning two vaults, or pointing both sides at one fabric, is accepted.
  • βœ… Both azure_to_azure_settings pairings ARE enforced by RequiredWith at plan time β€” the one place in this resource where a documented rule is backed by the schema.
  • ⚠️ failover_recovery_group and shutdown_recovery_group are MinItems: 1, MaxItems: 1 and carry no protected items β€” actions only.
  • ⚠️ There is no SchemaVersion and no state upgrader.

πŸ”‘ Required Azure RBAC Roles / Permissions

Principal Scope Requirement
The principal running Terraform The vault or its resource group Site Recovery Contributor β€” it carries Microsoft.RecoveryServices/vaults/replicationRecoveryPlans/*
The same principal The replicated protected items named in the boot groups Read. They are referenced by Resource ID and must be resolvable
The Automation account behind any runbook action Wherever the runbook acts Whatever the runbook needs. A runbook action names a runbook by ID; what it can do when the plan runs is decided by the Automation account's own identity, not by this plan's permissions or the vault's
Whoever runs the failover The vault Site Recovery Operator or higher, plus the ability to acknowledge manual actions
Alternatively The vault or its resource group Contributor or Owner

The third row is the one to read twice. A recovery plan is a mechanism for running arbitrary automation at the worst possible moment, and nothing in Azure RBAC on the vault bounds what a named runbook does. Reviewing a plan means reviewing the runbooks it names β€” runbook_actions emits their names for exactly that purpose.

Backup Contributor is the wrong role and it is a plausible mistake, because one Recovery Services vault serves both Azure Backup and Site Recovery. It grants nothing over replicationRecoveryPlans.

Plan access is not credential access. This module reads and holds no secret, and no output is sensitive.


Azure Prerequisites

  1. Microsoft.RecoveryServices registered in the subscription.
  2. An existing Recovery Services vault, and a fabric on each side, referenced here by Resource ID.
  3. Replicated protected items that already exist and are protected. A plan references them; it creates no protection. Each needs a protection container on both sides, a container mapping and a replication policy.
  4. An Azure Automation account and runbooks, for any runbook action. Microsoft's guidance lists the Azure modules the account must carry.
  5. A System Center VMM server, for any script action β€” Microsoft states a script can be added on the primary site only where VMM is deployed, so an Azure-to-Azure plan uses runbooks instead.
  6. At most 7 boot groups and 100 protected items, and no machine in two groups. Documented by Microsoft, enforced by this module, absent from the provider's schema.

πŸ“ Module Structure

terraform-azurerm-site-recovery-replication-recovery-plan/
β”œβ”€β”€ providers.tf   # required_version and the pinned azurerm; no provider block
β”œβ”€β”€ variables.tf   # 9 variables, 36 validations
β”œβ”€β”€ main.tf        # the keystone `this`, five dynamic block families, and the plan's derived shape
β”œβ”€β”€ outputs.tf     # 45 outputs: 5 passthrough, 31 derived, 9 constant
β”œβ”€β”€ README.md      # this file
β”œβ”€β”€ SCOPE.md       # the cross-module contract
β”œβ”€β”€ LICENSE        # MIT
└── .gitignore

βš™οΈ Quick Start

module "dr_rg" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-dr-prod"
  location = "eastus"
}

module "dr_vault" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-recovery-services-vault.git?ref=v1.0.0"

  name                = "rsv-dr-prod"
  resource_group_name = module.dr_rg.name
  location            = module.dr_rg.location
  sku                 = "Standard"
}

module "fabric_primary" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-fabric.git?ref=v1.0.0"

  name                = "fabric-eastus"
  location            = "eastus"
  recovery_vault_name = module.dr_vault.name
  resource_group_name = module.dr_rg.name
}

module "fabric_secondary" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-fabric.git?ref=v1.0.0"

  name                = "fabric-westus2"
  location            = "westus2"
  recovery_vault_name = module.dr_vault.name
  resource_group_name = module.dr_rg.name
}

module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-app-tier"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  # ORDER IS THE SEQUENCE. The database boots first, then everything that depends on it.
  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_db.id] },
    { replicated_protected_items = [module.protected_web.id] },
  ]
}

πŸ’‘ The smallest legal call needs one boot group and nothing else β€” the failover_recovery_group and shutdown_recovery_group blocks the provider requires are rendered empty by this module's {} defaults. Such a plan automates nothing and still does the most valuable thing a plan does: it orders the failover.

ℹ️ The caller configures the provider, its authentication and its features {} block. This module declares none of those.


πŸ”Œ Cross-Module Contract

Consumes

Input Type Source module
recovery_vault_id string terraform-azurerm-recovery-services-vault output id
source_recovery_fabric_id string terraform-azurerm-site-recovery-fabric output id
target_recovery_fabric_id string terraform-azurerm-site-recovery-fabric output id (recovery region)
boot_recovery_groups[*].replicated_protected_items list(string) terraform-azurerm-site-recovery-replicated-vm output id
action runbook_id string an azurerm_automation_runbook resource β€” there is no runbook module in this library

Emits

Output Description Consumed by
id The plan's Resource ID (first) a management lock, an import, or a report
recovery_plan_path_within_the_subscription This record's full plan-time path check blocks asserting two plans are distinct
source_fabric_path_within_the_subscription, target_fabric_path_within_the_subscription Both fabrics' paths check blocks against either fabric-creating module
vault_path_within_the_subscription The vault's path check blocks asserting one topology, one vault
protected_items_per_boot_group The sequence shape, in order plan review; the value a diff cannot show
manual_actions, runbook_actions, script_actions Action names by type review of what runs, and of what pauses
actions_that_do_not_run_during_a_test_failover The steps a drill never exercises check blocks, drill review
nine constants The facts that produce no error when they bite documentation, check blocks

Nothing in this library consumes this module. A recovery plan is the end of the chain.


πŸ“š Example Library

1 Β· The smallest real call
module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-app-tier"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_db.id] },
  ]
}

πŸ’‘ One group, one machine, no automation. plan_has_no_actions reports it as a decision rather than an unfinished configuration β€” a pure sequencing plan is a legitimate and common thing to own.

ℹ️ The failover_recovery_group and shutdown_recovery_group blocks are required by the provider and rendered empty here by this module's {} defaults, so the smallest call does not have to mention them.

2 Β· A three-tier application, in order
locals {
  # Resource IDs of machines already protected by terraform-azurerm-site-recovery-replicated-vm.
  # Example 15 shows that wiring end to end; the examples here keep the focus on the sequence.
  middleware_item_ids = [
    "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod/replicationFabrics/fabric-eastus/replicationProtectionContainers/container-eastus/replicationProtectedItems/vm-app-01",
    "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod/replicationFabrics/fabric-eastus/replicationProtectionContainers/container-eastus/replicationProtectedItems/vm-app-02",
  ]
  cache_item_id = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod/replicationFabrics/fabric-eastus/replicationProtectionContainers/container-eastus/replicationProtectedItems/vm-cache-01"
}

module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-three-tier"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    # Group 1: the database. Nothing else starts until it has.
    { replicated_protected_items = [module.protected_db.id] },

    # Group 2: middleware, in parallel with each other.
    { replicated_protected_items = local.middleware_item_ids },

    # Group 3: the web tier, last, so nobody reaches the app before it is ready.
    { replicated_protected_items = [module.protected_web.id] },
  ]
}

output "the_sequence" {
  value = module.app_recovery_plan.protected_items_per_boot_group
}

πŸ”΄ The list order IS the behaviour. Microsoft: "Machines in different groups fail over in group order, so that Group 2 machines start their failover only after all the machines in Group 1 have failed over and started", and "machines in a single group fail over in parallel".

⚠️ So reordering this list changes what the plan does while changing nothing else about it β€” and a Terraform diff will show three groups changing without saying that. protected_items_per_boot_group emits [1, 2, 1] so a review can read the shape at a glance.

3 Β· A runbook pre-action, and what bounds it
resource "azurerm_automation_account" "dr" {
  name                = "aa-dr-prod"
  resource_group_name = module.dr_rg.name
  location            = module.dr_rg.location
  sku_name            = "Basic"
}

resource "azurerm_automation_runbook" "failover_ag" {
  name                    = "ASR-SQL-FailoverAG"
  resource_group_name     = module.dr_rg.name
  location                = module.dr_rg.location
  automation_account_name = azurerm_automation_account.dr.name
  runbook_type            = "PowerShell"
  log_verbose             = true
  log_progress            = true
  content                 = file("${path.module}/runbooks/failover-ag.ps1")
}

module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-three-tier"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    {
      replicated_protected_items = [module.protected_db.id]

      pre_actions = [
        {
          name                 = "failover-sql-availability-group"
          type                 = "AutomationRunbookActionDetails"
          fail_over_directions = ["PrimaryToRecovery", "RecoveryToPrimary"]
          fail_over_types      = ["PlannedFailover", "TestFailover", "UnplannedFailover"]
          fabric_location      = "Recovery"
          runbook_id           = azurerm_automation_runbook.failover_ag.id
        },
      ]
    },
    { replicated_protected_items = [module.protected_web.id] },
  ]
}

πŸ”’ runbook_id and fabric_location are both documented as required for this action type, and the provider enforces neither. Omit runbook_id and an empty string is sent to Azure. This module rejects both omissions at plan time β€” and anchors runbook_id to an Automation runbook path, because the provider's own validator accepts any Resource ID at all.

⚠️ What this runbook can do is not bounded by the vault's RBAC. It runs with the Automation account's own identity. Reviewing a recovery plan means reviewing the runbooks it names; runbook_actions emits their names for that.

πŸ’‘ There is no runbook module in this library, so the runbook is a plain resource here β€” which is what a caller's root module would hold anyway.

4 Β· A manual gate, and the pause it creates
module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-with-gate"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_db.id] },
  ]

  failover_recovery_group = {
    pre_actions = [
      {
        name                      = "confirm-dns-cutover"
        type                      = "ManualActionDetails"
        fail_over_directions      = ["PrimaryToRecovery"]
        fail_over_types           = ["PlannedFailover", "UnplannedFailover"]
        manual_action_instruction = "Confirm with the network team that the DNS TTL has expired before continuing."
        # NOTE: fabric_location must NOT be set for a manual action.
      },
    ]
  }
}

check "the_plan_can_run_unattended" {
  assert {
    condition     = !module.app_recovery_plan.plan_pauses_for_a_manual_action
    error_message = "This plan contains a manual action, so a failover will stop and wait for a human. Remove the manual action, or accept that the recovery time objective includes however long someone takes to notice."
  }
}

πŸ”΄ A manual action stops the failover job. Microsoft: "if you add a manual action, when the recovery plan runs, it stops at the point at which you inserted the manual action. A dialog box prompts you to specify that the manual action was completed."

⚠️ So the plan's real recovery time objective includes human reaction time, and a drill that measured it with someone watching will understate the 3am case. The assertion above is for plans that must run unattended; delete it for plans where the gate is deliberate.

πŸ”’ manual_action_instruction is documented as required for this type and enforced nowhere β€” an omitted value sends an empty description, leaving an operator paused at a step with nothing to read. This module rejects that. And fabric_location must not be set here: that is the one rule the provider enforces, at apply time, without documenting it.

5 Β· The four requirements the provider documents and does not enforce
# EVERY ONE of these actions is accepted by the provider at plan time and rejected by this module.
module "app_recovery_plan_bad_actions" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-bad"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    {
      replicated_protected_items = [module.protected_db.id]

      pre_actions = [
        # 1. a runbook action with no runbook_id -> an empty string is sent to Azure
        {
          name                 = "no-runbook"
          type                 = "AutomationRunbookActionDetails"
          fail_over_directions = ["PrimaryToRecovery"]
          fail_over_types      = ["TestFailover"]
          fabric_location      = "Recovery"
        },
        # 2. a script action with no script_path -> an empty path is sent
        {
          name                 = "no-path"
          type                 = "ScriptActionDetails"
          fail_over_directions = ["PrimaryToRecovery"]
          fail_over_types      = ["TestFailover"]
          fabric_location      = "Primary"
        },
        # 3. a manual action with no instruction -> an empty description is sent
        {
          name                 = "no-instruction"
          type                 = "ManualActionDetails"
          fail_over_directions = ["PrimaryToRecovery"]
          fail_over_types      = ["TestFailover"]
        },
        # 4. a runbook action with no fabric_location
        {
          name                 = "no-location"
          type                 = "AutomationRunbookActionDetails"
          fail_over_directions = ["PrimaryToRecovery"]
          fail_over_types      = ["TestFailover"]
          runbook_id           = azurerm_automation_runbook.failover_ag.id
        },
      ]
    },
  ]
}

πŸ”΄ All four are documented requirements in the provider's own notes, and all four fields are plain Optional in the schema with no RequiredWith, no CustomizeDiff and no check in the create or update body. An omitted value is transmitted rather than refused.

⚠️ The inverse is the only thing the provider does check: fabric_location must not be set for a manual action, raised as a Go error inside its expand β€” at apply, and documented nowhere. So it documents four requirements it does not enforce and enforces one prohibition it does not document. This module covers all five at plan time.

6 Β· The three Azure limits the schema does not express
locals {
  # Eight already-protected items, by Resource ID. The point of this example is the group COUNT.
  eight_item_ids = [
    for n in range(1, 9) :
    format(
      "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod/replicationFabrics/fabric-eastus/replicationProtectionContainers/container-eastus/replicationProtectedItems/vm-%02d",
      n
    )
  ]
}

# Each of these is accepted by the provider and rejected by this module.
module "plan_too_many_groups" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-eight-groups"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  # EIGHT groups. Microsoft documents a maximum of seven; the provider sets no maximum.
  boot_recovery_groups = [for id in local.eight_item_ids : { replicated_protected_items = [id] }]
}

module "plan_duplicated_machine" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-duplicate"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  # The SAME machine in two groups: two contradictory positions in one sequence.
  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_db.id] },
    { replicated_protected_items = [module.protected_db.id] },
  ]
}

πŸ”’ Three documented limits, none in the schema: at most seven groups, at most 100 protected items in one plan, and "a machine or replication group can only belong to one group in a recovery plan."

⚠️ The duplicate rule is the sharpest of the three, because group order is the failover sequence: a machine listed twice has two contradictory positions in it, and neither the provider nor the plan output would say so.

πŸ’‘ If you genuinely need more than seven stages, model the extra ones as pre-actions or post-actions on an existing group rather than as groups.

7 Β· An empty boot group is a pause point, not a mistake
module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-with-pause"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_db.id] },

    # No machines: a stage that exists purely to run an action between two tiers.
    {
      pre_actions = [
        {
          name                      = "verify-database-is-serving"
          type                      = "ManualActionDetails"
          fail_over_directions      = ["PrimaryToRecovery"]
          fail_over_types           = ["PlannedFailover", "TestFailover", "UnplannedFailover"]
          manual_action_instruction = "Confirm the database accepts connections before the app tier starts."
        },
      ]
    },

    { replicated_protected_items = local.middleware_item_ids },
  ]
}

output "which_groups_are_empty" {
  value = module.app_recovery_plan.boot_groups_without_protected_items
}

πŸ’‘ replicated_protected_items is optional, so an empty group is legal β€” and it is how a plan carries a stage that is purely an action. boot_groups_without_protected_items emits the indexes rather than a count, because in an ordered sequence which group is empty is the point: index 1 here is a deliberate gate between two tiers, whereas an empty last group is usually a leftover.

8 Β· What a drill never exercises
output "untested_steps" {
  value = {
    never_in_a_drill   = module.app_recovery_plan.actions_that_do_not_run_during_a_test_failover
    only_in_a_drill    = module.app_recovery_plan.actions_that_only_run_during_a_test_failover
    shutdown_drill_only = module.app_recovery_plan.shutdown_actions_that_only_run_during_a_test_failover
  }
}

check "every_action_is_exercised_by_a_drill" {
  assert {
    condition     = length(module.app_recovery_plan.actions_that_do_not_run_during_a_test_failover) == 0
    error_message = "At least one action omits TestFailover from fail_over_types, so a test failover never exercises it. That is how a plan passes its rehearsal and fails the real event."
  }
}

⚠️ A test failover is the only way to verify a plan without an outage β€” Microsoft recommends measuring a plan's recovery time objective by running one. An action excluded from it is untested by construction, and the exclusion is a single missing string in a set.

πŸ”΄ The third value is narrower and sharper: the shutdown step is skipped during a test failover. Microsoft: "A shutdown step attempts to turn off the on-premises machines. The exception is if you run a test failover, in which case the primary site continues to run." So a shutdown-group action scoped only to TestFailover belongs to a phase the drill does not perform.

9 Β· Failback is a direction, not a second plan
output "failback_coverage" {
  value = {
    automates_failback  = module.app_recovery_plan.plan_automates_failback
    one_way_actions     = module.app_recovery_plan.actions_that_do_not_run_on_failback
  }
}

πŸ’‘ Microsoft notes that recovery plans "can be used for both failover to and failback from Azure", and the direction is RecoveryToPrimary. So the same plan serves the return journey β€” and a plan whose every action omits that direction automates the outward trip only, while looking complete.

ℹ️ plan_automates_failback is true when at least one action runs on failback. It is deliberately not an assertion: a plan that only ever fails forward is a legitimate design, and the flag exists so that choice is visible rather than accidental.

10 Β· Azure-to-Azure zones, and the pairing the provider does enforce
module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-zonal"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_db.id] },
  ]

  azure_to_azure_settings = {
    primary_zone  = "1"
    recovery_zone = "2"
  }
}

βœ… This is the one place in this resource where a documented rule is backed by the schema. primary_zone and recovery_zone each carry a RequiredWith naming the other, as do primary_edge_zone and recovery_edge_zone β€” so the provider rejects a half-specified pair at plan time with a message that names the missing attribute. This module therefore adds no pairing rule, and records that absence as examined.

⚠️ The contrast with the action rules is the point: four documented requirements on the action block are enforced nowhere, and these two are enforced properly. Same resource, same documentation style, opposite outcomes.

ℹ️ All four fields are force-new, and the force-new is not visible on the arguments. The module's rules here check the zone shape only β€” which zones exist depends on region and capacity, and this suite does not approximate an unpublished set.

11 Β· A script action in an Azure-to-Azure plan
module "app_recovery_plan_contradiction" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-script-in-a2a"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    {
      replicated_protected_items = [module.protected_db.id]

      pre_actions = [
        {
          name                 = "vmm-script"
          type                 = "ScriptActionDetails"
          fail_over_directions = ["PrimaryToRecovery"]
          fail_over_types      = ["PlannedFailover"]
          fabric_location      = "Primary"
          script_path          = "C:\\ASR\\failover.ps1"
        },
      ]
    },
  ]

  azure_to_azure_settings = {
    primary_zone  = "1"
    recovery_zone = "2"
  }
}

output "contradictory_actions" {
  value = module.app_recovery_plan_contradiction.script_actions_in_an_azure_to_azure_plan
}

⚠️ Microsoft's scenario table gives Runbook for Azure-to-Azure failover and failback, and Script only where System Center VMM is involved β€” the same article stating you can add a script on the primary site "only if you have a VMM server deployed".

πŸ’‘ So a script action beside azure_to_azure_settings references a capability that topology does not have. Reported rather than rejected: the provider accepts the combination, and a module cannot see which scenario a vault is actually configured for beyond what these arguments say. Use AutomationRunbookActionDetails for an Azure-to-Azure plan.

12 Β· The comparisons the provider never makes
check "the_plan_is_wired_to_its_own_topology" {
  assert {
    condition = alltrue([
      module.app_recovery_plan.source_fabric_is_in_the_plans_vault,
      module.app_recovery_plan.target_fabric_is_in_the_plans_vault,
      !module.app_recovery_plan.source_and_target_fabrics_are_the_same,
    ])
    error_message = "The plan's fabrics do not line up: check that both are in the plan's own vault, and that the source and target are not the same fabric."
  }
}

check "the_plan_orders_the_items_this_composition_declared" {
  assert {
    condition = (
      module.app_recovery_plan.source_fabric_path_within_the_subscription
      == module.fabric_primary.fabric_path_within_the_subscription
    )
    error_message = "The plan's source fabric is not the one this composition declares."
  }
}

⚠️ Nothing in the provider compares either fabric to the vault, or the two fabrics to each other. A plan spanning two vaults is accepted; so is one whose source and target are the same fabric β€” which cannot survive the event the plan exists for.

πŸ”’ The second assertion works because the fabric path is emitted in an identical form by both fabric-creating modules and by this one; this batch re-proved it byte-identical across four modules from three argument conventions.

13 Β· Refining the sequence after a drill
module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-three-tier"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  # The drill showed the cache had to be up before the app tier. Moving it is an in-place UPDATE.
  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_db.id, local.cache_item_id] },
    { replicated_protected_items = local.middleware_item_ids },
    { replicated_protected_items = [module.protected_web.id] },
  ]
}

βœ… The groups and their actions are editable in place. That is what makes a recovery plan practical to own in Terraform: the sequence is exactly the part that gets refined after every drill, and refining it neither replaces the plan nor disturbs the protected items it references.

⚠️ The other half is the trap. Correcting the plan's name, its vault or either fabric replaces the plan β€” and Microsoft's guidance is that Azure's backend uses the recovery plan name to identify a failover operation, so a replacement after a failover is not a cosmetic change.

14 Β· What a destroy does here β€” and what it does not
output "what_destroy_means" {
  value = module.app_recovery_plan.a_plan_orders_a_failover_and_creates_no_protection
}

πŸ’‘ Destroying a recovery plan removes the ORDERING and leaves every machine still protected. The plan references replicated protected items; it does not own them.

⚠️ That is the exact opposite of terraform-azurerm-site-recovery-replicated-vm, where a destroy disables protection and discards the replica. Two adjacent modules in one family, two opposite meanings for terraform destroy β€” worth knowing before running one against a composition that contains both.

15 Β· πŸ—οΈ End-to-end composition β€” an ordered Azure-to-Azure recovery
locals {
  platform_tags = {
    workload = "disaster-recovery"
    managed  = "terraform"
  }
}

module "dr_rg" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-dr-prod"
  location = "eastus"

  tags = local.platform_tags
}

module "app_rg_secondary" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-app-westus2"
  location = "westus2"

  tags = local.platform_tags
}

module "dr_vault" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-recovery-services-vault.git?ref=v1.0.0"

  name                = "rsv-dr-prod"
  resource_group_name = module.dr_rg.name
  location            = module.dr_rg.location
  sku                 = "Standard"

  tags = local.platform_tags
}

module "fabric_primary" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-fabric.git?ref=v1.0.0"

  name                = "fabric-eastus"
  location            = "eastus"
  recovery_vault_name = module.dr_vault.name
  resource_group_name = module.dr_rg.name
}

module "fabric_secondary" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-fabric.git?ref=v1.0.0"

  name                = "fabric-westus2"
  location            = "westus2"
  recovery_vault_name = module.dr_vault.name
  resource_group_name = module.dr_rg.name
}

module "container_primary" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-protection-container.git?ref=v1.0.0"

  name                 = "container-eastus"
  recovery_fabric_name = module.fabric_primary.name
  recovery_vault_name  = module.dr_vault.name
  resource_group_name  = module.dr_rg.name
}

module "container_secondary" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-protection-container.git?ref=v1.0.0"

  name                 = "container-westus2"
  recovery_fabric_name = module.fabric_secondary.name
  recovery_vault_name  = module.dr_vault.name
  resource_group_name  = module.dr_rg.name
}

module "dr_policy" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-policy.git?ref=v1.0.0"

  name                = "policy-24h"
  recovery_vault_name = module.dr_vault.name
  resource_group_name = module.dr_rg.name

  recovery_point_retention_in_minutes                  = 1440
  application_consistent_snapshot_frequency_in_minutes = 240
}

module "dr_mapping" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-protection-container-mapping.git?ref=v1.0.0"

  name                 = "mapping-eastus-to-westus2"
  recovery_fabric_name = module.fabric_primary.name
  recovery_vault_name  = module.dr_vault.name
  resource_group_name  = module.dr_rg.name

  recovery_source_protection_container_name = module.container_primary.name
  recovery_target_protection_container_id   = module.container_secondary.id
  recovery_replication_policy_id            = module.dr_policy.id
}

# The protected machines. Each is a module of its own; two shown, both abbreviated to the required arguments.
module "protected_db" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replicated-vm.git?ref=v1.0.0"

  name                = "vm-db-01"
  resource_group_name = module.dr_rg.name
  recovery_vault_name = module.dr_vault.name

  source_recovery_fabric_name               = module.fabric_primary.name
  source_recovery_protection_container_name = module.container_primary.name
  source_vm_id                              = module.db_vm.id

  recovery_replication_policy_id           = module.dr_policy.id
  target_recovery_fabric_id               = module.fabric_secondary.id
  target_recovery_protection_container_id = module.container_secondary.id
  target_resource_group_id                = module.app_rg_secondary.id

  depends_on = [module.dr_mapping]
}

module "protected_web" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replicated-vm.git?ref=v1.0.0"

  name                = "vm-web-01"
  resource_group_name = module.dr_rg.name
  recovery_vault_name = module.dr_vault.name

  source_recovery_fabric_name               = module.fabric_primary.name
  source_recovery_protection_container_name = module.container_primary.name
  source_vm_id                              = module.web_vm.id

  recovery_replication_policy_id           = module.dr_policy.id
  target_recovery_fabric_id               = module.fabric_secondary.id
  target_recovery_protection_container_id = module.container_secondary.id
  target_resource_group_id                = module.app_rg_secondary.id

  depends_on = [module.dr_mapping]
}

module "db_vm" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-linux-virtual-machine.git?ref=v1.0.0"

  name                  = "vm-db-01"
  resource_group_name   = module.dr_rg.name
  location              = "eastus"
  size                  = "Standard_D4s_v3"
  admin_username        = "azureadmin"
  network_interface_ids = [module.db_vm_nic.id]
}

module "web_vm" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-linux-virtual-machine.git?ref=v1.0.0"

  name                  = "vm-web-01"
  resource_group_name   = module.dr_rg.name
  location              = "eastus"
  size                  = "Standard_D2s_v3"
  admin_username        = "azureadmin"
  network_interface_ids = [module.web_vm_nic.id]
}

module "db_vm_nic" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-network-interface.git?ref=v1.0.0"

  name                = "nic-db-01"
  resource_group_name = module.dr_rg.name
  location            = "eastus"

  ip_configurations = {
    primary = { subnet_id = module.vnet_primary.subnet_ids["snet-app"] }
  }
}

module "web_vm_nic" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-network-interface.git?ref=v1.0.0"

  name                = "nic-web-01"
  resource_group_name = module.dr_rg.name
  location            = "eastus"

  ip_configurations = {
    primary = { subnet_id = module.vnet_primary.subnet_ids["snet-app"] }
  }
}

module "vnet_primary" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-virtual-network.git?ref=v1.0.0"

  name                = "vnet-app-eastus"
  resource_group_name = module.dr_rg.name
  location            = "eastus"
  address_space       = ["10.10.0.0/16"]

  subnets = {
    "snet-app" = { address_prefixes = ["10.10.1.0/24"] }
  }
}

# And the plan that orders them.
module "app_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-app-tier"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_primary.id
  target_recovery_fabric_id = module.fabric_secondary.id

  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_db.id] },
    { replicated_protected_items = [module.protected_web.id] },
  ]
}

check "the_plan_is_wired_to_its_own_topology" {
  assert {
    condition = alltrue([
      module.app_recovery_plan.source_fabric_is_in_the_plans_vault,
      module.app_recovery_plan.target_fabric_is_in_the_plans_vault,
      !module.app_recovery_plan.source_and_target_fabrics_are_the_same,
      module.app_recovery_plan.source_fabric_path_within_the_subscription == module.fabric_primary.fabric_path_within_the_subscription,
    ])
    error_message = "The recovery plan's fabrics do not match the topology this composition declares. Nothing in the provider compares them."
  }
}

output "the_plan_automates_nothing_yet" {
  description = "Named for the gap. The sequence is recorded; no automation and no drill verification exist."
  value = {
    no_actions          = module.app_recovery_plan.plan_has_no_actions
    no_failback         = !module.app_recovery_plan.plan_automates_failback
    the_sequence        = module.app_recovery_plan.protected_items_per_boot_group
    destroy_leaves_them = module.app_recovery_plan.a_plan_orders_a_failover_and_creates_no_protection
  }
}

πŸ”’ This is the first composition in the family that is genuinely complete for its scenario β€” vault, both fabrics, both containers, a policy, the mapping, two protected machines and the plan that orders them. Nothing here is waiting on a resource this provider lacks.

⚠️ It is still not finished as a disaster-recovery posture, and the closing output says why: the plan automates nothing, nothing runs on failback, and no drill has verified any of it. Those are choices to make deliberately, not gaps in the code.

ℹ️ depends_on on the two protected-VM modules is correct and necessary: the container mapping must exist before an item can replicate, and no argument references it.


πŸ“₯ Inputs

Group Variables
Identity name, recovery_vault_id, source_recovery_fabric_id, target_recovery_fabric_id
The sequence boot_recovery_groups (ordered, required)
Bracketing actions failover_recovery_group, shutdown_recovery_group
Scenario settings azure_to_azure_settings
Universal tail timeouts (no tags β€” the resource has none)
Full input schemas
Variable Type Default Notes
name string β€” Force-new. One anchored rule; the provider's own regex lacks a leading anchor.
recovery_vault_id string β€” Force-new. One anchored rule, load-bearing for the derivation.
source_recovery_fabric_id string β€” Force-new. One anchored rule.
target_recovery_fabric_id string β€” Force-new. One anchored rule. Not compared to the source β€” reported instead.
boot_recovery_groups list(object({…})) β€” (required) Ordered. 12 rules: three Azure limits, the item shape, and the seven action rules.
failover_recovery_group object({ pre_actions, post_actions }) {} Exactly one block, actions only. The seven action rules.
shutdown_recovery_group object({ pre_actions, post_actions }) {} Exactly one block, actions only. The seven action rules.
azure_to_azure_settings object({…}) null All fields force-new. Two zone-shape rules; no pairing rules β€” the provider's RequiredWith covers both.
timeouts object({ create, read, update, delete }) {} β†’ 30m / 5m / 30m / 30m Four keys, because an Update function exists.

36 validations. The seven action rules appear three times β€” once per group variable β€” because a Terraform validation block can only reference its own variable. That repetition is forced by the language, not a style choice, and each copy is genuine coverage of a documented requirement the provider does not enforce.

Action fields use explicit null tests, not try(). A try() on a defaultless optional() returns null rather than the fallback, because the attribute exists β€” so trimspace(try(a.script_path, "")) evaluates trimspace(null) and throws, and a rule whose condition throws is indistinguishable from a rule that never fired.


🧾 Outputs

45 outputs: 5 passthrough, 31 derived, 9 constant. None is sensitive; this module accepts and emits no secret.

The derived count is high because a sequence is not readable from a diff. Terraform can show that boot_recovery_groups changed; it cannot show that the change alters the order an application recovers in, which actions pause it, or which steps a drill never touches. Those are the plan's actual content, and they have to be derived to be reviewable.

Output Description
id The plan's Resource ID.
name The plan's name β€” which Azure uses to identify a failover operation.
recovery_vault_id, source_recovery_fabric_id, target_recovery_fabric_id As configured.
vault_name, vault_resource_group_name, vault_subscription_id Parsed from the vault ID.
source_fabric_name, target_fabric_name Parsed from the two fabric IDs.
vault_path_within_the_subscription Comparable across every authored site-recovery module.
source_fabric_path_within_the_subscription, target_fabric_path_within_the_subscription Comparable with both fabric-creating modules.
recovery_plan_path_within_the_subscription This record's own full path at plan time.
source_fabric_is_in_the_plans_vault, target_fabric_is_in_the_plans_vault Comparisons the provider never makes.
source_and_target_fabrics_are_the_same A plan that cannot survive its own event.
boot_group_count Against Microsoft's documented maximum of seven.
protected_items_per_boot_group The sequence shape, in order.
protected_item_count Against Microsoft's documented maximum of 100.
boot_groups_without_protected_items The indexes of action-only stages.
action_count Across all three group kinds.
manual_actions, runbook_actions, script_actions Names by type.
plan_pauses_for_a_manual_action Whether the plan can run unattended.
actions_that_do_not_run_during_a_test_failover The steps a drill never exercises.
actions_that_only_run_during_a_test_failover The mirror.
shutdown_actions_that_only_run_during_a_test_failover A phase the drill does not perform.
actions_that_do_not_run_on_failback The return journey's gaps.
plan_automates_failback Whether any action runs on the way back.
plan_has_no_actions A pure sequencing plan, named as a decision.
configured_for_azure_to_azure Whether the optional settings block was supplied.
script_actions_in_an_azure_to_azure_plan A contradiction the provider accepts.
zones_are_pinned, edge_zones_are_pinned Both force-new.
boot_group_order_is_the_failover_sequence Constant. Why the groups are a list.
the_provider_documents_four_conditional_requirements_and_enforces_none_of_them Constant.
three_documented_azure_limits_are_unenforced_by_the_provider Constant.
the_name_validator_in_the_provider_is_unanchored_at_the_start Constant.
a_manual_action_stops_the_failover_until_a_human_acknowledges_it Constant.
the_shutdown_step_is_skipped_during_a_test_failover Constant.
a_plan_orders_a_failover_and_creates_no_protection Constant. And what a destroy does.
runbook_actions_run_with_the_automation_accounts_own_permissions Constant. Unbounded by vault RBAC.
the_groups_are_editable_in_place_and_the_identity_is_not Constant.

🧠 Architecture Notes

Order is the resource. Microsoft's model is a shutdown step, then a parallel failover of everything in the plan, then the boot groups in sequence β€” group 2 starting only once every machine in group 1 has started, with machines inside a group failing over in parallel. So the list index of boot_recovery_groups is the plan's entire behaviour. That is why this module models the groups as a list rather than as the keyed map it uses for repeated blocks elsewhere in this family, and why the provider models the block as a TypeList rather than a TypeSet. A map would lose the ordering that is the point.

The provider's documentation and its enforcement diverge, in both directions. Four -> notes on the action block state conditional requirements β€” fabric_location for runbook and script actions, runbook_id for runbooks, script_path for scripts, manual_action_instruction for manual actions β€” and all four fields are plain Optional with no RequiredWith, no CustomizeDiff and no body check. An omitted value is transmitted as an empty string. Meanwhile the one thing the provider does enforce, that fabric_location must not be set for a manual action, is documented nowhere and raised only at apply. This module covers all five at plan time, which is coverage in the strongest sense: the provider published the rules and left them to Azure.

Three Azure limits live only in Microsoft's prose. Seven groups, 100 protected items per plan, and no machine in two groups. The schema expresses none of them. The duplicate rule is the sharpest, because in an ordered sequence a machine listed twice has two contradictory positions β€” and nothing in the provider or the plan output would say so.

Two facts decide whether a plan actually works, and neither is an argument. A manual action stops the failover job until a human acknowledges it, so a plan containing one cannot run unattended and its real recovery time objective includes reaction time. And the shutdown step is skipped during a test failover, so a drill never exercises that phase β€” which makes an action placed there and scoped to TestFailover alone belong to a phase the drill does not perform.

Nothing compares the fabrics to the vault, or to each other. Both fabric arguments are independent Resource IDs. A plan spanning two vaults is accepted; so is one whose source and target are the same fabric, which cannot survive the event the plan exists for. All three comparisons are reported rather than enforced, matching the judgement the network-mapping and replicated-VM modules make about their own equal-sides cases.

A runbook action is unbounded by this resource's permissions. It names an Automation runbook by ID, and what that runbook does when the plan runs is decided by the Automation account's identity. A recovery plan is therefore a mechanism for running arbitrary automation at the worst possible moment, and reviewing one means reviewing the runbooks it names.

The editable/immutable split is well chosen. The groups and their actions update in place β€” the parts refined after every drill β€” while the name, the vault and both fabrics force replacement. The trap is that the name is one of the immutable ones, and Azure's backend uses it to identify a failover operation.

Action fields use explicit null tests rather than try(). On a defaultless optional() the attribute exists with value null, so try(a.script_path, "") yields null and trimspace(null) throws β€” turning a validation into a condition-evaluation error that looks exactly like a rule that never fired. The proof harness caught six of these before they shipped.


🧱 Design Principles

There is no secure-defaults table on this module, and the reason is worth stating. Four required identity arguments and three blocks of sequencing and automation. Nothing here gates network exposure, and nothing has a public/private toggle. What a recovery plan governs is whether a recovery works β€” correctness and preparedness, which this suite's secure-by-default rule does not cover.

The choices that are made, and why:

Decision Value Reasoning
Boot groups as a list Yes, deliberately Order is the failover sequence. This is the one place in this suite where a keyed map would be wrong.
Pre- and post-actions as lists Yes Their order is meaningful too β€” Microsoft's portal offers Move Up and Move Down for them.
The four documented conditional requirements Enforced Documented by the provider and enforced nowhere by it. Coverage in the plainest sense.
The one undocumented prohibition Enforced too The provider raises it at apply as a Go error; restating it at plan is coverage with a nameable boundary.
The three documented Azure limits Enforced Published by Microsoft, absent from the schema. A caller cannot discover them from the provider.
The azure_to_azure_settings pairings Not restated β€” none, deliberately The provider enforces both with RequiredWith at plan time, with messages naming the missing attribute. Recorded as examined so the contrast with the action rules reads as a decision.
Zone membership Shape only Which zones exist depends on region and capacity. This suite enforces shape and refuses to approximate an unpublished set.
The three fabric/vault comparisons Reported, not enforced The provider documents no relationship. A cross-vault plan is unusual, not illegal.
A script action in an A2A plan Reported The provider accepts it; a module cannot see which scenario the vault serves.
failover_recovery_group / shutdown_recovery_group as objects Yes Their cardinality is exactly one. A one-element list would let a caller write zero or two and find out inside the resource.
Action-field null handling Explicit == null tests try() on a defaultless optional() returns null, so a trimspace() on it throws β€” and a throwing condition looks like an unfired rule.
Emitting the subscription in path outputs Omitted Matching the form used by the modules that take a vault by name keeps every path output in this family comparable.

πŸš€ Runbook

terraform init -backend=false
terraform fmt -check
terraform validate

Pin the module with ?ref=v1.0.0 and never a branch. This library is plan-only: a human applies from CI.


πŸ§ͺ Testing

terraform validate proves the configuration parses, all five nested block families exist and the types line up. It does not fire a variable validation when this module is called from another configuration β€” terraform console with a -var-file is the offline harness for that.

What was proven offline for this module:

  • All 36 validations fired, each by a fixture built to fail it, with zero condition-evaluation errors. Every bad fixture is the valid base with exactly one thing replaced, so no fixture can fire a rule other than the one it targets.
  • The first run reported 30/36 with SIX condition-evaluation errors, all one bug: try() on a defaultless optional() returns null, so trimspace(try(a.script_path, "")) threw. Rewritten as explicit null tests; a rule whose condition throws is indistinguishable from one that never fired, so this was a real defect caught before it shipped.
  • All 45 locals were driven to more than one value across five good fixtures: the minimal call, a three-tier plan with an empty middle group and pinned zones, an edge-zone plan with a script action and drill/failback gaps, a plan with both fabrics misplaced and identical, and a seven-group plan.
  • source_fabric_path_within_the_subscription was evaluated in four modules β€” this one, both fabric-creating modules and the Hyper-V policy association β€” from three different argument conventions, and came out byte-identical.
  • The terraform validate result was confirmed with the working directory printed, its .tf files listed, and a zz_negctl.tf negative control in the module's own directory: it fails with the control present and passes with it removed. That matters because terraform validate reports "Success!" in a directory containing no .tf files at all.

What only an apply exercises: whether Azure accepts the group composition, and every behaviour that happens at failover rather than at create.


πŸ’¬ Example Output

id = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod/replicationRecoveryPlans/rp-app-tier"
name = "rp-app-tier"
vault_name = "rsv-dr-prod"
source_fabric_name = "fabric-eastus"
target_fabric_name = "fabric-westus2"
recovery_plan_path_within_the_subscription = "/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod/replicationRecoveryPlans/rp-app-tier"
boot_group_count = 3
protected_items_per_boot_group = [1, 2, 1]
protected_item_count = 4
boot_groups_without_protected_items = []
action_count = 2
manual_actions = ["confirm-dns-cutover"]
runbook_actions = ["failover-sql-availability-group"]
script_actions = []
plan_pauses_for_a_manual_action = true
actions_that_do_not_run_during_a_test_failover = ["confirm-dns-cutover"]
plan_automates_failback = true
plan_has_no_actions = false
source_fabric_is_in_the_plans_vault = true
source_and_target_fabrics_are_the_same = false
boot_group_order_is_the_failover_sequence = true
a_manual_action_stops_the_failover_until_a_human_acknowledges_it = true

Read protected_items_per_boot_group as the plan's design: one database, then two middleware machines in parallel, then one web server.


πŸ” Troubleshooting

Symptom Cause Fix
boot_recovery_groups must contain at most 7 groups Microsoft's documented maximum, which the provider does not enforce Merge stages, or model extra ones as pre-actions or post-actions on an existing group.
the total number of replicated_protected_items ... must be at most 100 Microsoft's documented per-plan maximum Split the application across several recovery plans.
no replicated protected item may appear in more than one boot recovery group A machine listed twice, giving it two positions in one sequence Remove the duplicate. Group order is the sequence.
an action of type AutomationRunbookActionDetails must set runbook_id ... The provider documents this and enforces nothing; an omitted value is sent as an empty string Set runbook_id to an Automation runbook Resource ID.
fabric_location must be Primary or Recovery ... and must NOT be set for a manual action The requirement is documented and unenforced; the prohibition is enforced at apply and undocumented Set it for runbook and script actions, omit it for manual ones.
name must start with a letter, end with a letter or a number ... The provider's own regex is missing its leading anchor, so it would have accepted this Use the rule the provider's message states; this module enforces it.
The apply fails on a group composition that passed the plan Something Azure checks and neither the provider nor this module can see Compare the plan against the portal's recovery-plan view; the group and action limits this module enforces cover the documented cases.
A failover stopped and never finished A manual action is waiting for acknowledgement Acknowledge it in the portal. Assert plan_pauses_for_a_manual_action if the plan must run unattended.
A drill passed and the real failover skipped a step The action omits TestFailover, or sits in the shutdown group which a drill does not run Read actions_that_do_not_run_during_a_test_failover and shutdown_actions_that_only_run_during_a_test_failover.
Destroying the plan did not stop replication Correct β€” a plan owns only the ordering Protection is owned by the replicated-item resources, where a destroy does disable it.
Renaming the plan replaced it name is force-new, and Azure uses it to identify a failover operation Avoid renaming a plan that has been failed over.

πŸ”— Related Docs


πŸ’™ "Infrastructure as Code should be standardized, consistent, and secure."