Skip to content

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

☁️ Azure Site Recovery VMware Replicated VM Terraform Module

Enrols an on-premises VMware machine into Azure Site Recovery's modernized experience, targeting hashicorp/azurerm ~> 4.0.

Terraform azurerm Module Type Resources Posture


🧩 Overview

  • 🖥️ Manages azurerm_site_recovery_vmware_replicated_vm — the record that starts replicating a VMware machine Azure has never seen into a region it can be failed over to.
  • 🔴 Three of its arguments are friendly names matched at apply time, against records an on-premises appliance discovered. appliance_name, physical_server_credential_name and source_vm_name are emitted by no module in any library and are unverifiable by the provider and by this module.
  • 🔴 The vault must hold exactly one protection container. The provider lists them and fails unless there is precisely one, then takes the fabric and container names from it. Neither is an argument, and a vault shared with a second source site cannot host this resource.
  • 🔴 physical_server_credential_name names a root or administrator credential on the source machine, stored on the appliance rather than in Azure. This module accepts the name and never the secret — and no Azure role bounds what that credential can do on-premises.
  • 🔴 Two disk-configuration styles exist and mixing them never converges. A configuration that sets both the default_* fields and managed_disks, or that sets any per-disk log storage account, proposes a replacement on every plan, forever — and replacement re-seeds replication.
  • 🔴 test_network_id is read back and never updated. Changing it produces a plan that applies cleanly, changes nothing, and reappears identically next time.
  • ⚠️ Create is three phases, not one call: enable protection, poll until Azure reports the machine Protected, then a second call to apply the network interfaces. The 120-minute default is a data-transfer budget.
  • ⚠️ terraform destroy disables protection. The source machine in VMware is untouched; the Azure replica and its recovery points are discarded.

💡 Why it matters: almost nothing about this resource is checkable. Four values are resolved by lookup during the apply, two of the sharpest constraints appear only in Microsoft's prose, and the disk arguments contain a non-converging combination the provider documents as legal. So the module's job is to reject what is knowably wrong, name what the value actually is in every error message, and report the rest as flags a check block can assert — including the flag that says this configuration will replace itself on every apply.


❤️ Support this project

If this module saved you time:


🗺️ Where this fits in the family

This module is terraform-azurerm-site-recovery-vmware-replicated-vm, one of the fourteen blue nodes — all fourteen authored site_recovery_* modules share this diagram, and with this module the family is complete. The three dark nodes are not modules: each is a single ARM resource type that several of the blue modules write.

flowchart TB
  RG["terraform-azurerm-resource-group"]
  VAULT["terraform-azurerm-recovery-services-vault"]
  FABRIC["terraform-azurerm-site-recovery-fabric"]
  HVS["terraform-azurerm-site-recovery-services-vault-hyperv-site"]
  PC["terraform-azurerm-site-recovery-protection-container"]
  POLICY["terraform-azurerm-site-recovery-replication-policy"]
  HVPOL["terraform-azurerm-site-recovery-hyperv-replication-policy"]
  VMPOL["terraform-azurerm-site-recovery-vmware-replication-policy"]
  PCM["terraform-azurerm-site-recovery-protection-container-mapping"]
  HVASSOC["terraform-azurerm-site-recovery-hyperv-replication-policy-association"]
  VMASSOC["terraform-azurerm-site-recovery-vmware-replication-policy-association"]
  NM["terraform-azurerm-site-recovery-network-mapping"]
  HVNM["terraform-azurerm-site-recovery-hyperv-network-mapping"]
  RVM["terraform-azurerm-site-recovery-replicated-vm"]
  RRP["terraform-azurerm-site-recovery-replication-recovery-plan"]
  ARMF["ONE ARM type, TWO resource types: vaults replicationFabrics. Same name plus same vault means one object, silently"]
  ARMP["ONE ARM type, THREE resource types: vaults replicationPolicies. A2A, Hyper-V and VMware policies are interchangeable at the API"]
  ARMM["ONE ARM type, THREE resource types: replicationProtectionContainerMappings. The hardest group to assert on"]
  VM["terraform-azurerm-linux-virtual-machine or windows-virtual-machine: wire its id, NOT its virtual_machine_id"]
  VNET["terraform-azurerm-virtual-network: the failover network, the test network and the source network"]
  RUNBOOK["an Azure Automation runbook: a plan action names it by id and it runs with the automation account's own permissions. NOT bounded by vault RBAC"]
  VMWVM["terraform-azurerm-site-recovery-vmware-replicated-vm"]
  HOSTS["the Hyper-V hosts themselves, registered with a five-day vault key. NO azurerm resource does this"]
  APPL["the VMware replication appliance, deployed on-premises and registered against the vault. NO azurerm resource does this either"]
  VMM["the System Center VMM server, registered by installing the Site Recovery Provider on it. NO azurerm resource does this either"]
  VCENTER["the vCenter machines and the credentials stored on the appliance. source_vm_name and physical_server_credential_name are FRIENDLY names matched at apply time, and that credential is root or admin ON the source machine"]

  RG -->|"resource_group_name and location"| VAULT
  VAULT -->|"name, NOT id, so there is no dependency edge from a literal"| FABRIC
  VAULT -->|"id, NOT name: this sibling takes the vault ID instead"| HVS
  VAULT -->|"name, and NO fabric: a policy belongs to the vault"| POLICY
  VAULT -->|"id, a second convention on the same parent"| HVPOL
  VAULT -->|"id, and the association takes the SAME vault argument"| VMPOL
  VAULT -->|"id: BOTH the fabric and the container are discovered under it at apply time"| VMASSOC
  VAULT -->|"name plus resource group, a third convention again"| RVM
  VAULT -->|"id: the ONLY Resource ID this module resolves anything from"| HVNM
  VAULT -->|"id, and BOTH fabrics as ids: a fourth convention"| RRP
  FABRIC -->|"hardcoded instance type: an Azure fabric"| ARMF
  HVS -->|"hardcoded instance type: HyperVSite, absent from Swagger"| ARMF
  POLICY -->|"hardcoded A2A payload"| ARMP
  HVPOL -->|"hardcoded Hyper-V to Azure payload"| ARMP
  VMPOL -->|"hardcoded InMageRcm payload: the MODERNIZED VMware provider"| ARMP
  PCM -->|"container taken by name and by id"| ARMM
  HVASSOC -->|"container discovered under the fabric at apply time"| ARMM
  VMASSOC -->|"fabric AND container both discovered under the vault"| ARMM
  FABRIC -->|"name, so a Hyper-V site's name is equally legal here"| PC
  FABRIC -->|"name, for BOTH sides: ARM names, not friendly names"| NM
  FABRIC -->|"id, for BOTH sides, and neither is compared to the vault"| RRP
  PC -->|"name for the source side, id for the target side"| PCM
  POLICY -->|"id, where the container beside it is taken by name"| PCM
  HVPOL -->|"id, the only consumer of a Hyper-V policy"| HVASSOC
  HVS -->|"id, and the protection container is discovered from it"| HVASSOC
  VMPOL -->|"id, the only consumer of a VMware policy"| VMASSOC
  VNET -->|"id, twice. Only the SOURCE id is parsed by the provider"| NM
  VNET -->|"id, the target only. The SOURCE network is a VMM friendly name"| HVNM
  FABRIC -->|"source fabric by NAME, target fabric by ID: one resource, two conventions"| RVM
  PC -->|"source container by NAME, target container by ID"| RVM
  POLICY -->|"id, and a Hyper-V or VMware policy id would also be accepted"| RVM
  VM -->|"the machine being protected"| RVM
  VNET -->|"the failover network and the test network, both Optional AND Computed"| RVM
  PCM -->|"makes a container pair eligible before any item can replicate"| RVM
  NM -->|"decides where an Azure-to-Azure failover lands"| RVM
  HVNM -->|"decides where a VMM Hyper-V failover lands, for every machine on the source network"| VMM
  RVM -->|"ids, into ORDERED boot groups: group 2 starts only after group 1 has finished"| RRP
  VMWVM -->|"ids, into the same ORDERED boot groups as the Azure-to-Azure items"| RRP
  VAULT -->|"id ALONE, a fifth convention: the fabric AND the container are discovered from it, and there must be exactly ONE container"| VMWVM
  VMPOL -->|"id, its second consumer: the policy this VMware item replicates under"| VMWVM
  VNET -->|"the failover network and the test network. test_network_id is read back and NEVER updated"| VMWVM
  VMASSOC -->|"makes the discovered container eligible before any VMware machine can replicate"| VMWVM
  RUNBOOK -->|"id, from a plan pre-action or post-action"| RRP
  APPL -->|"appliance_name is its FRIENDLY name, matched against the fabric process servers. Microsoft: it cannot be changed once set"| VMWVM
  VCENTER -->|"source_vm_name and physical_server_credential_name, both matched by name against records the appliance discovered"| VMWVM
  HOSTS -->|"register into the site out of band, invisible to Terraform"| HVS
  HOSTS -->|"registration is what creates the container this record needs"| HVASSOC
  APPL -->|"required before any VMware machine can be enrolled"| VMASSOC
  VMM -->|"the fabric AND the source network are found under it by FRIENDLY NAME at apply time"| HVNM

  classDef me fill:#0078D4,stroke:#004578,color:#ffffff
  classDef keystone fill:#004578,stroke:#002438,color:#ffffff
  classDef sibling fill:#F3F6F9,stroke:#8A9BA8,color:#1B1F23
  class FABRIC,HVS,PC,POLICY,HVPOL,VMPOL,PCM,HVASSOC,VMASSOC,NM,HVNM,RVM,RRP,VMWVM me
  class ARMF,ARMP,ARMM keystone
  class RG,VAULT,VM,VNET,RUNBOOK,VCENTER,HOSTS,APPL,VMM sibling
Loading

Read the grey nodes carefully on this one. APPL and VCENTER are where three of this module's arguments come from, and no Terraform resource in any provider creates either.


🧬 What this module builds

flowchart TB
  IN_VAULT["recovery_vault_id: the ONLY Resource ID resolved from. The fabric and container are discovered under it, and there must be exactly ONE container"]
  IN_POLICY["recovery_replication_policy_id"]
  IN_NAMES["appliance_name, physical_server_credential_name, source_vm_name: THREE FRIENDLY NAMES, matched case-insensitively at apply time, emitted by no module"]
  IN_TARGET["target_vm_name, target_resource_group_id, target_vm_size, license_type, multi_vm_group_name"]
  IN_NET["target_network_id, test_network_id, and the network_interfaces list keyed by SOURCE MAC ADDRESS"]
  IN_PLACE["target_zone or target_availability_set_id, target_proximity_placement_group_id, target_boot_diagnostics_storage_account_id"]
  IN_DISK["EITHER the three default_ disk fields OR the managed_disks list. Never both: that combination can never converge"]
  IN_TIME["timeouts: four keys, create 120m by default because it is a data-transfer budget"]

  THIS["azurerm_site_recovery_vmware_replicated_vm.this"]

  PHASE1["1. enable protection, sending the modernized InMageRcm payload"]
  PHASE2["2. POLL every 15 seconds until Azure reports the machine Protected, for whatever remains of the create timeout"]
  PHASE3["3. a SECOND call that applies the network interfaces. NICs are never applied by step 1"]

  OUT_ID["id and name. Two ID segments are discovered during the apply, so the full path is NOT knowable at plan time"]
  OUT_PATH["vault_path_within_the_subscription and policy_path_within_the_subscription: byte-identical to the form every authored module in this family emits"]
  OUT_TRAP["disk_configuration_never_converges: TRUE when this configuration proposes a REPLACEMENT on every plan, forever"]
  OUT_NET["the NIC flags, including which interfaces are NOT failed over. is_primary alone decides that, with no separate argument"]
  OUT_DRILL["test_network_is_the_same_as_the_failover_network, which Microsoft advises against and the provider does not check"]
  OUT_CONST["ten constants: the one-container rule, the three name lookups, the root credential, the unupdatable test network, the three-phase create, the stuck-replication timeout, the read-modify-write update, disable-not-delete, the modernized experience, and the MAC as identifier"]

  IN_VAULT --> THIS
  IN_POLICY --> THIS
  IN_NAMES --> THIS
  IN_TARGET --> THIS
  IN_NET --> THIS
  IN_PLACE --> THIS
  IN_DISK --> THIS
  IN_TIME --> THIS

  THIS --> PHASE1
  PHASE1 -->|"then"| PHASE2
  PHASE2 -->|"then"| PHASE3

  PHASE3 --> OUT_ID
  THIS --> OUT_PATH
  THIS --> OUT_TRAP
  THIS --> OUT_NET
  THIS --> OUT_DRILL
  THIS --> OUT_CONST

  classDef me fill:#0078D4,stroke:#004578,color:#ffffff
  classDef keystone fill:#004578,stroke:#002438,color:#ffffff
  classDef sibling fill:#F3F6F9,stroke:#8A9BA8,color:#1B1F23
  class THIS keystone
  class PHASE1,PHASE2,PHASE3 me
  class IN_VAULT,IN_POLICY,IN_NAMES,IN_TARGET,IN_NET,IN_PLACE,IN_DISK,IN_TIME,OUT_ID,OUT_PATH,OUT_TRAP,OUT_NET,OUT_DRILL,OUT_CONST sibling
Loading
Resource Cardinality Note
azurerm_site_recovery_vmware_replicated_vm single (this) The keystone. Two nested block families, and a three-phase create.

A standalone module: one keystone resource, no for_each children. 20 arguments plus two block families.


✅ Provider / Versions

Item Value
Terraform >= 1.12.0
Provider hashicorp/azurerm ~> 4.0 (validated against 4.81.0)
Provider block None here. The caller configures provider "azurerm" { features {} }, auth and subscription.
ARM API version Microsoft.RecoveryServices 2024-04-01, replicationProtectedItems
tags Not supported — this resource has no tags and no location. Tag the vault instead.

Schema notes that bite

  • 🔴 The vault must hold EXACTLY ONE protection container. fetchSiteRecoveryContainerId lists them and returns there should be only 1 protection container in Recovery Vault, get: N for any other count. The fabric and container names then come out of that container's ID, so neither appears as an argument anywhere on this resource.
  • 🔴 Three arguments are matched by name, case-insensitively, at apply time. appliance_name against the fabric's process-server names, source_vm_name against every machine the appliance discovered in the VMware site, and physical_server_credential_name against the site's Run As accounts — that last one matched together with appliance_name, so the same credential name under a different appliance is a different account. The create path performs five such lookups in total.
  • 🔴 The docs and the schema enforce OPPOSITE directions of one pairing. The note says "target_network_id is required when network_interface is specified"; the RequiredWith sits on target_network_id and names network_interface. So a network without NICs is rejected at plan time and NICs without a network are accepted. This module covers the documented direction.
  • 🔴 The three default_* disk fields are write-only — "not returned by API, and only used in creation", in the provider's own comment — and the read populates neither them nor each disk's log_storage_account_id. The conditional force-new check compares prior-state disks against the configured defaults, so the comparison can never match after the first apply.
  • 🔴 test_network_id is read back, is not force-new, and is absent from the update payload. A change to it never converges.
  • ⚠️ Force-new: name, appliance_name, physical_server_credential_name, source_vm_name, target_vm_name, plus multi_vm_group_name, and managed_disk once a non-empty list has been supplied (removing it entirely does not). timeouts has four keys.
  • ⚠️ The create timeout is a budget for initial replication. 120 minutes by default, of which the polling phase consumes whatever remains after the enable-protection call.
  • ⚠️ The poll's failure test is a string match — a protection state ending in Failed or beginning with Cancel. Anything else unhealthy reads as pending, so the apply waits out the budget and reports a timeout instead of Azure's state.
  • ⚠️ The update is a read-modify-write. Unchanged fields are re-sent from the record Azure currently holds, so an out-of-band change is preserved rather than corrected — and an update to any field fails outright if Azure reports no NICs.
  • ⚠️ Three optional IDs do not map an empty value to unset on update (target_network_id, target_proximity_placement_group_id, target_boot_diagnostics_storage_account_id), unlike target_availability_set_id and target_zone, which do.
  • ⚠️ license_type is sent unconditionally, so an omission is transmitted as an empty string.
  • ⚠️ A NIC is identified by the SOURCE MAC ADDRESS, sent as the ARM record's NIC id. A non-primary NIC is created not-selected-for-failover, derived from is_primary with no separate argument.
  • ⚠️ There is no SchemaVersion and no state upgrader.

🔑 Required Azure RBAC Roles / Permissions

Principal Scope Requirement
The principal running Terraform The vault or its resource group Site Recovery Contributor — it carries Microsoft.RecoveryServices/vaults/replicationProtectedItems/*, and the create additionally READS the vault's fabrics, protection containers, VMware site machines and Run As accounts
The same principal The target resource group, network, disk encryption set, availability set, proximity placement group and storage accounts Read on each referenced resource. They are named by ID and must resolve
The same principal The vault Read on Run As accounts. Every refresh resolves physical_server_credential_name from the record's Run As account id with an extra API call, so a plan that cannot make that call fails
The credential named by physical_server_credential_name The source machine, on-premises Root (Linux) or an account with administrator privileges (Windows) — see below
Whoever runs the failover The vault Site Recovery Operator or higher
Alternatively The vault or its resource group Contributor or Owner

Row four is the one to read twice, and no Azure role limits it. Microsoft documents that these credentials are used to push-install the Mobility agent onto the source machine, that Linux requires root credentials and Windows an account with admin privileges, and that the username and password are encrypted and stored on the appliance rather than in Azure. This module accepts only the credential's name and never the secret — but naming it in a configuration means the replication will use privileged access to a machine, and reviewing this module means knowing which account that is.

Backup Contributor is the wrong role and it is a plausible mistake, because one Recovery Services vault serves both Azure Backup and Site Recovery. It grants nothing over replicationProtectedItems. The scaffold this module replaced named exactly that role.

Plan access is not credential access here. No secret is read, held or emitted, and no output is sensitive.


Azure Prerequisites

  1. Microsoft.RecoveryServices registered in the subscription.
  2. A Recovery Services vault NOT configured for the classic VMware experience. Microsoft retired that experience on 30 March 2026; this resource sends the modernized (InMageRcm) payload. terraform-azurerm-recovery-services-vault defaults classic_vmware_replication_enabled to false, which is correct — and that field is immutable, so a vault created with it enabled cannot be converted.
  3. Exactly ONE protection container in that vault. Not zero, not two. The provider discovers the fabric and container from it, and refuses any other count.
  4. An Azure Site Recovery replication appliance, deployed on-premises and registered against the vault. No Terraform resource in any provider does this. Its friendly name becomes appliance_name, and Microsoft documents that the name cannot be changed once set.
  5. vCenter registered with that appliance, and the machine discovered. source_vm_name is matched against discovered machines' display names, so the machine must already appear under the vault's Site Recovery infrastructure.
  6. A credential registered in the appliance configurator whose friendly name becomes physical_server_credential_name, holding root or administrator access to the source machine.
  7. A VMware replication policy and its container association, from terraform-azurerm-site-recovery-vmware-replication-policy and -vmware-replication-policy-association.
  8. A target resource group and, if you want to control placement, a target virtual network — plus a separate test-failover network, which Microsoft recommends be different from the failover network.

📁 Module Structure

terraform-azurerm-site-recovery-vmware-replicated-vm/
├── providers.tf   # required_version and the pinned azurerm; no provider block
├── variables.tf   # 23 variables, 41 validations
├── main.tf        # the keystone `this`, two dynamic block families, and 51 derived locals
├── outputs.tf     # 59 outputs: 11 passthrough, 38 derived, 10 constant
├── README.md      # this file
├── SCOPE.md       # the cross-module contract
├── LICENSE        # MIT
└── .gitignore

⚙️ Quick Start

module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name              = "vm-erp-01"
  recovery_vault_id = module.dr_vault.id

  recovery_replication_policy_id = module.vmware_policy.id

  # THREE FRIENDLY NAMES. None comes from a module output; all three are matched at apply time.
  appliance_name                  = "asr-appliance-dc1" # the appliance's friendly name
  physical_server_credential_name = "erp-root"          # a credential registered in the appliance
  source_vm_name                  = "ERP-PROD-01"       # the machine's vCenter display name

  target_vm_name           = "vm-erp-01-dr"
  target_resource_group_id = module.app_rg_recovery.id

  depends_on = [module.vmware_assoc]
}

💡 The smallest legal call takes eight arguments and configures no disks, no NICs and no networks. That is a real and common shape: Azure replicates every disk and chooses the failover networking itself.

🔒 Note what is not here: no default_* disk field and no managed_disks. That combination is the only one that converges — see example 6.

ℹ️ depends_on is correct and necessary: the container association must exist before a machine can replicate, and no argument on this resource references it.


🔌 Cross-Module Contract

Consumes

Input Type Source
recovery_vault_id string terraform-azurerm-recovery-services-vault output id
recovery_replication_policy_id string terraform-azurerm-site-recovery-vmware-replication-policy output id
target_resource_group_id string terraform-azurerm-resource-group output id
target_network_id, test_network_id string terraform-azurerm-virtual-network output id
target_boot_diagnostics_storage_account_id, default_log_storage_account_id, per-disk log_storage_account_id string terraform-azurerm-storage-account output id
appliance_name, physical_server_credential_name, source_vm_name string nothing. Read from the vault's Site Recovery infrastructure blade and the appliance configurator
target_availability_set_id, target_proximity_placement_group_id, default_target_disk_encryption_set_id string plain azurerm_* resources in the caller's configuration
ordering — terraform-azurerm-site-recovery-vmware-replication-policy-association, via depends_on

Emits

Output Description Consumed by
id The protected item's Resource ID (first) terraform-azurerm-site-recovery-replication-recovery-plan, as a boot-group member
vault_path_within_the_subscription The vault's path, in the family's shared form check blocks asserting one topology, one vault
policy_path_within_the_subscription The policy's path check blocks against the policy module
disk_configuration_never_converges Whether this configuration replaces itself on every plan check blocks — the highest-value assertion here
test_network_is_the_same_as_the_failover_network A drill that touches the network a real failover needs check blocks
network_interfaces_not_selected_for_failover Which interfaces will not come back drill review
ten constants Facts that produce no error when they bite documentation, check blocks

📚 Example Library

1 · The smallest real call
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-erp-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"
  target_vm_name                  = "vm-erp-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  depends_on = [module.vmware_assoc]
}

output "converges" {
  value = !module.protected_erp.disk_configuration_never_converges
}

💡 Eight required arguments and nothing else. Azure replicates every disk, picks the VM size from the source machine's shape, and chooses the failover networking. no_disk_configuration_at_all reports this as a state rather than a gap — it is the only disk shape that converges.

2 · The three names, and where each one comes from
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name              = "vm-erp-01"
  recovery_vault_id = module.dr_vault.id

  recovery_replication_policy_id = module.vmware_policy.id

  # 1. The appliance's FRIENDLY name, as shown under the vault's Site Recovery infrastructure.
  #    Matched against the fabric's process-server names. Microsoft: it cannot be changed once set.
  appliance_name = "asr-appliance-dc1"

  # 2. A credential's FRIENDLY name, as entered in the appliance configurator. Matched against the
  #    site's Run As accounts TOGETHER WITH appliance_name -- the same name under a different
  #    appliance is a different account.
  physical_server_credential_name = "erp-root"

  # 3. The machine's vCenter DISPLAY name, as discovered by the appliance.
  source_vm_name = "ERP-PROD-01"

  target_vm_name           = "vm-erp-01-dr"
  target_resource_group_id = module.app_rg_recovery.id

  depends_on = [module.vmware_assoc]
}

🔴 None of the three is an Azure identifier, and none is emitted by any module. Each is matched case-insensitively during the apply, so a typo is a runtime failure rather than a plan error. The module rejects a blank and rejects a Resource ID — the mistake a reader makes when every other reference on this resource is an ID — and each error message says what the value actually is.

⚠️ The three failures are distinguishable by their wording alone: process server ... not found means appliance_name, machine ... not found means source_vm_name, and run as account ... not found means physical_server_credential_name (or the wrong appliance beside it).

3 · The credential's blast radius
output "what_that_credential_can_do" {
  description = "Named for the fact no Azure role controls."
  value = {
    is_privileged_on_the_source = module.protected_erp.the_credential_named_here_is_root_or_administrator_on_the_source_machine
    credential_name             = module.protected_erp.physical_server_credential_name
  }
}

🔒 Microsoft documents that this credential push-installs the Mobility agent onto the source machine, that Linux requires root credentials and Windows an account with administrator privileges, and that the username and password are encrypted and stored on the appliance rather than in Azure.

🔒 So this module holds no secret and emits none — it accepts a name. What that name refers to is privileged access to a production machine, granted outside Azure, governed by nothing in this configuration. If you need to avoid it entirely, Microsoft documents installing the Mobility service manually and enabling replication without registered credentials.

⚠️ Note the second-order cost of the name: every refresh resolves it from the record's Run As account id with an extra API call, so a principal that can plan this module can read that account's display name, and a principal that cannot will see the refresh fail.

4 · One vault, one container — the constraint with no argument
check "the_vault_can_host_this_resource" {
  assert {
    condition     = module.protected_erp.the_vault_must_hold_exactly_one_protection_container
    error_message = "Always true -- this assertion exists to put the rule in front of a reviewer. Confirm the vault holds exactly ONE protection container before adding a second source site to it."
  }
}

check "the_policy_is_in_the_plans_own_vault" {
  assert {
    condition     = module.protected_erp.policy_is_in_the_same_vault_as_this_record
    error_message = "The replication policy is in a different vault from the one this record is created in. The provider compares the two nowhere, and a policy from another vault cannot apply."
  }
}

🔴 The provider lists the vault's protection containers and fails unless there is exactly one, with there should be only 1 protection container in Recovery Vault, get: N. It then takes the fabric name and container name out of that container's ID — which is why neither is an argument on this resource at all.

⚠️ The practical consequence is a design constraint, not a syntax one: a vault that serves two VMware sites, or that also serves Azure-to-Azure replication with its own container, cannot host this resource. Use one vault per source site. Nothing in the schema shows this and the failure is at apply.

💡 The second assertion covers a comparison the provider genuinely never makes, and the module can make it because both IDs are supplied.

5 · Disk configuration by defaults — the style that works
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-erp-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"
  target_vm_name                  = "vm-erp-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  # ONE style, applied to every disk. Do NOT combine these with managed_disks.
  default_recovery_disk_type            = "Premium_LRS"
  default_log_storage_account_id        = module.asr_log_storage.id
  default_target_disk_encryption_set_id = azurerm_disk_encryption_set.recovery.id

  depends_on = [module.vmware_assoc]
}

output "style" {
  value = module.protected_erp.disk_configuration_style
}

💡 disk_configuration_style reports default_fields here. The three fields apply to every replicated disk, which is what most callers want and what avoids naming individual VMware disk identifiers in a configuration.

⚠️ These three are write-only — the provider's own comment says they are "not returned by API, and only used in creation" — so they never appear in a refresh and never drift-detect. Their only effect is at creation, plus the conditional force-new described in example 6.

6 · The combination that never converges
# ACCEPTED by the provider, DOCUMENTED as legal, and it replaces the resource on every plan.
module "protected_erp_trap" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-erp-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"
  target_vm_name                  = "vm-erp-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  # BOTH styles at once.
  default_log_storage_account_id = module.asr_log_storage.id

  managed_disks = [
    {
      disk_id                = "disk-6000C29-os"
      target_disk_type       = "Premium_LRS"
      log_storage_account_id = module.asr_log_storage.id # the same account, as the docs require
    },
  ]

  depends_on = [module.vmware_assoc]
}

check "this_configuration_settles" {
  assert {
    condition     = !module.protected_erp_trap.disk_configuration_never_converges
    error_message = "This disk configuration proposes a REPLACEMENT on every plan, forever. Replacing a replicated item disables protection and re-seeds replication over the network. Use EITHER the three default_* fields OR managed_disks, and do not set a per-disk log_storage_account_id."
  }
}

🔴 The provider's notes say three things about this and the third contradicts the first: "Only one of default_log_storage_account_id or managed_disk must be specified"; "Changing it forces a new resource to be created. But removing it does not"; and "When it co-exists with managed_disk, the value must be the same as the log_storage_account_id of every managed_disk or it forces a new resource to be created."

🔴 The code implements co-existence and cannot reach the documented escape hatch. The force-new check compares each disk in prior state against the configured default — and the read populates neither the default_* fields nor each disk's log_storage_account_id. So after the first apply it compares an empty string against a set value, and matching them in the configuration changes nothing. The same applies to managed_disks on its own: any entry setting log_storage_account_id never round-trips.

⚠️ This module reports it rather than rejecting it, because the provider accepts the combination and its own documentation blesses it. Rejecting legal input is the worse error — but shipping the check block is not optional, which is why it is here.

7 · Per-disk configuration, done safely
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-erp-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"
  target_vm_name                  = "vm-erp-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  # Per-disk, and NO log_storage_account_id anywhere -- that field is what breaks convergence.
  managed_disks = [
    {
      disk_id                       = "disk-6000C29-os"
      target_disk_type              = "Premium_LRS"
      target_disk_encryption_set_id = azurerm_disk_encryption_set.recovery.id
    },
    {
      disk_id          = "disk-6000C29-data"
      target_disk_type = "Premium_ZRS"
    },
  ]

  depends_on = [module.vmware_assoc]
}

output "disk_review" {
  value = {
    mixed_types     = module.protected_erp.managed_disks_mixing_more_than_one_disk_type
    cmk             = module.protected_erp.managed_disks_using_customer_managed_keys
    zonal           = module.protected_erp.managed_disks_on_a_zonal_disk_type
    constrained     = module.protected_erp.managed_disks_on_a_constrained_disk_type
    never_converges = module.protected_erp.disk_configuration_never_converges
  }
}

💡 Per-disk configuration is the right style when disks genuinely differ — an OS disk on Premium_LRS and a data disk on Premium_ZRS, say. managed_disks_mixing_more_than_one_disk_type reports that as a state, because it is usually deliberate.

⚠️ Once a non-empty managed_disks list has been supplied, changing it forces a new resource — except removing it entirely. So this list is close to immutable in practice, and adding a disk later replaces the record.

🔒 disk_id is the VMware disk identifier discovered by the appliance, not an Azure ID. Nothing in Terraform can verify it, so the module checks only that it is non-empty and not duplicated.

ℹ️ managed_disks_on_a_constrained_disk_type names disks on UltraSSD_LRS or PremiumV2_LRS. Both carry zone, region and host-caching constraints Azure applies at failover, and neither the provider nor this module can check them from a configuration.

8 · Network interfaces, keyed by MAC address
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-erp-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"
  target_vm_name                  = "vm-erp-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  # The DOCUMENTED requirement: a NIC list needs a target network. This module enforces it.
  target_network_id = module.vnet_recovery.id
  test_network_id   = module.vnet_drill.id

  network_interfaces = [
    {
      source_mac_address = "00:50:56:9a:bc:de" # read off the SOURCE machine's interface
      is_primary         = true
      target_subnet_name = "snet-app"   # selected BY NAME inside target_network_id
      test_subnet_name   = "snet-drill" # selected BY NAME inside test_network_id
      target_static_ip   = "10.20.1.40"
    },
    {
      source_mac_address = "00:50:56:9a:bc:df"
      is_primary         = false
      target_subnet_name = "snet-data"
    },
  ]

  depends_on = [module.vmware_assoc]
}

output "which_nics_come_back" {
  value = {
    primary  = module.protected_erp.primary_network_interface_mac_addresses
    excluded = module.protected_erp.network_interfaces_not_selected_for_failover
  }
}

🔴 The MAC address IS the identifier. The provider sends source_mac_address as the ARM record's NIC id and reads it back from the same field, which is why it is validated as a MAC address. It is read off the source machine, corresponds to no Azure identifier, and cannot be wired from any module output.

🔴 The second NIC will not be failed over. The provider derives the ARM "selected for failover" flag from is_primary alone, with no separate argument — so every non-primary NIC is created not-selected, and there is no configuration that changes it. network_interfaces_not_selected_for_failover names them so the omission is visible rather than assumed.

⚠️ The provider enforces the CONVERSE of what it documents here. Its note says target_network_id is required when network_interface is set; the schema's RequiredWith makes target_network_id require a NIC instead. So a network with no NICs is rejected at plan time, NICs with no network are accepted, and this module covers the documented direction.

ℹ️ NICs are applied by a second API call after the initial protection completes, never by the call that enables it.

9 · The test network Microsoft tells you to keep separate
check "a_drill_does_not_touch_the_real_failover_network" {
  assert {
    condition     = !module.protected_erp.test_network_is_the_same_as_the_failover_network
    error_message = "The test-failover network is the same as the failover network. Microsoft's guidance is to keep them apart, 'to make sure the failover network is readily available in case of an actual disaster'. The provider does not check this."
  }
}

output "drill_readiness" {
  value = {
    test_network_configured = module.protected_erp.test_network_is_configured
    same_as_failover        = module.protected_erp.test_network_is_the_same_as_the_failover_network
    unupdatable             = module.protected_erp.test_network_id_is_read_back_but_never_updated
  }
}

🔴 test_network_id cannot be corrected after creation. The update payload omits it entirely, the read populates it from Azure, and the schema does not mark it force-new. So changing it yields a plan that applies cleanly, changes nothing in Azure, and shows the identical diff on the next plan. Get it right at creation; the only real fix is replacing the record, which re-seeds replication.

⚠️ That is precisely why the assertion above belongs in a first apply rather than a later review: this is the field where "we will tidy it up later" does not work.

ℹ️ Omitting test_network_id altogether is legal and means a drill runs against whatever Azure chooses.

10 · Placement after failover — zone or set, never both
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-erp-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"
  target_vm_name                  = "vm-erp-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  target_zone   = "2"
  target_vm_size = "Standard_E16s_v5"

  target_proximity_placement_group_id        = azurerm_proximity_placement_group.recovery.id
  target_boot_diagnostics_storage_account_id = module.asr_boot_diagnostics.id

  depends_on = [module.vmware_assoc]
}

output "placement" {
  value = {
    model              = module.protected_erp.resilience_model_after_failover
    size_chosen_by_you = !module.protected_erp.target_vm_size_is_chosen_by_azure
    boot_diagnostics   = module.protected_erp.boot_diagnostics_are_configured
  }
}

✅ The zone-versus-availability-set exclusion is enforced by the provider and restated here, so the error names both fields and says which to keep. Both halves of the rule sit on target_zone, because a Terraform validation may reference another variable only if it also references its own.

⚠️ Only the zone's SHAPE is checked — "1", "2", "3" rather than "eastus-1" or "az1". Which zones exist depends on the recovery region, and whether one has capacity at the moment of a failover is not knowable from a configuration at all.

💡 Configure boot diagnostics. It is how you find out why a failed-over machine did not start, which is the first question a real incident asks. Omitting it is legal and leaves you without that evidence.

ℹ️ resilience_model_after_failover returns availability_zone, availability_set or none. none is legal and common.

11 · Multi-VM consistency, and what it costs
module "protected_app_01" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-app-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "APP-PROD-01"
  target_vm_name                  = "vm-app-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  # Shared, crash-consistent recovery points across the group. Spelling is the only thing joining them.
  multi_vm_group_name = "erp-tier"

  depends_on = [module.vmware_assoc]
}

check "the_consistency_group_is_spelled_consistently" {
  assert {
    condition     = module.protected_app_01.multi_vm_group_name == module.protected_erp.multi_vm_group_name
    error_message = "Two machines that should share a multi-VM consistency group have different group names, so each is in a group of one and they have no shared recovery point."
  }
}

💡 A multi-VM consistency group gives its members shared, crash-consistent recovery points, which is what a distributed application needs to fail over to a single point in time.

⚠️ It is not free. Microsoft documents that multi-VM consistency is CPU-intensive on the source machine, and the group replicates as a unit — so one unhealthy member holds back the whole group's recovery points.

🔴 The name is a free-form string and nothing validates the group. A typo silently creates a second group of one, with no error anywhere, which is exactly the failure this argument is most likely to produce. The check block above is the only thing that catches it, and it works because the module emits the name back.

ℹ️ multi_vm_group_name is force-new, so moving a machine between groups replaces the record.

12 · Licensing is a billing assertion, so the module picks nothing
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-erp-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"
  target_vm_name                  = "vm-erp-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  # Azure Hybrid Benefit: this ASSERTS you hold Software Assurance or a qualifying subscription.
  license_type = "WindowsServer"

  depends_on = [module.vmware_assoc]
}

output "licensing" {
  value = module.protected_erp.claims_azure_hybrid_benefit
}

🔒 This module supplies no default for license_type, and the reason is a boundary rather than an oversight: WindowsServer is Azure Hybrid Benefit, and choosing it asserts a licence entitlement the caller holds and the module cannot see. Defaulting either way would either overstate an entitlement or spend money unnecessarily — and this suite's secure-by-default rule covers exposure, not licensing.

⚠️ The provider sends this field unconditionally, so an omitted value is transmitted as an empty string rather than left unset.

13 · Timeouts, and why the default is 120 minutes
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name                            = "vm-erp-01"
  recovery_vault_id               = module.dr_vault.id
  recovery_replication_policy_id  = module.vmware_policy.id
  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"
  target_vm_name                  = "vm-erp-01-dr"
  target_resource_group_id        = module.app_rg_recovery.id

  # A 4 TB file server over a constrained link needs more than the 120-minute default, not less.
  timeouts = {
    create = "300m"
    delete = "90m"
  }

  depends_on = [module.vmware_assoc]
}

check "the_create_budget_was_not_shortened_by_accident" {
  assert {
    condition     = !module.protected_erp.create_timeout_is_shorter_than_the_provider_default
    error_message = "The create timeout is below the provider's 120-minute default. That budget covers initial replication of the machine's disks over the network, so shortening it makes a timeout more likely rather than the apply faster."
  }
}

🔴 The create is three phases and the timeout covers all of them. It enables protection, then polls every 15 seconds until Azure reports the machine Protected for whatever remains of the budget, then makes a second call to apply the network interfaces. So the figure is a data-transfer budget that scales with the machine's disk volume, not an API timeout.

⚠️ A stuck replication consumes the whole budget. The poll treats a state as failed only if it ends in Failed or begins with Cancel, against a field the provider's own comment calls "pretty much enums". Anything else unhealthy reads as pending, so the apply waits out the timeout and reports a timeout rather than the state Azure returned.

⚠️ A create that fails after protection succeeded leaves a protected machine whose NICs were never applied. A further apply retries the update rather than recreating, which is the right behaviour and worth knowing before assuming a failed apply changed nothing.

ℹ️ Object types silently discard undeclared keys, so a fifth timeouts key would vanish without an error. All four declared here are real.

14 · What a destroy does — and what an unrelated update does
output "lifecycle_facts" {
  value = {
    destroy_disables_rather_than_deletes = module.protected_erp.destroying_this_disables_protection_and_discards_the_replica
    update_rewrites_current_values       = module.protected_erp.an_unrelated_update_rewrites_azures_current_values
    modernized_experience                = module.protected_erp.this_is_the_modernized_vmware_experience
  }
}

🔴 terraform destroy disables protection; it does not delete a machine. The source VM in VMware is left untouched, and the Azure-side replica and all its recovery points are discarded. So a destroy is not reversible by re-applying — replication starts again from scratch, over the network, on the create timeout budget.

⚠️ The update is a read-modify-write. Every field the plan did not change is re-sent from the record Azure currently holds, so a value changed out of band is preserved rather than corrected by an unrelated apply. And if Azure reports no NICs at all, an update to any other field fails outright with retrieving network_interface: VMNics was nil.

ℹ️ this_is_the_modernized_vmware_experience records the payload this resource sends. Microsoft retired the classic VMware experience on 30 March 2026, so the modernized one is the only correct choice — and a vault configured for classic will not serve this resource.

15 · 🏗️ End-to-end composition — a VMware topology, complete for its scenario
locals {
  platform_tags = {
    workload = "disaster-recovery"
    managed  = "terraform"
  }
}

module "dr_rg" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-dr-prod"
  location = "eastus"

  tags = local.platform_tags
}

module "app_rg_recovery" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-app-recovery"
  location = "eastus"

  tags = local.platform_tags
}

module "dr_vault" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-recovery-services-vault.git?ref=v1.0.0"

  name                = "rsv-dr-prod"
  resource_group_name = module.dr_rg.name
  location            = module.dr_rg.location
  sku                 = "Standard"

  # LEAVE classic_vmware_replication_enabled AT ITS DEFAULT OF false. It is immutable, and this
  # module sends the MODERNIZED payload -- Microsoft retired the classic experience in March 2026.

  tags = local.platform_tags
}

module "fabric_vmware" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-fabric.git?ref=v1.0.0"

  name                = "fabric-eastus"
  location            = "eastus"
  recovery_vault_name = module.dr_vault.name
  resource_group_name = module.dr_rg.name
}

# EXACTLY ONE protection container in this vault. Not two -- the provider refuses any other count.
module "container_vmware" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-protection-container.git?ref=v1.0.0"

  name                 = "container-eastus"
  recovery_fabric_name = module.fabric_vmware.name
  recovery_vault_name  = module.dr_vault.name
  resource_group_name  = module.dr_rg.name
}

module "vmware_policy" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replication-policy.git?ref=v1.0.0"

  name              = "policy-vmware-24h"
  recovery_vault_id = module.dr_vault.id

  recovery_point_retention_in_minutes                  = 1440
  application_consistent_snapshot_frequency_in_minutes = 240
}

module "vmware_assoc" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replication-policy-association.git?ref=v1.0.0"

  name              = "assoc-vmware"
  recovery_vault_id = module.dr_vault.id
  policy_id         = module.vmware_policy.id
}

module "vnet_recovery" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-virtual-network.git?ref=v1.0.0"

  name                = "vnet-recovery"
  resource_group_name = module.app_rg_recovery.name
  location            = module.app_rg_recovery.location
  address_space       = ["10.20.0.0/16"]

  subnets = {
    "snet-app"  = { address_prefixes = ["10.20.1.0/24"] }
    "snet-data" = { address_prefixes = ["10.20.2.0/24"] }
  }
}

# A SEPARATE drill network, per Microsoft's guidance.
module "vnet_drill" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-virtual-network.git?ref=v1.0.0"

  name                = "vnet-drill"
  resource_group_name = module.app_rg_recovery.name
  location            = module.app_rg_recovery.location
  address_space       = ["10.30.0.0/16"]

  subnets = {
    "snet-drill" = { address_prefixes = ["10.30.1.0/24"] }
  }
}

module "asr_boot_diagnostics" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-storage-account.git?ref=v1.0.0"

  name                = "stasrbootdiag001"
  resource_group_name = module.app_rg_recovery.name
  location            = module.app_rg_recovery.location

  # Microsoft documents that only STANDARD account types are accepted for boot diagnostics here.
  account_tier             = "Standard"
  account_replication_type = "LRS"

  tags = local.platform_tags
}

module "asr_log_storage" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-storage-account.git?ref=v1.0.0"

  name                = "stasrlogs001"
  resource_group_name = module.dr_rg.name
  location            = module.dr_rg.location

  account_tier             = "Standard"
  account_replication_type = "LRS"

  tags = local.platform_tags
}

resource "azurerm_disk_encryption_set" "recovery" {
  name                = "des-recovery"
  resource_group_name = module.app_rg_recovery.name
  location            = module.app_rg_recovery.location
  key_vault_key_id    = azurerm_key_vault_key.recovery.id

  identity {
    type = "SystemAssigned"
  }
}

resource "azurerm_proximity_placement_group" "recovery" {
  name                = "ppg-recovery"
  resource_group_name = module.app_rg_recovery.name
  location            = module.app_rg_recovery.location
}

# The protected machines. Note the disk style: per-disk, with NO log_storage_account_id anywhere.
module "protected_erp" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name              = "vm-erp-01"
  recovery_vault_id = module.dr_vault.id

  recovery_replication_policy_id = module.vmware_policy.id

  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "ERP-PROD-01"

  target_vm_name           = "vm-erp-01-dr"
  target_resource_group_id = module.app_rg_recovery.id
  target_vm_size           = "Standard_E16s_v5"
  target_zone              = "2"

  target_network_id = module.vnet_recovery.id
  test_network_id   = module.vnet_drill.id

  target_boot_diagnostics_storage_account_id = module.asr_boot_diagnostics.id

  multi_vm_group_name = "erp-tier"
  license_type        = "WindowsServer"

  managed_disks = [
    {
      disk_id                       = "disk-6000C29-os"
      target_disk_type              = "Premium_LRS"
      target_disk_encryption_set_id = azurerm_disk_encryption_set.recovery.id
    },
    {
      disk_id          = "disk-6000C29-data"
      target_disk_type = "Premium_ZRS"
    },
  ]

  network_interfaces = [
    {
      source_mac_address = "00:50:56:9a:bc:de"
      is_primary         = true
      target_subnet_name = "snet-app"
      test_subnet_name   = "snet-drill"
      target_static_ip   = "10.20.1.40"
    },
  ]

  timeouts = {
    create = "300m"
  }

  depends_on = [module.vmware_assoc]
}

module "protected_app_01" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-vmware-replicated-vm.git?ref=v1.0.0"

  name              = "vm-app-01"
  recovery_vault_id = module.dr_vault.id

  recovery_replication_policy_id = module.vmware_policy.id

  appliance_name                  = "asr-appliance-dc1"
  physical_server_credential_name = "erp-root"
  source_vm_name                  = "APP-PROD-01"

  target_vm_name           = "vm-app-01-dr"
  target_resource_group_id = module.app_rg_recovery.id
  target_zone              = "3"

  target_network_id = module.vnet_recovery.id
  test_network_id   = module.vnet_drill.id

  multi_vm_group_name = "erp-tier"

  network_interfaces = [
    {
      source_mac_address = "00:50:56:9a:bc:e0"
      is_primary         = true
      target_subnet_name = "snet-app"
      test_subnet_name   = "snet-drill"
    },
  ]

  depends_on = [module.vmware_assoc]
}

# And the plan that orders them into a failover sequence.
module "erp_recovery_plan" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-site-recovery-replication-recovery-plan.git?ref=v1.0.0"

  name                      = "rp-erp"
  recovery_vault_id         = module.dr_vault.id
  source_recovery_fabric_id = module.fabric_vmware.id
  target_recovery_fabric_id = module.fabric_vmware.id

  boot_recovery_groups = [
    { replicated_protected_items = [module.protected_erp.id] },
    { replicated_protected_items = [module.protected_app_01.id] },
  ]
}

check "every_protected_item_settles" {
  assert {
    condition = alltrue([
      !module.protected_erp.disk_configuration_never_converges,
      !module.protected_app_01.disk_configuration_never_converges,
      !module.protected_erp.test_network_is_the_same_as_the_failover_network,
      module.protected_erp.policy_is_in_the_same_vault_as_this_record,
      module.protected_erp.vault_path_within_the_subscription == module.protected_app_01.vault_path_within_the_subscription,
    ])
    error_message = "One of the protected items either replaces itself on every plan, drills against its own failover network, uses a policy from another vault, or sits in a different vault from its sibling."
  }
}

output "what_terraform_cannot_reach" {
  description = "Named for the gap rather than the achievement. Everything below happens outside this configuration."
  value = {
    appliance_is_deployed_out_of_band = "no Terraform resource in any provider deploys or registers an ASR replication appliance"
    three_names_are_matched_at_apply  = module.protected_erp.three_arguments_are_matched_by_name_at_apply_time
    the_credential_is_privileged      = module.protected_erp.the_credential_named_here_is_root_or_administrator_on_the_source_machine
    one_container_only                = module.protected_erp.the_vault_must_hold_exactly_one_protection_container
    destroy_disables_protection       = module.protected_erp.destroying_this_disables_protection_and_discards_the_replica
  }
}

🔒 This composition is complete for its scenario in Terraform terms, and the closing output says exactly where Terraform stops: the appliance, the vCenter registration and the privileged credential are all out of band, and three of this module's arguments are strings matched against what that out-of-band work produced.

⚠️ Note the single protection container. Adding a second source site to this vault would break every protected item in it, because the provider refuses any container count other than one.

⚠️ Both fabric arguments on the recovery plan are the same fabric here, which is correct for VMware-to-Azure — the source side is the on-premises fabric and the failover target is Azure itself. The plan module reports the equal-sides case rather than rejecting it, which is what makes this composition expressible.

ℹ️ depends_on on both protected items is correct and necessary: the container association must exist before a machine can replicate, and no argument references it.


📥 Inputs

Group Variables
Identity and parents name, recovery_vault_id, recovery_replication_policy_id
The three friendly names appliance_name, physical_server_credential_name, source_vm_name
Failover target target_vm_name, target_resource_group_id, target_vm_size, license_type, multi_vm_group_name
Failover networking target_network_id, test_network_id, network_interfaces
Failover placement target_zone, target_availability_set_id, target_proximity_placement_group_id, target_boot_diagnostics_storage_account_id
Disks — pick ONE style default_log_storage_account_id, default_recovery_disk_type, default_target_disk_encryption_set_id or managed_disks
Universal tail timeouts (no tags — the resource has none)
Full input schemas
Variable Type Default Notes
name string — Force-new. Blank and Resource-ID rules.
recovery_vault_id string — Anchored, plus a rule rejecting a fabric ID. The whole record path derives from it.
recovery_replication_policy_id string — Force-new. Anchored to replicationPolicies.
appliance_name string — Force-new. A friendly name, matched at apply.
physical_server_credential_name string — Force-new. A friendly name for a privileged credential.
source_vm_name string — Force-new. A vCenter display name, matched at apply.
target_vm_name string — Force-new. Not used until a failover happens.
target_resource_group_id string — Anchored to a resource group exactly. Editable.
license_type string null Closed set. No default on purpose — it is a licensing assertion.
multi_vm_group_name string null Force-new. Free-form; nothing validates the group.
target_vm_size string null Shape only. Omitting it lets Azure choose.
target_network_id string null Anchored to a virtual network.
test_network_id string null Anchored. Read back and never updated.
target_availability_set_id string null Anchored. Mutually exclusive with target_zone.
target_zone string null Shape only, plus the exclusion rule (carried here).
target_proximity_placement_group_id string null Anchored.
target_boot_diagnostics_storage_account_id string null Anchored. Standard tiers only, which an ID cannot express.
default_log_storage_account_id string null Anchored. Write-only; see the convergence trap.
default_recovery_disk_type string null Closed set. Write-only.
default_target_disk_encryption_set_id string null Anchored. Write-only.
managed_disks list(object({…})) [] 5 rules: duplicate disk_id, blank disk_id, disk type set, and both nested IDs.
network_interfaces list(object({…})) [] 8 rules, including the MAC format, duplicate MACs, exactly one primary, and the documented target_network_id pairing.
timeouts object({ create, read, update, delete }) {} → 120m / 5m / 90m / 90m Four keys; create is a data-transfer budget.

41 validations. Both nested collections are LISTS rather than keyed maps, matching the provider's block type — and unlike a recovery plan's boot groups, the provider treats neither list's order as meaningful, so the list is a faithful mirror rather than a semantic choice.


🧾 Outputs

59 outputs: 11 passthrough, 38 derived, 10 constant. None is sensitive; this module accepts and emits no secret.

The derived count is high because most of this resource is unobservable. Four values are resolved during the apply, two disk-configuration styles interact in a way that produces a permanent replacement, and several of the most consequential behaviours have no argument at all. Each of those is emitted as something a check block can assert.

Output Description
id, name Identity. The full ID is not knowable before an apply.
recovery_vault_id, recovery_replication_policy_id, target_resource_group_id As configured.
source_vm_name, target_vm_name, appliance_name, physical_server_credential_name Read back after an apply.
target_vm_size The resource attribute, so it reports Azure's choice when the argument is omitted.
vault_name, vault_resource_group_name, vault_subscription_id, policy_name Parsed from the supplied IDs.
vault_path_within_the_subscription, policy_path_within_the_subscription Comparable across the family.
replicated_item_path_prefix_within_the_subscription As much of this record's path as is honestly knowable at plan time.
policy_is_in_the_same_vault_as_this_record A comparison the provider never makes.
target_resource_group_is_in_the_vaults_subscription, target_resource_group_is_the_vaults_resource_group, target_resource_group_shares_the_vaults_blast_radius Topology facts; the second is normally false and correct.
disk_configuration_style, disk_configuration_never_converges, default_disk_fields_and_managed_disks_are_both_set, managed_disks_setting_a_log_storage_account, no_disk_configuration_at_all The convergence trap, from four angles.
managed_disk_count, managed_disk_target_disk_types, managed_disks_mixing_more_than_one_disk_type, managed_disks_using_customer_managed_keys, managed_disks_on_a_zonal_disk_type, managed_disks_on_a_constrained_disk_type Disk review.
network_interface_count, primary_network_interface_mac_addresses, network_interfaces_not_selected_for_failover, network_interfaces_without_a_target_subnet, network_interfaces_without_a_test_subnet, network_interfaces_with_a_static_target_ip NIC review.
failover_network_is_configured, test_network_is_configured, test_network_is_the_same_as_the_failover_network Networking and drill readiness.
resilience_model_after_failover, boot_diagnostics_are_configured, target_vm_size_is_chosen_by_azure, claims_azure_hybrid_benefit, joins_a_multi_vm_consistency_group, multi_vm_group_name, proximity_placement_group_is_configured Failover configuration.
create_timeout_is_shorter_than_the_provider_default Whether the data-transfer budget was cut.
the_vault_must_hold_exactly_one_protection_container Constant. The constraint with no argument.
three_arguments_are_matched_by_name_at_apply_time Constant.
the_credential_named_here_is_root_or_administrator_on_the_source_machine Constant. The security fact.
test_network_id_is_read_back_but_never_updated Constant.
create_polls_for_protection_then_makes_a_second_call_for_the_nics Constant.
a_stuck_replication_consumes_the_whole_create_budget Constant.
an_unrelated_update_rewrites_azures_current_values Constant.
destroying_this_disables_protection_and_discards_the_replica Constant.
this_is_the_modernized_vmware_experience Constant.
the_mac_address_is_the_nic_identifier Constant.

🧠 Architecture Notes

Almost nothing here is checkable, and that shapes every decision. Three arguments are friendly names matched case-insensitively against records an on-premises appliance discovered — the appliance's own name against the fabric's process servers, the source machine's name against discovered VMware machines, and a credential's name against the site's Run As accounts, that last one matched together with the appliance name. A fourth value, the protection container, is not an argument at all: the provider lists the vault's containers, refuses any count other than one, and takes the fabric and container names out of the single result. So the create path performs five lookups, and a configuration that passes terraform validate can still fail five different ways at apply. The module's answer is to reject what is knowably wrong (a blank, a Resource ID where a friendly name belongs), say in each message what the value actually is, and name the failure wording so a reader can tell which lookup failed.

The one-container rule is a topology constraint disguised as an error message. A vault that serves two VMware sites, or that also serves Azure-to-Azure replication with its own container, cannot host this resource — and nothing in the schema, the documentation or a plan says so. One vault per source site.

The disk arguments contain a combination that can never converge, and the provider documents it as legal. The three default_* fields are write-only — "not returned by API, and only used in creation" in the provider's own comment — and the read populates neither them nor each disk's log_storage_account_id. The conditional force-new check compares each disk in prior state against the configured default. After the first apply that comparison is an empty string against a set value, so it forces replacement, and the documented escape hatch ("make the values match") cannot be reached through Terraform because the per-disk value never round-trips. managed_disks has the same problem on its own whenever an entry sets a log storage account. Replacement of a replicated item disables protection and re-seeds replication over the network, so this is the most expensive trap in the module — reported as disk_configuration_never_converges, with the assertion shipped in the README, because rejecting input the provider accepts and blesses is the worse error.

test_network_id is a third kind of non-convergence. It is read back from Azure, it is not force-new, and the update payload omits it entirely. So a change produces a plan that applies cleanly, changes nothing, and reappears identically — and the only real correction is replacing the record. That makes the drill-network decision a first-apply decision, which is worth saying out loud because Microsoft's guidance (keep the test network separate from the failover network) is the kind of thing teams intend to tidy up later.

Create is three phases and the timeout is a data-transfer budget. Enable protection, poll every 15 seconds until Azure reports Protected for whatever remains of the 120-minute budget, then a second call to apply the NICs. Two consequences follow. A create that fails after protection succeeded leaves a protected machine whose networking was never applied, and a further apply retries the update rather than recreating. And a stuck replication consumes the entire budget, because the poll's failure test is a string match on states ending in Failed or beginning with Cancel — anything else unhealthy reads as pending, so the apply times out and reports a timeout rather than Azure's actual state.

The update is a read-modify-write, which is unusual enough to state. Every field the plan did not change is re-sent from the record Azure currently holds, so an out-of-band change is preserved rather than corrected by an unrelated apply. Three optional IDs also fail to map an empty value to unset, unlike the availability set and zone, which do — so clearing them transmits an empty string. And an update to any field fails outright if Azure reports no NICs.

The documentation and the schema enforce opposite directions of one pairing. The note says target_network_id is required when network_interface is specified; the RequiredWith makes target_network_id require a NIC. So the provider rejects a network without NICs and accepts NICs without a network, which is the reverse of what it documents. This module covers the documented direction.

A NIC is identified by the source machine's MAC address, sent as the ARM record's NIC id and read back from the same field. And is_primary alone decides whether a NIC is failed over at all — the provider derives the ARM selected-for-failover flag from it with no separate argument, so every secondary interface is created not-selected and no configuration changes that.

A destroy is not a deletion. The delete sends a disable-protection request: the source machine is untouched and the Azure replica plus its recovery points are discarded. That is the same shape as the Azure-to-Azure replicated VM in this family, and the opposite of the recovery plan, where a destroy removes only an ordering.


🧱 Design Principles

There is no secure-defaults table on this module, and the scaffold that claimed one was wrong. Eight required arguments and no empty call; no public-access toggle, no TLS setting, no protection switch. What this resource governs is whether a disaster recovery works — correctness and preparedness — which this suite's secure-by-default rule does not cover. The scaffold this module replaced asserted "the empty call yields the hardened resource" and offered public access, weaker TLS, disabled protections as the opt-outs; none of those exists here, and it also named a resource_group_name argument the resource does not have and Backup Contributor as the role.

The choices that are made, and why:

Decision Value Reasoning
The three friendly names Shape validated, existence not No rule can confirm a name exists in an environment Terraform never created. Each rule rejects a blank and a Resource ID, and each message says what the value is and where to read it.
The documented network_interfaces → target_network_id pairing Enforced Documented by the provider and enforced in the opposite direction by its schema.
The zone / availability-set exclusion Restated The provider enforces it too; restating it lets the message name both fields and recommend one.
Both mutual-exclusion notes on the disk styles Reported, not enforced The provider's notes both forbid and permit the combination, and it accepts it. Rejecting legal input is the worse error — so the flag and the check block carry it.
disk_configuration_never_converges Emitted, and asserted in the README The failure is a permanent replacement with no error message anywhere. This is the highest-value thing the module can say.
license_type No default It is a licensing assertion with a billing consequence. Secure-by-default covers exposure, not entitlement.
target_vm_size Emitted as the RESOURCE attribute Omitting the argument lets Azure choose from the source machine's shape, and the attribute is the only window onto that choice.
Zone membership, VM size availability, disk-type constraints Shape only All three depend on region and capacity at failover time and are unpublished. This suite does not approximate them.
Both nested collections as lists Yes The provider's blocks are lists and it treats neither order as meaningful, so a list is a faithful mirror. Contrast a recovery plan's boot groups, where order is the behaviour.
policy_is_in_the_same_vault_as_this_record Reported The provider compares them nowhere and the module has both IDs.
physical_server_credential_name Accepted as a NAME, never a secret The credential lives on the appliance. The module's contribution is documenting what it can do, since no Azure role bounds it.
Path outputs omit the subscription Yes Keeps every path output in this family byte-comparable; proved across four modules and three vault conventions.
No tags variable Correct The resource exposes neither tags nor location.

🚀 Runbook

terraform init -backend=false
terraform fmt -check
terraform validate

Pin the module with ?ref=v1.0.0 and never a branch. This library is plan-only: a human applies from CI.


🧪 Testing

terraform validate proves the configuration parses, both nested block families exist and the types line up. It does not fire a variable validation when this module is called from another configuration — terraform console with a -var-file is the offline harness for that.

What was proven offline for this module:

  • All 41 validations fired, each by a fixture built to fail it, with zero condition-evaluation errors on the first run. Every bad fixture is the valid base with exactly one thing replaced, so no fixture can fire a rule other than the one it targets.
  • All 51 locals were driven to more than one value across six good fixtures: the minimal call, a full per-disk-and-networking case, the defaults style with an availability set and a shared test network, both disk styles together, a policy in another vault, and a second vault entirely.
  • The sixth fixture was added because of the proof, not before it. The first five shared one vault, which left the six vault-derived locals never-varying — and since each references var.recovery_vault_id, that was a fixture gap rather than a constant in the wrong place.
  • vault_path_within_the_subscription was evaluated in four modules — this one, the Azure-to-Azure replicated VM, the fabric and the recovery plan — from three different vault conventions (an ID alone, a name plus resource group, and an ID), and came out byte-identical. Without that, every cross-module check comparing them would be silently useless.
  • One output name baked in a wrong count and was renamed before shipping: it claimed four arguments were resolved by name lookup, when three are arguments and the fourth is the protection container, which is not an argument at all.
  • The terraform validate result was confirmed with the working directory printed, its .tf files listed, and a zz_negctl.tf negative control in the module's own directory: it fails with the control present and passes with it removed. That matters because terraform validate reports "Success!" in a directory containing no .tf files at all.

What only an apply exercises: all five name lookups, whether the vault really holds one container, the three-phase create, and every behaviour that happens at failover rather than at create.


💬 Example Output

id = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod/replicationFabrics/fabric-eastus/replicationProtectionContainers/container-eastus/replicationProtectedItems/vm-erp-01"
name = "vm-erp-01"
source_vm_name = "ERP-PROD-01"
target_vm_name = "vm-erp-01-dr"
appliance_name = "asr-appliance-dc1"
physical_server_credential_name = "erp-root"
target_vm_size = "Standard_E16s_v5"
vault_path_within_the_subscription = "/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod"
replicated_item_path_prefix_within_the_subscription = "/resourceGroups/rg-dr-prod/providers/Microsoft.RecoveryServices/vaults/rsv-dr-prod"
policy_is_in_the_same_vault_as_this_record = true
disk_configuration_style = "per_disk"
disk_configuration_never_converges = false
managed_disk_count = 2
managed_disks_mixing_more_than_one_disk_type = true
managed_disks_on_a_zonal_disk_type = ["disk-6000C29-data"]
network_interface_count = 1
primary_network_interface_mac_addresses = ["00:50:56:9a:bc:de"]
network_interfaces_not_selected_for_failover = []
test_network_is_the_same_as_the_failover_network = false
resilience_model_after_failover = "availability_zone"
create_timeout_is_shorter_than_the_provider_default = false
the_vault_must_hold_exactly_one_protection_container = true
the_credential_named_here_is_root_or_administrator_on_the_source_machine = true

Note replicated_item_path_prefix_within_the_subscription stopping at the vault: the fabric and container segments visible in id were discovered during the apply and were not knowable when the plan was made.


🔍 Troubleshooting

Symptom Cause Fix
there should be only 1 protection container in Recovery Vault, get: 2 The provider discovers the container by listing the vault and accepts exactly one Use one vault per source site. Nothing in this resource can select among containers.
fetch process server id: ... process server <name> not found appliance_name does not match a process server on the vault's fabric Read the appliance's friendly name from the vault's Site Recovery infrastructure blade. Microsoft: the name cannot be changed once set.
fetch discovery machine id: ... machine <name> not found source_vm_name does not match a discovered machine's vCenter display name Confirm the machine appears under the vault's discovered machines; the appliance must have discovered it first.
fetch run as account id: ... run as account <name> not found physical_server_credential_name does not match, or appliance_name beside it does not The pair is matched together. Check both before assuming the credential name is wrong.
retrieving ...: Detail Type mismatch The vault's fabric is not a modernized VMware fabric The vault is configured for another scenario, most likely the classic VMware experience Microsoft retired in March 2026. Use a vault with classic_vmware_replication_enabled = false.
Every plan proposes a replacement and nothing in the configuration changed The disk-configuration trap: both styles set, or a per-disk log_storage_account_id Read disk_configuration_never_converges. Use one style, and do not set a per-disk log storage account.
A change to test_network_id applies but nothing happens, and the diff returns The update payload omits the field entirely Set it correctly at creation. Correcting it later requires replacing the record.
The apply times out after 120 minutes with no useful message The poll only recognises states ending in Failed or beginning with Cancel; anything else reads as pending Check the protection state in the portal. Raise the create timeout for a large machine — it is a data-transfer budget.
retrieving network_interface: VMNics was nil on an unrelated change The update reads existing NICs when network_interfaces has not changed, and Azure reported none Configure network_interfaces explicitly, or investigate why the record has no NICs.
at most one network_interfaces entry may set is_primary = true Two primaries; the provider does not check this Keep one primary. Note that is_primary also decides whether a NIC is failed over.
A secondary NIC did not come back after a failover Expected: the selected-for-failover flag is derived from is_primary alone Read network_interfaces_not_selected_for_failover. There is no argument that changes this.
Destroying the module did not remove the source VM Correct — the delete disables protection The source machine is untouched; the Azure replica and its recovery points are gone and replication must start again.

🔗 Related Docs


💙 "Infrastructure as Code should be standardized, consistent, and secure."