Skip to content

Latest commit

Β 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

☁️ Azure Local Instance (Stack HCI Cluster) Terraform Module

Registers one Azure Local (Azure Stack HCI) instance in Azure Resource Manager β€” the root of the Azure Local family β€” targeting hashicorp/azurerm ~> 4.0.

Terraform azurerm Module Type Resources Posture


🧩 Overview

  • πŸ›οΈ Creates one azurerm_stack_hci_cluster named this β€” the ARM record for an Azure Local instance, and the root every other module in this family descends from.
  • 🚨 States plainly, in an output, that creating this resource deploys nothing: no nodes joined, no storage, no networking, and no custom location. That is azurerm_stack_hci_deployment_setting's job.
  • 🧭 Emits arc_setting_id β€” the <id>/arcSettings/default string the extension module needs β€” composed once, correctly, instead of by hand.
  • πŸ”‘ Emits resource_provider_object_id and explains that this, not the instance's own managed identity, is the principal that reads the deployment key vault. Both are GUIDs.
  • πŸͺͺ Takes the instance's Entra client_id as resource data, argues why that does not violate this suite's auth-is-the-caller's rule, and never accepts the application secret.
  • 🏷️ Carries tags β€” and notes it is the only taggable resource in most of this family.

πŸ’‘ Why it matters: the gap between "the apply went green" and "the instance works" is wider here than anywhere else in this library. A successful apply gives you a registered, empty cluster record. Everything a security review would want to see β€” BitLocker, SMB signing, Credential Guard, HVCI, drift control β€” lives on the deployment setting, and so does the custom location five sibling modules cannot function without.


❀️ Support this project

If this module saved you time:


πŸ—ΊοΈ Where this fits in the family

flowchart TB
  RG["terraform-azurerm-resource-group"]
  ARCM["Azure Arc machine resources (registered per node)"]
  CLU["terraform-azurerm-stack-hci-cluster"]
  DS["terraform-azurerm-stack-hci-deployment-setting"]
  ARC["arcSettings/default (created by the deployment)"]
  EXT["terraform-azurerm-stack-hci-extension"]
  CL["custom location (created by the deployment)"]
  CHILD["five custom-location children: logical-network, storage-path, marketplace-gallery-image, virtual-hard-disk, network-interface"]

  RG -->|"resource_group_name"| CLU
  ARCM -->|"arc_resource_ids"| DS
  CLU -->|"id as stack_hci_cluster_id"| DS
  DS -->|"deployment creates"| CL
  DS -->|"deployment creates"| ARC
  CLU -->|"id plus /arcSettings/default"| EXT
  ARC -->|"must exist first"| EXT
  CL -->|"custom_location_id"| CHILD

  style CLU fill:#0078D4,stroke:#004578,color:#ffffff
  style DS fill:#004578,stroke:#00335c,color:#ffffff
  style RG fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style ARCM fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style ARC fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style EXT fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style CL fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style CHILD fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
Loading

Read the diagram outward from the highlighted node and the ownership boundaries fall out. This module creates the instance record. The deployment setting consumes its id and β€” as a side effect of deploying β€” brings the custom location and the Arc setting into existence. Only then can the five custom-location children and the extension be created. This module is the only one in the family that consumes nothing from it.


🧬 What this module builds

flowchart TB
  IN["required: name, resource_group_name, location"]
  IDENT["identity object: type -- only SystemAssigned is documented"]
  ENTRA["entra_application object: client_id, tenant_id -- the CLUSTER's app registration, not Terraform's credential"]
  AM["automanage_configuration_id -- hands part of the config to another service"]
  THIS["azurerm_stack_hci_cluster.this"]
  ARC["Azure Arc"]
  NODES["physical machines in your own datacentre"]
  OUTID["id -- Microsoft.AzureStackHCI/clusters"]
  COMPUTED["read back after apply: cloud_id, service_endpoint, resource_provider_object_id, identity principal_id, effective tenant_id"]
  DERIVED["derived: arc_setting_id, has_identity, uses_entra_application, entra_tenant_was_left_to_the_provider, is_automanaged"]
  NOTDEPLOYED["NOT created here: nodes joined, storage, networking, custom location"]

  IN --> THIS
  IDENT --> THIS
  ENTRA --> THIS
  AM --> THIS
  THIS -->|"registers through"| ARC
  ARC -->|"represents"| NODES
  THIS --> OUTID
  THIS --> COMPUTED
  THIS --> DERIVED
  THIS -.->|"a deployment setting does this, not this resource"| NOTDEPLOYED

  style THIS fill:#0078D4,stroke:#004578,color:#ffffff
  style OUTID fill:#004578,stroke:#00335c,color:#ffffff
  style IN fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style IDENT fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style ENTRA fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style AM fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style ARC fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style NODES fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style COMPUTED fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style DERIVED fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
  style NOTDEPLOYED fill:#f2f4f7,stroke:#98a2b3,color:#1d2939
Loading

Resource inventory

Resource Count Notes
azurerm_stack_hci_cluster 1 (this) The keystone. No owned children.

The dotted edge is the important one: it points at what this resource does not create.


βœ… Provider / Versions

Requirement Value
Terraform >= 1.12.0
hashicorp/azurerm ~> 4.0
Provider block None in this module. The caller configures provider "azurerm" { features {} }, authentication and subscription.
Module version v1.0.0

Schema notes that bite β€” argument names, types and block cardinality from the pinned provider's binary schema; force-new facts from the provider's own resource documentation, because the binary schema carries no force-new marker at all; deployment behaviour from Microsoft's Azure Local documentation:

  • πŸ”΄ Creating this resource does not deploy anything. It is a registration. Microsoft documents the real deployment as taking 2.5–3 hours, with one step ("Deploy Moc and ARB Stack") accounting for 40–45 minutes β€” and none of that happens here.
  • πŸ”΄ The custom location the family's five children need is created by the deployment, not here. This resource has no custom-location argument and no custom-location output, and the ordering is not expressible as a Terraform dependency because the custom location is not a resource Terraform creates.
  • πŸ”΄ client_id and tenant_id are the instance's Entra application, not provider authentication. The provider's own wording is "used by" the cluster.
  • ⚠️ tenant_id is optional and computed β€” omitted, it defaults to the provider's tenant. That is the one place provider configuration determines this resource's data, and it is force-new, so a pipeline authenticated against a different tenant replaces the instance rather than updating it.
  • ⚠️ identity is max = 1 with a required type inside, and SystemAssigned is the only documented value. Modelled here as one optional object β€” but the resource attribute is still a list, so reading principal_id needs try(...identity[0].principal_id, null).
  • ⚠️ Microsoft documents that the instance name must differ from every node name β€” unenforceable here, because the node names are inputs to the deployment setting. That module enforces it.
  • ⚠️ A destroy is not a decommission. It removes the registration and leaves deployed software and Arc machine resources behind.
  • ℹ️ cloud_id is not the ARM id β€” it is an immutable platform UUID, and the Resource ID is not immutable across a recreation.
  • ℹ️ resource_provider_object_id is the resource provider's service principal, not the instance's managed identity. It is the one that reads the deployment key vault.
  • ℹ️ Timeouts default to 30 minutes create/update/delete, 5 minutes read. Ample β€” the long work is elsewhere.
  • ℹ️ The product is Azure Local; ARM and Terraform still say Stack HCI. A rename, not a deprecation, and this is the resource where the confusion costs most because it is what a policy definition targets.

πŸ”‘ Required Azure RBAC Roles / Permissions

Scope Role / permission Why
The target resource group Azure Stack HCI Administrator The role for creating and managing the instance record itself.
The target resource group Azure Connected Machine Onboarding + Azure Connected Machine Resource Administrator Held by whoever registers the physical machines with Azure Arc. Not needed to create this resource β€” but the machines must already be registered for the deployment that follows, so the grants are usually made together.
The Automanage configuration profile (only when assigned) Reader A principal that cannot see the profile cannot confirm the ID it is passing.

πŸ”’ Neither plan nor state access is credential access here. This resource has no secret-bearing attribute at all. client_id is an application ID β€” an identifier that appears in sign-in logs, deliberately not sensitive β€” and the application's secret is never accepted by this module, so it never enters this state file. resource_provider_object_id, identity.principal_id, cloud_id and tenant_id are likewise identifiers. That is a per-resource conclusion; the deployment setting, one step away, reaches a different one for a different reason.

⚠️ resource_provider_object_id versus identity_principal_id. Two different service principals are involved in an Azure Local deployment, both are GUIDs, and the key vault grant belongs to the resource provider's. Granting the wrong one produces a permissions failure that names neither.

πŸ”΄ What no Azure role constrains: a destroy here is not a decommission. Deleting this resource removes the Azure registration; it does not un-deploy software on the nodes, un-join them, or clean up the Arc machine resources. Whoever can delete it can strand a working instance in a state that is neither registered nor cleanly reusable.


Azure Prerequisites

  • The Microsoft.AzureStackHCI resource provider registered on the subscription. Microsoft.ExtendedLocation and Microsoft.HybridCompute are needed for what follows.
  • The physical machines registered with Azure Arc, all on the same OS version, with the same network adapter configuration, and with deployment permissions assigned. Microsoft treats this as a prerequisite of the deployment rather than of the record β€” but there is little point creating the record without it.
  • A region where Azure Local is available. Not enforced here: Microsoft adds to that list, and a closed check would reject a newly supported region.
  • An Entra application registration, if the instance is to use one. Created out of band; only its client ID comes in here.
  • The caller configures provider "azurerm" { features {} }; this module declares no provider block.

πŸ“ Module Structure

terraform-azurerm-stack-hci-cluster/
β”œβ”€β”€ providers.tf    # required_version >= 1.12.0, azurerm ~> 4.0. No provider block.
β”œβ”€β”€ variables.tf    # 8 typed inputs, 13 validations, the tags + timeouts tail.
β”œβ”€β”€ main.tf         # locals for every derived fact, then the single keystone resource.
β”œβ”€β”€ outputs.tf      # id first, then the computed trio, the derived flags, and the constants.
β”œβ”€β”€ README.md       # this file.
β”œβ”€β”€ SCOPE.md        # the cross-module contract.
β”œβ”€β”€ LICENSE         # MIT.
└── .gitignore      # the canonical library ignore file.

βš™οΈ Quick Start

provider "azurerm" {
  features {}
}

module "azlocal" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                = "azlocal-cluster-01"
  resource_group_name = "rg-azlocal"
  location            = "eastus2"

  identity = {
    type = "SystemAssigned"
  }
}

ℹ️ The caller owns the provider block, authentication and the mandatory features {} β€” this module declares none of them.

πŸ”΄ That is a registration, not a deployment. After a green apply you have an empty cluster record. Add terraform-azurerm-stack-hci-deployment-setting to make it an instance β€” see example 12.


πŸ”Œ Cross-Module Contract

Consumes

Input Type Typical source
resource_group_name string module.rg.name from terraform-azurerm-resource-group
name, location string Caller
identity object({ type }) (optional) Caller
entra_application object({ client_id, tenant_id }) (optional) An Entra app registration created out of band
automanage_configuration_id string (optional) An Automanage configuration profile
tags, timeouts map(string), object Caller

Emits

Output Consumed by
id terraform-azurerm-stack-hci-deployment-setting β†’ stack_hci_cluster_id
arc_setting_id terraform-azurerm-stack-hci-extension β†’ arc_setting_id
resource_provider_object_id A role assignment on the deployment key vault
identity_principal_id Role assignments, diagnostic settings
cloud_id, service_endpoint, tenant_id, client_id Support cases, firewall allow-lists, review
The derived flags and constants Change, access and security review

πŸ“š Example Library

1 Β· The smallest registration
module "azlocal" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                = "azlocal-cluster-01"
  resource_group_name = "rg-azlocal"
  location            = "eastus2"
}

ℹ️ Three arguments is the floor β€” the provider requires all three, so there is no empty call.

πŸ”΄ And it is only a registration. creating_this_resource_does_not_deploy_anything reads true for every call this module can make.

2 Β· With a system-assigned identity, which you almost always want
module "azlocal" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                = "azlocal-cluster-01"
  resource_group_name = "rg-azlocal"
  location            = "eastus2"

  identity = {
    type = "SystemAssigned"
  }
}

output "instance_principal" {
  value = module.azlocal.identity_principal_id
}

πŸ’‘ The identity is what a diagnostic setting, a policy assignment or a key-vault grant targets. Adding it later is a real change to the resource, not a no-op, and anything waiting on the principal ID has to be re-applied.

ℹ️ SystemAssigned is the only value the pinned provider documents β€” there is no user-assigned option on an Azure Local instance. The module models it as an object with one field anyway, so a second type would be a new value rather than a breaking interface change.

3 Β· The instance's own Entra application
module "azlocal" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                = "azlocal-cluster-01"
  resource_group_name = "rg-azlocal"
  location            = "eastus2"

  entra_application = {
    client_id = var.azlocal_app_client_id
    tenant_id = var.entra_tenant_id
  }
}

πŸ”’ This is resource data, not Terraform's credential, and the distinction is worth being precise about. This suite's rule is that authentication is entirely the caller's provider configuration β€” and that rule is about the credential Terraform uses. These two fields identify the application registration the cluster uses for its own Azure communication. Two different principals; only one of them is Terraform's.

πŸ”΄ The module accepts the client ID and will never accept the secret. An application ID is an identifier, not a credential. The secret is provisioned out of band β€” Microsoft's guidance puts it in the key vault the deployment setting points at.

⚠️ Both fields are force-new. Rotating to a different app registration replaces the instance's ARM record.

4 Β· Letting the provider's tenant govern β€” and knowing that it did
module "azlocal" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                = "azlocal-cluster-01"
  resource_group_name = "rg-azlocal"
  location            = "eastus2"

  entra_application = {
    client_id = var.azlocal_app_client_id
    # tenant_id deliberately omitted
  }
}

output "tenant_came_from_the_provider" {
  value = module.azlocal.entra_tenant_was_left_to_the_provider # true
}

output "effective_tenant" {
  value = module.azlocal.tenant_id # read back from the resource
}

⚠️ This is the one place provider configuration determines this resource's data. tenant_id is optional and computed, and omitted it defaults to the provider's tenant. Usually right β€” and it means a configuration that is correct today silently changes meaning if applied by a pipeline authenticated against a different tenant.

πŸ”΄ And tenant_id is force-new, so that change is a replacement, not an update.

πŸ’‘ Read tenant_id for the effective value; the flag says only where it came from.

5 Β· Getting the extension's `arc_setting_id` right
module "monitor_agent" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-extension.git?ref=v1.0.0"

  name = "AzureMonitorWindowsAgent"

  # Composed once, correctly, by the cluster module.
  arc_setting_id = module.azlocal.arc_setting_id

  publisher = "Microsoft.Azure.Monitor"
  type      = "AzureMonitorWindowsAgent"
}

πŸ’‘ The extension does not take a bare cluster ID β€” it takes a child record of it, conventionally named default. Appending /arcSettings/default by hand is the most common mistake in this family, so the module does it once.

⚠️ default is a convention, not a guarantee. Azure Local creates the Arc setting under that name during deployment; this module does not create it. The output is a correctly-formed ID for the setting that normally exists, not proof that it does.

6 Β· Granting the resource provider access to the deployment key vault
resource "azurerm_role_assignment" "rp_reads_deployment_secrets" {
  scope                = var.deployment_key_vault_id
  role_definition_name = "Key Vault Secrets User"

  # The RESOURCE PROVIDER's service principal -- not the instance's identity.
  principal_id = module.azlocal.resource_provider_object_id
}

πŸ”‘ This is the grant that is easiest to miss entirely. Microsoft's deployment guidance stores the instance's deployment secrets β€” the local administrator credentials among them β€” in a key vault, and the resource provider is the principal that reads them.

πŸ”΄ Not identity_principal_id. Two different service principals, both GUIDs, and granting the wrong one produces a permissions failure that names neither. The output names are deliberately different lengths so they do not look interchangeable in a plan.

ℹ️ It is an object ID, so it belongs in principal_id and nowhere an application (client) ID is wanted.

7 Β· Assigning an Automanage configuration profile
module "azlocal" {
  source                      = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                        = "azlocal-cluster-01"
  resource_group_name         = "rg-azlocal"
  location                    = "eastus2"

  automanage_configuration_id = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-governance/providers/Microsoft.Automanage/configurationProfiles/prod-baseline"
}

output "also_managed_elsewhere" {
  value = module.azlocal.is_automanaged # true
}

⚠️ A true means something other than Terraform is also configuring this instance. That is the purpose of Automanage β€” and it means a later drift may originate from the profile rather than from a change here, with nothing in a Terraform plan to say so.

πŸ’‘ The ID is anchored at both ends, and both plausible wrong values are rejected by name: this instance's own ID, and a custom location ID. See example 11.

8 Β· Tagging, and why it matters more here than it looks
module "azlocal" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                = "azlocal-cluster-01"
  resource_group_name = "rg-azlocal"
  location            = "eastus2"

  tags = {
    environment  = "production"
    platform     = "azure-local"
    cost_centre  = "infra-4102"
    data_class   = "internal"
  }
}

⚠️ This is the only taggable resource in most of this family. Neither azurerm_stack_hci_deployment_setting nor azurerm_stack_hci_extension exposes a tags attribute at all, so a governance query expecting to find every Azure Local component by tag will find this instance and the custom-location children, and nothing else.

πŸ’‘ It is also the only argument on this resource that is not force-new.

9 Β· Reading the computed trio
output "instance_facts" {
  value = {
    arm_id           = module.azlocal.id           # changes if recreated
    platform_uuid    = module.azlocal.cloud_id     # immutable
    data_endpoint    = module.azlocal.service_endpoint
    rp_principal     = module.azlocal.resource_provider_object_id
  }
}

ℹ️ cloud_id is not the ARM id. The Resource ID changes if the instance is recreated in a different resource group; cloud_id is the platform's own immutable identifier and is what Azure Local tooling and support cases key off.

πŸ’‘ service_endpoint is the outbound endpoint that has to be reachable when the instance sits behind a proxy or a restrictive firewall. It is determined by location rather than chosen.

⚠️ All three are computed, so all three are unknown at plan time.

10 Β· Several instances, keyed by site
module "azlocal_sites" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  for_each = {
    dc1 = "eastus2"
    dc2 = "centralus"
  }

  name                = "azlocal-${each.key}"
  resource_group_name = "rg-azlocal-${each.key}"
  location            = each.value

  identity = { type = "SystemAssigned" }

  tags = { site = each.key }
}

output "sites_missing_an_identity" {
  value = [for k, m in module.azlocal_sites : k if !m.has_identity]
}

πŸ’‘ for_each over a keyed map, never count. Removing dc1 later must not re-index dc2 β€” and here re-indexing would mean destroying and recreating a cluster registration, which strands a deployed instance.

ℹ️ Each site needs its own deployment setting; the registration is per-instance and so is everything below it.

11 Β· The rejections, and what each message says instead of "invalid"
# Rejected: the instance's own ID, pasted into the Automanage field.
automanage_configuration_id = module.azlocal.id

# Rejected: a custom location -- which is not an input to this resource at all.
automanage_configuration_id = var.custom_location_id

# Rejected: only "SystemAssigned" is documented, and the casing is checked.
identity = { type = "systemassigned" }

# Rejected: a tenant's domain name is not a tenant ID.
entra_application = { tenant_id = "contoso.onmicrosoft.com" }

# Rejected: a display name, and upper case.
location = "East US 2"

πŸ”’ Each message names the wrong thing rather than reporting a bad string. The custom-location rejection is the one worth reading twice: it says the custom location is not an input to this resource because the deployment creates it β€” which is the fact a reader actually needs, not a note about ID formats.

πŸ’‘ The Automanage regex is anchored at both ends, so a child ID or a truncated one is rejected too.

ℹ️ A client secret pasted into client_id is also rejected, because a secret is not a GUID β€” and the message says the secret does not belong in this module at any price.

12 Β· Registration is not deployment: the two-module minimum
module "azlocal" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                = "azlocal-cluster-01"
  resource_group_name = "rg-azlocal"
  location            = "eastus2"
  identity            = { type = "SystemAssigned" }
}

# Without this, the instance above is an empty record forever.
module "deployment" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-deployment-setting.git?ref=v1.0.0"

  stack_hci_cluster_id = module.azlocal.id
  arc_resource_ids     = var.arc_machine_ids
  deployment_version   = "10.0.0.0"

  scale_unit = [{
    active_directory_organizational_unit_path = "OU=azlocal,DC=contoso,DC=com"
    domain_fqdn                               = "contoso.com"
    name_prefix                               = "azl"
    secrets_location                          = "https://kv-azlocal-01.vault.azure.net/"

    cluster = {
      name                  = module.azlocal.name
      azure_service_endpoint = "core.windows.net"
      cloud_account_name     = "stazlocalwitness01"
      witness_type           = "Cloud"
      witness_path           = "\\\\fileserver\\witness"
    }

    optional_service = {
      # The custom location the five sibling modules need comes from HERE.
      custom_location = "cl-azlocal-eastus2"
    }

    storage = { configuration_mode = "Express" }

    physical_node = [
      { name = "azl-node-1", ipv4_address = "10.60.0.11" },
      { name = "azl-node-2", ipv4_address = "10.60.0.12" },
    ]

    infrastructure_network = [{
      gateway     = "10.60.0.1"
      subnet_mask = "255.255.255.0"
      dns_server  = ["10.60.1.10"]
      ip_pool     = [{ starting_address = "10.60.0.20", ending_address = "10.60.0.40" }]
    }]

    host_network = {
      intent = [{
        name         = "converged"
        adapter      = ["Port1", "Port2"]
        traffic_type = ["Compute", "Storage", "Management"]
      }]
      storage_network = [{
        name                 = "Storage1"
        network_adapter_name = "Port1"
        vlan_id              = "711"
      }]
    }
  }]
}

πŸ”΄ This is the minimum viable Azure Local configuration, and it is two modules. The first is a record; the second is a three-hour deployment.

πŸ’‘ optional_service.custom_location is where the custom location gets its name β€” the value five sibling modules then need as custom_location_id. It does not exist until the deployment completes.

⚠️ cluster.name is deliberately wired from module.azlocal.name. Microsoft documents that the instance name must differ from every node name, and the deployment setting module enforces exactly that β€” because there the cluster name and the node names are inputs to the same variable.

13 Β· πŸ—οΈ End-to-end composition
provider "azurerm" {
  features {}
}

variable "arc_machine_ids" {
  type        = list(string)
  description = "Microsoft.HybridCompute/machines IDs for the nodes, registered with Azure Arc beforehand."
}

variable "deployment_key_vault_id" {
  type        = string
  description = "The key vault holding the deployment secrets. Created out of band; only its ID and URI come in here."
}

variable "deployment_key_vault_uri" {
  type        = string
  description = "The vault URI, e.g. https://kv-azlocal-01.vault.azure.net/"
}

locals {
  # The resource group module takes tags as an input but does not emit them, so the
  # composition owns the canonical map. Neither the deployment setting nor the
  # extension accepts tags at all.
  common_tags = { environment = "production", platform = "azure-local" }
}

module "rg" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-azlocal"
  location = "eastus2"

  tags = local.common_tags
}

# --- 1. The registration: this module. Fast, and deploys nothing. ---
module "azlocal" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-cluster.git?ref=v1.0.0"

  name                = "azlocal-cluster-01"
  resource_group_name = module.rg.name
  location            = "eastus2"

  identity = { type = "SystemAssigned" }

  tags = local.common_tags
}

# --- 2. The grant the deployment cannot proceed without. ---
resource "azurerm_role_assignment" "rp_reads_deployment_secrets" {
  scope                = var.deployment_key_vault_id
  role_definition_name = "Key Vault Secrets User"
  principal_id         = module.azlocal.resource_provider_object_id
}

# --- 3. The deployment: 2.5-3 hours, and the source of the custom location. ---
module "deployment" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-deployment-setting.git?ref=v1.0.0"

  stack_hci_cluster_id = module.azlocal.id
  arc_resource_ids     = var.arc_machine_ids
  deployment_version   = "10.0.0.0"

  scale_unit = [{
    active_directory_organizational_unit_path = "OU=azlocal,DC=contoso,DC=com"
    domain_fqdn                               = "contoso.com"
    name_prefix                               = "azl"
    secrets_location                          = var.deployment_key_vault_uri

    # Security defaults, all explicit.
    bitlocker_boot_volume_enabled  = true
    bitlocker_data_volume_enabled  = true
    credential_guard_enabled       = true
    drtm_protection_enabled        = true
    hvci_protection_enabled        = true
    side_channel_mitigation_enabled = true
    smb_cluster_encryption_enabled = true
    smb_signing_enabled            = true
    wdac_enabled                   = true
    drift_control_enabled          = true

    cluster = {
      name                   = module.azlocal.name
      azure_service_endpoint = "core.windows.net"
      cloud_account_name     = "stazlocalwitness01"
      witness_type           = "Cloud"
      witness_path           = "\\\\fileserver\\witness"
    }

    optional_service = { custom_location = "cl-azlocal-eastus2" }
    storage          = { configuration_mode = "Express" }

    physical_node = [
      { name = "azl-node-1", ipv4_address = "10.60.0.11" },
      { name = "azl-node-2", ipv4_address = "10.60.0.12" },
    ]

    infrastructure_network = [{
      gateway     = "10.60.0.1"
      subnet_mask = "255.255.255.0"
      dns_server  = ["10.60.1.10", "10.60.1.11"]
      ip_pool     = [{ starting_address = "10.60.0.20", ending_address = "10.60.0.40" }]
    }]

    host_network = {
      intent = [{
        name         = "converged"
        adapter      = ["Port1", "Port2"]
        traffic_type = ["Compute", "Storage", "Management"]
      }]
      storage_network = [{
        name                 = "Storage1"
        network_adapter_name = "Port1"
        vlan_id              = "711"
      }]
    }
  }]

  timeouts = { create = "8h" }

  depends_on = [azurerm_role_assignment.rp_reads_deployment_secrets]
}

# --- 4. A cluster-level agent, on the Arc setting the deployment created. ---
module "monitor_agent" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-stack-hci-extension.git?ref=v1.0.0"

  name           = "AzureMonitorWindowsAgent"
  arc_setting_id = module.azlocal.arc_setting_id
  publisher      = "Microsoft.Azure.Monitor"
  type           = "AzureMonitorWindowsAgent"

  depends_on = [module.deployment]
}

output "azure_local_platform" {
  value = {
    instance_id   = module.azlocal.id
    platform_uuid = module.azlocal.cloud_id
    deployment_id = module.deployment.id

    # The name of a custom location that does not exist until the deployment finishes.
    custom_location_name = module.deployment.custom_location_name

    node_count        = module.deployment.node_count
    witness_in_use    = module.deployment.witness_in_use
    security_posture  = module.deployment.security_features_disabled

    agent = module.monitor_agent.extension_identifier
  }
}

πŸ—οΈ The wiring worth noticing is the ordering, and one of the three edges is not a data dependency. The role assignment must exist before the deployment starts, and nothing in the deployment's arguments references it β€” so it is an explicit depends_on, which is exactly the case that rule exists for. The extension likewise waits on the deployment, because the Arc setting it targets does not exist until the deployment creates it.

πŸ”΄ The five custom-location children are deliberately absent from this composition. They need module.deployment.custom_location_name resolved into a Resource ID for a custom location that Terraform never creates β€” so it is not an attribute of anything in the graph. Apply them in a later run, or look the ID up with a data block once the deployment has finished. A composition that tries to do it all in one apply fails on the children, not on the deployment.

⚠️ witness_path needs doubled backslashes twice over β€” "\\\\fileserver\\witness" is the UNC path \\fileserver\witness. A single backslash is an invalid escape sequence and fails to parse before any validation runs.

πŸ’‘ timeouts.create = "8h" against a provider default of six hours and a documented 2.5–3 hour deployment. The default is generous; the override is for a large scale unit or a slow link.


πŸ“₯ Inputs

Required β€” name, resource_group_name, location.

Optional β€” identity, entra_application, automanage_configuration_id, plus the universal tags and timeouts tail.

Full input schemas
Name Type Default Notes
name string β€” Required, force-new. Must differ from every node name β€” enforced in the deployment setting module. Rejects a Resource ID.
resource_group_name string β€” Required, force-new. Rejects a Resource ID.
location string β€” Required, force-new. Places the ARM projection, not the hardware. Rejects a display name and upper case. Region availability is deliberately not enforced.
identity object({ type }) null SystemAssigned only β€” the single documented value, enforced as a closed set. One object rather than a boolean, so a future second type is not a breaking change.
entra_application object({ client_id, tenant_id }) {} The instance's app registration, not Terraform's credential. Both bare GUIDs, both force-new. tenant_id defaults to the provider's tenant. A client secret is never accepted.
automanage_configuration_id string null Force-new. Microsoft.Automanage/configurationProfiles, anchored at both ends. This instance's own ID and a custom location ID are each rejected by name.
tags map(string) {} The only non-force-new argument β€” and the only tags in most of this family.
timeouts object({ create, read, update, delete }) null All four exist; provider defaults 30m/5m/30m/30m. Ample, because the long work is on the deployment setting. Unknown keys discarded silently.

🧾 Outputs

Output Type Notes
id string The instance's Resource ID. Emitted first.
name string
resource_group_name string
location string The ARM record's region.
arc_setting_id string Derived <id>/arcSettings/default. What the extension consumes.
cloud_id string Computed. Immutable platform UUID β€” not the ARM id.
service_endpoint string Computed. The regional data-path endpoint.
resource_provider_object_id string Computed. The principal that reads the deployment key vault.
tenant_id string Computed-or-supplied effective tenant.
client_id string The instance's Entra application ID, or null. Deliberately not sensitive.
automanage_configuration_id string Or null.
automanage_configuration_name string Derived. null when unassigned.
identity object The whole block, or null. Read with [0] internally.
identity_principal_id string Computed. Not the RP principal.
tags map(string)
has_identity bool Derived.
uses_entra_application bool Derived.
entra_tenant_was_left_to_the_provider bool Derived. Where provider config sets resource data.
is_automanaged bool Derived.
creating_this_resource_does_not_deploy_anything bool Constant true. The headline.
the_custom_location_its_children_consume_is_created_by_the_deployment_not_by_this_resource bool Constant true.
the_node_names_live_on_the_deployment_setting_so_the_collision_check_lives_there bool Constant true.
this_is_an_arc_projected_resource_so_azure_state_can_diverge_from_the_hardware bool Constant true.
the_product_is_now_called_azure_local_but_arm_and_terraform_still_say_stack_hci string The ARM type.
secure_by_default_has_almost_nothing_to_act_on_here bool Constant true, with the compensations named.

No output carries key material β€” the resource has none.


🧠 Architecture Notes

The most important property of this resource is what it does not do. azurerm_stack_hci_cluster creates an ARM record. After a successful apply there are no nodes joined, no storage pool, no networking, and no custom location. Microsoft documents the real deployment as taking 2.5 to 3 hours, with a single step accounting for 40 to 45 minutes of it β€” and none of that happens here. The provider's own 30-minute create default is the tell: the deployment setting's defaults to six hours for exactly this reason. So a green apply on this module proves the registration succeeded and essentially nothing about the instance.

The custom location is created by the deployment, and that breaks a dependency chain people expect to exist. Five modules in this family require a custom_location_id, and none of them can obtain it from this resource β€” this resource has no such argument and no such output. The custom location comes into being as a side effect of the deployment, whose scale_unit.optional_service.custom_location names it. Because Terraform never creates that resource, its ID is not an attribute of anything in the graph, so the ordering cannot be expressed as a dependency. In practice the five children are applied in a later run, or the ID is looked up with a data block once the deployment has finished. A composition that attempts everything in one apply fails on the children rather than on the deployment, which is a confusing place for it to fail.

Two of this resource's fields look like provider authentication and are not. client_id and tenant_id identify the Entra application the cluster uses to talk to Azure β€” the provider's wording is "used by" the cluster. This suite's rule that authentication is never a module variable is about the credential Terraform uses, and neither of these is it. The rule still bites in one place, though, and this module honours it absolutely: the application's secret is never accepted here. An application ID is an identifier and is deliberately not marked sensitive; the secret belongs in the key vault the deployment points at. The one genuine leak of provider configuration into resource data is tenant_id's default β€” omit it and the provider's tenant governs, which is usually right and is force-new, so a pipeline authenticated elsewhere replaces the instance rather than updating it.

Two service principals matter and both are GUIDs. identity_principal_id is the instance's own managed identity β€” what a diagnostic setting or policy assignment targets. resource_provider_object_id is the Azure Stack HCI resource provider's service principal, and it is what needs read access to the key vault holding the deployment secrets. Granting the wrong one produces a permissions failure that names neither, which is why the two outputs are named as differently as they are.

Two cross-field rules Microsoft publishes are enforced elsewhere, and the module says so rather than staying quiet. The instance name must differ from every node name β€” but the node names are inputs to azurerm_stack_hci_deployment_setting, so that module enforces the rule and this one reports it. This is the same asymmetry the network interface module documents about address containment: a rule gets enforced wherever both operands happen to be inputs, and reported wherever one of them is not. Saying it explicitly is what stops the weaker module reading as careless rather than differently placed.

Arc projection, and a destroy that is not a decommission. The record represents physical machines reached through Azure Arc, so ARM state can outlive the hardware β€” and on this resource that matters more than on its children, because the record carries the billing relationship. Deleting it removes the Azure registration; it does not un-deploy software on the nodes, un-join them from the cluster, or clean up the Arc machine resources. Running those steps in the wrong order leaves an instance that is neither registered nor cleanly reusable.


🧱 Design Principles

Secure-by-default has almost nothing to act on here, and pretending otherwise would be the misleading choice. Three arguments are required, so there is no meaningful empty call β€” and this resource has no exposure toggle, no network surface, no encryption argument and no authentication setting of its own. Everything a security review wants to see about an Azure Local instance β€” BitLocker on boot and data volumes, SMB signing and cluster encryption, Credential Guard, HVCI, WDAC, side-channel mitigation, drift control β€” lives on azurerm_stack_hci_deployment_setting.

The rule is therefore compensated:

Compensation How it appears here
Enforce the closed value sets One exists: identity.type, the single value the pinned provider documents. The casing is named in the message, along with the fact that there is no user-assigned option.
Name the restrictive fact inside the error message That a client secret does not belong in this module at any price; that a tenant's domain name is not a tenant ID; that a custom location is not an input because the deployment creates it; that an Automanage profile ID is not this instance's own ID.
Emit every unverifiable decision That creating this resource deploys nothing; that the custom location its children need comes from the deployment; that the name-versus-node-name rule is enforced elsewhere; that a destroy is not a decommission; and which of the two service principals reads the key vault.

The defaults that do exist:

Concern This module's default Opt-out
identity Unset, so no identity β€” and has_identity reports false and says what that costs Supply { type = "SystemAssigned" }
entra_application {} β€” no application, and the tenant left to the provider Supply a client ID, and a tenant ID if the app is in another tenant
automanage_configuration_id Unset β€” Terraform is the only thing configuring the instance Assign a profile, and read is_automanaged
tags {} Supply tags

πŸ”’ The one posture decision this module declines to make is the identity. Assigning a system-assigned identity is almost always right, but creating a principal the caller did not ask for is not this suite's habit β€” so the default stands and has_identity reports the false with its cost attached.


πŸš€ Runbook

terraform init -backend=false
terraform validate
terraform fmt -check

Pin the module with ?ref=v1.0.0 β€” never a branch. This library is plan-only; a human applies from CI.

πŸ”΄ Do not read a green apply here as a working instance. Confirm the deployment setting applied, and confirm the custom location exists, before applying anything that depends on either.

⚠️ Review any plan touching this module for # forces replacement. Everything except tags is force-new, and replacing this record strands a deployed instance.


πŸ§ͺ Testing

The offline gate proves the configuration and the type system, and nothing beyond it:

Covered offline Only exercised by plan/apply
HCL parses; the identity and entra_application objects reject unknown attributes. Whether Azure Local is available in the chosen region.
All 13 validations, provable with terraform console and a .tfvars file. Whether the machines are registered with Azure Arc.
Every derived local, in both directions β€” with and without an identity, an application, a tenant, a profile. Whether the Entra application exists or has the right permissions.
The Automanage regex anchored at both ends, plus both by-name rejections. Whether the deployment that follows succeeds. Nothing here tests that.
The arc_setting_id string composition. Whether the Arc setting is actually named default.

πŸ’‘ terraform console fires root-module variable validations β€” terraform validate on a calling configuration does not. Feed it a deliberately bad .tfvars.

πŸ’‘ Prove validations by line number, not by message text. Terraform hard-wraps error messages, so grepping for a phrase from an error_message silently matches nothing. Collect the variables.tf:<line> references instead and compare their union against the count of validation blocks.

⚠️ A short error list is not proof a check is missing. Terraform skips a validation whose referenced variable has itself failed. Every validation here references only its own variable, which is why client_id and tenant_id share one object.


πŸ’¬ Example Output

Outputs:

id                            = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-azlocal/providers/Microsoft.AzureStackHCI/clusters/azlocal-cluster-01"
name                          = "azlocal-cluster-01"
resource_group_name           = "rg-azlocal"
location                      = "eastus2"
arc_setting_id                = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-azlocal/providers/Microsoft.AzureStackHCI/clusters/azlocal-cluster-01/arcSettings/default"

cloud_id                      = "7f3a1c88-2b45-4d6e-9f01-5c8a2e7b4d90"
service_endpoint              = "https://eastus2.dp.stackhci.azure.com"
resource_provider_object_id   = "1412d0e3-3f4b-4a1c-9d2e-6b7a8c9d0e1f"
tenant_id                     = "cccccccc-4444-5555-6666-dddddddddddd"
client_id                     = null

identity                      = {
  principal_id = "9a8b7c6d-5e4f-3a2b-1c0d-9e8f7a6b5c4d"
  tenant_id    = "cccccccc-4444-5555-6666-dddddddddddd"
  type         = "SystemAssigned"
}
identity_principal_id         = "9a8b7c6d-5e4f-3a2b-1c0d-9e8f7a6b5c4d"

automanage_configuration_id   = null
automanage_configuration_name = null
tags                          = { "environment" = "production", "platform" = "azure-local" }

has_identity                          = true
uses_entra_application                = false
entra_tenant_was_left_to_the_provider = true
is_automanaged                        = false

creating_this_resource_does_not_deploy_anything = true
the_custom_location_its_children_consume_is_created_by_the_deployment_not_by_this_resource = true
the_node_names_live_on_the_deployment_setting_so_the_collision_check_lives_there = true
this_is_an_arc_projected_resource_so_azure_state_can_diverge_from_the_hardware = true
the_product_is_now_called_azure_local_but_arm_and_terraform_still_say_stack_hci = "Microsoft.AzureStackHCI/clusters"
secure_by_default_has_almost_nothing_to_act_on_here = true

πŸ”‘ Note resource_provider_object_id and identity_principal_id side by side. Two different principals, both GUIDs β€” the key vault grant belongs to the first.


πŸ” Troubleshooting

Symptom Cause Fix
The apply succeeded but the instance shows no nodes and has no custom location Expected. This resource is a registration; it deploys nothing. Apply terraform-azurerm-stack-hci-deployment-setting against this instance's id.
A sibling module fails because the custom location does not exist The custom location is created by the deployment, and it is not a Terraform resource β€” so nothing in the graph can be depended on. Apply the deployment first, then the children in a later run, or look the ID up with a data block.
identity.type must be exactly "SystemAssigned"… A casing slip, or an attempt at UserAssigned / SystemAssigned, UserAssigned. Use SystemAssigned. There is no user-assigned option on this resource.
entra_application.client_id must be a bare GUID… Braces around the GUID, surrounding text β€” or a client secret in the client ID field. Pass the application (client) ID. The secret does not belong in this module at all.
entra_application.tenant_id must be a bare GUID. A tenant domain name such as contoso.onmicrosoft.com. Use the tenant's GUID.
automanage_configuration_id is an Azure Local (Stack HCI) CLUSTER Resource ID… This instance's own ID pasted into the Automanage field. Use the Microsoft.Automanage/configurationProfiles ID.
automanage_configuration_id is a CUSTOM LOCATION Resource ID. A custom location passed here β€” understandable, since five siblings take one. This resource never does. Use the Automanage profile ID.
location must be lower case / contains a space A region display name such as "East US 2". Use eastus2.
Apply fails saying the region is not supported Azure Local is available in a subset of regions, which this module deliberately does not enforce. Check Microsoft's current region list.
The deployment later fails reading its secrets The resource provider's service principal has no access to the key vault. Grant resource_provider_object_id β€” not identity_principal_id β€” on the vault.
A role assignment against the instance silently affects nothing The wrong principal was used. Both are GUIDs and nothing reports the mismatch. identity_principal_id for the instance itself; resource_provider_object_id for the RP.
A plan shows # forces replacement after changing tenant_id Expected. tenant_id is force-new β€” and it may have changed because the pipeline authenticated against a different tenant. Check entra_tenant_was_left_to_the_provider; set the tenant explicitly if it must be pinned.
Terraform reports the instance exists but the hardware is gone ARM state outlived the hardware. Reconcile against the machines; do not trust state alone for Arc-projected resources.
Destroying the module left the nodes deployed Expected. A destroy is not a decommission. Un-deploy on the nodes and clean up the Arc machine resources separately, in that order.
Settings on the instance change without a Terraform plan An Automanage configuration profile is assigned and is re-applying its own settings. Check is_automanaged.

πŸ”— Related Docs


πŸ’™ "Infrastructure as Code should be standardized, consistent, and secure."