Skip to content

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

☁️ Azure Cassandra Datacenter Terraform Module

The nodes of an Azure Managed Instance for Apache Cassandra cluster — where the provider and the ARM contract disagree about two defaults, and two arguments look editable but are not. Targets hashicorp/azurerm ~> 4.0.

Terraform azurerm Module Type Resources Posture


🧩 Overview

  • 🗄️ Manages azurerm_cosmosdb_cassandra_datacenter — a datacenter inside an Azure Managed Instance for Apache Cassandra cluster: the virtual machine scale set that actually runs Cassandra, in its own region, on its own delegated subnet.
  • 🔴 This is not the Cosmos DB Cassandra API. Managed Instance runs pure open-source Cassandra on virtual machines in your own virtual network. Microsoft states there is no architectural dependency between it and the Cosmos DB backend, even though both live under Microsoft.DocumentDB and in this module family.
  • 🔴 Omitting disk_count sends 0 to Azure. The provider declares it optional with no default and then sends its value unconditionally — so an absent argument transmits a number the provider's own validator rejects when typed. ARM documents the default as 4; this module supplies 4.
  • 🔴 disk_count and availability_zones_enabled cannot be changed. The update payload omits both properties, and neither argument is force-new — so a change plans cleanly, applies without error, and does nothing.
  • 🔴 The provider and the ARM contract document different default VM SKUs — Standard_E16s_v5 here, Standard_DS14_v2 in ARM. And Microsoft's own pages publish mutually inconsistent lists of allowed sizes, with the provider's default appearing in only one of them.
  • ⚠️ Two prerequisites are governance acts, not resources — a tenant-wide service principal, and a role assignment on the virtual network. Neither is anything this module can create, and their absence fails the apply.
  • ⚠️ node_count is a desired state Azure converges to asynchronously, and the provider exposes no way to see actual node status.
  • ⚠️ A datacenter has its own location, separate from the cluster's — which is how a multi-region managed Cassandra cluster is expressed.

💡 Why it matters: almost everything consequential about this resource is invisible in a plan. Two arguments accept changes and discard them. One argument sends a value nobody chose when you leave it out. The SKU you get by omission depends on which API surface created the resource. And two of the deployment's hard requirements are performed by a tenant administrator, outside Terraform entirely. So the module supplies the default the provider drops, reports the rest, and emits the operands for the cross-resource checks it cannot make itself.


❤️ Support this project

If this module saved you time:


🗺️ Where this fits in the family

This is the shared cosmosdb family DAG. The fourteen blue nodes are the authored modules; the single dark node is not a module at all — it is the constraint that decides which children are even legal on a given account.

flowchart TB
  RG["terraform-azurerm-resource-group"]
  KV["terraform-azurerm-key-vault: the customer-managed key, if the account uses one"]
  VNET["terraform-azurerm-virtual-network: a delegated subnet, for the MANAGED CASSANDRA service only"]

  ACCT["terraform-azurerm-cosmosdb-account: the account, PLUS its sql_database and sql_container children"]
  KEYSPACE["terraform-azurerm-cosmosdb-cassandra-keyspace"]
  CASSTABLE["terraform-azurerm-cosmosdb-cassandra-table"]
  MONGODB["terraform-azurerm-cosmosdb-mongo-database"]
  MONGOCOLL["terraform-azurerm-cosmosdb-mongo-collection"]
  GREMDB["terraform-azurerm-cosmosdb-gremlin-database"]
  GREMGRAPH["terraform-azurerm-cosmosdb-gremlin-graph"]
  PGCLUSTER["terraform-azurerm-cosmosdb-postgresql-cluster"]
  MICLUSTER["terraform-azurerm-cosmosdb-cassandra-cluster: MANAGED INSTANCE for Apache Cassandra, a DIFFERENT service from the Cassandra API"]
  MICHILD["terraform-azurerm-cosmosdb-cassandra-datacenter: the NODES of a managed cluster, with its own region and its own delegated subnet"]
  SQLTRIGGER["terraform-azurerm-cosmosdb-sql-trigger: JavaScript that the client must ASK FOR on each request, so registering it does not put it into effect"]
  SQLFUNC["terraform-azurerm-cosmosdb-sql-function: a SIDE-EFFECT-FREE query extension, invoked only from query text as udf.name, with no context object at all"]
  SQLSPROC["terraform-azurerm-cosmosdb-sql-stored-procedure: TRANSACTIONAL JavaScript on the primary replica, scoped to ONE logical partition, and the only container child taking FOUR NAMES instead of a container id"]
  SQLROLEDEF["terraform-azurerm-cosmosdb-sql-role-definition: ENTRA ID data-plane RBAC. Defines what may be granted, and grants nothing by itself"]
  SQLROLEASSIGN["terraform-azurerm-cosmosdb-sql-role-assignment: the half that HANDS THE ROLE TO AN ENTRA PRINCIPAL. Creating one is the moment data access begins"]
  MONGOROLEDEF["terraform-azurerm-cosmosdb-mongo-role-definition: MONGO-NATIVE RBAC, a wholly different mechanism from the Entra ID RBAC above. Roles live INSIDE a database"]
  MONGOUSERDEF["terraform-azurerm-cosmosdb-mongo-user-definition: the MONGO USER that holds a role, authenticating with SCRAM rather than Entra ID. The ONLY module in this family carrying a PASSWORD, and Azure never returns it"]
  SQLGATEWAY["terraform-azurerm-cosmosdb-sql-dedicated-gateway: a SINGLETON SERVICE on the account, not a child record. Provisions the integrated cache, is BILLED HOURLY whether used, and does nothing until a client switches to gateway mode"]
  TABLE["terraform-azurerm-cosmosdb-table: the Table API, and the ONLY child with NO intermediate layer -- a table sits directly on the account, so two names are the whole path"]

  API["ONE ACCOUNT SERVES ONE API, chosen by capabilities: EnableCassandra, EnableMongo, EnableGremlin, EnableTable, or none for the SQL API. This is why each API's children are separate modules and not more children of the account composite"]

  PGKIDS["not yet authored, and DELIBERATELY SO: azurerm_cosmosdb_postgresql_role, azurerm_cosmosdb_postgresql_firewall_rule, azurerm_cosmosdb_postgresql_node_configuration and azurerm_cosmosdb_postgresql_coordinator_configuration. Microsoft documents Cosmos DB for PostgreSQL as on a retirement path, so these await a maintainer decision rather than effort"]

  RG -->|"resource_group_name and location"| ACCT
  RG -->|"resource_group_name"| MICLUSTER
  RG -->|"resource_group_name"| PGCLUSTER
  KV -->|"key id, for account encryption"| ACCT
  VNET -->|"a DELEGATED subnet id, required by the managed service and by nothing else here"| MICLUSTER
  VNET -->|"a SECOND delegated subnet, which must sit in the datacenter's own region and must be able to ROUTE to the cluster's management subnet"| MICHILD
  KV -->|"two VERSIONED key uris, for backup storage and for the nodes' managed disks"| MICHILD

  ACCT -->|"capabilities decide which of the children below are even legal"| API

  ACCT -->|"NAME plus resource group, so the subscription comes from the provider and Terraform builds no edge from a literal"| KEYSPACE
  ACCT -->|"name plus resource group, the same convention"| MONGODB
  ACCT -->|"name plus resource group, the same convention again"| GREMDB
  ACCT -->|"THREE names and NO intermediate layer: the Table API has no database, so account plus table name is the entire hierarchy. Every other API here interposes a database, keyspace or graph"| TABLE

  KEYSPACE -->|"id: the child takes an ID where the parent takes names, and the provider rebuilds that ID using the PROVIDER's subscription, so a cross-subscription id is silently relocated"| CASSTABLE
  MONGODB -->|"THREE names at once: this parent emits name, account_name AND resource_group_name, so the four-name child wires entirely from one module"| MONGOCOLL
  GREMDB -->|"three names, and this parent emits all three for exactly that reason. The child cannot express a HIERARCHICAL partition key at all"| GREMGRAPH
  MICLUSTER -->|"cluster id: the provider takes the subscription, resource group AND cluster name from this id and nothing from its own configuration, so a cross-subscription id is honored rather than relocated. The exact opposite of the keyspace edge above"| MICHILD
  MONGODB -->|"the DATABASE id: the provider reads the subscription, resource group, account AND database name out of it and takes nothing from its own configuration, so a cross-subscription id is HONORED rather than relocated. The same convention as the cluster to datacenter edge, and the opposite of the keyspace to table edge"| MONGOROLEDEF
  MONGODB -->|"the same DATABASE id, on the same convention. This database is also the authSource a client must authenticate against"| MONGOUSERDEF
  MONGOROLEDEF -->|"the ROLE NAME, from its role_name output. A role confers nothing until a user holds it, so this edge is where Mongo-native access actually begins"| MONGOUSERDEF
  ACCT -->|"container id: the provider parses the resource group, account, database and container out of it and then takes the SUBSCRIPTION FROM ITS OWN CONFIGURATION, so a cross-subscription container id is silently relocated. The same trap as the keyspace-to-table edge, and the opposite of the cluster-to-datacenter edge"| SQLTRIGGER
  ACCT -->|"container id: the SAME subscription substitution as the trigger edge. Its child is a QUERY EXTENSION rather than an operation hook, so it shares the id trap and none of the invocation story"| SQLFUNC
  ACCT -->|"FOUR NAMES: resource group, account, database and container. So the caller never supplies a subscription and the provider has nothing to override -- the relocation hazard its two siblings carry is IMPOSSIBLE here rather than silent"| SQLSPROC
  ACCT -->|"THREE names, and an allow-list of data actions. Nothing here is a container child: a role definition belongs to the ACCOUNT and names the databases and containers it may be assigned over, using the DATA-PLANE scope grammar with dbs and colls rather than sqlDatabases and containers"| SQLROLEDEF
  ACCT -->|"TWO names only, resource group and account, and BOTH are force-new. The assignment's own name is an optional GUID rather than a label, so it is not a third name from the parent"| SQLROLEASSIGN
  SQLROLEDEF -->|"the definition's RESOURCE ID, not its bare GUID: Terraform validates this argument as a role-definition id, so the id output is the one to wire and the role_definition_id output is for things outside Terraform"| SQLROLEASSIGN
  ACCT -->|"the ACCOUNT id, and the only child here that takes one. The gateway appends services/SqlDedicatedGateway to it, so there is no name to supply and exactly ONE gateway per account"| SQLGATEWAY
  SQLGATEWAY -.->|"routes through, and caches for, whichever containers a client reads via the dedicated endpoint"| ACCT
  PGCLUSTER --> PGKIDS

  classDef me fill:#0078D4,stroke:#004578,color:#ffffff
  classDef keystone fill:#004578,stroke:#002438,color:#ffffff
  classDef sibling fill:#F3F6F9,stroke:#8A9BA8,color:#1B1F23
  class KEYSPACE,CASSTABLE,MONGOCOLL,GREMDB,GREMGRAPH,ACCT,MONGODB,PGCLUSTER,MICLUSTER,MICHILD,SQLTRIGGER,SQLFUNC,SQLSPROC,SQLROLEDEF,SQLROLEASSIGN,MONGOROLEDEF,MONGOUSERDEF,SQLGATEWAY,TABLE me
  class API keystone
  class RG,KV,VNET,PGKIDS sibling
Loading

Read the two edges into this module together with the one above them. The Cassandra API keyspace takes an account ID and has the provider quietly substitute its own subscription; the managed-instance cluster ID feeds this datacenter every segment it needs and substitutes nothing. Same family, opposite behavior.


🧬 What this module builds

flowchart TB
  IN_CLUSTER["cassandra_cluster_id: a cassandraClusters Resource ID. FORCE-NEW, and it supplies the subscription, resource group and cluster name for this datacenter's own id"]
  IN_NAME["name and location: FORCE-NEW. The location is the DATACENTER's region, independent of the cluster's, which is how a multi-region cluster is expressed"]
  IN_SUBNET["delegated_management_subnet_id: FORCE-NEW. Must be in the datacenter's region and must ROUTE to the cluster's management subnet, neither of which is checkable at plan time"]
  IN_CAPACITY["node_count defaults to 3, disk_count defaults to 4 BECAUSE THE PROVIDER SENDS ZERO OTHERWISE, disk_sku defaults to P30, sku_name defaults to Standard_E16s_v5 where ARM documents Standard_DS14_v2"]
  IN_ZONES["availability_zones_enabled defaults to true, and the deployment FAILS in a region without zone support"]
  IN_CMK["two optional VERSIONED Key Vault key uris: backup storage and managed disks"]
  IN_YAML["base64_encoded_yaml_fragment: only an unpublished subset of cassandra.yaml keys is accepted"]

  PROV["the provider contributes NO id segment here, unlike elsewhere in this family: every segment comes from the cluster id"]

  THIS["azurerm_cosmosdb_cassandra_datacenter.this"]

  UPDATE["the update payload OMITS disk capacity and availability zone, and neither argument is force-new, so a change to either plans cleanly, applies without error, and does nothing"]

  OUT_ID["id and name"]
  OUT_PATHS["cluster_path_within_the_subscription, datacenter_path_within_the_subscription, delegated_subnet_vnet_path: the operands a check block compares, since the module cannot compare them itself"]
  OUT_CAPACITY["total_data_disks, disk_count_is_below_microsofts_recommended_minimum, disk_sku_is_not_the_p30_the_disk_count_guidance_assumes"]
  OUT_SKU["sku_is_a_preview_write_through_caching_tier, sku_appears_in_no_microsoft_published_list, sku_is_the_arm_documented_default"]
  OUT_SEEDS["seed_node_ip_addresses and seed_node_count, known only after apply: what a hybrid ring needs in its own seed provider configuration"]
  OUT_CONST["sixteen constants: the two silently dropped updates, the zero-disk send, the two governance acts a tenant administrator must perform, and the node count Azure converges to asynchronously"]

  IN_CLUSTER --> THIS
  IN_NAME --> THIS
  IN_SUBNET --> THIS
  IN_CAPACITY --> THIS
  IN_ZONES --> THIS
  IN_CMK --> THIS
  IN_YAML --> THIS
  PROV -.->|"nothing"| THIS

  THIS --> UPDATE
  THIS --> OUT_ID
  THIS --> OUT_PATHS
  THIS --> OUT_CAPACITY
  THIS --> OUT_SKU
  THIS --> OUT_SEEDS
  THIS --> OUT_CONST

  classDef me fill:#0078D4,stroke:#004578,color:#ffffff
  classDef keystone fill:#004578,stroke:#002438,color:#ffffff
  classDef sibling fill:#F3F6F9,stroke:#8A9BA8,color:#1B1F23
  class THIS keystone
  class OUT_ID,OUT_PATHS,OUT_CAPACITY,OUT_SKU,OUT_SEEDS,OUT_CONST me
  class IN_CLUSTER,IN_NAME,IN_SUBNET,IN_CAPACITY,IN_ZONES,IN_CMK,IN_YAML,PROV,UPDATE sibling
Loading
Resource Cardinality Note
azurerm_cosmosdb_cassandra_datacenter single (this) The keystone. One nested block, timeouts; every other argument is a scalar.

A standalone module: one keystone resource, no for_each children.


✅ Provider / Versions

Item Value
Terraform >= 1.12.0
Provider hashicorp/azurerm ~> 4.0 (validated against 4.81.0)
Provider block None here. The caller configures provider "azurerm" { features {} }, auth and subscription.
ARM API type Microsoft.DocumentDB/cassandraClusters/dataCenters (API version 2023-04-15)
tags Not supported — this resource has no tags. Tag the cluster or the resource group instead.
Timeouts Provider defaults are 60 minutes for create, update and delete; 5 minutes for read.

Schema notes that bite

  • 🔴 disk_count has no default and is sent unconditionally. The schema is Optional with IntBetween(1, 10) and no Default; the create path sends DiskCapacity: int64(d.Get("disk_count").(int)) regardless. An omitted argument therefore transmits 0 — below the validator's own floor. ARM documents the property's default as 4, and Microsoft's portal guidance calls four "strongly recommended", so this module defaults it to 4.
  • 🔴 The update payload omits disk_count and availability_zones_enabled. It carries the delegated subnet, node count, SKU, location and disk SKU — and nothing else. Neither omitted argument is force-new, so Terraform plans a change, the apply succeeds, Azure ignores it, and the following read pulls the old value back into state.
  • 🔴 Two defaults for one property. The provider defaults sku_name to Standard_E16s_v5 (a v4.0 change it documents in a note); the ARM contract documents its own default as Standard_DS14_v2. The same omitted argument gives a different virtual machine size, at a different price, depending on the surface.
  • 🔴 Microsoft's published SKU lists disagree with each other. One page lists only v5 sizes; another lists v4, DS and D sizes; a third adds the L series. The provider's default appears in only one. So this module enforces no closed set — it rejects a case-only near miss and lets an unlisted SKU through.
  • ⚠️ The Standard_L* sizes are a public preview with no SLA. Microsoft introduced them for write-through caching with locally attached disks and does not recommend them for production. Nothing in a plan says so.
  • ⚠️ VM family transitions are refused by the service. Microsoft: it "doesn't support transitioning across VM families", giving Standard_DS13_v2 → Standard_DS14_v2 as an example that needs a support ticket. sku_name is not force-new and is sent on update, so Terraform will happily plan a change the service declines.
  • ⚠️ Force-new: name, cassandra_cluster_id, location, delegated_management_subnet_id. Four, and location is force-new because it comes from commonschema.Location().
  • ⚠️ The ID constructor takes every segment from cassandra_cluster_id — subscription included — and nothing from provider configuration. That is the opposite of the Cassandra API table in this family, whose constructor substitutes the provider's subscription for the supplied one.
  • ⚠️ The update waits out a known API defect. After the long-running operation completes, the provider polls provisioning state with a one-minute initial delay, citing Azure/azure-rest-api-specs#19078: "The API cannot update the property after WaitForCompletionRef is returned. It has to wait a while after that."
  • ⚠️ Both customer-key arguments require a VERSIONED Key Vault key ID — keyvault.ValidateNestedItemID(VersionTypeVersioned, NestedItemTypeKey). A versionless URI, a secret URI and a certificate URI are all rejected.
  • ⚠️ seed_node_ip_addresses is computed only. The read's flattener appends each IPAddress pointer rather than the dereferenced string, which is unlike the surrounding code; and the read sets disk_count via int(*props.DiskCapacity) with no nil guard, where every neighboring field is set through the pointer.
  • ⚠️ There is no SchemaVersion and no state upgrader.

🔑 Required Azure RBAC Roles / Permissions

Principal Scope Requirement
The principal running Terraform The managed Cassandra cluster, or its resource group Microsoft.DocumentDB/cassandraClusters/dataCenters/* — carried by Contributor at the cluster's resource group
The Azure Cosmos DB application (a232010e-820c-4083-83bb-3ace5fc29d0b) The virtual network holding the delegated subnet Network Contributor (role definition 4d97b98b-1d4f-4787-a291-c67834d212e7), assigned with the command Microsoft publishes. Not the caller's identity — the service's
The cluster's system-assigned identity The Key Vault key, if either customer-managed key is used get, wrap and unwrap on the key. The identity belongs to the cluster, not to this datacenter
Alternatively, for the first row The cluster or its resource group Owner

Control plane only. Creating a datacenter grants no access to the Cassandra data in it — that runs over the cluster's own certificates and credentials, and this module touches neither. No key is read, no secret is held, and no output is sensitive, so plan access is not credential access here.

The second and third rows are the ones that get missed, because neither is about the person running Terraform. The scaffold this module replaced named plain Contributor and "the service-specific SQL/DB Contributor role" — the wrong service, and silent on both of the grants that actually block a deployment.


Azure Prerequisites

  1. Microsoft.DocumentDB registered in the subscription.
  2. 🔴 The Azure Cosmos DB service principal present in the tenant. The provider documents that the application ID a232010e-820c-4083-83bb-3ace5fc29d0b must exist as a service principal, added by a tenant administrator:
    New-AzADServicePrincipal -ApplicationId a232010e-820c-4083-83bb-3ace5fc29d0b
    This is tenant-wide and is not a resource any module here creates.
  3. 🔴 A Network Contributor role assignment for that application on the virtual network, before the deployment is attempted. terraform-azurerm-role-assignments can make it; this module cannot, and nothing here checks it.
  4. An existing managed Cassandra cluster, from terraform-azurerm-cosmosdb-cassandra-cluster.
  5. An existing delegated subnet in a region that matches this datacenter's location, able to route to the cluster's own management subnet. Both requirements are Microsoft's and neither is checkable at plan time.
  6. Outbound internet access from that subnet. Microsoft: deployment "requires internet access" and fails where it is restricted — Azure Storage, Azure Key Vault, Azure Virtual Machine Scale Sets, Azure Monitor, Microsoft Entra ID and Microsoft Defender for Cloud must all be reachable.
  7. A region that supports availability zones, if availability_zones_enabled is left at its default of true, with the chosen sku_name available in every zone.
  8. A Key Vault key with a version, plus the cluster identity's permissions on it, if either customer-managed key is used.

There is no resource_group_name argument on this resource — it inherits the cluster's, parsed out of cassandra_cluster_id.


📁 Module Structure

terraform-azurerm-cosmosdb-cassandra-datacenter/
├── providers.tf   # required_version and the pinned azurerm; no provider block
├── variables.tf   # 13 variables, 27 validations
├── main.tf        # the keystone `this`, one dynamic timeouts block, and 35 derived locals
├── outputs.tf     # 50 outputs: 11 passthrough, 23 derived, 16 constant
├── README.md      # this file
├── SCOPE.md       # the cross-module contract
├── LICENSE        # MIT
└── .gitignore

⚙️ Quick Start

module "dc_eastus" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                 = "dc-eastus"
  cassandra_cluster_id = module.cassandra_cluster.id
  location             = "eastus"

  # Must be in this datacenter's region and route to the cluster's management subnet.
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]
}

💡 That empty call gets three nodes, four P30 disks each, Standard_E16s_v5, and availability zones on. The four disks come from this module, not the provider — see example 2.

⚠️ It will still fail without the tenant service principal and the virtual network role assignment from Azure Prerequisites. Neither is expressible as an argument.


🔌 Cross-Module Contract

Consumes

Input Type Source
cassandra_cluster_id string terraform-azurerm-cosmosdb-cassandra-cluster output id
delegated_management_subnet_id string terraform-azurerm-virtual-network, a delegated subnet id
location string the caller — this datacenter's own region
the subscription — the cluster ID. Not the caller's provider, unusually for this family

Emits

Output Description Consumed by
id The datacenter's Resource ID (first) a management lock, an import, a report
seed_node_ip_addresses The seed node IPs Azure assigned a hybrid ring's own cassandra.yaml
delegated_subnet_vnet_path The subnet's virtual network check blocks comparing against the cluster's network
cluster_path_within_the_subscription, datacenter_path_within_the_subscription Exact but for the subscription check blocks
total_data_disks node_count times disk_count cost review
sku_is_a_preview_write_through_caching_tier An L-series size, no SLA check blocks
disk_count_changes_are_silently_dropped_on_update One of the two no-op arguments documentation
sixteen constants Facts that produce no error when they bite check blocks, runbooks

📚 Example Library

1 · The smallest real call
module "dc_eastus" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]
}

output "what_the_defaults_gave_us" {
  value = {
    nodes      = module.dc_eastus.node_count
    disks      = module.dc_eastus.disk_count
    all_disks  = module.dc_eastus.total_data_disks
    vm_size    = module.dc_eastus.sku_name
    zones      = module.dc_eastus.availability_zones_enabled
  }
}

💡 Three nodes, four disks each — twelve P30 disks — on Standard_E16s_v5, zones on. total_data_disks exists because neither node_count nor disk_count shows that number on its own, and disks are attached per node.

⚠️ Four required arguments and no resource_group_name: this resource has none. The resource group comes out of the cluster ID, which is why cluster_resource_group_name is a derived output rather than a passthrough.

2 · The disk count the provider sends when you omit it
# This module supplies 4. The provider supplies nothing and sends 0.
module "dc_default_disks" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]

  # disk_count omitted -> this module applies 4, the ARM-documented default.
}

# REJECTED by the module and by the provider: 0 is below the validator's floor.
module "dc_zero_disks" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-broken"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]

  disk_count = 0
}

output "the_provider_would_have_sent_zero" {
  value = module.dc_default_disks.omitting_disk_count_would_send_zero_disks_to_azure
}

🔴 The provider declares disk_count optional with no default, and then sends it unconditionally: DiskCapacity: pointer.To(int64(d.Get("disk_count").(int))). An absent integer reads as its zero value, so omitting the argument transmits 0 — a number the provider's own IntBetween(1, 10) rejects the moment a caller types it, as the second module above shows.

🔒 So this module defaults it to 4, and that is a correction rather than a preference. ARM documents the property's default as 4; Microsoft's portal calls four disks per node "strongly recommended". Supplying 4 sends the documented default instead of a zero nobody chose. This is the one place the module overrides the provider's own idea of an omission.

ℹ️ Microsoft's sizing guidance: plan for a maximum sustained utilization of 50 percent, because tombstones and system services consume space and backups use local disk before being persisted to Blob storage.

3 · The two arguments that accept a change and discard it
module "dc_eastus" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]

  node_count = 6 # editable: the provider sends it and Microsoft documents scaling this way
  disk_count = 8 # NOT editable after creation, and nothing marks it force-new
}

check "the_no_op_arguments_are_understood" {
  assert {
    condition = alltrue([
      module.dc_eastus.disk_count_changes_are_silently_dropped_on_update,
      module.dc_eastus.availability_zone_changes_are_silently_dropped_on_update,
    ])
    error_message = "Both of these are always true, so this assertion never fails -- it exists to put the fact in the plan output. The provider's update payload omits the disk capacity and availability zone properties, and neither argument is force-new, so changing either one produces a plan that shows the change, an apply that reports success, and no change in Azure. Delete this assertion once the team knows."
  }
}

🔴 The update payload carries the delegated subnet, node count, SKU, location and disk SKU — and nothing else. disk_count and availability_zones_enabled are simply absent from it, and neither argument is marked force-new. So Terraform believes both are editable in place, Azure never hears about the change, and the read that follows pulls the unchanged value back into state.

⚠️ Treat both as create-time decisions, because that is what they are — the schema just does not say so. Changing the disk count for real means replacing the datacenter, which means moving its data first.

💡 node_count, by contrast, genuinely is editable: Microsoft documents scaling a datacenter by changing it, and the provider sends it. See example 4 for what "editable" means here.

4 · Scaling nodes, and why a successful apply is not a running cluster
module "dc_eastus" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]

  node_count = 9

  # The provider's own update default is 60 minutes, and the update path additionally waits out a
  # known API defect before it can confirm the change.
  timeouts = {
    update = "2h"
  }
}

output "what_apply_success_does_not_mean" {
  value = {
    desired_not_actual = module.dc_eastus.node_count_is_a_desired_state_azure_converges_to_asynchronously
    no_status_here     = module.dc_eastus.node_status_is_not_visible_from_terraform
    requested          = module.dc_eastus.node_count
  }
}

🔴 Microsoft: node_count is "the desired number. After it is set, it may take some time for the data center to be scaled to match." So a successful apply means Azure accepted the target, not that nine nodes are serving traffic. A pipeline that provisions and then immediately connects can find fewer nodes than it asked for.

🔴 And there is no Terraform-native way to check. Microsoft's route is the cluster's fetchNodeStatus operation, which this provider exposes as neither a resource nor a data source. Use az managed-cassandra cluster status or Azure Monitor.

⚠️ The provider's update timeout defaults to 60 minutes, and the update path polls provisioning state with a one-minute initial delay because of Azure/azure-rest-api-specs#19078 — "The API cannot update the property after WaitForCompletionRef is returned. It has to wait a while after that." Scaling is slow; raise the timeout rather than fighting it.

5 · The VM SKU, where three sources disagree
# The provider's default, applied when sku_name is omitted.
module "dc_provider_default" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]
  # sku_name omitted -> Standard_E16s_v5
}

# What ARM would have defaulted to for the same property.
module "dc_arm_default" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus2"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus2"
  delegated_management_subnet_id = var.eastus2_cassandra_subnet_id

  sku_name = "Standard_DS14_v2"
}

# REJECTED: matches a published SKU except in case.
#   sku_name = "standard_e16s_v5"
# ALLOWED, and reported: in none of Microsoft's lists.
#   sku_name = "Standard_E48s_v5"

output "sku_facts" {
  value = {
    provider_default = module.dc_provider_default.sku_is_the_providers_default
    arm_default      = module.dc_arm_default.sku_is_the_arm_documented_default
    lists_disagree   = module.dc_provider_default.microsofts_published_vm_sku_lists_disagree_with_each_other
  }
}

🔴 The provider defaults this to Standard_E16s_v5; the ARM contract documents its default as Standard_DS14_v2. Two datacenters created from the same omission, through different surfaces, get different virtual machines at different prices. That matters when comparing a Terraform-managed datacenter against one an ARM template or the CLI created.

🔴 Microsoft publishes several different lists of allowed sizes across different pages — one v5-only, one v4-plus-DS, one adding the L series — and the provider's own default appears in only one of them. So no closed set is enforced here. The module rejects a case-only near miss, which is a real mistake with a real fix, and lets an unlisted SKU through, because refusing a legal size is the worse error.

⚠️ Microsoft: the service "doesn't support transitioning across VM families", with Standard_DS13_v2 → Standard_DS14_v2 given as an example needing a support ticket. sku_name is not force-new and is sent on update, so Terraform will plan a change the service may decline. vm_family_transitions_require_a_support_ticket is emitted for that reason.

6 · The preview tier with no SLA
module "dc_cached" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]

  # An L-series size: locally attached disks, write-through caching.
  sku_name = "Standard_L16s_v3"
}

check "no_preview_tiers_in_production" {
  assert {
    condition     = !module.dc_cached.sku_is_a_preview_write_through_caching_tier
    error_message = "This datacenter is on an L-series VM size. Microsoft introduced those as write-through caching tiers with locally attached disks, ships them in PUBLIC PREVIEW without a service-level agreement, and does not recommend them for production workloads. Nothing in the plan says so, and the SKU name does not look different from any other. Choose an E or D series size, or delete this assertion having accepted the preview terms."
  }
}

🔴 Microsoft: write-through caching "is provided without a service-level agreement" and "We don't recommend it for production workloads." The tier is genuinely useful for read-heavy workloads — locally attached disks give higher read IOPS and lower tail latency — but that trade is a decision, and Standard_L16s_v3 looks exactly as production-ready as Standard_E16s_v5 in a plan diff.

💡 The flag is derived from the name rather than a list, so a new L-series size Microsoft adds later is caught without a module change.

7 · Availability zones, where the provider's own default can fail the apply
# The default: zones on.
module "dc_zoned" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]
  # availability_zones_enabled omitted -> true
}

# A region without zone support: the argument must be turned OFF explicitly.
module "dc_unzoned_region" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-secondary"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "westcentralus"
  delegated_management_subnet_id = var.westcentralus_cassandra_subnet_id

  availability_zones_enabled = false
}

output "zone_hazard" {
  value = module.dc_zoned.availability_zones_are_unsupported_in_some_regions_and_fail_the_deployment
}

🔴 Microsoft: "Availability zones aren't supported in all regions. Deployments fail if you select a region where availability zones aren't supported." The provider defaults this argument to true, so the provider's own default fails the apply in a region without zone support — and the module cannot see which regions those are, so it does not guess.

⚠️ Even in a zone-enabled region, success depends on capacity in every zone. Microsoft: deployment is "subject to the availability of compute resources in all the zones", so a large sku_name can fail where a smaller one succeeds.

🔒 The default is kept rather than flipped. Zone spreading raises the service's availability SLA, and this suite treats resilience as the caller's decision rather than something to reduce quietly — the same boundary that keeps secure-by-default about exposure rather than spend or redundancy.

⚠️ And per example 3, turning this off after creation does nothing at all.

8 · The two grants no argument can express
# 1. Tenant-wide, and a tenant administrator's job. Not Terraform's:
#
#      New-AzADServicePrincipal -ApplicationId a232010e-820c-4083-83bb-3ace5fc29d0b
#
# 2. Network Contributor for that same application, on the VNET holding the subnet.
#    This one IS expressible, with a sibling module:

module "cassandra_service_network_access" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-role-assignments.git?ref=v1.0.0"

  scope = module.cassandra_network.id

  role_assignments = {
    cosmos_db_service_network_contributor = {
      role_definition_name = "Network Contributor"

      # The Azure Cosmos DB application's object id in THIS tenant -- resolve it out of band
      # from the application id a232010e-820c-4083-83bb-3ace5fc29d0b.
      principal_id = var.cosmos_db_service_principal_object_id
    }
  }
}

module "dc_eastus" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]

  # The grant must exist BEFORE the deployment, and nothing here creates the ordering edge.
  depends_on = [module.cassandra_service_network_access]
}

🔴 Both prerequisites are governance acts rather than resources, and neither is visible from this module. The tenant service principal is created once per tenant by an administrator; the role assignment is scoped to a virtual network that this module only ever sees as a subnet ID. Their absence fails the apply with a networking error that does not name either of them.

💡 depends_on is the right tool here and is valid in a module block — unlike lifecycle, which is not. There is no attribute reference between the role assignment and the datacenter, so the ordering edge has to be stated.

⚠️ Microsoft's own command uses the role definition GUID 4d97b98b-1d4f-4787-a291-c67834d212e7 rather than a name; the sibling module takes the name, which resolves to the same definition.

ℹ️ principal_id is the service principal's object ID in your tenant, which is not the application ID above — resolve it once and pass it in. The module does not accept an application ID because Azure role assignments do not take one.

9 · Customer-managed keys, and whose identity needs the grant
module "cassandra_vault" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-key-vault.git?ref=v1.0.0"

  name                = "kv-cassandra-prod"
  resource_group_name = module.cassandra_rg.name
  location            = module.cassandra_rg.location
  tenant_id           = var.tenant_id

  keys = {
    # key_opts is REQUIRED by that module, and these three are exactly the operations
    # Microsoft says the cluster's identity needs on the key.
    cassandra-backup-cmk = {
      key_type = "RSA"
      key_size = 2048
      key_opts = ["get", "wrapKey", "unwrapKey"]
    }
    cassandra-disk-cmk = {
      key_type = "RSA"
      key_size = 2048
      key_opts = ["get", "wrapKey", "unwrapKey"]
    }
  }
}

module "dc_eastus" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]

  # key_ids, NOT key_versionless_ids -- see the callout below.
  backup_storage_customer_key_uri = module.cassandra_vault.key_ids["cassandra-backup-cmk"]
  managed_disk_customer_key_uri   = module.cassandra_vault.key_ids["cassandra-disk-cmk"]
}

# ALL THREE REJECTED at plan time:
#   "https://kv.vault.azure.net/keys/cmk"                 -- versionless
#   "https://kv.vault.azure.net/secrets/cmk/<version>"    -- a secret, not a key
#   "my-cmk"                                              -- not a URI at all

output "encryption" {
  value = {
    backup_cmk        = module.dc_eastus.backup_storage_uses_a_customer_managed_key
    disk_cmk          = module.dc_eastus.managed_disks_use_a_customer_managed_key
    identity_grant    = module.dc_eastus.customer_managed_keys_need_key_permissions_on_the_clusters_identity
  }
}

🔴 The identity that needs key permissions belongs to the CLUSTER, not to this datacenter. Microsoft: "Ensure the system assigned identity of the cluster has been assigned appropriate permissions (key get/wrap/unwrap permissions) on the key." So the grant is made against the cluster module's identity output, one level up from the resource that references the key.

🔴 The vault module emits both forms and only ONE of them works here. key_ids carries each key's versioned ID; key_versionless_ids carries the versionless one, and that module's own description recommends the versionless form "for CMK wiring that should follow rotation". This provider validates with VersionTypeVersioned and rejects it, so the general advice is wrong for this specific resource — wire key_ids and rotate by updating the argument.

⚠️ That is why the versionless case gets its own error message rather than falling through the generic one: omitting the version usually intends automatic rotation, and here it achieves nothing at all.

🔒 Neither argument is a secret and neither output emits one. A key URI is a reference; the key material never enters Terraform, and marking the URI sensitive would redact a value that exists to be referenced while protecting nothing.

ℹ️ Omitting both leaves Microsoft-managed encryption at rest in place. It is on either way — the choice is who holds the key, not whether there is one.

10 · The cassandra.yaml override
module "dc_eastus" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                           = "dc-eastus"
  cassandra_cluster_id           = module.cassandra_cluster.id
  location                       = "eastus"
  delegated_management_subnet_id = module.cassandra_network.subnet_ids["cassandra-nodes"]

  # Base64, not raw YAML -- and only an unpublished subset of keys is accepted.
  base64_encoded_yaml_fragment = base64encode(file("${path.module}/cassandra-fragment.yaml"))
}

# REJECTED by the module: raw YAML does not decode as base64.
#   base64_encoded_yaml_fragment = "num_tokens: 256"

output "config" {
  value = module.dc_eastus.cassandra_yaml_is_overridden
}

💡 The fragment is merged into cassandra.yaml on every node in this datacenter, which is the flexibility the managed-instance service offers over the platform Cassandra API — real open-source configuration, not an interoperability layer.

⚠️ Microsoft states that "only a subset of keys are allowed" and does not publish which subset. So the module validates the encoding and says nothing about the contents; a rejected key surfaces at apply, not at plan. The decode check is a heuristic on the likely mistake — passing the YAML instead of its encoding — and its message says so rather than implying a rule about the keys.

ℹ️ cassandra_yaml_is_overridden exists because a base64 blob in a plan diff tells a reviewer nothing. The flag at least says a configuration override is in play and that the module cannot inspect it.

11 · The paths, and the comparisons the module cannot make itself
check "the_datacenter_belongs_to_the_cluster_this_composition_created" {
  assert {
    condition     = module.dc_eastus.cassandra_cluster_id == module.cassandra_cluster.id
    error_message = "The datacenter names a different cluster from the one this composition created. Nothing in the provider compares them."
  }
}

check "the_nodes_share_the_clusters_virtual_network" {
  assert {
    condition = (
      module.dc_eastus.delegated_subnet_vnet_path ==
      module.cassandra_network.id
    )
    error_message = "This datacenter's delegated subnet is in a different virtual network from the one this composition created. Microsoft requires the datacenter's subnet to be able to ROUTE to the cluster's management subnet -- legal across peered networks, but then the peering and the service's role assignment on THAT network both have to exist. Assert deliberately or delete this check."
  }
}

output "path_operands" {
  value = {
    cluster_path      = module.dc_eastus.cluster_path_within_the_subscription
    datacenter_path   = module.dc_eastus.datacenter_path_within_the_subscription
    subnet_vnet       = module.dc_eastus.delegated_subnet_vnet_path
    cross_subscription = module.dc_eastus.subnet_crosses_a_subscription_boundary
    rules_unchecked   = module.dc_eastus.the_subnet_region_and_routing_rules_are_documented_and_unchecked
  }
}

🔴 Microsoft states two constraints on the subnet that no configuration can verify: it must be in the same region as this datacenter's location, and it must be able to route to the cluster's management subnet. A subnet Resource ID carries no region, and routing is not a plan-time fact. So the module emits the operands instead of pretending to check — delegated_subnet_vnet_path is the value a check block can actually compare.

💡 Same-virtual-network is a sufficient condition, not a necessary one. Microsoft documents datacenters reached across peered networks, so a failing second assertion above is a prompt to confirm the peering exists, not proof of a mistake. subnet_crosses_a_subscription_boundary is emitted for the same reason: legal, and it means the service's Network Contributor grant has to be made in the other subscription.

ℹ️ Unlike the Cassandra API modules in this family, there is no shared account_path_within_the_subscription to compare here — this is the managed-instance branch, and the cluster module was authored before that convention existed. So the first assertion compares the cluster ID directly, which is what the cluster module emits.

12 · 🏗️ End-to-end composition — a multi-region managed Cassandra cluster
locals {
  platform_tags = {
    workload = "cassandra"
    managed  = "terraform"
  }

  # A datacenter per region. Each needs its own subnet, in its own region.
  regions = {
    eastus = {
      vnet_cidr   = "10.20.0.0/16"
      subnet_cidr = "10.20.1.0/24"
      nodes       = 3
    }
    westus2 = {
      vnet_cidr   = "10.21.0.0/16"
      subnet_cidr = "10.21.1.0/24"
      nodes       = 3
    }
  }
}

module "cassandra_rg" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-cassandra-prod"
  location = "eastus"

  tags = local.platform_tags
}

# One virtual network per region, each with a subnet delegated to the managed service.
module "cassandra_network" {
  for_each = local.regions
  source   = "git::https://github.com/microsoftexpert/terraform-azurerm-virtual-network.git?ref=v1.0.0"

  name                = "vnet-cassandra-${each.key}"
  resource_group_name = module.cassandra_rg.name
  location            = each.key
  address_space       = [each.value.vnet_cidr]

  subnets = {
    cassandra-nodes = {
      address_prefixes = [each.value.subnet_cidr]
    }
  }

  tags = local.platform_tags
}

# The service's OWN access to each network. Required before any datacenter deploys.
module "cassandra_service_network_access" {
  for_each = local.regions
  source   = "git::https://github.com/microsoftexpert/terraform-azurerm-role-assignments.git?ref=v1.0.0"

  scope = module.cassandra_network[each.key].id

  role_assignments = {
    cosmos_db_service_network_contributor = {
      role_definition_name = "Network Contributor"
      principal_id         = var.cosmos_db_service_principal_object_id
    }
  }
}

module "cassandra_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-cluster.git?ref=v1.0.0"

  name                = "cassandra-prod"
  resource_group_name = module.cassandra_rg.name
  location            = "eastus"

  # The cluster's MANAGEMENT subnet. Every datacenter's subnet must be able to route to it.
  delegated_management_subnet_id = module.cassandra_network["eastus"].subnet_ids["cassandra-nodes"]

  # REQUIRED by the cluster module, sensitive, and immutable -- provisioned out of band and never
  # committed. It is required even when password authentication is not the chosen method.
  default_admin_password = var.cassandra_default_admin_password

  tags = local.platform_tags
}

# One datacenter per region. All four of its parents come from the modules above.
module "cassandra_datacenter" {
  for_each = local.regions
  source   = "git::https://github.com/microsoftexpert/terraform-azurerm-cosmosdb-cassandra-datacenter.git?ref=v1.0.0"

  name                 = "dc-${each.key}"
  cassandra_cluster_id = module.cassandra_cluster.id
  location             = each.key

  delegated_management_subnet_id = module.cassandra_network[each.key].subnet_ids["cassandra-nodes"]

  node_count = each.value.nodes

  # Both are create-time decisions in practice, so they are stated rather than left to drift.
  disk_count                 = 4
  availability_zones_enabled = true

  # Scaling and the known update defect both want more than the default hour.
  timeouts = {
    create = "90m"
    update = "2h"
  }

  # The service reaches the network through a grant this composition makes, not an argument.
  depends_on = [module.cassandra_service_network_access]
}

check "every_datacenter_uses_its_own_regional_network" {
  assert {
    condition = alltrue([
      for region, dc in module.cassandra_datacenter :
      dc.delegated_subnet_vnet_path == module.cassandra_network[region].id
    ])
    error_message = "A datacenter's delegated subnet is not in that region's own virtual network. Microsoft requires the subnet to be in the same region as the datacenter, and nothing in the provider checks it."
  }
}

check "no_datacenter_is_on_a_preview_tier" {
  assert {
    condition = alltrue([
      for dc in module.cassandra_datacenter : !dc.sku_is_a_preview_write_through_caching_tier
    ])
    error_message = "A datacenter is on an L-series VM size, which Microsoft ships in public preview with no SLA and does not recommend for production."
  }
}

output "what_this_composition_accepts" {
  description = "Named for the decisions that cannot be revisited, and the ones made outside Terraform."
  value = {
    total_disks_billed    = { for r, dc in module.cassandra_datacenter : r => dc.total_data_disks }
    seed_nodes            = { for r, dc in module.cassandra_datacenter : r => dc.seed_node_ip_addresses }
    disks_are_create_only = module.cassandra_datacenter["eastus"].disk_count_changes_are_silently_dropped_on_update
    zones_are_create_only = module.cassandra_datacenter["eastus"].availability_zone_changes_are_silently_dropped_on_update
    nodes_converge_later  = module.cassandra_datacenter["eastus"].node_count_is_a_desired_state_azure_converges_to_asynchronously
    tenant_principal      = module.cassandra_datacenter["eastus"].the_tenant_must_carry_the_azure_cosmos_db_service_principal
    egress_required       = module.cassandra_datacenter["eastus"].the_service_requires_outbound_internet_access
  }
}

🔒 A multi-region managed Cassandra cluster is one cluster and several datacenters, each with its own location and its own regional subnet — which is why location is a datacenter argument rather than something inherited from the cluster. The for_each over local.regions keys the datacenters, their networks and their role assignments by region, so all three stay aligned by construction.

🔴 The role assignment is a real module here, not a comment, and depends_on carries the ordering edge because there is no attribute reference between a role assignment and a datacenter. The tenant service principal is the one prerequisite that stays outside the configuration entirely.

💡 disk_count and availability_zones_enabled are written out explicitly even though both match the defaults. They are create-time decisions that nothing marks as such, so stating them keeps the intent in the file rather than in a plan diff that would silently accept a later edit.

⚠️ Twelve P30 disks per datacenter, twenty-four in total, and total_data_disks reports it per region because neither node_count nor disk_count shows the product. Combined with Standard_E16s_v5 at three nodes per region, this is not a small default.

ℹ️ seed_node_ip_addresses is known only after apply. It is what an existing on-premises or other-cloud Cassandra ring needs in its own seed provider configuration to form a hybrid cluster with these datacenters.


📥 Inputs

Group Variables
Identity and placement name, cassandra_cluster_id, location, delegated_management_subnet_id
Capacity and cost node_count, disk_count, disk_sku, sku_name
Resilience availability_zones_enabled
Encryption backup_storage_customer_key_uri, managed_disk_customer_key_uri
Cassandra configuration base64_encoded_yaml_fragment
Universal tail timeouts (no tags — the resource has none)
Full input schemas
Variable Type Default Notes
name string — Force-new. Blank and Resource-ID rules.
cassandra_cluster_id string — Force-new. Anchored at cassandraClusters/<name> with a terminator, plus a rule rejecting a databaseAccounts ID — the wrong service in the same family.
location string — Force-new. The datacenter's region. Blank and Resource-ID rules.
delegated_management_subnet_id string — Force-new. Anchored at subnets/<name>; a virtual network ID is the near miss.
node_count number 3 >= 3, whole. Editable, and a desired state.
disk_count number 4 1–10, whole. The default is this module's correction — the provider has none and sends 0. Not editable in practice.
disk_sku string "P30" Blank rule, plus a heuristic rejecting a Standard_-prefixed VM SKU swapped in.
sku_name string "Standard_E16s_v5" Blank rule, a case-only near-miss rule against the union of Microsoft's published lists, and a heuristic rejecting a P30-style disk designator. No closed set.
availability_zones_enabled bool true No rules — a bool with two legal values. Not editable in practice.
base64_encoded_yaml_fragment string null Blank-when-set, plus a Base64-decodability heuristic.
backup_storage_customer_key_uri string null Versioned key URI; versionless and secret/certificate URIs each rejected with their own message.
managed_disk_customer_key_uri string null The same three rules.
timeouts object({ create, read, update, delete }) {} → 60m / 5m / 60m / 60m Four keys, all meaningful here.

27 validations. Three are coverage the provider has no equivalent for (the wrong-service cluster ID, and the two SKU-swap heuristics); the rest either anchor an ID the provider validates loosely or restate a bound so the message names the argument. Six of the twenty-seven are on the two customer-key URIs, because a key URI has three distinct ways of being wrong and each deserves its own message.

No cross-field validations. Unusually for this family, there are none to make: the region-versus-subnet and subnet-versus-cluster relationships Microsoft documents are not derivable from the values supplied, so they are prerequisites and emitted operands instead of rules.


🧾 Outputs

50 outputs: 11 passthrough, 23 derived, 16 constant. None is sensitive; this module accepts and emits no secret.

Sixteen constants is a lot, and it is proportionate here. Every one of them describes something that produces no error when it bites: an update that succeeds and changes nothing, an omission that sends a value the validator forbids, a grant performed by a tenant administrator, a node count Azure is still converging toward.

Output Description
id, name, location, cassandra_cluster_id, delegated_management_subnet_id Identity and placement, as applied.
node_count, disk_count, disk_sku, sku_name, availability_zones_enabled Capacity, as applied.
seed_node_ip_addresses, seed_node_count The seed nodes Azure assigned. After apply only, and not the same as node_count.
cluster_name, cluster_resource_group_name, cluster_subscription_id Parsed out of the cluster ID — including the resource group, because this resource has no argument for one.
cluster_path_within_the_subscription, datacenter_path_within_the_subscription Exact but for the subscription.
delegated_subnet_name, delegated_subnet_virtual_network_name, delegated_subnet_vnet_path The subnet, and the network a check block compares.
subnet_is_in_the_clusters_subscription, subnet_is_in_the_clusters_resource_group, subnet_crosses_a_subscription_boundary Legal either way; reported.
total_data_disks node_count times disk_count — what actually bills.
disk_count_is_below_microsofts_recommended_minimum, disk_sku_is_not_the_p30_the_disk_count_guidance_assumes Where Microsoft's sizing guidance stops describing what is deployed.
sku_is_the_providers_default, sku_is_the_arm_documented_default, sku_is_a_preview_write_through_caching_tier, sku_appears_in_no_microsoft_published_list The four things worth knowing about a VM size here.
backup_storage_uses_a_customer_managed_key, managed_disks_use_a_customer_managed_key, any_customer_managed_key_is_configured Encryption posture.
cassandra_yaml_is_overridden A configuration the module cannot inspect.
sixteen constants Enumerated in Architecture Notes.

🧠 Architecture Notes

The disk count is the sharpest thing in this resource. disk_count is Optional with IntBetween(1, 10) and no Default, and the create path sends DiskCapacity: pointer.To(int64(d.Get("disk_count").(int))) unconditionally. An absent integer reads as zero in the plugin SDK, so omitting the argument transmits 0 — a value the same schema rejects the instant a caller types it. ARM documents the property's default as 4 and Microsoft's portal calls four "strongly recommended", so this module supplies 4. That is the only place the module overrides the provider's notion of an omission, and it does so because every documented source agrees on a value the provider does not send.

Two arguments accept a change and discard it. The update payload carries the delegated subnet, node count, SKU, location and disk SKU. It omits disk capacity and availability zone entirely — and neither disk_count nor availability_zones_enabled is force-new. So Terraform plans the change, the apply reports success, Azure never hears about it, and the read that follows overwrites state with the unchanged value. Both are create-time decisions that nothing in the schema marks as such, which is why each gets its own constant output: the failure is silent in both directions.

Three sources disagree about the virtual machine SKU. The provider defaults sku_name to Standard_E16s_v5; the ARM contract documents Standard_DS14_v2; and Microsoft's own pages publish several different lists of allowed sizes, with the provider's default in only one of them. That is why no closed set is enforced — a list assembled from any one page would refuse sizes another page publishes. The module rejects a case-only near miss, because that is a real mistake with an unambiguous fix, and allows everything else while reporting whether the value appears in any published list. Layered on top, Microsoft says the service "doesn't support transitioning across VM families", so even a valid change may be declined at apply.

Two prerequisites are governance acts. The Azure Cosmos DB application a232010e-820c-4083-83bb-3ace5fc29d0b must exist as a service principal in the tenant, added once by an administrator, and it must hold Network Contributor on the virtual network holding the delegated subnet. The first is not a resource in any provider's vocabulary; the second is a role assignment on a network this module sees only as part of a subnet ID. Both fail the apply when absent, and neither failure names itself clearly, so both are constants and both are in the prerequisites.

This resource is in the wrong family, in a sense worth stating. Azure Managed Instance for Apache Cassandra runs pure open-source Cassandra on virtual machine scale sets in the caller's own virtual network. Azure Cosmos DB for Apache Cassandra is a platform service exposing the Cassandra wire protocol over the Cosmos DB backend. Microsoft says there is no architectural dependency between them. They share the Microsoft.DocumentDB resource provider, the cosmosdb_cassandra_ prefix and this module family, and nothing else — which is exactly the confusion the cassandra_cluster_id validation catches by rejecting a databaseAccounts ID.

The ID constructor here does the opposite of its sibling's. NewDataCenterID takes the subscription, resource group and cluster name from the parsed cassandra_cluster_id and nothing from provider configuration, so a cross-subscription cluster ID is honored. The Cassandra API table in this family parses its parent's ID and then substitutes the provider's subscription, silently relocating the resource. Same family, same shape of argument, opposite semantics — which is why this module emits cluster_subscription_id as a fact that is used rather than merely recorded.

node_count is a target, not a state. Microsoft: "the desired number. After it is set, it may take some time for the data center to be scaled to match." A successful apply means the target was accepted. The way to see reality is the cluster's fetchNodeStatus operation, which this provider exposes as neither a resource nor a data source — so there is no Terraform-native confirmation that a scale-up finished. The provider's 60-minute create, update and delete defaults, and the update path's one-minute polling delay for Azure/azure-rest-api-specs#19078, are all consequences of the same asynchrony.

What the module cannot compare, it emits. Microsoft requires the delegated subnet to be in this datacenter's region and to route to the cluster's management subnet. A subnet ID carries no region and routing is not a plan-time fact, so there are no cross-field validations on this resource at all — unusually for this family. Instead delegated_subnet_vnet_path, cluster_path_within_the_subscription and subnet_crosses_a_subscription_boundary are emitted as operands, and a check block in the caller's configuration does the comparing it is actually able to do.

Two observations about the read worth knowing before debugging one. It sets disk_count via int(*props.DiskCapacity) with no nil guard, where every neighboring field is set through the pointer; and the seed-node flattener appends each IPAddress pointer rather than the dereferenced string, unlike the surrounding code. Neither is something a module can work around — they are recorded so that odd behavior in those two fields is recognized rather than re-derived.


🧱 Design Principles

There is no secure-defaults table on this module, and the scaffold that claimed one was describing a different resource. Nothing here gates network exposure, there is no TLS setting, and encryption at rest is on regardless — the only choice is who holds the key. The scaffold asserted "the empty call yields the hardened resource" with public access, weaker TLS, disabled protections as the opt-outs, named a resource_group_name argument this resource does not have, listed only name and that non-existent argument as immutable while four fields are force-new, and required merely "an existing resource group in a supported US Azure region" for a deployment that needs a tenant service principal, a role assignment on a virtual network, and outbound internet access.

Decision Value Reasoning
disk_count default 4, supplied by this module The provider has no default and sends 0, below its own validator's floor. ARM documents 4; Microsoft recommends at least four. Restoring a documented default is not inventing one.
node_count default 3 The provider's default and its documented floor. Nothing to add.
sku_name closed set Not enforced Microsoft's published lists disagree with each other and the provider's default is in only one. A closed set would refuse legal sizes.
sku_name case near miss Rejected A case-only mismatch is a real mistake with one obvious fix, and the value is case-sensitive.
The disk/VM SKU swap Rejected in both directions Two adjacent arguments, neither validated against a list. Stated as a heuristic on the naming convention in both messages.
availability_zones_enabled default The provider's true, kept Zone spreading raises the availability SLA. Secure-by-default covers exposure, not resilience — reducing it silently is not this module's call, even though the default can fail an apply in a region without zone support.
The region/routing rules Prerequisites, plus emitted operands Not derivable from the values supplied. A validation would have to guess.
base64_encoded_yaml_fragment Encoding validated, contents not Microsoft does not publish which cassandra.yaml keys are accepted. The decode check names itself a heuristic.
Customer-key URIs Three rules each The provider requires a versioned key ID; versionless, secret and certificate URIs each fail differently and each earn their own message.
Key URIs marked sensitive No A key URI is a reference that exists to be passed around; the key material never enters Terraform. Redaction would obscure plan review and protect nothing.
location as an argument Correct, not inherited A datacenter carries its own region. That is how a multi-region cluster is expressed.
No tags variable Correct The resource exposes none. Tag the cluster or the resource group.
The two governance acts Constants and prerequisites Neither is a resource this module can create, and both fail the apply with an error that does not name them.

🚀 Runbook

terraform init -backend=false
terraform fmt -check
terraform validate

Pin the module with ?ref=v1.0.0 and never a branch. This library is plan-only: a human applies from CI.

Budget real time for an apply. The provider's create, update and delete timeouts all default to 60 minutes, and Microsoft notes a cluster alone can take up to 15 minutes before its first datacenter starts.


🧪 Testing

terraform validate proves the configuration parses. It does not fire a variable validation when this module is called from another configuration — terraform console with a -var-file is the offline harness for that.

What was proven offline for this module:

  • All 27 validations fired, each by a fixture built to fail it, and zero condition-evaluation errors in the final run.
  • The first run reported 27/27 reached with TWO condition-evaluation errors, and both were a FIXTURE defect rather than a module defect. The raw-YAML fixture for base64_encoded_yaml_fragment contained an embedded newline, and the generator escapes only the double quote — so the .tfvars had a raw line break inside a quoted string and failed to parse. A parse failure reads exactly like a rule that throws, which reads exactly like a rule that never fires, so it was worth chasing rather than accepting. Rewritten as a single-line fragment; the rule then fired on the encoding, as intended.
  • All 35 locals were driven to more than one value across four good fixtures: the bare minimum, the ARM-default SKU with both customer keys and a non-P30 disk size and zones off, a preview L-series tier with a subnet in a different subscription and resource group, and an unlisted SKU against an entirely different cluster.
  • One local was never-varying and was inlined rather than fixture-covered. published_sku_names held the union of Microsoft's SKU lists and referenced no variable, so it could not vary by construction — the standing discriminator for "constant in the wrong place" rather than "fixture gap". It was folded into the expression that used it, and the module was then grepped for the deleted name.
  • The output tuple was computed mechanically before this README was written, so the counts here were never a hand tally.
  • The terraform validate result was confirmed with the working directory printed, its .tf files listed, and a zz_negctl.tf negative control in the module's own directory: it fails with the control present and passes with it removed.
  • The .tf files were swept for non-ASCII immediately after writing, and again after every edit.

What only an apply exercises: whether the tenant carries the Cosmos DB service principal, whether the virtual network role assignment exists, whether the subnet's region matches and routes to the cluster, whether the region supports availability zones, whether the chosen SKU has capacity in every zone, and whether the cassandra.yaml fragment's keys are accepted.


💬 Example Output

id = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-cassandra-prod/providers/Microsoft.DocumentDB/cassandraClusters/cassandra-prod/dataCenters/dc-eastus"
name = "dc-eastus"
location = "eastus"
node_count = 3
disk_count = 4
disk_sku = "P30"
sku_name = "Standard_E16s_v5"
availability_zones_enabled = true
total_data_disks = 12
seed_node_ip_addresses = ["10.20.1.4", "10.20.1.5", "10.20.1.6"]
seed_node_count = 3
cluster_name = "cassandra-prod"
cluster_resource_group_name = "rg-cassandra-prod"
datacenter_path_within_the_subscription = "/resourceGroups/rg-cassandra-prod/providers/Microsoft.DocumentDB/cassandraClusters/cassandra-prod/dataCenters/dc-eastus"
delegated_subnet_vnet_path = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-cassandra-prod/providers/Microsoft.Network/virtualNetworks/vnet-cassandra-eastus"
subnet_is_in_the_clusters_subscription = true
disk_count_is_below_microsofts_recommended_minimum = false
disk_sku_is_not_the_p30_the_disk_count_guidance_assumes = false
sku_is_the_providers_default = true
sku_is_a_preview_write_through_caching_tier = false
sku_appears_in_no_microsoft_published_list = false
any_customer_managed_key_is_configured = false
cassandra_yaml_is_overridden = false

Read total_data_disks = 12 as the bill: three nodes, four P30 disks each. Neither node_count nor disk_count shows that number.


🔍 Troubleshooting

Symptom Cause Fix
The apply fails with a networking or permission error naming nothing recognizable The Azure Cosmos DB service principal is missing from the tenant, or lacks Network Contributor on the virtual network Both are in Azure Prerequisites. They are the most common cause of a first-deployment failure and neither is an argument.
cassandra_cluster_id must be a full managed Cassandra CLUSTER Resource ID A datacenter ID, or an ID with extra segments The check is anchored with a terminator because a datacenter ID extends the same path. Pass the cluster module's id.
cassandra_cluster_id looks like a Cosmos DB ACCOUNT Resource ID The Cosmos DB Cassandra API was confused with the managed-instance service Different services. This resource needs a cassandraClusters ID; a databaseAccounts ID has no datacenters.
A plan proposes changing disk_count and the apply succeeds, but nothing changes Expected: the update payload omits the property and the argument is not force-new No fix. Treat the disk count as create-time; changing it for real means replacing the datacenter.
Toggling availability_zones_enabled does nothing The same defect Same answer. Decide before the first apply.
The deployment fails as soon as zones are involved The region does not support availability zones, or the SKU lacks capacity in every zone Set availability_zones_enabled = false, or choose a region or SKU that supports it.
An apply reports success but the cluster has fewer nodes than requested node_count is a desired state Azure converges to asynchronously Expected. Check with az managed-cassandra cluster status; the provider exposes no node-status surface.
A SKU change fails at apply The service refuses transitions across VM families Open a support ticket, per Microsoft. Terraform cannot tell in advance because the module cannot see the deployed value.
sku_name matches a SKU Microsoft publishes for this service except in case A lowercased SKU name Use the published spelling; the value is case-sensitive.
An unlisted SKU is accepted and later fails Deliberate: Microsoft's lists disagree, so unlisted values pass Read sku_appears_in_no_microsoft_published_list and confirm the size is offered in the target region.
backup_storage_customer_key_uri is a VERSIONLESS key URI A key URI without its version segment The provider requires a specific version. Append it.
Encryption with a customer key fails on the key operation The cluster's system-assigned identity lacks get/wrap/unwrap Grant it on the key. The identity is the cluster's, not the datacenter's.
base64_encoded_yaml_fragment must be Base64-encoded Raw YAML Wrap it: base64encode(file(...)).
The apply fails on an unrecognized cassandra.yaml key Only an unpublished subset is accepted Not checkable at plan time. Remove the key or consult Microsoft.
The update runs for a very long time Scaling is a scale-set operation, and the provider waits out Azure/azure-rest-api-specs#19078 after the operation returns Raise timeouts.update; the provider's default is already 60 minutes.
A caller wants prevent_destroy on this module lifecycle is not valid in a module block Use a CanNotDelete management lock — which prevents deletion, not replacement.

🔗 Related Docs


💙 "Infrastructure as Code should be standardized, consistent, and secure."