Skip to content

Latest commit

Β 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

☁️ Azure Data Factory Dataset β€” Parquet Terraform Module

A typed Data Factory dataset describing Parquet files in Blob Storage, Data Lake Gen2 or on an HTTP server, targeting hashicorp/azurerm ~> 4.0.

Terraform azurerm Module Type Resources Caveat


🧩 Overview

  • πŸ—œοΈ Creates one azurerm_data_factory_dataset_parquet β€” a named description of Parquet files that pipelines read from and write to.
  • πŸ“ Accepts exactly one of three location blocks β€” HTTP server, Blob Storage, or Data Lake Gen2 β€” enforced by the provider at terraform plan, offline and without credentials, even though this resource's documentation never says so.
  • 🚫 Accepts eight compression codecs. "None" is not one of them β€” and the delimited-text dataset accepts it. To configure no compression here, omit the argument.
  • πŸ”€ Makes http_server_location.path optional, where the delimited-text and JSON datasets both require it. One block name, three contracts across the family.
  • 🎚️ Carries up to three dynamic flags per location and reports six distinct mismatches plus a roll-up.

πŸ’‘ Why it matters: this resource shares all three location blocks, both compression arguments and the whole schema_column block with the delimited-text dataset β€” and still refuses a configuration copied from it, in two independent ways. Two resources can share almost everything and still disagree.


❀️ Support this project

If this module saved you time:


πŸ—ΊοΈ Where this fits in the family

flowchart TB
  RG["terraform-azurerm-resource-group"]
  SA["terraform-azurerm-storage-account"]
  ADF["terraform-azurerm-data-factory"]
  LS["a Data Factory linked service"]
  THIS["terraform-azurerm-data-factory-dataset-parquet"]
  DTEXT["terraform-azurerm-data-factory-dataset-delimited-text"]
  JSONDS["terraform-azurerm-data-factory-dataset-json"]
  SQLT["terraform-azurerm-data-factory-dataset-azure-sql-table"]
  PIPE["a Data Factory pipeline"]
  FILES["a container, a Data Lake file system, or an HTTP server"]

  RG -->|"name"| ADF
  RG -->|"name"| SA
  SA -->|"named inside the location block, never read here"| THIS
  ADF -->|"id"| THIS
  ADF -->|"id"| DTEXT
  ADF -->|"id"| JSONDS
  ADF -->|"id"| SQLT
  LS -->|"linked_service_name, a bare NAME"| THIS
  LS -->|"linked_service_id, a full Resource ID"| SQLT
  THIS -->|"codec set of EIGHT, and None is not one of them"| PIPE
  DTEXT -->|"codec set of NINE, including None"| PIPE
  JSONDS -->|"no compression arguments at all"| PIPE
  PIPE -->|"reads only at run time, never here"| FILES

  classDef this fill:#0078D4,stroke:#004578,color:#ffffff,stroke-width:2px
  classDef keystone fill:#004578,stroke:#00243d,color:#ffffff,stroke-width:2px
  classDef sibling fill:#eef3f8,stroke:#b9c8d8,color:#1b2733
  class THIS this
  class ADF keystone
  class RG,SA,LS,DTEXT,JSONDS,SQLT,PIPE,FILES sibling
Loading

The edge labels carry the difference most likely to survive a copy-paste: this resource's codec set has eight members and the delimited-text dataset's has nine, the extra one being exactly "None".


🧬 What this module builds

flowchart TB
  subgraph INPUTS["Inputs"]
    NAME["name (force-new)"]
    ADFID["data_factory_id (force-new)"]
    LSNAME["linked_service_name, a bare NAME"]
    LOC["EXACTLY ONE of three location blocks"]
    COMP["compression_codec, eight values, and compression_level"]
    COLS["schema_column list"]
    META["parameters, additional_properties, annotations, description, folder"]
  end

  THIS["azurerm_data_factory_dataset_parquet.this"]

  subgraph OUTPUTS["Outputs"]
    OID["id and name"]
    OKIND["location_kind, normalised path, filename, container_or_file_system"]
    OMIS["dynamic_flags_look_inconsistent, plus its six components"]
    ONONE["location_names_nothing and describes_no_specific_file"]
    OCOMP["the codec set EXCLUDES None, and compression_level_without_codec"]
    OFACTS["the constant facts, including the documentation gap"]
  end

  NAME --> THIS
  ADFID --> THIS
  LSNAME --> THIS
  LOC --> THIS
  COMP --> THIS
  COLS --> THIS
  META --> THIS

  THIS --> OID
  LOC --> OKIND
  LOC --> OMIS
  LOC --> ONONE
  META --> ONONE
  COMP --> OCOMP
  THIS --> OFACTS

  classDef this fill:#0078D4,stroke:#004578,color:#ffffff,stroke-width:2px
  classDef sibling fill:#eef3f8,stroke:#b9c8d8,color:#1b2733
  class THIS this
  class NAME,ADFID,LSNAME,LOC,COMP,COLS,META,OID,OKIND,OMIS,ONONE,OCOMP,OFACTS sibling
Loading
Resource Count Notes
azurerm_data_factory_dataset_parquet.this 1 The keystone. A real ARM child of the factory.
http_server_location dynamic, 0..1 relative_url + filename required; path optional.
azure_blob_storage_location dynamic, 0..1 Only container required.
azure_blob_fs_location dynamic, 0..1 Nothing required.
schema_column dynamic, 0..n Optional column definitions.
timeouts dynamic, 0..1 All four keys exist.

βœ… Provider / Versions

Requirement Value
Terraform >= 1.12.0
hashicorp/azurerm ~> 4.0
Provider block None. The caller configures the provider, including the mandatory features {} block.

Schema notes that bite β€” verified against the live provider source, not inferred from the schema:

  • πŸ”΄ compression_codec accepts EIGHT values here and NINE on the delimited-text dataset. The missing one is exactly "None". compression_codec = "None" plans cleanly on a delimited-text dataset and is refused here; omit the argument instead. Everything else about the two resources' compression arguments is identical.
  • πŸ”΄ http_server_location.path is OPTIONAL here and REQUIRED on the delimited-text and JSON datasets. One block name, three contracts.
  • πŸ”΄ ExactlyOneOf is declared across the three location blocks β€” and this resource's documentation never mentions it. The delimited-text page states "(exactly one of them must be set)"; this page just lists the blocks. The code is identical on both, so the rule holds and only the documentation is silent. Reading the page alone would suggest all three are optional.
  • πŸ”΄ The three location blocks do not share a shape. http_server_location requires relative_url and filename; azure_blob_storage_location requires only container; azure_blob_fs_location requires nothing β€” so azure_blob_fs_location = {} satisfies the rule while naming no file.
  • ⚠️ Each block names its dynamic flags after the field they govern β€” dynamic_container_enabled on the blob block, dynamic_file_system_enabled on the Data Lake block.
  • ⚠️ compression_level does not require compression_codec. A ratio with nothing to compress is accepted.
  • ⚠️ Only name and data_factory_id force replacement. The linked service and the entire location block update in place.
  • ⚠️ Timeouts default to 30m / 5m / 30m / 30m, all four keys present.
  • ℹ️ The requires-import error names this resource correctly β€” unlike the JSON dataset, whose guard names the delimited-text type.

πŸ”‘ Required Azure RBAC Roles / Permissions

Permission Scope Why
Microsoft.DataFactory/factories/datasets/write the Data Factory Create and update.
Microsoft.DataFactory/factories/datasets/read the Data Factory Refresh and plan.
Microsoft.DataFactory/factories/datasets/delete the Data Factory Destroy β€” see the warning below.
Data Factory Contributor the Data Factory The built-in role containing all three.

πŸ”’ No permission on the storage account, the Data Lake file system or the HTTP server is required, requested or used. The dataset names a file and its compression; it never opens one. Access belongs to the linked service, so whoever can write datasets can point one at any container the factory's identity can already reach β€” including data they have no direct permission on.

⚠️ Delete is the operation to review, not create. Removing a dataset breaks every pipeline that references it by name.


Azure Prerequisites

  • An existing Data Factory, and its Resource ID.
  • A linked service already configured in that factory, and its name. Nothing here verifies it.
  • A container, Data Lake file system or HTTP endpoint for the location block to name β€” or parameters.
  • A dataset name meeting Microsoft's Data Factory naming rules. The provider checks only non-emptiness.
  • The Microsoft.DataFactory resource provider registered in the subscription.

πŸ“ Module Structure

terraform-azurerm-data-factory-dataset-parquet/
β”œβ”€β”€ providers.tf    # required_version + the pinned azurerm; no provider block
β”œβ”€β”€ variables.tf    # 15 typed inputs, 19 validations
β”œβ”€β”€ main.tf         # the keystone, the location normalisation, five dynamic blocks
β”œβ”€β”€ outputs.tf      # 73 outputs; id first
β”œβ”€β”€ README.md       # this file
β”œβ”€β”€ SCOPE.md        # the cross-module contract
β”œβ”€β”€ LICENSE         # MIT
└── .gitignore

βš™οΈ Quick Start

provider "azurerm" {
  features {}
}

module "orders_parquet" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_orders_parquet"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.adls_linked_service_name

  azure_blob_fs_location = {
    file_system = "curated"
    path        = "orders"
    filename    = "orders.parquet"
  }

  compression_codec = "snappy"
}

ℹ️ Exactly one location block is required β€” a rule the provider enforces and this resource's documentation omits.


πŸ”Œ Cross-Module Contract

Consumes

Input Type Typical source
data_factory_id string terraform-azurerm-data-factory β†’ id
linked_service_name string a linked service's name
one location block object the caller

Emits (selected β€” 73 in total)

Output Description
id The dataset's Resource ID.
location_kind Which of the three blocks was supplied. Never null.
dynamic_flags_look_inconsistent The one output to assert on.
the_compression_codec_set_EXCLUDES_None_unlike_the_delimited_text_dataset Constant. Eight here, nine there.
the_http_location_block_is_looser_here_than_on_its_siblings Constant. path is optional only here.
the_provider_documentation_omits_the_exactly_one_rule Constant. A doc gap.

πŸ“š Example Library

1 Β· The minimum call β€” Data Lake Gen2
module "orders" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_orders"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.adls_linked_service_name

  azure_blob_fs_location = {
    file_system = "curated"
    path        = "orders"
    filename    = "orders.parquet"
  }
}

πŸ’‘ Parquet on Data Lake Gen2 is the common shape. Every field on this block is optional.

2 Β· Omitting the location is refused β€” a rule the docs do not state
# ❌ NOT ACCEPTED β€” rejected at terraform plan, offline and without credentials
module "broken" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_broken"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.adls_linked_service_name
  # no location block at all
}

⚠️ The provider declares ExactlyOneOf across the three blocks, so this fails offline. This resource's registry page never mentions the rule β€” only the delimited-text page does. the_provider_documentation_omits_the_exactly_one_rule records that.

3 Β· `"None"` is a legal codec on delimited-text and is refused here
# ❌ NOT ACCEPTED β€” the Parquet codec set has eight members and None is not one of them
module "copied_codec" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_copied"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.blob_linked_service_name

  azure_blob_storage_location = { container = "curated" }

  compression_codec = "None" # legal on the delimited-text dataset, REFUSED here
}

⚠️ To configure no compression on this resource, omit the argument. This is the difference most likely to survive a copy-paste, because everything else about the two resources' compression arguments is identical.

4 Β· An HTTP location without a path β€” legal only here
module "vendor_feed" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_vendor_feed"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.http_linked_service_name

  http_server_location = {
    relative_url = "/exports/daily"
    filename     = "orders.parquet"
    # path omitted -- OPTIONAL on this resource
  }
}

πŸ”’ The delimited-text and JSON datasets both require path on their block of the same name. A configuration valid here fails on both of them.

5 Β· A Blob Storage location, container only
module "curated" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_curated"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.blob_linked_service_name

  azure_blob_storage_location = {
    container = "curated"
  }
}

ℹ️ container is the only required field β€” the same rule as delimited-text, and not the same as the JSON dataset, which requires container, path and filename.

6 Β· The empty location that names nothing
module "resolves_to_nothing" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_empty"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.adls_linked_service_name

  azure_blob_fs_location = {} # satisfies ExactlyOneOf, names no file
}

check "dataset_points_somewhere" {
  assert {
    condition     = !module.resolves_to_nothing.describes_no_specific_file
    error_message = "The dataset names no file and exposes no parameters; no pipeline can resolve it."
  }
}

⚠️ This applies cleanly and reports nothing. location_names_nothing is true.

7 Β· All eight codecs, at their exact casing
locals {
  # The provider's own list mixes capitalised and lowercase names. Copy it; do not guess.
  codecs = ["bzip2", "gzip", "deflate", "ZipDeflate", "TarGzip", "Tar", "snappy", "lz4"]
}

module "by_codec" {
  for_each = toset(local.codecs)

  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_${lower(replace(each.key, "-", "_"))}"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.blob_linked_service_name

  azure_blob_storage_location = { container = "curated" }

  compression_codec = each.key
}

πŸ”’ Compared case-sensitively: "Gzip" is rejected where "gzip" is accepted.

8 Β· Compression level, and the pairing nothing enforces
module "compressed" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_compressed"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.blob_linked_service_name

  azure_blob_storage_location = { container = "curated", filename = "orders.parquet" }

  compression_codec = "snappy"
  compression_level = "Optimal" # capitalised -- "optimal" is REJECTED
}

check "compression_is_coherent" {
  assert {
    condition     = !module.compressed.compression_level_without_codec
    error_message = "A compression level was set with no codec; it configures nothing."
  }
}

⚠️ The provider accepts a level with no codec. The module reports that rather than refusing it.

9 Β· A dynamic path, flag and value agreeing
module "daily" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_daily"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.adls_linked_service_name

  azure_blob_fs_location = {
    file_system          = "curated"
    path                 = "@concat('orders/', formatDateTime(utcnow(), 'yyyy/MM/dd'))"
    dynamic_path_enabled = true
    filename             = "orders.parquet"
  }
}

πŸ”’ dynamic_flags_look_inconsistent is false β€” the value is an expression and the flag is set.

10 Β· The mismatch this module exists to catch
module "silently_wrong" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_silently_wrong"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.adls_linked_service_name

  azure_blob_fs_location = {
    file_system = "@pipeline().parameters.fs"
    # dynamic_file_system_enabled left at false
  }
}

check "dynamic_flags_agree" {
  assert {
    condition     = !module.silently_wrong.dynamic_flags_look_inconsistent
    error_message = "A dynamic flag disagrees with the value it governs."
  }
}

⚠️ Note the flag is dynamic_file_system_enabled on this block and dynamic_container_enabled on the blob block β€” the provider names it after the field it governs. The module unifies them into one output, dynamic_container_or_file_system_enabled.

11 Β· When the heuristic is wrong, and why it only reports
module "literal_at" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_literal_at"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.blob_linked_service_name

  azure_blob_storage_location = {
    container = "curated"
    path      = "@archive/2026" # a real folder that starts with @
  }
}

ℹ️ path_expression_without_dynamic_flag is true and the configuration is correct. Refusing would reject legal input, and a failing validation {} would also block terraform destroy.

12 Β· A typed column schema
module "typed" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_typed"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.adls_linked_service_name

  azure_blob_fs_location = { file_system = "curated", filename = "orders.parquet" }

  schema_column = [
    { name = "OrderId", type = "Int64", description = "Primary key." },
    { name = "Placed", type = "DateTimeOffset" },
    { name = "Notes" }, # untyped -- legal
  ]
}

check "every_column_is_typed" {
  assert {
    condition     = length(module.typed.untyped_schema_column_names) == 0
    error_message = "Untyped columns: ${join(", ", module.typed.untyped_schema_column_names)}"
  }
}

ℹ️ A Parquet file carries its own schema in its footer, so omitting schema_column is especially normal here. The fifteen types are case-sensitive and identical across the family.

13 Β· Many datasets from one map
locals {
  feeds = ["orders", "customers", "ledger"]
}

module "feeds" {
  for_each = toset(local.feeds)

  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_${each.key}"
  data_factory_id     = var.data_factory_id
  linked_service_name = var.adls_linked_service_name

  azure_blob_fs_location = {
    file_system = "curated"
    path        = each.key
    filename    = "${each.key}.parquet"
  }

  compression_codec = "snappy"
  folder            = "silver/parquet"
}

check "no_feed_has_a_flag_mismatch" {
  assert {
    condition     = alltrue([for m in module.feeds : !m.dynamic_flags_look_inconsistent])
    error_message = "At least one feed's dynamic flags disagree with the values they govern."
  }
}

πŸ’‘ The provider takes no lock on the Data Factory for this resource, so these are created concurrently and safely.

14 Β· Re-pointing a dataset is an in-place update
module "curated" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_curated"          # changing this REPLACES
  data_factory_id     = var.data_factory_id   # changing this REPLACES
  linked_service_name = var.other_linked_svc  # changing this updates IN PLACE

  azure_blob_fs_location = {                  # changing any of this updates IN PLACE
    file_system = "archive"
    path        = "orders"
  }
}

output "what_forces_replacement" {
  value = module.curated.force_new_fields # ["name", "data_factory_id"]
}

⚠️ Everything describing what the dataset reads β€” including which storage account, by way of the linked service β€” changes in place. A review scanning plans for destroys will not see it.

15 Β· πŸ—οΈ End-to-end composition
provider "azurerm" {
  features {}
}

module "rg" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-analytics-eastus"
  location = "eastus"
}

module "storage" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-storage-account.git?ref=v1.0.0"

  name                = "stanalyticseastus01"
  resource_group_name = module.rg.name
  location            = module.rg.location

  is_hns_enabled = true # Data Lake Gen2

  containers = {
    curated = {}
  }
}

module "adf" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory.git?ref=v1.0.0"

  name                = "adf-analytics-eastus"
  resource_group_name = module.rg.name
  location            = module.rg.location

  identity = {
    type = "SystemAssigned"
  }
}

# No linked-service module exists in this suite yet, so the linked service is
# declared directly. It authenticates with the factory's managed identity, so
# no secret appears anywhere in this configuration.
resource "azurerm_data_factory_linked_service_azure_blob_storage" "curated" {
  name            = "ls_adls_curated"
  data_factory_id = module.adf.id

  service_endpoint     = module.storage.primary_blob_endpoint
  use_managed_identity = true
}

module "orders_parquet" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"

  name                = "ds_orders_parquet"
  data_factory_id     = module.adf.id
  linked_service_name = azurerm_data_factory_linked_service_azure_blob_storage.curated.name

  azure_blob_fs_location = {
    file_system          = "curated"
    path                 = "@concat('orders/', formatDateTime(utcnow(), 'yyyy/MM/dd'))"
    dynamic_path_enabled = true
    filename             = "orders.parquet"
  }

  compression_codec = "snappy" # NOT "None" -- that value does not exist on this resource
  compression_level = "Optimal"
  folder            = "silver/parquet"

  schema_column = [
    { name = "OrderId", type = "Int64" },
    { name = "Placed", type = "DateTimeOffset" },
  ]

  annotations = ["silver", "daily"]
}

check "wiring_is_consistent" {
  assert {
    condition     = !module.orders_parquet.dynamic_flags_look_inconsistent
    error_message = "A dynamic flag disagrees with the value it governs."
  }
  assert {
    condition     = !module.orders_parquet.describes_no_specific_file
    error_message = "The dataset resolves to nothing."
  }
  assert {
    condition     = !module.orders_parquet.compression_level_without_codec
    error_message = "A compression level was set with no codec."
  }
}

output "dataset_id" {
  value = module.orders_parquet.id
}

output "reads_from" {
  value = {
    kind        = module.orders_parquet.location_kind
    file_system = module.orders_parquet.container_or_file_system
    path        = module.orders_parquet.location_path
    filename    = module.orders_parquet.location_filename
  }
}

πŸ”’ The factory's system-assigned identity authenticates to storage. Grant it Storage Blob Data Reader on the container out of band β€” no credential is stored in Terraform.


πŸ“₯ Inputs

Required: name, data_factory_id, linked_service_name, and exactly one of http_server_location / azure_blob_storage_location / azure_blob_fs_location. Compression: compression_codec (eight values, no "None"), compression_level. Shape: schema_column, parameters, additional_properties. Metadata: description, folder, annotations. Tail: timeouts. There is no tags variable β€” the resource supports none.

Full input schemas
Name Type Default Notes
name string β€” Force-new. Non-empty only; a leading / is refused.
data_factory_id string β€” Force-new. Anchored Resource-ID validator.
linked_service_name string β€” A bare name. A Resource ID is refused.
http_server_location object({ relative_url, filename, path, dynamic_path_enabled, dynamic_filename_enabled }) null relative_url + filename required; path optional. Carries the exactly-one check.
azure_blob_storage_location object({ container, path, filename, dynamic_container_enabled, dynamic_path_enabled, dynamic_filename_enabled }) null Only container required.
azure_blob_fs_location object({ file_system, path, filename, dynamic_file_system_enabled, dynamic_path_enabled, dynamic_filename_enabled }) null Nothing required.
compression_codec string null Closed, case-sensitive set of eight. "None" is absent.
compression_level string null Optimal or Fastest, case-sensitive.
schema_column list(object({ name, type, description })) [] Closed, case-sensitive set of fifteen types.
parameters map(string) {} Supplied per pipeline run.
additional_properties map(string) {} Unvalidated top-level keys.
annotations list(string) [] Not tags.
description string null Empty string refused; omission accepted.
folder string null The authoring tree, not a storage path.
timeouts object({ create, read, update, delete }) null 30m / 5m / 30m / 30m.

🧾 Outputs

Output Description Notes
id The dataset's Resource ID. First, by convention.
name, data_factory_id, data_factory_name Identity and parent. Name parsed from the end of the ID.
resource_group_name, subscription_id Where the factory lives.
linked_service_name The linked service, as supplied. A name β€” nothing can verify it.
location_kind Which block was supplied. Never null.
supplied_location_count Always 1 on a config that plans. Assert on the invariant.
location_path, location_filename, container_or_file_system, relative_url Normalised across three blocks.
location_names_nothing The location addresses no file.
describes_no_specific_file Names nothing and no parameters. Assert on this.
dynamic_path_enabled, dynamic_filename_enabled, dynamic_container_or_file_system_enabled How the strings are read.
path_looks_like_an_expression, filename_looks_like_an_expression, container_looks_like_an_expression The @ heuristic.
path_expression_without_dynamic_flag, filename_expression_without_dynamic_flag, container_expression_without_dynamic_flag Expression used literally.
path_marked_dynamic_but_looks_literal, filename_marked_dynamic_but_looks_literal, container_marked_dynamic_but_looks_literal The opposite mismatch. Different remedy.
dynamic_flags_look_inconsistent Any of the six. Assert on this.
compression_codec, compression_level The case-sensitive enums.
the_compression_codec_set_EXCLUDES_None_unlike_the_delimited_text_dataset Constant. Eight here, nine there.
compression_level_without_codec A ratio with nothing to compress. Assert on this.
the_compression_codec_set_is_closed_and_case_sensitive Constant. Mixed casing.
exactly_one_location_is_REQUIRED_on_this_resource Constant. The JSON dataset differs.
the_provider_documentation_omits_the_exactly_one_rule Constant. A doc gap.
the_three_location_blocks_do_not_share_a_shape Constant.
the_http_location_block_is_looser_here_than_on_its_siblings Constant. path optional only here.
schema_column_count, has_schema_columns, schema_column_names, untyped_schema_column_names The column schema.
parameters, parameter_count The run-time inputs.
uses_additional_properties, additional_property_count The escape hatch.
annotations, annotation_count, has_annotations A list, not tags.
description, has_description, folder, has_folder Metadata.
folder_and_the_location_path_mean_different_things Constant.
force_new_fields, fields_that_can_change_after_creation, fields_azure_returns_on_read The change surface.
this_resource_takes_a_linked_service_NAME_not_an_ID Constant. A 4-to-1 split.
the_dataset_family_is_not_a_clone_cluster Constant. This resource is the evidence.
the_linked_service_is_not_verified_to_exist Constant.
the_dynamic_flags_change_the_meaning_of_their_neighbours Constant. The trap, named.
the_schema_column_type_set_is_closed_and_case_sensitive Constant.
the_schema_column_block_is_identical_on_ten_of_the_twelve_datasets Constant. Identical on ten of twelve; binary has no block and snowflake's differs.
this_dataset_moves_no_data_by_itself Constant.
the_module_cannot_see_which_pipelines_use_this Constant. Read before destroying.
the_module_cannot_verify_the_file_exists, no_credential_is_configured_here Constants.
destroying_the_factory_destroys_this_dataset_too, destroying_this_does_not_touch_the_files Constants.
lifecycle_prevent_destroy_is_not_available_to_a_module_caller Constant.
this_is_a_real_azure_resource_not_a_composite, the_provider_takes_no_lock_on_the_data_factory Constants.
this_resource_supports_no_azure_resource_tags Constant. Why there is no tags variable.
no_secret_is_accepted_or_emitted_by_this_module Constant.

No output is sensitive, and none can be.


🧠 Architecture Notes

This resource is the family's best evidence against templating. It shares all three location blocks, both compression arguments and the entire schema_column block with the delimited-text dataset, and it still refuses a configuration copied from it in two independent ways. compression_codec accepts eight values here and nine there, the extra one being exactly "None" β€” so a working delimited-text configuration fails at terraform validate when the resource type is swapped. And http_server_location.path is optional here and required on both the delimited-text and JSON datasets. Two resources can share almost everything and still disagree.

The exactly-one rule holds, and the documentation does not say so. The provider declares ExactlyOneOf across the three location blocks on this resource exactly as it does on delimited-text. Only the delimited-text registry page states it. Reading this resource's page alone would suggest all three blocks are optional, and the failure appears at plan rather than in review β€” which is the good direction, but only if you know to expect it.

Three blocks, three shapes. http_server_location requires relative_url and filename; azure_blob_storage_location requires only container; azure_blob_fs_location requires nothing at all. That last one is why location_names_nothing exists: azure_blob_fs_location = {} satisfies the exactly-one rule and addresses no file.

The check is placed, not gathered. A validation {} condition must reference its own variable, and two variables validating each other is rejected as a cycle. The exactly-one check therefore lives on http_server_location and reads the other two one-directionally, which is why its message names blocks the caller may not have been editing.

Six mismatches, one roll-up. Each location names its dynamic flags after the field they govern, so the three blocks do not share a flag set β€” dynamic_container_enabled on one, dynamic_file_system_enabled on another. The module unifies them into dynamic_container_or_file_system_enabled, compares each flag against whether its value begins with @, emits all six mismatches individually, and rolls them up. All of it reports; none of it refuses.

Only identity forces replacement. The linked service and the entire location block update in place.

The module cannot see downstream. Pipelines reference a dataset by name and nothing points back. lifecycle is not valid inside a module block, so a caller cannot add prevent_destroy; a CanNotDelete lock prevents deletion but not the replacement that editing name would cause.


🧱 Design Principles

Concern This module's default Opt-out
Secrets None accepted, none emitted. Not available β€” by design.
Credentials Live on the linked service, never here. β€”
Storage access No permission required or used. β€”
A missing or duplicated location Refused at plan, by the provider and by this module. None.
Compression codec Enforced against the provider's eight, "None" excluded, with the exclusion named in the error message. None β€” the set is closed.
No compression Omit the argument. There is no value that means it. β€”
A compression level with no codec Reported, not refused. Ignore the output.
Dynamic flags All default to false β€” literal, the reading that cannot surprise you. Set one to true.
A flag disagreeing with its value Reported six ways plus a roll-up. Never refused. Ignore the outputs.
A location naming nothing Reported via location_names_nothing. Ignore the output.
Column types Enforced against the provider's case-sensitive fifteen. None β€” the set is closed.
tags Not offered β€” the resource supports none. Tag the factory.

πŸ”’ There is no risky toggle to close here: the resource holds no credential and touches no data. The exposures are the delete and the silent mismatches, so the module documents the first and reports the second.


πŸš€ Runbook

terraform init -backend=false
terraform validate
terraform fmt -check

Pin the module with ?ref=v1.0.0 β€” never a branch. This library is plan-only: a human applies from CI.


πŸ§ͺ Testing

terraform plan and the module's own validation {} blocks cover, offline and without credentials:

  • the exactly-one-location rule, in both failing directions;
  • each location block's own requiredness, which differs per block β€” including that an HTTP location without a path is accepted here;
  • the anchored Resource-ID shape of data_factory_id, and the refusal of an ID passed as linked_service_name;
  • all eight codecs accepted at their exact casing, and "None" refused β€” the case that distinguishes this resource from delimited-text;
  • the case-sensitive column-type set;
  • every derived flag β€” all six dynamic-flag mismatches, their roll-up, location_names_nothing, describes_no_specific_file and compression_level_without_codec β€” each with one fixture per branch through terraform console, which does fire root-module variable validations.

The harness also drives the same fixture against the delimited-text and JSON modules and asserts they disagree where the provider source says they must β€” so the contrast this README draws is executed, not merely asserted.

⚠️ terraform validate reaches none of that through a module call. Validate evaluates no module variable values, so neither this module's validation {} blocks nor the provider's ExactlyOneOf is reached and it reports success. The refusals above land at plan β€” still offline and without credentials.

Only a real apply can tell you whether the linked service exists, whether the file exists, or whether a pipeline still depends on the dataset.


πŸ’¬ Example Output

dataset_id = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-analytics-eastus/providers/Microsoft.DataFactory/factories/adf-analytics-eastus/datasets/ds_orders_parquet"

reads_from = {
  "file_system" = "curated"
  "filename"    = "orders.parquet"
  "kind"        = "azure_blob_fs_location"
  "path"        = "@concat('orders/', formatDateTime(utcnow(), 'yyyy/MM/dd'))"
}

supplied_location_count         = 1
dynamic_path_enabled            = true
path_looks_like_an_expression   = true
dynamic_flags_look_inconsistent = false
location_names_nothing          = false
describes_no_specific_file      = false
compression_codec               = "snappy"
compression_level_without_codec = false
force_new_fields                = ["name", "data_factory_id"]

πŸ” Troubleshooting

Symptom Cause Fix
compression_codec must be one of bzip2, gzip, deflate, ... on "None" "None" is not in this resource's set, though the delimited-text dataset accepts it. Omit the argument to configure no compression.
compression_codec must be one of bzip2, gzip, ... on "Gzip" The set is case-sensitive and mixes casings. Use "gzip". Copy the list; do not guess.
EXACTLY ONE of http_server_location, azure_blob_storage_location or azure_blob_fs_location must be set. No location, or more than one. Supply exactly one. This resource's documentation omits the rule; the provider enforces it.
http_server_location requires a non-empty relative_url AND filename. One of the two required fields was blank. path is optional here; the other two are not.
A config valid here fails on a delimited-text or JSON dataset Those require http_server_location.path. Add path.
azure_blob_storage_location.container must not be empty The one required field on that block. Supply a container.
A dataset applies but no pipeline can use it azure_blob_fs_location = {} with no parameters. Check describes_no_specific_file.
A pipeline looks for a folder literally named @concat(...) An expression with its dynamic flag unset. Set the flag. path_expression_without_dynamic_flag reports it.
A pipeline fails evaluating a plain folder name A literal marked dynamic. Unset the flag. path_marked_dynamic_but_looks_literal reports it.
dynamic_flags_look_inconsistent is true on a correct config The @ heuristic β€” a real folder may begin with @. Read the six component outputs and ignore the roll-up.
Setting dynamic_container_enabled on the Data Lake block does nothing That block's flag is dynamic_file_system_enabled. Use the flag named for the field.
Compression seems not to happen A level was set with no codec. Check compression_level_without_codec.
A timeouts key seems to have no effect Terraform silently discards an undeclared object key. Compare against object({ create, read, update, delete }).

πŸ”— Related Docs


πŸ’™ "Infrastructure as Code should be standardized, consistent, and secure."