A typed Data Factory dataset describing Parquet files in Blob Storage, Data Lake Gen2 or on an HTTP server, targeting
hashicorp/azurerm ~> 4.0.
- ποΈ Creates one
azurerm_data_factory_dataset_parquetβ a named description of Parquet files that pipelines read from and write to. - π Accepts exactly one of three location blocks β HTTP server, Blob Storage, or Data Lake Gen2 β
enforced by the provider at
terraform plan, offline and without credentials, even though this resource's documentation never says so. - π« Accepts eight compression codecs.
"None"is not one of them β and the delimited-text dataset accepts it. To configure no compression here, omit the argument. - π Makes
http_server_location.pathoptional, where the delimited-text and JSON datasets both require it. One block name, three contracts across the family. - ποΈ Carries up to three dynamic flags per location and reports six distinct mismatches plus a roll-up.
π‘ Why it matters: this resource shares all three location blocks, both compression arguments and the whole
schema_columnblock with the delimited-text dataset β and still refuses a configuration copied from it, in two independent ways. Two resources can share almost everything and still disagree.
If this module saved you time:
- β Star the repository β it is the cheapest signal that this work is worth continuing.
- πΌ Connect on LinkedIn: linkedin.com/in/microsoftexpert
- β Buy me a coffee: buymeacoffee.com/microsoftexpert
flowchart TB
RG["terraform-azurerm-resource-group"]
SA["terraform-azurerm-storage-account"]
ADF["terraform-azurerm-data-factory"]
LS["a Data Factory linked service"]
THIS["terraform-azurerm-data-factory-dataset-parquet"]
DTEXT["terraform-azurerm-data-factory-dataset-delimited-text"]
JSONDS["terraform-azurerm-data-factory-dataset-json"]
SQLT["terraform-azurerm-data-factory-dataset-azure-sql-table"]
PIPE["a Data Factory pipeline"]
FILES["a container, a Data Lake file system, or an HTTP server"]
RG -->|"name"| ADF
RG -->|"name"| SA
SA -->|"named inside the location block, never read here"| THIS
ADF -->|"id"| THIS
ADF -->|"id"| DTEXT
ADF -->|"id"| JSONDS
ADF -->|"id"| SQLT
LS -->|"linked_service_name, a bare NAME"| THIS
LS -->|"linked_service_id, a full Resource ID"| SQLT
THIS -->|"codec set of EIGHT, and None is not one of them"| PIPE
DTEXT -->|"codec set of NINE, including None"| PIPE
JSONDS -->|"no compression arguments at all"| PIPE
PIPE -->|"reads only at run time, never here"| FILES
classDef this fill:#0078D4,stroke:#004578,color:#ffffff,stroke-width:2px
classDef keystone fill:#004578,stroke:#00243d,color:#ffffff,stroke-width:2px
classDef sibling fill:#eef3f8,stroke:#b9c8d8,color:#1b2733
class THIS this
class ADF keystone
class RG,SA,LS,DTEXT,JSONDS,SQLT,PIPE,FILES sibling
The edge labels carry the difference most likely to survive a copy-paste: this resource's codec set has
eight members and the delimited-text dataset's has nine, the extra one being exactly "None".
flowchart TB
subgraph INPUTS["Inputs"]
NAME["name (force-new)"]
ADFID["data_factory_id (force-new)"]
LSNAME["linked_service_name, a bare NAME"]
LOC["EXACTLY ONE of three location blocks"]
COMP["compression_codec, eight values, and compression_level"]
COLS["schema_column list"]
META["parameters, additional_properties, annotations, description, folder"]
end
THIS["azurerm_data_factory_dataset_parquet.this"]
subgraph OUTPUTS["Outputs"]
OID["id and name"]
OKIND["location_kind, normalised path, filename, container_or_file_system"]
OMIS["dynamic_flags_look_inconsistent, plus its six components"]
ONONE["location_names_nothing and describes_no_specific_file"]
OCOMP["the codec set EXCLUDES None, and compression_level_without_codec"]
OFACTS["the constant facts, including the documentation gap"]
end
NAME --> THIS
ADFID --> THIS
LSNAME --> THIS
LOC --> THIS
COMP --> THIS
COLS --> THIS
META --> THIS
THIS --> OID
LOC --> OKIND
LOC --> OMIS
LOC --> ONONE
META --> ONONE
COMP --> OCOMP
THIS --> OFACTS
classDef this fill:#0078D4,stroke:#004578,color:#ffffff,stroke-width:2px
classDef sibling fill:#eef3f8,stroke:#b9c8d8,color:#1b2733
class THIS this
class NAME,ADFID,LSNAME,LOC,COMP,COLS,META,OID,OKIND,OMIS,ONONE,OCOMP,OFACTS sibling
| Resource | Count | Notes |
|---|---|---|
azurerm_data_factory_dataset_parquet.this |
1 | The keystone. A real ARM child of the factory. |
http_server_location |
dynamic, 0..1 |
relative_url + filename required; path optional. |
azure_blob_storage_location |
dynamic, 0..1 |
Only container required. |
azure_blob_fs_location |
dynamic, 0..1 |
Nothing required. |
schema_column |
dynamic, 0..n |
Optional column definitions. |
timeouts |
dynamic, 0..1 |
All four keys exist. |
| Requirement | Value |
|---|---|
| Terraform | >= 1.12.0 |
hashicorp/azurerm |
~> 4.0 |
| Provider block | None. The caller configures the provider, including the mandatory features {} block. |
Schema notes that bite β verified against the live provider source, not inferred from the schema:
- π΄
compression_codecaccepts EIGHT values here and NINE on the delimited-text dataset. The missing one is exactly"None".compression_codec = "None"plans cleanly on a delimited-text dataset and is refused here; omit the argument instead. Everything else about the two resources' compression arguments is identical. - π΄
http_server_location.pathis OPTIONAL here and REQUIRED on the delimited-text and JSON datasets. One block name, three contracts. - π΄
ExactlyOneOfis declared across the three location blocks β and this resource's documentation never mentions it. The delimited-text page states "(exactly one of them must be set)"; this page just lists the blocks. The code is identical on both, so the rule holds and only the documentation is silent. Reading the page alone would suggest all three are optional. - π΄ The three location blocks do not share a shape.
http_server_locationrequiresrelative_urlandfilename;azure_blob_storage_locationrequires onlycontainer;azure_blob_fs_locationrequires nothing β soazure_blob_fs_location = {}satisfies the rule while naming no file. β οΈ Each block names its dynamic flags after the field they govern βdynamic_container_enabledon the blob block,dynamic_file_system_enabledon the Data Lake block.β οΈ compression_leveldoes not requirecompression_codec. A ratio with nothing to compress is accepted.β οΈ Onlynameanddata_factory_idforce replacement. The linked service and the entire location block update in place.β οΈ Timeouts default to 30m / 5m / 30m / 30m, all four keys present.- βΉοΈ The requires-import error names this resource correctly β unlike the JSON dataset, whose guard names the delimited-text type.
| Permission | Scope | Why |
|---|---|---|
Microsoft.DataFactory/factories/datasets/write |
the Data Factory | Create and update. |
Microsoft.DataFactory/factories/datasets/read |
the Data Factory | Refresh and plan. |
Microsoft.DataFactory/factories/datasets/delete |
the Data Factory | Destroy β see the warning below. |
Data Factory Contributor |
the Data Factory | The built-in role containing all three. |
π No permission on the storage account, the Data Lake file system or the HTTP server is required, requested or used. The dataset names a file and its compression; it never opens one. Access belongs to the linked service, so whoever can write datasets can point one at any container the factory's identity can already reach β including data they have no direct permission on.
- An existing Data Factory, and its Resource ID.
- A linked service already configured in that factory, and its name. Nothing here verifies it.
- A container, Data Lake file system or HTTP endpoint for the location block to name β or
parameters. - A dataset name meeting Microsoft's Data Factory naming rules. The provider checks only non-emptiness.
- The
Microsoft.DataFactoryresource provider registered in the subscription.
terraform-azurerm-data-factory-dataset-parquet/
βββ providers.tf # required_version + the pinned azurerm; no provider block
βββ variables.tf # 15 typed inputs, 19 validations
βββ main.tf # the keystone, the location normalisation, five dynamic blocks
βββ outputs.tf # 73 outputs; id first
βββ README.md # this file
βββ SCOPE.md # the cross-module contract
βββ LICENSE # MIT
βββ .gitignore
provider "azurerm" {
features {}
}
module "orders_parquet" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_orders_parquet"
data_factory_id = var.data_factory_id
linked_service_name = var.adls_linked_service_name
azure_blob_fs_location = {
file_system = "curated"
path = "orders"
filename = "orders.parquet"
}
compression_codec = "snappy"
}βΉοΈ Exactly one location block is required β a rule the provider enforces and this resource's documentation omits.
Consumes
| Input | Type | Typical source |
|---|---|---|
data_factory_id |
string |
terraform-azurerm-data-factory β id |
linked_service_name |
string |
a linked service's name |
| one location block | object |
the caller |
Emits (selected β 73 in total)
| Output | Description |
|---|---|
id |
The dataset's Resource ID. |
location_kind |
Which of the three blocks was supplied. Never null. |
dynamic_flags_look_inconsistent |
The one output to assert on. |
the_compression_codec_set_EXCLUDES_None_unlike_the_delimited_text_dataset |
Constant. Eight here, nine there. |
the_http_location_block_is_looser_here_than_on_its_siblings |
Constant. path is optional only here. |
the_provider_documentation_omits_the_exactly_one_rule |
Constant. A doc gap. |
1 Β· The minimum call β Data Lake Gen2
module "orders" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_orders"
data_factory_id = var.data_factory_id
linked_service_name = var.adls_linked_service_name
azure_blob_fs_location = {
file_system = "curated"
path = "orders"
filename = "orders.parquet"
}
}π‘ Parquet on Data Lake Gen2 is the common shape. Every field on this block is optional.
2 Β· Omitting the location is refused β a rule the docs do not state
# β NOT ACCEPTED β rejected at terraform plan, offline and without credentials
module "broken" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_broken"
data_factory_id = var.data_factory_id
linked_service_name = var.adls_linked_service_name
# no location block at all
}
β οΈ The provider declaresExactlyOneOfacross the three blocks, so this fails offline. This resource's registry page never mentions the rule β only the delimited-text page does.the_provider_documentation_omits_the_exactly_one_rulerecords that.
3 Β· `"None"` is a legal codec on delimited-text and is refused here
# β NOT ACCEPTED β the Parquet codec set has eight members and None is not one of them
module "copied_codec" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_copied"
data_factory_id = var.data_factory_id
linked_service_name = var.blob_linked_service_name
azure_blob_storage_location = { container = "curated" }
compression_codec = "None" # legal on the delimited-text dataset, REFUSED here
}
β οΈ To configure no compression on this resource, omit the argument. This is the difference most likely to survive a copy-paste, because everything else about the two resources' compression arguments is identical.
4 Β· An HTTP location without a path β legal only here
module "vendor_feed" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_vendor_feed"
data_factory_id = var.data_factory_id
linked_service_name = var.http_linked_service_name
http_server_location = {
relative_url = "/exports/daily"
filename = "orders.parquet"
# path omitted -- OPTIONAL on this resource
}
}π The delimited-text and JSON datasets both require
pathon their block of the same name. A configuration valid here fails on both of them.
5 Β· A Blob Storage location, container only
module "curated" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_curated"
data_factory_id = var.data_factory_id
linked_service_name = var.blob_linked_service_name
azure_blob_storage_location = {
container = "curated"
}
}βΉοΈ
containeris the only required field β the same rule as delimited-text, and not the same as the JSON dataset, which requirescontainer,pathandfilename.
6 Β· The empty location that names nothing
module "resolves_to_nothing" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_empty"
data_factory_id = var.data_factory_id
linked_service_name = var.adls_linked_service_name
azure_blob_fs_location = {} # satisfies ExactlyOneOf, names no file
}
check "dataset_points_somewhere" {
assert {
condition = !module.resolves_to_nothing.describes_no_specific_file
error_message = "The dataset names no file and exposes no parameters; no pipeline can resolve it."
}
}
β οΈ This applies cleanly and reports nothing.location_names_nothingistrue.
7 Β· All eight codecs, at their exact casing
locals {
# The provider's own list mixes capitalised and lowercase names. Copy it; do not guess.
codecs = ["bzip2", "gzip", "deflate", "ZipDeflate", "TarGzip", "Tar", "snappy", "lz4"]
}
module "by_codec" {
for_each = toset(local.codecs)
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_${lower(replace(each.key, "-", "_"))}"
data_factory_id = var.data_factory_id
linked_service_name = var.blob_linked_service_name
azure_blob_storage_location = { container = "curated" }
compression_codec = each.key
}π Compared case-sensitively:
"Gzip"is rejected where"gzip"is accepted.
8 Β· Compression level, and the pairing nothing enforces
module "compressed" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_compressed"
data_factory_id = var.data_factory_id
linked_service_name = var.blob_linked_service_name
azure_blob_storage_location = { container = "curated", filename = "orders.parquet" }
compression_codec = "snappy"
compression_level = "Optimal" # capitalised -- "optimal" is REJECTED
}
check "compression_is_coherent" {
assert {
condition = !module.compressed.compression_level_without_codec
error_message = "A compression level was set with no codec; it configures nothing."
}
}
β οΈ The provider accepts a level with no codec. The module reports that rather than refusing it.
9 Β· A dynamic path, flag and value agreeing
module "daily" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_daily"
data_factory_id = var.data_factory_id
linked_service_name = var.adls_linked_service_name
azure_blob_fs_location = {
file_system = "curated"
path = "@concat('orders/', formatDateTime(utcnow(), 'yyyy/MM/dd'))"
dynamic_path_enabled = true
filename = "orders.parquet"
}
}π
dynamic_flags_look_inconsistentisfalseβ the value is an expression and the flag is set.
10 Β· The mismatch this module exists to catch
module "silently_wrong" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_silently_wrong"
data_factory_id = var.data_factory_id
linked_service_name = var.adls_linked_service_name
azure_blob_fs_location = {
file_system = "@pipeline().parameters.fs"
# dynamic_file_system_enabled left at false
}
}
check "dynamic_flags_agree" {
assert {
condition = !module.silently_wrong.dynamic_flags_look_inconsistent
error_message = "A dynamic flag disagrees with the value it governs."
}
}
β οΈ Note the flag isdynamic_file_system_enabledon this block anddynamic_container_enabledon the blob block β the provider names it after the field it governs. The module unifies them into one output,dynamic_container_or_file_system_enabled.
11 Β· When the heuristic is wrong, and why it only reports
module "literal_at" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_literal_at"
data_factory_id = var.data_factory_id
linked_service_name = var.blob_linked_service_name
azure_blob_storage_location = {
container = "curated"
path = "@archive/2026" # a real folder that starts with @
}
}βΉοΈ
path_expression_without_dynamic_flagistrueand the configuration is correct. Refusing would reject legal input, and a failingvalidation {}would also blockterraform destroy.
12 Β· A typed column schema
module "typed" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_typed"
data_factory_id = var.data_factory_id
linked_service_name = var.adls_linked_service_name
azure_blob_fs_location = { file_system = "curated", filename = "orders.parquet" }
schema_column = [
{ name = "OrderId", type = "Int64", description = "Primary key." },
{ name = "Placed", type = "DateTimeOffset" },
{ name = "Notes" }, # untyped -- legal
]
}
check "every_column_is_typed" {
assert {
condition = length(module.typed.untyped_schema_column_names) == 0
error_message = "Untyped columns: ${join(", ", module.typed.untyped_schema_column_names)}"
}
}βΉοΈ A Parquet file carries its own schema in its footer, so omitting
schema_columnis especially normal here. The fifteen types are case-sensitive and identical across the family.
13 Β· Many datasets from one map
locals {
feeds = ["orders", "customers", "ledger"]
}
module "feeds" {
for_each = toset(local.feeds)
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_${each.key}"
data_factory_id = var.data_factory_id
linked_service_name = var.adls_linked_service_name
azure_blob_fs_location = {
file_system = "curated"
path = each.key
filename = "${each.key}.parquet"
}
compression_codec = "snappy"
folder = "silver/parquet"
}
check "no_feed_has_a_flag_mismatch" {
assert {
condition = alltrue([for m in module.feeds : !m.dynamic_flags_look_inconsistent])
error_message = "At least one feed's dynamic flags disagree with the values they govern."
}
}π‘ The provider takes no lock on the Data Factory for this resource, so these are created concurrently and safely.
14 Β· Re-pointing a dataset is an in-place update
module "curated" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_curated" # changing this REPLACES
data_factory_id = var.data_factory_id # changing this REPLACES
linked_service_name = var.other_linked_svc # changing this updates IN PLACE
azure_blob_fs_location = { # changing any of this updates IN PLACE
file_system = "archive"
path = "orders"
}
}
output "what_forces_replacement" {
value = module.curated.force_new_fields # ["name", "data_factory_id"]
}
β οΈ Everything describing what the dataset reads β including which storage account, by way of the linked service β changes in place. A review scanning plans for destroys will not see it.
15 Β· ποΈ End-to-end composition
provider "azurerm" {
features {}
}
module "rg" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"
name = "rg-analytics-eastus"
location = "eastus"
}
module "storage" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-storage-account.git?ref=v1.0.0"
name = "stanalyticseastus01"
resource_group_name = module.rg.name
location = module.rg.location
is_hns_enabled = true # Data Lake Gen2
containers = {
curated = {}
}
}
module "adf" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory.git?ref=v1.0.0"
name = "adf-analytics-eastus"
resource_group_name = module.rg.name
location = module.rg.location
identity = {
type = "SystemAssigned"
}
}
# No linked-service module exists in this suite yet, so the linked service is
# declared directly. It authenticates with the factory's managed identity, so
# no secret appears anywhere in this configuration.
resource "azurerm_data_factory_linked_service_azure_blob_storage" "curated" {
name = "ls_adls_curated"
data_factory_id = module.adf.id
service_endpoint = module.storage.primary_blob_endpoint
use_managed_identity = true
}
module "orders_parquet" {
source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-dataset-parquet.git?ref=v1.0.0"
name = "ds_orders_parquet"
data_factory_id = module.adf.id
linked_service_name = azurerm_data_factory_linked_service_azure_blob_storage.curated.name
azure_blob_fs_location = {
file_system = "curated"
path = "@concat('orders/', formatDateTime(utcnow(), 'yyyy/MM/dd'))"
dynamic_path_enabled = true
filename = "orders.parquet"
}
compression_codec = "snappy" # NOT "None" -- that value does not exist on this resource
compression_level = "Optimal"
folder = "silver/parquet"
schema_column = [
{ name = "OrderId", type = "Int64" },
{ name = "Placed", type = "DateTimeOffset" },
]
annotations = ["silver", "daily"]
}
check "wiring_is_consistent" {
assert {
condition = !module.orders_parquet.dynamic_flags_look_inconsistent
error_message = "A dynamic flag disagrees with the value it governs."
}
assert {
condition = !module.orders_parquet.describes_no_specific_file
error_message = "The dataset resolves to nothing."
}
assert {
condition = !module.orders_parquet.compression_level_without_codec
error_message = "A compression level was set with no codec."
}
}
output "dataset_id" {
value = module.orders_parquet.id
}
output "reads_from" {
value = {
kind = module.orders_parquet.location_kind
file_system = module.orders_parquet.container_or_file_system
path = module.orders_parquet.location_path
filename = module.orders_parquet.location_filename
}
}π The factory's system-assigned identity authenticates to storage. Grant it
Storage Blob Data Readeron the container out of band β no credential is stored in Terraform.
Required: name, data_factory_id, linked_service_name, and exactly one of
http_server_location / azure_blob_storage_location / azure_blob_fs_location.
Compression: compression_codec (eight values, no "None"), compression_level.
Shape: schema_column, parameters, additional_properties.
Metadata: description, folder, annotations.
Tail: timeouts. There is no tags variable β the resource supports none.
Full input schemas
| Name | Type | Default | Notes |
|---|---|---|---|
name |
string |
β | Force-new. Non-empty only; a leading / is refused. |
data_factory_id |
string |
β | Force-new. Anchored Resource-ID validator. |
linked_service_name |
string |
β | A bare name. A Resource ID is refused. |
http_server_location |
object({ relative_url, filename, path, dynamic_path_enabled, dynamic_filename_enabled }) |
null |
relative_url + filename required; path optional. Carries the exactly-one check. |
azure_blob_storage_location |
object({ container, path, filename, dynamic_container_enabled, dynamic_path_enabled, dynamic_filename_enabled }) |
null |
Only container required. |
azure_blob_fs_location |
object({ file_system, path, filename, dynamic_file_system_enabled, dynamic_path_enabled, dynamic_filename_enabled }) |
null |
Nothing required. |
compression_codec |
string |
null |
Closed, case-sensitive set of eight. "None" is absent. |
compression_level |
string |
null |
Optimal or Fastest, case-sensitive. |
schema_column |
list(object({ name, type, description })) |
[] |
Closed, case-sensitive set of fifteen types. |
parameters |
map(string) |
{} |
Supplied per pipeline run. |
additional_properties |
map(string) |
{} |
Unvalidated top-level keys. |
annotations |
list(string) |
[] |
Not tags. |
description |
string |
null |
Empty string refused; omission accepted. |
folder |
string |
null |
The authoring tree, not a storage path. |
timeouts |
object({ create, read, update, delete }) |
null |
30m / 5m / 30m / 30m. |
| Output | Description | Notes |
|---|---|---|
id |
The dataset's Resource ID. | First, by convention. |
name, data_factory_id, data_factory_name |
Identity and parent. | Name parsed from the end of the ID. |
resource_group_name, subscription_id |
Where the factory lives. | |
linked_service_name |
The linked service, as supplied. | A name β nothing can verify it. |
location_kind |
Which block was supplied. | Never null. |
supplied_location_count |
Always 1 on a config that plans. | Assert on the invariant. |
location_path, location_filename, container_or_file_system, relative_url |
Normalised across three blocks. | |
location_names_nothing |
The location addresses no file. | |
describes_no_specific_file |
Names nothing and no parameters. | Assert on this. |
dynamic_path_enabled, dynamic_filename_enabled, dynamic_container_or_file_system_enabled |
How the strings are read. | |
path_looks_like_an_expression, filename_looks_like_an_expression, container_looks_like_an_expression |
The @ heuristic. |
|
path_expression_without_dynamic_flag, filename_expression_without_dynamic_flag, container_expression_without_dynamic_flag |
Expression used literally. | |
path_marked_dynamic_but_looks_literal, filename_marked_dynamic_but_looks_literal, container_marked_dynamic_but_looks_literal |
The opposite mismatch. | Different remedy. |
dynamic_flags_look_inconsistent |
Any of the six. | Assert on this. |
compression_codec, compression_level |
The case-sensitive enums. | |
the_compression_codec_set_EXCLUDES_None_unlike_the_delimited_text_dataset |
Constant. | Eight here, nine there. |
compression_level_without_codec |
A ratio with nothing to compress. | Assert on this. |
the_compression_codec_set_is_closed_and_case_sensitive |
Constant. | Mixed casing. |
exactly_one_location_is_REQUIRED_on_this_resource |
Constant. | The JSON dataset differs. |
the_provider_documentation_omits_the_exactly_one_rule |
Constant. | A doc gap. |
the_three_location_blocks_do_not_share_a_shape |
Constant. | |
the_http_location_block_is_looser_here_than_on_its_siblings |
Constant. | path optional only here. |
schema_column_count, has_schema_columns, schema_column_names, untyped_schema_column_names |
The column schema. | |
parameters, parameter_count |
The run-time inputs. | |
uses_additional_properties, additional_property_count |
The escape hatch. | |
annotations, annotation_count, has_annotations |
A list, not tags. | |
description, has_description, folder, has_folder |
Metadata. | |
folder_and_the_location_path_mean_different_things |
Constant. | |
force_new_fields, fields_that_can_change_after_creation, fields_azure_returns_on_read |
The change surface. | |
this_resource_takes_a_linked_service_NAME_not_an_ID |
Constant. | A 4-to-1 split. |
the_dataset_family_is_not_a_clone_cluster |
Constant. | This resource is the evidence. |
the_linked_service_is_not_verified_to_exist |
Constant. | |
the_dynamic_flags_change_the_meaning_of_their_neighbours |
Constant. | The trap, named. |
the_schema_column_type_set_is_closed_and_case_sensitive |
Constant. | |
the_schema_column_block_is_identical_on_ten_of_the_twelve_datasets |
Constant. Identical on ten of twelve; binary has no block and snowflake's differs. |
|
this_dataset_moves_no_data_by_itself |
Constant. | |
the_module_cannot_see_which_pipelines_use_this |
Constant. | Read before destroying. |
the_module_cannot_verify_the_file_exists, no_credential_is_configured_here |
Constants. | |
destroying_the_factory_destroys_this_dataset_too, destroying_this_does_not_touch_the_files |
Constants. | |
lifecycle_prevent_destroy_is_not_available_to_a_module_caller |
Constant. | |
this_is_a_real_azure_resource_not_a_composite, the_provider_takes_no_lock_on_the_data_factory |
Constants. | |
this_resource_supports_no_azure_resource_tags |
Constant. | Why there is no tags variable. |
no_secret_is_accepted_or_emitted_by_this_module |
Constant. |
No output is sensitive, and none can be.
This resource is the family's best evidence against templating. It shares all three location blocks, both
compression arguments and the entire schema_column block with the delimited-text dataset, and it still
refuses a configuration copied from it in two independent ways. compression_codec accepts eight values here
and nine there, the extra one being exactly "None" β so a working delimited-text configuration fails at
terraform validate when the resource type is swapped. And http_server_location.path is optional here and
required on both the delimited-text and JSON datasets. Two resources can share almost everything and still
disagree.
The exactly-one rule holds, and the documentation does not say so. The provider declares ExactlyOneOf
across the three location blocks on this resource exactly as it does on delimited-text. Only the
delimited-text registry page states it. Reading this resource's page alone would suggest all three blocks are
optional, and the failure appears at plan rather than in review β which is the good direction, but only
if you know to expect it.
Three blocks, three shapes. http_server_location requires relative_url and filename;
azure_blob_storage_location requires only container; azure_blob_fs_location requires nothing at all.
That last one is why location_names_nothing exists: azure_blob_fs_location = {} satisfies the
exactly-one rule and addresses no file.
The check is placed, not gathered. A validation {} condition must reference its own variable, and two
variables validating each other is rejected as a cycle. The exactly-one check therefore lives on
http_server_location and reads the other two one-directionally, which is why its message names blocks the
caller may not have been editing.
Six mismatches, one roll-up. Each location names its dynamic flags after the field they govern, so the
three blocks do not share a flag set β dynamic_container_enabled on one, dynamic_file_system_enabled on
another. The module unifies them into dynamic_container_or_file_system_enabled, compares each flag against
whether its value begins with @, emits all six mismatches individually, and rolls them up. All of it
reports; none of it refuses.
Only identity forces replacement. The linked service and the entire location block update in place.
The module cannot see downstream. Pipelines reference a dataset by name and nothing points back.
lifecycle is not valid inside a module block, so a caller cannot add prevent_destroy; a CanNotDelete
lock prevents deletion but not the replacement that editing name would cause.
| Concern | This module's default | Opt-out |
|---|---|---|
| Secrets | None accepted, none emitted. | Not available β by design. |
| Credentials | Live on the linked service, never here. | β |
| Storage access | No permission required or used. | β |
| A missing or duplicated location | Refused at plan, by the provider and by this module. | None. |
| Compression codec | Enforced against the provider's eight, "None" excluded, with the exclusion named in the error message. |
None β the set is closed. |
| No compression | Omit the argument. There is no value that means it. | β |
| A compression level with no codec | Reported, not refused. | Ignore the output. |
| Dynamic flags | All default to false β literal, the reading that cannot surprise you. |
Set one to true. |
| A flag disagreeing with its value | Reported six ways plus a roll-up. Never refused. | Ignore the outputs. |
| A location naming nothing | Reported via location_names_nothing. |
Ignore the output. |
| Column types | Enforced against the provider's case-sensitive fifteen. | None β the set is closed. |
tags |
Not offered β the resource supports none. | Tag the factory. |
π There is no risky toggle to close here: the resource holds no credential and touches no data. The exposures are the delete and the silent mismatches, so the module documents the first and reports the second.
terraform init -backend=false
terraform validate
terraform fmt -checkPin the module with ?ref=v1.0.0 β never a branch. This library is plan-only: a human applies from CI.
terraform plan and the module's own validation {} blocks cover, offline and without credentials:
- the exactly-one-location rule, in both failing directions;
- each location block's own requiredness, which differs per block β including that an HTTP location without a path is accepted here;
- the anchored Resource-ID shape of
data_factory_id, and the refusal of an ID passed aslinked_service_name; - all eight codecs accepted at their exact casing, and
"None"refused β the case that distinguishes this resource from delimited-text; - the case-sensitive column-type set;
- every derived flag β all six dynamic-flag mismatches, their roll-up,
location_names_nothing,describes_no_specific_fileandcompression_level_without_codecβ each with one fixture per branch throughterraform console, which does fire root-module variable validations.
The harness also drives the same fixture against the delimited-text and JSON modules and asserts they disagree where the provider source says they must β so the contrast this README draws is executed, not merely asserted.
terraform validate reaches none of that through a module call. Validate evaluates no module
variable values, so neither this module's validation {} blocks nor the provider's ExactlyOneOf is
reached and it reports success. The refusals above land at plan β still offline and without credentials.
Only a real apply can tell you whether the linked service exists, whether the file exists, or
whether a pipeline still depends on the dataset.
dataset_id = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-analytics-eastus/providers/Microsoft.DataFactory/factories/adf-analytics-eastus/datasets/ds_orders_parquet"
reads_from = {
"file_system" = "curated"
"filename" = "orders.parquet"
"kind" = "azure_blob_fs_location"
"path" = "@concat('orders/', formatDateTime(utcnow(), 'yyyy/MM/dd'))"
}
supplied_location_count = 1
dynamic_path_enabled = true
path_looks_like_an_expression = true
dynamic_flags_look_inconsistent = false
location_names_nothing = false
describes_no_specific_file = false
compression_codec = "snappy"
compression_level_without_codec = false
force_new_fields = ["name", "data_factory_id"]
| Symptom | Cause | Fix |
|---|---|---|
compression_codec must be one of bzip2, gzip, deflate, ... on "None" |
"None" is not in this resource's set, though the delimited-text dataset accepts it. |
Omit the argument to configure no compression. |
compression_codec must be one of bzip2, gzip, ... on "Gzip" |
The set is case-sensitive and mixes casings. | Use "gzip". Copy the list; do not guess. |
EXACTLY ONE of http_server_location, azure_blob_storage_location or azure_blob_fs_location must be set. |
No location, or more than one. | Supply exactly one. This resource's documentation omits the rule; the provider enforces it. |
http_server_location requires a non-empty relative_url AND filename. |
One of the two required fields was blank. | path is optional here; the other two are not. |
| A config valid here fails on a delimited-text or JSON dataset | Those require http_server_location.path. |
Add path. |
azure_blob_storage_location.container must not be empty |
The one required field on that block. | Supply a container. |
| A dataset applies but no pipeline can use it | azure_blob_fs_location = {} with no parameters. |
Check describes_no_specific_file. |
A pipeline looks for a folder literally named @concat(...) |
An expression with its dynamic flag unset. | Set the flag. path_expression_without_dynamic_flag reports it. |
| A pipeline fails evaluating a plain folder name | A literal marked dynamic. | Unset the flag. path_marked_dynamic_but_looks_literal reports it. |
dynamic_flags_look_inconsistent is true on a correct config |
The @ heuristic β a real folder may begin with @. |
Read the six component outputs and ignore the roll-up. |
Setting dynamic_container_enabled on the Data Lake block does nothing |
That block's flag is dynamic_file_system_enabled. |
Use the flag named for the field. |
| Compression seems not to happen | A level was set with no codec. | Check compression_level_without_codec. |
A timeouts key seems to have no effect |
Terraform silently discards an undeclared object key. | Compare against object({ create, read, update, delete }). |
azurerm_data_factory_dataset_parquetazurerm_data_factory_dataset_delimited_textβ accepts"None"as a codecazurerm_data_factory_dataset_jsonβ permits no location- Azure Data Factory datasets
- Parquet format in Data Factory
SCOPE.mdβ this module's cross-module contract- Sibling modules:
terraform-azurerm-data-factory,terraform-azurerm-data-factory-dataset-delimited-text,terraform-azurerm-data-factory-dataset-json
π "Infrastructure as Code should be standardized, consistent, and secure."