Skip to content

Latest commit

Β 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

☁️ Azure Data Factory Flowlet Terraform Module

Manages one azurerm_data_factory_flowlet_data_flow β€” the reusable fragment of mapping-data-flow logic that other data flows embed. Targets hashicorp/azurerm ~> 4.0.

Terraform azurerm Module Type Resources Blast radius

🧩 Overview

  • 🎯 Creates one flowlet inside an existing Data Factory.
  • 🧩 A flowlet is a fragment, not a flow: it runs as part of whichever mapping data flow embeds it.
  • πŸ“₯ flow_source and sink are optional here β€” a flowlet may take both from the embedding flow, and usually does.
  • πŸ“ Takes the Data Flow Script as either a single script string or an ordered script_lines list; at least one is required.
  • πŸ”— May embed other flowlets by name β€” and nothing prevents a cycle.
  • 🚫 Carries no tags β€” the resource exposes none.

πŸ’‘ Why it matters: a flowlet exists to be reused, so its blast radius is larger than a data flow's, not smaller. Several flows may embed it, each naming it as a plain string, and none of those references is a Terraform dependency. Renaming or destroying this resource applies cleanly and leaves every embedding flow failing at run time with no plan-time signal anywhere. That reuse is also why the provider makes source and sink optional here and required on the mapping data flow β€” the one schema difference between two otherwise near-identical resources.

❀️ Support this project

If this module saves you time:

πŸ—ΊοΈ Where this fits in the family

flowchart TB
  RG["azurerm_resource_group"]
  ADF["azurerm_data_factory"]
  DS["azurerm_data_factory_dataset_*"]
  LS["azurerm_data_factory_linked_service_*"]
  PL["azurerm_data_factory_pipeline"]
  DF["data_flow"]
  FL["flowlet_data_flow"]

  RG -->|"resource group"| ADF
  ADF -->|"data_factory_id"| DF
  ADF -->|"data_factory_id"| FL
  ADF -->|"data_factory_id"| DS
  ADF -->|"data_factory_id"| LS
  ADF -->|"data_factory_id"| PL
  DS -.->|"dataset by NAME"| DF
  LS -.->|"linked_service by NAME"| DF
  FL -.->|"flowlet by NAME, no dependency edge"| DF
  FL -.->|"a flowlet may embed another, CYCLE possible"| FL
  DF -.->|"invoked by NAME from Execute Data Flow"| PL

  classDef me fill:#0078D4,stroke:#004578,color:#ffffff
  classDef key fill:#004578,stroke:#002b47,color:#ffffff
  classDef ext fill:#F0F3F6,stroke:#9AA5B1,color:#1F2933
  class DF,FL me
  class ADF key
  class RG,DS,LS,PL ext
Loading

Every dashed edge is a name, not a reference β€” including the self-edge, because a flowlet may embed another flowlet and nothing prevents a cycle.

🧬 What this module builds

flowchart TB
  VN["name, data_factory_id -- the ONLY two force-new fields"]
  VS["script OR script_lines -- AtLeastOneOf, both legal together"]
  VSRC["flow_source -- OPTIONAL here, REQUIRED on the mapping data flow"]
  VSNK["sink -- OPTIONAL here"]
  VT["transformation -- OPTIONAL, and only THREE reference blocks"]
  R["azurerm_data_factory_flowlet_data_flow.this"]
  EMB["Embedded BY NAME by every consuming data flow"]
  OID["id, name, script_form"]
  OFLAG["defines_no_source_or_sink, referenced_flowlet_names"]
  OWARN["destroying_this_flowlet_breaks_every_data_flow_that_embeds_it"]

  VN --> R
  VS --> R
  VSRC --> R
  VSNK --> R
  VT --> R
  R -->|"no dependency edge in either direction"| EMB
  R --> OID
  R --> OFLAG
  R --> OWARN

  classDef me fill:#0078D4,stroke:#004578,color:#ffffff
  classDef key fill:#004578,stroke:#002b47,color:#ffffff
  classDef ext fill:#F0F3F6,stroke:#9AA5B1,color:#1F2933
  class R key
  class EMB me
  class VN,VS,VSRC,VSNK,VT,OID,OFLAG,OWARN ext
Loading
Resource Count Role
azurerm_data_factory_flowlet_data_flow.this 1 The keystone. No child resources β€” source, sink and transformation are rendered inline.

βœ… Provider / Versions

Item Value
Terraform >= 1.12.0
hashicorp/azurerm ~> 4.0
Provider block None in this module β€” the caller configures the provider, its authentication, and the mandatory features {} block
Module type standalone

Schema notes that bite β€” each verified against the provider source at the pinned line:

  • πŸ”΄ source and sink are OPTIONAL here and REQUIRED on the mapping data flow. The provider uses SchemaForDataFlowletSourceAndSink here and SchemaForDataFlowSourceAndSink there β€” the same Elem, a different cardinality. That is the entire schema difference between the two resources.
  • A flowlet with neither is legal and usual β€” it takes both from the embedding flow.
  • Everything is referenced by name, in both directions, including this flowlet from the flows that embed it.
  • Nothing prevents a cycle between two flowlets that embed each other.
  • Nothing validates the script beyond StringIsNotEmpty.
  • script and script_lines carry a mutual AtLeastOneOf; supplying both is legal and unresolved.
  • script_lines is ordered and the order is the program.
  • πŸ”΄ source/sink and transformation are NOT the same shape β€” five reference blocks versus three.
  • πŸ”΄ An undeclared key on transformation is silently discarded β€” no error at validate, plan or apply.
  • Only two force-new fields: name, data_factory_id.
  • source is a RESERVED Terraform variable name, so this module's input is flow_source.
  • No CustomizeDiff, no version gate, no tags.

πŸ”‘ Required Azure RBAC Roles / Permissions

Role Scope Why
Data Factory Contributor the Data Factory Create, read, update and delete flowlet definitions.

⚠️ This module grants nothing and runs nothing. A flowlet is a definition; it executes with the Data Factory's managed identity as part of an embedding flow.

πŸ”’ Write access to a shared flowlet is write access to every flow that embeds it. That is the difference from a data flow, whose blast radius is one flow. A flowlet exists to be reused, so a change here propagates to consumers that nothing in Terraform records.

Azure Prerequisites

  • Microsoft.DataFactory registered in the subscription.
  • An existing Data Factory.
  • The datasets, linked services and other flowlets named anywhere in this flowlet must already exist. Nothing checks this at any Terraform stage.
  • At least one mapping data flow that embeds this flowlet, if it is ever to run. A flowlet nothing embeds is inert and produces no error of any kind.

πŸ“ Module Structure

terraform-azurerm-data-factory-flowlet-data-flow/
β”œβ”€β”€ providers.tf    # required_version + the pinned azurerm; no provider block
β”œβ”€β”€ variables.tf    # 11 typed inputs, 14 validations
β”œβ”€β”€ main.tf         # locals + the single keystone resource
β”œβ”€β”€ outputs.tf      # 47 outputs; id first, then name, then the derived facts
β”œβ”€β”€ README.md       # this file
β”œβ”€β”€ SCOPE.md        # the cross-module contract
β”œβ”€β”€ LICENSE         # MIT
└── .gitignore

βš™οΈ Quick Start

provider "azurerm" {
  features {}
}

module "standardise_address" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-standardise-address"
  data_factory_id = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-data-eastus/providers/Microsoft.DataFactory/factories/adf-corp"

  script = file("${path.module}/flows/standardise_address.dfs")

  # No source and no sink: this fragment takes both from the flow that embeds it.
}

πŸ’‘ That really is the whole call. A flowlet with neither a source nor a sink is the common case, and the mapping data flow sibling cannot be created that way at all.

⚠️ The input for sources is flow_source, not source β€” source is a meta-argument of the module block itself.

πŸ”Œ Cross-Module Contract

Consumes

Input Type Source
data_factory_id Resource ID terraform-azurerm-data-factory β†’ id
…dataset.name dataset names terraform-azurerm-data-factory-dataset-* β†’ name
…linked_service.name linked service names terraform-azurerm-data-factory-linked-service-* β†’ name
…flowlet.name another flowlet's name another instance of this module β†’ name

Emits

Output Note
id The ARM Resource ID
name The name every embedding data flow refers to
defines_no_source_or_sink The common flowlet shape, at plan
referenced_flowlet_names Other flowlets this one embeds β€” where a cycle would show
referenced_dataset_names, referenced_linked_service_names For review by eye
script_form, supplies_both_script_forms Which form is in use
force_new_fields Just two

πŸ“š Example Library

1 Β· Minimal β€” a fragment with no source and no sink
module "standardise_address" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-standardise-address"
  data_factory_id = var.data_factory_id

  script = file("${path.module}/flows/standardise_address.dfs")
}

πŸ’‘ The common case. defines_no_source_or_sink reports true at plan. The mapping data flow sibling requires both and cannot be created this way.

2 Β· A flowlet with its own source
module "with_source" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-reference-lookup"
  data_factory_id = var.data_factory_id

  script = file("${path.module}/flows/reference_lookup.dfs")

  flow_source = {
    "countryCodes" = { dataset = { name = "ds-country-codes" } }
  }
}

πŸ’‘ Legal and useful: a lookup fragment can bring its own reference data while still taking the main input from the embedding flow.

3 Β· Transformations only
module "dedupe" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-dedupe"
  data_factory_id = var.data_factory_id

  script = file("${path.module}/flows/dedupe.dfs")

  transformation = {
    "aggregateByKey" = {}
    "pickLatest"     = {}
  }
}

ℹ️ A transformation entry with no reference block is normal β€” the step is defined in the script, and the entry exists so the name is declared.

4 Β· The script as an ordered list of lines
module "line_form" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-line-form"
  data_factory_id = var.data_factory_id

  script_lines = [
    "input1 derive(country = upper(country)) ~> upperCountry",
    "upperCountry output() ~> output1",
  ]
}

⚠️ The order of this list is the program. Reordering it changes what the fragment does.

5 Β· Composing flowlets
module "address_pipeline" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-address-pipeline"
  data_factory_id = var.data_factory_id

  script = file("${path.module}/flows/address_pipeline.dfs")

  transformation = {
    "standardise" = {
      flowlet = { name = "fl-standardise-address" }
    }
    "validate" = {
      flowlet = { name = "fl-validate-address" }
    }
  }
}

πŸ”΄ Nothing prevents a cycle. These are names, not references β€” two flowlets that embed each other apply cleanly and fail only when a data flow using either is executed. referenced_flowlet_names emits what this one embeds so a reviewer can trace it; the module cannot see the other direction.

6 Β· A source resolved by linked service
module "inline_source" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-inline-source"
  data_factory_id = var.data_factory_id

  script = file("${path.module}/flows/inline.dfs")

  flow_source = {
    "rawFiles" = { linked_service = { name = "ls-datalake" } }
  }
}

πŸ’‘ dataset, linked_service and flowlet are alternative resolutions, so this module enforces at most one per block, mirroring the provider's MaxItems: 1 on each.

7 Β· Schema and rejected-row references
module "with_extra_refs" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-strict"
  data_factory_id = var.data_factory_id

  script = file("${path.module}/flows/strict.dfs")

  flow_source = {
    "incoming" = {
      dataset               = { name = "ds-incoming" }
      schema_linked_service = { name = "ls-schema-store" }
    }
  }

  sink = {
    "checked" = {
      dataset                 = { name = "ds-checked" }
      rejected_linked_service = { name = "ls-quarantine" }
    }
  }
}

πŸ’‘ These two are not alternatives to dataset β€” they answer a different question β€” so combining them is normal and this module does not treat it as ambiguous.

⚠️ Both exist on flow_source and sink only. transformation does not declare them, and because Terraform silently discards undeclared object keys, writing one there produces no error at all.

8 Β· Both script forms at once
module "both_forms" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-both-forms"
  data_factory_id = var.data_factory_id

  script       = file("${path.module}/flows/base.dfs")
  script_lines = ["input1 output() ~> output1"]
}

⚠️ Legal β€” AtLeastOneOf, not ExactlyOneOf β€” but which one the service honours when they disagree is not documented. Reported through supplies_both_script_forms rather than resolved.

9 Β· Metadata: description, folder and annotations
module "documented" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-standardise-address"
  data_factory_id = var.data_factory_id

  script = file("${path.module}/flows/standardise_address.dfs")

  description = "Uppercases country codes and trims postal codes."
  folder      = "shared/fragments"
  annotations = ["owner:data-platform", "shared:true"]
}

πŸ’‘ Filing shared flowlets in their own folder is worth doing precisely because nothing records who embeds them β€” the folder is the only organisational signal a reviewer gets.

⚠️ folder is a label, not a resource; a typo files the flowlet elsewhere silently.

10 Β· Many fragments from one map
locals {
  fragments = ["standardise-address", "validate-email", "dedupe"]
}

module "shared" {
  source   = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"
  for_each = toset(local.fragments)

  name            = "fl-${each.value}"
  data_factory_id = var.data_factory_id
  folder          = "shared/fragments"

  script = file("${path.module}/flows/${each.value}.dfs")
}

πŸ’‘ Each fragment gets its own script file, which matters more here than for a data flow β€” a shared fragment is reviewed by people who did not write it.

11 Β· Reviewing what the flowlet points at
output "flowlet_review" {
  value = {
    takes_from_parent = module.standardise_address.defines_no_source_or_sink
    sources           = module.standardise_address.source_names
    embeds_flowlets   = module.standardise_address.referenced_flowlet_names
    datasets          = module.standardise_address.referenced_dataset_names
    script_form       = module.standardise_address.script_form
    replace_on        = module.standardise_address.force_new_fields
  }
}

πŸ’‘ Every value is known at plan. referenced_flowlet_names is the one to read twice β€” it is where a composition cycle would first become visible, and the module can see only this half of it.

12 Β· What this module deliberately cannot tell you
# There is NO output for "which data flows embed this flowlet", because this
# module cannot know. Embedding flows name the flowlet as a plain string, so
# nothing in Terraform records the relationship in that direction.
#
# The honest substitute is a naming and annotation convention the caller owns:

module "standardise_address" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-standardise-address"
  data_factory_id = var.data_factory_id
  folder          = "shared/fragments"

  script = file("${path.module}/flows/standardise_address.dfs")

  annotations = [
    "shared:true",
    "embedded-by:df-clean-customers",
    "embedded-by:df-clean-suppliers",
  ]
}

⚠️ Those annotations are documentation, not a mechanism β€” nothing keeps them current and nothing validates them. They are suggested because the alternative is no record at all, not because they are enforced.

πŸ”’ This is the module declining to invent a protective claim: a "consumers" output would have to be fabricated, and a reader would reasonably trust it.

13 Β· πŸ—οΈ End-to-end composition
provider "azurerm" {
  features {}
}

module "resource_group" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-resource-group.git?ref=v1.0.0"

  name     = "rg-data-eastus"
  location = "eastus"
}

module "data_factory" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory.git?ref=v1.0.0"

  name                = "adf-corp"
  resource_group_name = module.resource_group.name
  location            = module.resource_group.location

  identity = {
    type = "SystemAssigned"
  }
}

module "standardise_address" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-flowlet-data-flow.git?ref=v1.0.0"

  name            = "fl-standardise-address"
  data_factory_id = module.data_factory.id
  folder          = "shared/fragments"

  script      = file("${path.module}/flows/standardise_address.dfs")
  description = "Uppercases country codes and trims postal codes."
  annotations = ["shared:true"]
}

module "clean_customers" {
  source = "git::https://github.com/microsoftexpert/terraform-azurerm-data-factory-data-flow.git?ref=v1.0.0"

  name            = "df-clean-customers"
  data_factory_id = module.data_factory.id

  script = file("${path.module}/flows/clean_customers.dfs")

  flow_source = { "customersRaw" = { dataset = { name = "ds-customers-raw" } } }

  transformation = {
    "standardise" = {
      flowlet = { name = module.standardise_address.name }
    }
  }

  sink = { "customersClean" = { dataset = { name = "ds-customers-clean" } } }
}

output "flowlet_name" {
  value = module.standardise_address.name
}

⚠️ module.standardise_address.name creates a Terraform ordering edge but NOT a Data Factory one. Terraform creates the flowlet first because the data flow reads its name β€” but Data Factory resolves by name at run time, so renaming or replacing this flowlet later leaves the flow pointing at whatever now holds that name, with no plan-time signal.

πŸ”’ A second flow embedding the same flowlet gets no edge at all unless it also reads the output. Prefer reading module.<flowlet>.name in every consumer rather than hardcoding the string β€” it is the only ordering Terraform can give you here.

πŸ“₯ Inputs

Required: name, data_factory_id, and at least one of script / script_lines.

Optional: flow_source, sink, transformation, description, folder, annotations, timeouts.

Full input schemas
Name Type Default Notes
name string β€” Force-new. Every embedding flow refers to this name.
data_factory_id string β€” Force-new. Anchored to Microsoft.DataFactory/factories/<name> with $.
script string null AtLeastOneOf with script_lines. Not validated beyond non-empty.
script_lines list(string) [] Ordered β€” the order is the program.
flow_source map(object({ description, dataset, linked_service, flowlet, schema_linked_service, rejected_linked_service })) {} Optional here, required on the mapping data flow. Named flow_source because source is reserved.
sink same shape {} Optional here.
transformation same shape minus schema_linked_service and rejected_linked_service {} A different shape β€” only three reference blocks.
description string null Non-empty if set.
folder string null A label, not a resource.
annotations list(string) [] Ordered. Not Azure tags.
timeouts object({create, read, update, delete}) null All four wired correctly.

Each dataset / linked_service / schema_linked_service / rejected_linked_service is { name, parameters }; flowlet adds dataset_parameters.

At most one of dataset, linked_service and flowlet may be set per block. schema_linked_service and rejected_linked_service are excluded from that rule.

🧾 Outputs

Output Type Notes
id string The ARM Resource ID
name string The name every embedding data flow refers to
data_factory_id / data_factory_name / resource_group_name / subscription_id string The owning factory
defines_no_source_or_sink bool The common flowlet shape, known at plan
source_names / sink_names / transformation_names list(string) The names the script must match
source_count / sink_count / transformation_count number Zero is legal and usual for the first two
referenced_flowlet_names list(string) Other flowlets this one embeds
referenced_dataset_names / referenced_linked_service_names list(string) For review by eye
sources_with_no_reference / sinks_with_no_reference list(string)
script_form / supplies_both_script_forms string / bool
force_new_fields / fields_that_can_change_after_creation list(string)
import_address string For terraform import
the CONSTANT outputs bool Facts that are consequential, invisible in state, and inferable from nothing else

Nothing is sensitive: this resource carries no credential. There is deliberately no "which flows embed this" output β€” see Example 12.

🧠 Architecture Notes

A flowlet's blast radius is larger than a data flow's. It exists to be reused, and every consumer names it as a plain string. Nothing in Terraform records that relationship, so destroying this resource β€” or renaming it, which is a destroy and a create β€” applies cleanly and leaves every embedding flow failing at run time. The module states this in destroying_this_flowlet_breaks_every_data_flow_that_embeds_it rather than leaving a reader to infer it from the data flow module's gentler equivalent.

The one schema difference is cardinality. The provider gives this resource SchemaForDataFlowletSourceAndSink (Optional) and the mapping data flow SchemaForDataFlowSourceAndSink (Required) β€” the same Elem, a different cardinality. The two provider resource files are otherwise near-identical, right down to sharing the transformation helper, which is exactly why this had to be read rather than assumed. A flowlet with neither source nor sink is the common shape; defines_no_source_or_sink reports it.

Composition is permitted and cycles are not prevented. A flowlet may embed another flowlet by name. Two that embed each other apply cleanly and fail only when a flow using either runs. This module emits referenced_flowlet_names so the outbound half is visible; it cannot see the inbound half, and does not pretend to β€” claiming cycle detection would be a protective claim that does not exist.

The three blocks look alike and are not the same shape. source and sink carry five reference blocks; transformation carries three. And because Terraform's object-type conversion silently discards undeclared keys, putting schema_linked_service on a transformation produces no error at validate, plan or apply β€” captured by running the case, not inferred.

Nothing validates the script, here or on the sibling. The failure surfaces when an embedding flow executes, which on a shared fragment may be in a pipeline nobody associated with this change.

source is reserved. It is a module block meta-argument, so the input is flow_source; the rendered resource still emits a source block exactly as the provider defines it.

features {} dependence. As with every module in this suite, the caller supplies provider "azurerm" { features {} }.

🧱 Design Principles

Concern Default in the empty call Opt-out
A fragment with no logic Impossible β€” AtLeastOneOf across script and script_lines, mirrored here none; the provider forbids it
A fragment with no source or sink Legal and usual β€” reported, not refused, because the provider permits it here declare either
Ambiguous source resolution At most one of dataset / linked_service / flowlet per block none
Composition cycles Documented, not detected β€” the module can see only its own half none possible
Existence of anything named Not checkable β€” emitted as lists for human review instead none
Who embeds this flowlet Deliberately not emitted β€” it cannot be known, and a fabricated output would be trusted use annotations as documentation

Two rules govern the validations. Never invent a constraint that could reject legal input β€” a validation {} failure blocks terraform destroy as well as apply. And enforce a configuration rule, report a service or runtime one β€” which is why this module refuses exactly where the provider refuses, and nowhere else. Notably it does not require a source or a sink, because the provider does not; the mapping data flow module does, because there it does.

πŸš€ Runbook

terraform init -backend=false
terraform validate
terraform fmt -check

Pin the module with ?ref=v1.0.0 β€” never a branch. This module is plan-only; a human applies from CI.

πŸ§ͺ Testing

Know which command reaches which check β€” the three stages are not interchangeable. terraform validate on a configuration that CALLS this module checks types and syntax only: it evaluates none of the module's variable values. The 14 validation {} blocks and the provider's AtLeastOneOf are reached at terraform plan β€” and neither needs credentials, because variable validation runs before the provider is configured. To exercise them without a plan at all, run the module as the root module and drive it through terraform console -var-file=....

Because this resource has no CustomizeDiff and no version gate, nothing is deferred to plan by the provider.

What no Terraform stage exercises at all: whether the script is valid, whether any referenced dataset, linked service or flowlet exists, whether a composition cycle exists, and whether anything embeds this flowlet at all. A flowlet nothing embeds applies cleanly, costs nothing and does nothing, and no error is produced anywhere.

πŸ’¬ Example Output

$ terraform output
id                              = "/subscriptions/.../factories/adf-corp/dataflows/fl-standardise-address"
name                            = "fl-standardise-address"
data_factory_name               = "adf-corp"
defines_no_source_or_sink       = true
source_names                    = []
sink_names                      = []
transformation_names            = ["standardise"]
referenced_flowlet_names        = []
referenced_dataset_names        = []
script_form                     = "script"
folder                          = "shared/fragments"
force_new_fields                = ["name", "data_factory_id"]

πŸ” Troubleshooting

Symptom Cause Fix
Invalid expression value: string required, but have object. at terraform init, pointing at the source line source is a module block meta-argument holding the module's ADDRESS, so a map written there is type-checked against that, not against any variable. Terraform reports Error: Incorrect value type. Use flow_source for the data flow's sources; leave source as the module address.
at least one of script and script_lines must be set -- a data flow with neither has no transformation logic at all. Neither script form was supplied. Supply one, or both.
The flowlet applies but nothing ever runs it Nothing embeds it. A flowlet is inert on its own and produces no error of any kind. Add a flowlet block naming it in a mapping data flow's flow_source, sink or transformation.
Renaming the flowlet broke several data flows at once Embedding flows name it as a plain string; that is not a Terraform dependency. Update every embedding flow. Prefer reading module.<flowlet>.name in consumers so at least Terraform orders them.
Two flowlets embed each other and the flow fails at run time Nothing prevents a cycle β€” these are names, not references. Check referenced_flowlet_names on both. The module can only see one direction.
a source may reference at most ONE of dataset, linked_service or flowlet. Two alternative resolutions on one block. Keep one. schema_linked_service and rejected_linked_service are exempt.
schema_linked_service on a transformation had no effect and no error transformation does not declare it, and Terraform silently discards undeclared object keys. Move it to the flow_source or sink entry. Nothing will ever report this.
The flowlet moved folders unexpectedly folder is a free-text label with no folder object behind it. Check folder.

πŸ”— Related Docs

πŸ’™ "Infrastructure as Code should be standardized, consistent, and secure."