martinpronk.com
Azure Firewall Policy Inheritance: Building a Parent/Child Model From Scratch, and Migrating an Existing Firewall Into It

Azure Firewall Policy Inheritance: Building a Parent/Child Model From Scratch, and Migrating an Existing Firewall Into It

How Azure Firewall Policy's parent/child inheritance lets an MSP keep a non-negotiable security baseline in code while a customer safely self-serves their own rules, built from scratch as a Terraform POC, and mapped out for migrating an existing customer firewall into the model.

· 10 min read

If you run firewalls for multiple customers, you’ll recognize this tension: you want a central, non-negotiable security baseline (deny known-bad traffic, enforce Threat Intelligence, IDPS), but you also don’t want to be the bottleneck every time a customer needs to open a port for their own application. Give customers direct access to the firewall and you lose the baseline. Keep it fully locked down and you become a ticket queue.

Azure Firewall Policy has a built-in answer to this: policy inheritance. A parent (“base”) policy holds the baseline. One or more child policies inherit from it and can add their own rules on top, evaluated after the parent’s, never able to override it. Scope a customer’s access to just their child policy, and they can self-serve without ever touching the baseline.

This post covers two things: building the model from scratch as a Terraform proof-of-concept, and what changes when you’re retrofitting it onto a firewall that’s already live at a customer.

How the hierarchy actually behaves

A few things about this feature aren’t obvious from the name alone, and matter a lot once you start relying on it.

Evaluation order is fixed, and priority numbers don’t cross policy boundaries. Azure Firewall always processes NAT rules, then network rules, then application rules, each is terminating, so a match stops further evaluation of the remaining types entirely. Within network and application rules specifically, the parent’s rule collections are always evaluated before the child’s of the same type, regardless of what priority number either side uses. A child rule with priority 100 still runs after every parent network rule. Priority only orders things within one policy, not between parent and child.

NAT rules are the one exception, they don’t inherit at all. Straight from Microsoft’s own docs: “NAT rule collections are not inherited, as they are specific to individual firewalls. If you want to use NAT rules, you must define them in the child policy.” The reasoning makes sense once you think about it: a policy is designed to be reused across many firewalls, but a DNAT rule translates traffic hitting one specific firewall’s public IP. If NAT rules inherited, a parent-defined DNAT rule would get blindly applied to every child firewall, most of which have a completely different public IP. So: NAT rules always go in the child, never the parent.

Threat Intelligence mode inherits, but only in one direction. A child can make it stricter than the parent (parent set to “Alert only” → child can set “Alert and deny”), but never looser or off. Same for the TI allowlist, the child can only add entries, not remove ones the parent set.

This has a sharp edge once the parent is already at the strictest setting. Set the parent to “Deny” and leave the child’s mode unconfigured, and the child falls back to its own default (“Alert”), which is looser than the parent’s. Azure rejects that outright with FirewallPolicyChildPolicyThreatIntelModeError. If your baseline is already maximally strict, every child has to say so explicitly, there’s no “inherit implicitly” option that satisfies the constraint, you have to set it to “Deny” yourself even though that’s already what you’d get by doing nothing in a world where the parent were less strict.

One firewall, one policy, but one policy, many firewalls. This is the detail with the biggest architectural consequence. A firewall can only ever have one policy attached. But the same policy can be attached to many firewalls at once. For a parent policy, that’s exactly what you want, one shared baseline, reused across every customer’s firewall, changed once and it takes effect everywhere. For a child policy it means the opposite: since each firewall gets exactly one policy, you can’t put two different customers behind a single shared firewall using separate child policies. Each customer needs their own firewall instance. That’s a direct cost input worth knowing before you promise “many customers, one policy model” to anyone holding the budget.

One caveat that only shows up once you’re running this across actual separate customers rather than one organization’s internal teams: parent and child have to live in the same Microsoft Entra ID tenant. A policy in your own management tenant can’t serve as the base_policy_id for a child sitting in a customer’s tenant, they’re simply not addressable across that boundary. So “one shared baseline” doesn’t mean one shared Azure resource spanning every customer, it means one shared Terraform definition, deployed as its own resource into each customer’s own tenant. Same content everywhere, separate objects. The customer still never touches their copy of the parent, and a change to the definition still rolls out to everyone on the next apply, it just does so as N deployments instead of one.

Scenario 1: building it from scratch

For a proof-of-concept, you don’t need Virtual WAN or a Secured Hub, a standalone Azure Firewall in a plain VNet is enough to prove the hierarchy mechanism works, and it’s cheaper to tear down between test sessions.

The Terraform structure ends up mirroring the ownership split conceptually, not just technically, which file something lives in should tell you who’s allowed to touch it:

# fw-parent-baseline.tf, fully managed by us, no exceptions.
resource "azurerm_firewall_policy" "parent" {
  name                = "fwpol-parent"
  resource_group_name = azurerm_resource_group.firewall.name
  location            = azurerm_resource_group.firewall.location
  sku                 = "Standard"
}

resource "azurerm_firewall_policy_rule_collection_group" "parent" {
  name               = "rcg-parent-baseline"
  firewall_policy_id = azurerm_firewall_policy.parent.id
  priority           = 200

  network_rule_collection {
    name     = "baseline-deny-known-bad"
    priority = 200
    action   = "Deny"
    # ... actual baseline rules
  }
}
# fw-child.tf, customer-editable content, deliberately separated.
resource "azurerm_firewall_policy" "child" {
  name                = "fwpol-child"
  resource_group_name = azurerm_resource_group.firewall_child.name # own resource group
  location            = azurerm_resource_group.firewall_child.location
  sku                 = "Standard"
  base_policy_id      = azurerm_firewall_policy.parent.id
}

resource "azurerm_firewall_policy_rule_collection_group" "child" {
  name               = "rcg-child"
  firewall_policy_id = azurerm_firewall_policy.child.id
  priority           = 300

  application_rule_collection {
    name     = "starter-rules"
    priority = 300
    action   = "Allow"
    # ... starter content, meant to be replaced via the Portal afterwards
  }

  lifecycle {
    ignore_changes = [application_rule_collection, network_rule_collection, nat_rule_collection]
  }
}

Two design choices worth calling out:

The child policy lives in its own resource group, separate from the parent and the firewall. That’s not accidental, it means RBAC scoping the customer’s access to “everything in this one resource group” already gets you most of the way to “customer can edit their child policy and nothing else,” without needing a custom role definition.

For finer-grained control than “the whole resource group,” Microsoft publishes a reference architecture for exactly this scenario, a custom role built around Microsoft.Network/firewallPolicies/ruleCollectionGroups/write. Their example pairs it with a blanket */read, though, read access to everything in the scope just so the customer can see what they’re editing. Checked against the resource provider’s actual list of available operations rather than assumed, the real minimum is four actions:

Microsoft.Resources/subscriptions/resourceGroups/read
Microsoft.Network/firewallPolicies/read
Microsoft.Network/firewallPolicies/ruleCollectionGroups/read
Microsoft.Network/firewallPolicies/ruleCollectionGroups/write

assigned at the scope of the child’s resource group, not the subscription. That’s enough to see the policy, see its rule collection groups, and edit them, nothing more. Two things are deliberately left out: Microsoft.Network/firewallPolicies/write, which would let the customer touch the policy object itself rather than just its rule collection groups, its DNS settings, its Threat Intelligence mode, its base_policy_id, exactly what the parent/child split exists to keep out of their hands, and both delete actions, since being able to create and edit a rule collection group doesn’t imply needing to remove one, or the policy, entirely.

ignore_changes on the child’s rule collections is what makes this survivable long-term. Terraform still refreshes real state on every plan/apply, so it always knows what the customer actually has configured, it just never treats a difference from the .tf file as something to fix. Without this, the next pipeline run from the MSP side would silently overwrite whatever the customer added through the Portal. The parent gets no such treatment, it stays fully Terraform-enforced, because the whole point is that the baseline is not up for negotiation outside of code review.

One real deployment error worth mentioning, because the error message alone doesn’t immediately point at the cause: if the firewall module defaults to zone-redundant (zones = ["1", "2", "3"]), its public IP has to be zone-redundant too, or you’ll hit ZonalAzureFirewallCannotReferenceNoZonePublicIp. Costs nothing extra to fix, zones are a resilience property, not a pricing tier, just add the same zones argument to the azurerm_public_ip resource.

A consequence you don’t get for free: monitoring what the customer changes

Because ignore_changes suppresses the diff at plan time, not just apply time, a normal Terraform-plan-based drift check will report “no changes” no matter what the customer does in their child policy, that’s the entire point of ignore_changes, but it also means it’s blind to exactly the thing you might want visibility into (is someone using priority ranges you didn’t agree on, did someone add an overly broad source_addresses = ["*"] rule). If you need that visibility, it has to come from something that reads the live resource directly, independent of Terraform’s plan/apply cycle, an Azure Policy definition auditing rule collection groups is a better fit here than a custom script, since it runs continuously without needing its own pipeline.

Scenario 2: retrofitting an existing, already-live firewall

The from-scratch version is the easy case. Real customers usually already have a firewall with rules on it, and the migration path depends entirely on what it’s currently running.

If the customer already uses a Firewall Policy (just not a hierarchical one yet), this is almost anticlimactically simple: create the new parent, then point the existing policy’s base_policy_id at it. Confirmed behavior, the existing rules are not touched, they just become subordinate in processing order to whatever the new parent defines. No rule collection groups need to be recreated, no rules re-entered by hand.

If the customer is still on Classic Rules (rules defined directly on the azurerm_firewall resource, no separate policy object at all), there’s a mandatory intermediate step, because classic rules can’t be given a parent directly. Microsoft publishes an official PowerShell script, AZFWMigrationScript.ps1, that reads the live firewall and recreates every rule collection (application, network, NAT) as an equivalent Firewall Policy. Once that’s done, you’re back in the first case: point the new policy’s base_policy_id at your parent.

A few things worth planning around rather than discovering mid-change:

  • This is a live change to a production firewall’s rule evaluation. Even though Azure Firewall handles most reconfiguration without a hard outage, treat it like a maintenance-window change, not something to run mid-afternoon on a whim.
  • If the firewall wasn’t already managed by Terraform, the migration and the IaC-adoption are two separate concerns, you’ll likely still need a terraform import pass afterwards to bring the resulting policy (and the firewall, VNet, public IP, etc.) under management, same as any other brownfield migration.
  • Test the exact sequence somewhere disposable first, and you don’t even need a firewall to do it. A Firewall Policy attached to zero or one firewall instances is free, and everything worth verifying before you touch a customer, inheritance, evaluation order, RBAC scoping, is visible directly on the policy objects in the Portal. Stand up a throwaway parent and child, confirm it behaves the way you expect, then tear it down. Much cheaper, in money and in risk, than finding out something doesn’t behave the way the docs describe on a customer’s live perimeter firewall.

Wrapping up

The mechanism itself is solid and well-documented once you know where to look, the surprises here weren’t bugs, they were details that don’t show up until you actually build the thing: priority numbers meaning nothing across a policy boundary, NAT rules deliberately breaking the inheritance pattern, a public IP’s zone configuration silently blocking a firewall deployment, and a ignore_changes block that solves the “don’t overwrite the customer” problem while quietly creating a “now you can’t see what they changed” problem of its own. None of it was hard to fix once identified, but all of it was worth verifying against Microsoft’s actual documentation rather than assuming, which is exactly what this post has tried to do throughout.