Summary

I built an OpenTofu module that manages Google Cloud Pub/Sub. You keep a single list of names, and for every name it creates a topic, a subscription, a dead-letter topic and a dead-letter subscription, plus exactly the IAM that dead-lettering needs. Adding a topic is one new line in the list instead of four resources and two IAM grants by hand. Changes go through a GitLab pipeline: validate → plan → manual apply.

The most useful part turned out to be the migration. Most topics had been created by hand in the console. While importing them, it turned out that every dead-letter topic except one had no subscription, so any message dead-lettered on them was dropped the moment it arrived. After the migration, the code, the state and GCP match (No changes), so there is no more drift between what is written down and what actually runs.

The problem

Topics clicked together in the console drift apart. Nobody reviews them, and nothing tells you when one of them is configured wrong. The worst case is quiet: a subscription with a dead-letter policy whose dead-letter topic has no subscription. Pub/Sub accepts the message, and since nobody is subscribed, drops it immediately.

What it builds

The only file you normally should edit is terraform.tfvars:

names = [
  "orders",
  "payments",
  "notifications",
]

Every name becomes four Pub/Sub resources and two IAM bindings:

orders     topic  ──►  orders     subscription
                         │
                         │  message fails 5 delivery attempts
                         ▼
orders-DL  topic  ──►  orders-DL  subscription   (keeps messages for 7 days)

Adding a topic is one new line pushed to main. The plan job shows 6 to add, 0 to destroy, and nothing happens in GCP until you read the plan and click apply.

How it works

The code is declarative: it describes what should exist, not the steps to get there. tofu plan compares that with reality and lists the difference, and tofu apply carries it out. That cuts both ways: adding a name to the list creates its resources, and removing a name destroys them.

The whole module is for_each over a set of names:

locals {
  names        = toset(var.names)
  pubsub_agent = "serviceAccount:service-${var.project_number}@gcp-sa-pubsub.iam.gserviceaccount.com"
}

resource "google_pubsub_subscription" "subscription" {
  for_each = local.names
  name     = each.key
  topic    = google_pubsub_topic.topic[each.key].name
  expiration_policy { ttl = "" }

  dead_letter_policy {
    dead_letter_topic     = google_pubsub_topic.dead_letter_topic[each.key].id
    max_delivery_attempts = var.max_delivery_attempts
  }
}

pubsub_agent is the address of Google’s Pub/Sub service agent, the account that actually moves failed messages to the dead-letter topic. Google builds it from the project number (not the ID), so the number is set once in terraform.tfvars, and the IAM bindings below grant their roles to this address.

State in a GCS bucket

Tofu keeps a state file: its record of which real GCP resource belongs to which block in the code, for example google_pubsub_topic.topic["orders"] → topic orders. Every plan compares three things: the code, the state and what actually exists in GCP.

The state lives in a GCS bucket, not on anyone’s laptop. A local state file only knows about applies run from that one machine; if it gets lost, or two people work from different copies, Tofu loses track of what it manages and tries to create resources that already exist.

backend "gcs" {
  bucket = "josipziva-tfstate"
  prefix = "pubsub-creator"
}
  • Shared: the pipeline and a local tofu plan see the same state.
  • Locked: while one apply runs, a lock object in the bucket blocks every other one.
  • Versioned: bucket versioning keeps every old version of the state, so a state broken by mistake can be restored in a minute. The pipeline’s service account has access to this bucket only and cannot turn versioning off, but it can still delete old versions: versioning protects against mistakes, not against a leaked key.

Minimal needed permissions

Every identity gets only the permissions its job needs, on the narrowest scope that works:

IdentityRoleScopeWhy
Pub/Sub service agentpubsub.publishereach -DL topicpublish a failed message to the dead-letter topic
Pub/Sub service agentpubsub.subscribereach source subscriptionacknowledge the message after forwarding it
Terraform service accountpubsub.adminprojectcreate topics and subscriptions, set IAM on them
Terraform service accountstorage.objectAdminstate bucket onlyread and write the state

This matters because the pipeline’s key is the one secret in this setup that can leak. With these roles, a leaked key reaches Pub/Sub and the state bucket, never the rest of the project: it cannot grant itself roles on the project or touch any other service.

The known trade-off is Pub/Sub itself. pubsub.admin has to be granted on the project, because a topic that doesn’t exist yet has no IAM policy of its own to grant on. So a leaked key could still delete any topic or subscription in the project, including the ones Tofu doesn’t manage. The planned answer is to remove the long-lived key altogether (see What’s next).

Bringing existing topics under Tofu

The state knew about one name (payments, with all four of its Pub/Sub resources), while the list asked for 21. The migration happened before the per-resource IAM from Minimal needed permissions existed: dead-lettering then used two project-wide bindings that were already in the state, so the module added only the four Pub/Sub resources per name. The first plan therefore said 80 to add (20 names × 4), and most of those would have failed with 409 ALREADY_EXISTS.

The fix for that is import: it writes an existing resource into the state under its address in the code, without creating or changing anything in GCP. From then on, Tofu manages it like any resource it created itself:

tofu import 'google_pubsub_topic.topic["orders"]' projects/PROJECT_ID/topics/orders

But first you need to know what actually exists. An inventory of GCP for the 20 names missing from the state showed:

ResourceExists in GCP?Action
topics✅ 20import
subscriptions✅ 20import
dead-letter topics✅ 20import
dead-letter subscriptions❌ nonecreate

After 60 tofu imports (20 names × 3 existing resources), the plan showed 20 to add, 0 to change, 0 to destroy: of the original 80, only the 20 missing dead-letter subscriptions were left. 0 to change confirmed that the code matched reality, and those 20 were a real fix, not paperwork.

The per-resource IAM bindings came in a later change. An *_iam_member only adds one grant, so there was nothing to import: the plan showed them as new and apply created them.

Pipeline

validate  (fmt, validate, tofu test, no credentials)
   ↓
plan      (prints only the CHANGES list)
   ↓
apply     (manual, main only, one at a time)

The service account’s JSON key is stored as a base64 string in a GitLab CI variable that is masked and hidden, so it never shows up in job logs or in the settings UI, and protected, so only pipelines on main get it. Base64 turns the multi-line JSON into a single line, which GitLab needs in order to mask it. Every job decodes it straight into an environment variable, so the key is never written to disk:

before_script:
  - export GOOGLE_CREDENTIALS="$(echo "$GOOGLE_APPLICATION_CREDENTIALS_base64" | base64 -d)"
  • apply runs the saved tfplan from the plan job, so what you reviewed is exactly what gets applied.
  • resource_group makes sure two apply jobs never even start at the same time, on top of the state lock.

Testing without GCP

tofu test with a mocked provider runs in CI without any project or credentials. The tests check things that are easy to break and hard to notice:

run "subscriptions_never_expire" {
  command = plan

  assert {
    condition     = google_pubsub_subscription.dead_letter_subscription["orders"].expiration_policy[0].ttl == ""
    error_message = "Dead-letter subscription must never expire (the 31-day default would silently drop it)."
  }
}

Gotchas

  • Dead-letter subscriptions expire after 31 days by default. A DL subscription is idle by nature, so without ttl = "" it would delete itself and the original problem would come back. Code review caught this, and now a test does.
  • Name rules are only checked by GCP at create time. t1 passed the plan and failed on apply (min. 3 characters), while t1-DL was created. State recorded what succeeded, so renaming fixed it.
  • Order matters when taking a role away. First apply the code that removes the old project-level bindings (the SA originally had Project IAM Admin to manage those bindings), then remove Project IAM Admin from the SA. The other way around, Tofu can’t even destroy the old bindings.
  • prevent_destroy is off on purpose. It made removing a topic a two-commit dance. The price: removing a name deletes the topic and its unprocessed messages, so the plan has to be read before every apply.

Lessons learned

  • Importing is mostly inventory work. The tofu import commands are the easy part.
  • A plan with 0 to change after an import is the proof you need.
  • Keep the import procedure as documentation, not as a script. A script next to the names list is a second source of truth. In hindsight, import blocks with for_each over the same names list would have done this declaratively: no second list, and every import shows up in the plan before it happens.

What’s next

  • An alert on the dead-letter backlog. This is the most important gap. A dead-lettered message now waits in its -DL subscription for 7 days and then expires, again without anyone noticing: the module turned “lost immediately” into “lost after a week”. The fix is an alert on subscription/num_undelivered_messages (or oldest_unacked_message_age) for every -DL subscription, created by Tofu as a google_monitoring_alert_policy in the same for_each, so every new name gets its alert automatically.
  • Workload Identity Federation instead of a long-lived key
  • A daily drift pipeline (plan -detailed-exitcode)

Source

Code, pipeline and the full write-up are on GitLab/josip.zivic93