Give a coding agent the same Terraform request on two different days and you get two dialects from the same model. On the first day the network is labelled main and the provider lives in main.tf. On the second the network is vpc_primary, the provider has moved to providers.tf, and the subnet IDs are read through an output name the first stack never published. Somewhere in one of the two sits an argument the provider has never had, which terraform validate will be the first to notice. And the session that was asked for a plan runs an apply, because nothing in its context said that was someone else's job. The model is the same on both days. What changed is where the conventions lived, and on both days that was the head of whoever typed the prompt.

This article is a golden path for running Terraform with an AI agent in a professional repository: one reference setup, written to be cloned and used as it stands. It is opinionated on purpose. The conventions live in the repository as files the agent reads: a CLAUDE.md of about eighty lines that is always in context, five skills and three subagents that are loaded on demand, three hooks that run outside the model, and an identity that can only plan in production. The same setup serves two jobs. It builds a new estate from zero on whichever cloud you point it at. It also translates an existing estate from AWS to Google Cloud (or back) by treating the finished stacks of one cloud as the specification for the other. The companion repository holds both flavors of the setup and a complete floor for a platform named atlas on each cloud. On AWS that is a state bucket in 00-remote-backend, a two-zone VPC with one NAT gateway in 01-networking, an EKS 1.36 cluster with a small managed node group in 02-eks, and a private PostgreSQL 18 instance plus an application bucket in 03-data. On Google Cloud it is the same four stacks as a GCS state bucket, a custom VPC with Cloud NAT, a regional GKE cluster and Cloud SQL on a private IP. Everything validates offline. Every stack plans and applies with the commands at the end of this article.

The thesis fits in five words. The conventions are the prompt. Numbered stacks with their own state, one object variable per domain, this and main as labels, upstream outputs read through locals, an identity that can only plan in production. Write those down where the agent reads them and it becomes dependable for infrastructure work on any cloud, because most of the rules never mention one.

The path goes in order. First the MCP servers and plugins the loop depends on. Then the anatomy of the setup, what stays in context and what is kept out of the model altogether, and the loop every stack goes through. Then the AWS floor built from zero, stack by stack, with the design decisions that make it safe to hand to an agent. Then the same floor on Google Cloud with the AWS stacks as the spec, the cost of both floors as Infracost reads them, what to change when your target is a third cloud, and the commands to run all of it yourself.

Step Zero Is Installing the Servers the Loop Depends On

The tooling rule in CLAUDE.md is a loop with a fixed order: schema, write, validate, plan, cost, review, and only then a human apply. Three of those steps call tools Claude Code does not carry on its own. The schema step needs HashiCorp's Terraform MCP server, whose registry tools return the provider's own documentation for the pinned version, so no resource argument is written from memory. The documentation step needs Context7 when the registry page is not enough. The cost step needs the Infracost plugin. Read-only checks against the account go through the AWS MCP or the gcloud MCP. Install them once, before the first prompt. The repository ships a project-scoped .mcp.json in each of aws/ and gcp/; the AWS one, aws/.mcp.json, is short:

{
  "mcpServers": {
    "terraform": {
      "type": "stdio",
      "command": "docker",
      "args": ["run", "-i", "--rm", "hashicorp/terraform-mcp-server"]
    },
    "context7": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@upstash/context7-mcp"]
    }
  }
}

gcp/.mcp.json adds a third entry, gcloud, running npx -y @google-cloud/gcloud-mcp, and Claude Code asks you to trust project-scoped servers on first start. To carry the same servers into every repository, use the user-scope commands from docs/mcp-setup.md, checked against each project's own README:

# Terraform MCP (official install command from github.com/hashicorp/terraform-mcp-server)
claude mcp add terraform -s user -t stdio -- docker run -i --rm hashicorp/terraform-mcp-server
# Context7 (github.com/upstash/context7)
claude mcp add context7 -s user -t stdio -- npx -y @upstash/context7-mcp
# gcloud MCP (github.com/googleapis/gcloud-mcp), GCP repositories only
claude mcp add gcloud -s user -t stdio -- npx -y @google-cloud/gcloud-mcp

-s user stores the server in your user config, -t stdio is the transport, and everything after -- goes to the server untouched. The Terraform server is a Docker image (hashicorp/terraform-mcp-server, 1.3.0 at the time of writing); the other two run through npx on Node 20 or later. The gcloud server inherits the permissions of the authenticated gcloud CLI, which is why its README recommends impersonating a service account with narrow roles, the plan-only and apply pair this repository defines. AWS and Infracost arrive as Claude Code plugins:

# AWS: official plugin from aws/agent-toolkit-for-aws (MCP proxy + skills). Needs `uv`.
/plugin install aws-core@claude-plugins-official
# if the official marketplace is not available in your install:
/plugin marketplace add aws/agent-toolkit-for-aws
# Infracost (www.infracost.io/docs/ai_editor_plugins/ai_skills/)
claude plugin marketplace add infracost/agent-skills
claude plugin install infracost@infracost
#   exposes /infracost:scan, /infracost:price-lookup, /infracost:iac-generation
#   first use prompts for dependencies and account authentication (free account)
# Context7 is also available as a plugin instead of an MCP entry:
/plugin install context7@claude-plugins-official

Finish with claude mcp list; every server should read ✔ Connected. A ✘ Failed to connect on terraform almost always means Docker is not running.

Know what the Terraform server does before you write rules around it. Its registry toolset is search_providers, get_provider_details and get_latest_provider_version, plus the module and policy equivalents. It also exposes HCP Terraform tools for workspaces and runs. It does not run init, validate or plan. Sample agent configurations on the internet often name tools such as get_schema or terraform_validate that the official server does not have. A model told to call a tool that does not exist will either stall or improvise. The setup in the repository names the real ones. The schema step is search_providers followed by get_provider_details, and init, validate and plan are CLI commands the agent runs through Bash.

In short, step zero is five installs and one claude mcp list: the Terraform MCP for schema, Context7 for documentation, Infracost for cost, the AWS or gcloud servers as read-only eyes on the account. The schema call is search_providers then get_provider_details; validate and plan stay on the CLI.

What Stays in Context, and What Stays Outside the Model

The setup has three layers.

None
Anatomy of the setup: CLAUDE.md (about eighty lines, always in context) in the centre; the five skills and three subagents on either side, loaded on demand; underneath, the two guardrails that live outside the model, the hooks wired in settings.json and the plan-only IAM role.

The first layer is CLAUDE.md, the one file always in context, 82 lines on AWS and 84 on GCP. It holds only what must be true on every turn: the identity, the safety rules, the order in which tools are used, the response format, what to delegate, and the bootstrap block with the account id, the state bucket and the two role ARNs every provider block keys off. I keep it that short on purpose. A rule on line 400 of a long instruction file is a rule the model half remembers in a long session, and every procedure belongs in a skill. The two sections doing the most work are the guardrails and the tooling strategy, from aws/CLAUDE.md:

## Safety Guardrails (ALWAYS-ON)
**EXECUTION PROHIBITED**:
- Never run `terraform apply` or `terraform destroy` when `terraform.workspace = "production"`.
- Always run `terraform plan` first; surface the diff and ask the human for explicit approval before any state-changing operation.
- The repo also ships a hard guardrail (a `terraform-plan-readonly` IAM role assumed only when workspace is `production`). The model-level rule above complements but does not replace it.
**WORKSPACE CHECK**:
- At session start, run `terraform workspace show`.
- Adapt all suggestions to the active workspace. Sandbox and staging allow full CRUD; production allows plan only.
## Tooling Strategy (ALWAYS-ON)
Four categories of tools, used in this strict order:
1. **Claude Code native** (Read, Write, Edit, Bash, Glob, Grep) for all file I/O. The Terraform MCP does NOT edit files.
2. **Terraform MCP** (official server, registry toolset) for the schema step, before writing any resource or data source block (anti-hallucination):
   - `search_providers` to locate the resource's documentation for the provider pinned in `main.tf`
   - `get_provider_details` to read its arguments, blocks and attributes
   - `get_latest_provider_version` when checking or bumping the pin
   The server does not run `init`, `validate` or `plan`; those are CLI commands.
3. **Terraform CLI via Bash** for the loop itself:
   - `terraform init -backend=false` when providers are not installed (no credentials needed)
   - `terraform fmt` and `terraform validate` immediately after Write/Edit (the PostToolUse hook runs both on every `.tf` write as well)
   - `terraform workspace show`, then `terraform plan -out tfplan` to simulate impact; never `terraform apply`
4. **Infracost plugin** (`/infracost:scan`) after every `terraform plan`. Flag any resource above $500/month and ask for confirmation.

The workspace check is what makes everything downstream conditional. The role map in every stack keys off terraform.workspace, so the agent has to know where it stands before it proposes anything. The third bullet under execution states, in the file the model reads, that the model-level rule is a complement to an identity in IAM that cannot write. That sentence is the design of the whole setup in one line.

The second layer is what gets called. Five skills, each a single-purpose procedure invoked by slash command or matched by its description. /tf-scaffold-stack lays out a numbered stack and validates the skeleton before any business resource goes in. /tf-variables-review enforces the one-object-per-domain shape, /tf-naming-review polices this against main and the Name tag, /tf-cross-stack wires a downstream stack to upstream outputs starting from the upstream outputs.tf, and /tf-outputs-review checks what a stack publishes. Three subagents carry a persona across turns. terraform-architect owns the thirteen conventions and never applies, terraform-cost-reviewer reads the plan and the Infracost output against a $500-a-month per-resource threshold, and terraform-security-reviewer audits IAM, encryption, network posture and drift. Splitting always-on from on-demand is what keeps the short file sharp.

The third layer sits outside the model, and if I had to keep one layer it would be this one. Three command hooks, PreToolUse on Bash, PostToolUse on Edit|Write and Stop, are wired in aws/.claude/settings.json, each a shell script under .claude/hooks/ reached through $CLAUDE_PROJECT_DIR. A hook receives its input as JSON on stdin and runs in the session's working directory, so every script reads the fields it needs with jq and changes into the right directory before touching Terraform. The guard, pre-tool-guard.sh, answers with a decision:

input="$(cat)"
command="$(printf '%s' "$input" | jq -r '.tool_input.command // empty')"
cwd="$(printf '%s' "$input" | jq -r '.cwd // empty')"
[ -n "$command" ] || exit 0
printf '%s' "$command" | grep -qE '(^|[[:space:];&|])terraform([[:space:]]|$)' || exit 0
deny() {
  jq -n --arg reason "$1" '{
    hookSpecificOutput: {
      hookEventName: "PreToolUse",
      permissionDecision: "deny",
      permissionDecisionReason: $reason
    }
  }'
  exit 0
}
# 1. destroy without -target
if printf '%s' "$command" | grep -qE 'terraform[[:space:]]+destroy' \
   && ! printf '%s' "$command" | grep -qE -- '-target[= ]'; then
  deny "Blocked: 'terraform destroy' without -target. Humans run destroy, outside the agent."
fi
# 2. apply -auto-approve, anywhere
if printf '%s' "$command" | grep -qE 'terraform[[:space:]]+apply' \
   && printf '%s' "$command" | grep -qE -- '-auto-approve'; then
  deny "Blocked: 'terraform apply -auto-approve'. Plan, then hand the apply to a human."
fi

A third rule follows the two shown. For apply, destroy, import, state mv and state rm it resolves the directory the command will run in, honouring a leading cd <dir> &&, asks terraform workspace show there, and denies when the answer is production. The hook is deterministic, testable offline with a sample payload. Its denial reaches the model as a reason it can act on. It cannot see a terraform apply typed by a human in another terminal, which is what the role is for.

The companion, post-tool-tf-validate.sh, runs after every write to a .tf file. It reads tool_input.file_path, changes into the file's directory, formats the file, initialises the stack without a backend if the providers are missing, and runs terraform validate. On success it returns additionalContext saying the stack is clean; on failure it exits 2 with the validate output on stderr. A PostToolUse hook cannot block, since the write already happened. What exit 2 buys is that stderr is fed back to the model, so an invented attribute such as cidr_blok lands in front of the agent on the same turn, with the line number, instead of waiting for someone to run validate by hand. The Stop hook runs fmt -check and validate over every initialised stack and greps the tree for credential-looking material before the turn ends, skipping its own hooks/ directory so the patterns it searches for do not match themselves.

Then the identity. The rule is that whatever is assumed when terraform.workspace is production can read everything and write nothing but the state lock, so that any write call fails at the cloud's boundary regardless of what the model decides. How that is written in IAM gets its own section, right after the AWS floor.

In short, the setup is one always-on file of about eighty lines, eight on-demand components that own the conventions and the reviews, and two guardrails that do not depend on the model at all. The hooks catch the agent; the role catches everyone. A prompt can be argued with. An AccessDenied cannot.

The Loop for One Stack

Every stack on both clouds goes through the same nine steps.

None
The loop for one stack, nine steps in a serpentine: prompt; the conventions apply; provider docs through the Terraform MCP; write the .tf files; the PostToolUse hook runs fmt and validate; plan under the workspace check; Infracost; the cost and security reviewers; a human applies. Callouts mark the loop back from a failed validate and what the PreToolUse hook blocks.

The prompt is a sentence, something like "scaffold the networking stack". CLAUDE.md is already in context, the request matches /tf-scaffold-stack. The skill's sequence takes over: next number, kebab-case name, skeleton, init and validate before a single business resource exists. Each resource block is written from the provider's documentation for the pinned version. The write fires the PostToolUse hook, so terraform fmt and terraform validate have run by the time the agent sees its own file, and anything wrong comes back as error text to fix. When the stack validates, the agent runs terraform workspace show and terraform plan -out tfplan, assuming the apply role on sandbox and the plan-only role on production, then Infracost and the two reviewers read the plan. The answer comes back in the shape CLAUDE.md prescribes, workspace, resource blocks, validate output, the first fifty lines of the plan, the cost estimate, and then a pause, because the apply is a human's command. The hook refuses -auto-approve and anything state-changing on production. The role refuses it again at the API.

Building From Zero: The AWS Floor, One Stack at a Time

A new estate starts with an empty directory, the setup copied in (CLAUDE.md and .claude/ at the root of aws/), and the bootstrap block filled with your values. The repository ships placeholders: atlas, account 123456789012, bucket atlas-us-east-1-bucket-terraform-state, us-east-1, and the two role ARNs. From there the agent builds one stack per prompt.

The first stack, 00-remote-backend, fixes the layout every later stack copies. main.tf holds only the terraform and provider blocks. Resources go in one kebab-case file per resource group, data sources in datasources.tf, and the state key follows <stack>/<stack>.tfstate. The provider is pinned at ~> 6.0. The stack itself is variables.tf with three objects (tags, assume_role, remote_backend), s3-bucket.tf with the bucket, versioning, encryption and public access block all labelled this, outputs.tf, and a README with the bootstrap sequence, since the bucket the backend points at does not exist on the first apply.

The provider block is where the workspace becomes a security decision. Each workspace maps to an identity. The lookup has to survive the workspace every fresh checkout sits in. terraform validate and a first terraform init both run in default, which is not in the map, so a direct index such as local.workspace_role_arn[terraform.workspace] fails validate on the first resource that uses the provider. The repository resolves it with a fallback that fails closed, from aws/terraform/00-remote-backend/main.tf:

# Workspace-aware role selection: production is plan-only for the agent.
# The hard guardrail lives in IAM (see iam/README.md), not here.
# Any workspace not listed here (including "default", the one a fresh
# `terraform init` lands in) falls back to var.assume_role.role_arn, which
# defaults to the plan-only role: unknown workspaces fail closed.
locals {
  workspace_role_arn = {
    sandbox    = "arn:aws:iam::123456789012:role/terraform-apply"
    staging    = "arn:aws:iam::123456789012:role/terraform-apply"
    production = "arn:aws:iam::123456789012:role/terraform-plan-readonly"
  }
}
provider "aws" {
  region = var.assume_role.region
  default_tags {
    tags = var.tags
  }
  assume_role {
    role_arn = lookup(local.workspace_role_arn, terraform.workspace, var.assume_role.role_arn)
  }
}

lookup falls back to var.assume_role.role_arn, whose default is the plan-only role. Any workspace not in the map, default included, gets the identity that cannot write, which is the direction you want an unknown to fail in. The provider block is identical in every stack from here on, with only the backend key changing.

01-networking is the vpc object with enable_dns_support and enable_dns_hostnames on, because EKS needs hostnames. Two public and two private subnets in us-east-1a and 1b are driven by count over the lists. The singletons (internet gateway, EIP, NAT gateway) are labelled this, the VPC is main, subnets and route tables carry their role, and names follow <type>-<project>-<condensed-region>, as in vpc-atlas-useast1. The private subnet file, subnet-private.tf, is ten lines and shows three of the conventions at once:

resource "aws_subnet" "private" {
  count             = length(var.vpc.private_subnets)
  vpc_id            = aws_vpc.main.id
  cidr_block        = var.vpc.private_subnets[count.index].cidr_block
  availability_zone = var.vpc.private_subnets[count.index].availability_zone
  tags = {
    Name = var.vpc.private_subnets[count.index].name
  }
}

count over a list inside the domain object, a role label, and a tags block that carries Name and nothing else because Project, Env and ManagedBy come from the provider's default_tags. Two decisions sit in this stack and are written in the file. One NAT gateway, the right trade-off for a sandbox floor; a per-AZ NAT is the production setting and the cost reviewer will price it for you. And the Name-only rule has one exception, the tags a managed service reads for discovery. The day an ingress controller lands, the subnets need kubernetes.io/role/elb and kubernetes.io/role/internal-elb, and the naming skill allows them as the one exception, listed in the stack README when they are added.

02-eks begins the way /tf-cross-stack insists, by reading 01-networking/outputs.tf before writing anything, then a datasources.tf with the remote state and the locals that normalise what the stack consumes, from aws/terraform/02-eks/datasources.tf:

# Named workspaces (sandbox/staging/production) store state under
# env:/<workspace>/<key>; without `workspace` this block would silently read
# the default workspace's state, which does not exist in this layout.
data "terraform_remote_state" "network" {
  backend   = "s3"
  workspace = terraform.workspace
  config = {
    bucket = "atlas-us-east-1-bucket-terraform-state"
    key    = "networking/networking.tfstate"
    region = "us-east-1"
  }
}
locals {
  vpc_id             = data.terraform_remote_state.network.outputs.vpc.id
  public_subnet_ids  = data.terraform_remote_state.network.outputs.public_subnets_ids
  private_subnet_ids = data.terraform_remote_state.network.outputs.private_subnets_ids
}

The workspace line is the one most cross-stack examples leave out, and in a workspace-per-environment layout it is mandatory. The S3 backend stores named workspaces at env:/<workspace>/<key>. A terraform_remote_state without workspace reads the default workspace's state, so in a sandbox session 02-eks would look for networking/networking.tfstate at the root of the bucket, an object this layout never writes. The stack would validate clean and fail at plan time with an error about a missing output. Every downstream stack in the repository carries the line.

The cluster is one eks object, with the managed policy ARNs as lists so that the attachments use count. AmazonEKSClusterPolicy goes on the control plane. The nodes get AmazonEKSWorkerNodePolicy, AmazonEC2ContainerRegistryPullOnly (which AWS now recommends over the ReadOnly policy) and AmazonEKS_CNI_Policy, kept on the node role for a floor with no add-ons stack. The cluster runs 1.36, the newest version in standard support, with authentication_mode = "API" and bootstrap_cluster_creator_admin_permissions = false, so the Terraform role gets no Kubernetes access and the only cluster-admin is an explicit access entry for a platform-admin principal bound to AmazonEKSClusterAdminPolicy. The node group is two t4g.small on AL2023_ARM_64_STANDARD, a value from the EKS API reference. The IAM policy documents live in datasources.tf with every other data source. The labels are cluster and node, lowercase and without hyphens, as HashiCorp's style guide asks.

03-data has two domains, so it has two objects, rds and app_bucket. The rule counts domains, and a stack can own more than one. RDS PostgreSQL 18 runs on db.t4g.micro with 20 GiB of gp3, encrypted, on a subnet group over the private subnets, behind a security group with one ingress rule from the VPC CIDR written with the standalone aws_vpc_security_group_ingress_rule, and publicly_accessible = false. The master password never enters state, because manage_master_user_password = true leaves it with Secrets Manager. The sandbox flags, deletion_protection = false and skip_final_snapshot = true, sit in the object with a comment saying what production flips. The bucket is the same four resources as the state bucket.

In short, the AWS floor is four stacks that share one layout and one provider block. The three decisions that make it safe to hand to an agent are a workspace lookup that fails closed, a remote state that names its workspace, and an identity for production that cannot write.

The Plan-Only Identity, Written in IAM's Own Grammar

The production identity has two requirements that pull against each other. It must not write, and terraform plan must still be able to lock the state. The S3 backend with use_lockfile = true acquires a lock by creating <key>.tflock and deletes it when done. The backend documentation lists s3:GetObject, s3:PutObject and s3:DeleteObject as required on the lock file. A role with only Get and List on the bucket stops terraform plan at lock acquisition, before it reads a single resource. There are two exits. Run terraform plan -lock=false on production and accept an unlocked plan that a concurrent apply can make stale, or let the role write the lock marker and nothing else. The repository takes the second, in the first two statements of aws/iam/plan-only-state.json:

{
      "Sid": "StateLockFileOnly",
      "Effect": "Allow",
      "Action": [
        "s3:PutObject",
        "s3:DeleteObject"
      ],
      "Resource": "arn:aws:s3:::atlas-us-east-1-bucket-terraform-state/*.tflock"
    },
    {
      "Sid": "DenyWritesOutsideLockFile",
      "Effect": "Deny",
      "Action": [
        "s3:PutObject",
        "s3:DeleteObject"
      ],
      "NotResource": "arn:aws:s3:::atlas-us-east-1-bucket-terraform-state/*.tflock"
    },

Allow on *.tflock, explicit deny on every other key. The role still cannot write state; it can only leave the marker that keeps a human's apply and the agent's plan from crossing. The suffix pattern holds for any future stack because the key convention is <stack>/<stack>.tfstate, so its lock is always <stack>/<stack>.tfstate.tflock. A third statement, ReadStateBucket, keeps the state read explicit. DynamoDB is not in the picture at all; the backend documentation calls DynamoDB-based locking deprecated and slated for removal.

The read side has its own detail. A read-everything policy is often written as *:Describe*, *:List* and *:Get*. IAM's Action element is <service>:<action>. The documented grammar allows a wildcard in the action name (ec2:Describe*), the whole-service form (s3:*) and the bare *, never in the service namespace. So the role is three pieces, documented in aws/iam/README.md. The AWS managed policy arn:aws:iam::aws:policy/ReadOnlyAccess enumerates the read actions of every service in the grammar's own words and is maintained by AWS. The state policy above needs no final deny, because IAM denies by default whatever no policy allows. And a permissions boundary, aws/iam/plan-only-boundary.json, is the ceiling that survives any policy attached to the role later. A boundary is a single managed policy and IAM does not compose them, so it cannot be ReadOnlyAccess plus a second document, and ReadOnlyAccess alone as the boundary would deny the lock write. It enumerates the read patterns the four stacks need, ec2:Describe* and ec2:Get*, eks:Describe* and eks:List*, iam:Get* and iam:List*, rds:Describe* and rds:List*, s3:Get* and s3:List*, plus the same .tflock allow. A new service in a future stack means a new line in the boundary. Until that line exists, the stack's plan fails closed with AccessDenied on production, the safe direction.

Translating the Floor to Google Cloud, With AWS as the Spec

The instruction for the second cloud is the thesis in the imperative. Do not design GCP; translate AWS. Each gcp/terraform/NN-* is written with the matching aws/terraform/NN-* open, resource by resource, and each GCP stack's README ends with the translation table. This is the workflow for a migration. It is also the fastest way to stand up a second cloud from a first one you already trust, because the structural decisions are already made.

The setup is translated first. gcp/CLAUDE.md, the three agents, the five skills and the hooks are the AWS files with the cloud-specific lines swapped and each one marked GCP: inline. In the always-on file 24 lines differ. The identity says Google Cloud and multi-project, the hard-guardrail sentence names a plan-only service account, terraform-plan@atlas-demo-123.iam.gserviceaccount.com, impersonated only on production, the Context7 library becomes /hashicorp/terraform-provider-google, and the bootstrap block holds a project id, a GCS bucket and two service account emails. Of the thirteen conventions, eight keep their rule text and change only the nouns in their examples. Three keep the idea and swap arguments: backend "gcs" with a prefix, impersonate_service_account with default_labels, and a remote state config of { bucket, prefix }. Two are different mechanisms, the name argument and lowercase labels against a Name tag, and a service account with a conditional .tflock binding against an IAM role. The hooks differ by four lines, the deny message and the secret patterns. The convention-by-convention table, with the design rules above, is in docs/conventions.md.

The backend and provider block of gcp/terraform/00-remote-backend/main.tf shows the translation of the backend and provider conventions in thirty lines:

# GCS locks natively and encrypts at rest by default: no lock table, no
  # `encrypt`/`use_lockfile` switches. The state object is
  # <prefix>/<workspace>.tfstate, so the workspace name, not the stack name,
  # ends the path. Backend access uses the ambient credentials (ADC) or
  # GOOGLE_BACKEND_IMPERSONATE_SERVICE_ACCOUNT; see iam/plan-only-role.md.
  backend "gcs" {
    bucket = "atlas-us-central1-bucket-terraform-state"
    prefix = "remote-backend"
  }
}
# Workspace-aware identity selection: production is plan-only for the agent.
# The hard guardrail lives in IAM (see iam/plan-only-role.md), not here.
# Any workspace not listed here (including "default", the one a fresh
# `terraform init` lands in) falls back to var.impersonation.service_account,
# which defaults to the plan-only account: unknown workspaces fail closed.
locals {
  workspace_service_account = {
    sandbox    = "terraform-apply@atlas-demo-123.iam.gserviceaccount.com"
    staging    = "terraform-apply@atlas-demo-123.iam.gserviceaccount.com"
    production = "terraform-plan@atlas-demo-123.iam.gserviceaccount.com"
  }
}
provider "google" {
  project        = var.impersonation.project_id
  region         = var.impersonation.region
  default_labels = var.labels
  impersonate_service_account = lookup(local.workspace_service_account, terraform.workspace, var.impersonation.service_account)
}

Same map, same fail-closed lookup, same comment, different nouns. The state key convention <stack>/<stack>.tfstate has no equivalent on GCS, because the backend names the object after the workspace and takes a prefix instead of a key; the stack README records it. Stack 00 also gains a resource with no AWS twin, google_project_service.this[*] over the six APIs the other stacks call, because a fresh project has them off. And four resources become one, since versioning, public access prevention and uniform bucket-level access are arguments of google_storage_bucket.

01-networking is where the two clouds disagree most. The translation table in its README is the one to read. Four subnets become one regional subnet with two secondary ranges for GKE pods and services. The internet gateway disappears, because the VPC has a default route. The EIP and NAT gateway become a Cloud Router and a Cloud NAT, regional and Google-managed, so the AWS stack's single-AZ trade-off has no counterpart. The route tables and associations disappear, because Cloud NAT applies to the subnet. Zone placement moves to the GKE stack's node_locations. Ten resources become four. The layout does not move at all: vpc.tf, subnetwork.tf, cloud-router.tf, cloud-nat.tf, one vpc object with inline defaults, main for the network, this for the singletons, outputs that expose whole objects plus the two range names downstream needs. The subnet shows how the count convention reads for nested blocks, in subnetwork.tf:

resource "google_compute_subnetwork" "this" {
  name                     = var.vpc.subnetwork.name
  network                  = google_compute_network.main.id
  region                   = var.impersonation.region
  ip_cidr_range            = var.vpc.subnetwork.ip_cidr_range
  private_ip_google_access = var.vpc.subnetwork.private_ip_google_access
  dynamic "secondary_ip_range" {
    for_each = var.vpc.subnetwork.secondary_ip_ranges
    content {
      range_name    = secondary_ip_range.value.range_name
      ip_cidr_range = secondary_ip_range.value.ip_cidr_range
    }
  }
}

count for plurals is a rule about resources. secondary_ip_range is a nested block list, nested blocks cannot take count, and so they take dynamic with for_each.

In 02-gke, resource by resource, the cluster IAM role has no GCP counterpart, since the control plane runs under a Google-managed service agent. The node IAM role becomes a dedicated google_service_account.node with four project roles attached through count. access_config and the access entry have no counterpart, because project IAM authorises. desired_size = 2 becomes node_count = 1 per zone across two node_locations, and AL2023_ARM_64_STANDARD becomes nothing, since Container-Optimized OS is the default. Workload Identity is on. Nodes are private with a public endpoint, and deletion_protection = false is written out, because on provider 5.0.0 and later it has to be set to false explicitly before a destroy will succeed. The datasources.tf is the AWS one with backend = "gcs", config = { bucket, prefix } and workspace = terraform.workspace, since GCS stores named workspaces as <prefix>/<name>.tfstate and the same rule applies.

03-data is where the schema step earns its place. The subnet group and security group pair become Private Services Access, a google_compute_global_address reserved for VPC peering plus a google_service_networking_connection, with ipv4_enabled = false on the instance; there is no firewall object because the peering is the boundary. The instance carries an explicit depends_on on the peering connection, because the provider documentation warns that Cloud SQL does not interpolate it and the private IP allocation fails on first apply without it. manage_master_user_password becomes IAM database authentication, a dedicated google_service_account.app with roles/cloudsql.instanceUser and a google_sql_user of type CLOUD_IAM_SERVICE_ACCOUNT whose name is the email with .gserviceaccount.com trimmed, as the provider documentation requires for PostgreSQL. No password anywhere. One argument deserves its comment, in gcp/terraform/03-data/variables.tf:

default = {
    instance_name    = "sql-postgres-atlas-uscentral1"
    database_version = "POSTGRES_18"
    tier             = "db-f1-micro"
    # PostgreSQL 16+ defaults to ENTERPRISE_PLUS, which rejects shared-core
    # tiers such as db-f1-micro; ENTERPRISE must be explicit.
    edition                = "ENTERPRISE"
    availability_type      = "ZONAL"
    disk_size              = 10
    disk_type              = "PD_SSD"

The provider documentation for google_sql_database_instance spells it out. When edition is unset the API picks it from database_version; POSTGRES_16 and later default to ENTERPRISE_PLUS, which supports only the db-perf-optimized-N-* machine types, so shared-core tiers such as db-f1-micro require edition = "ENTERPRISE". Without that line the stack validates and the apply fails at create time. A model writing from memory of an older provider has no reason to know it. The pinned documentation does.

The plan-only identity crosses over with a different syntax. The GCS backend documentation asks for Storage Object Admin on the bucket. A roles/viewer account can read state but cannot create the lock object. The answer in gcp/iam/plan-only-role.md is roles/viewer on the project, roles/storage.objectViewer on the bucket, and roles/storage.objectUser on the bucket restricted by an IAM condition to object names ending in .tflock:

PROJECT=atlas-demo-123
BUCKET=atlas-us-central1-bucket-terraform-state
SA=terraform-plan@${PROJECT}.iam.gserviceaccount.com
gcloud iam service-accounts create terraform-plan --project "$PROJECT" \
  --display-name "Terraform plan-only (production)"
gcloud projects add-iam-policy-binding "$PROJECT" \
  --member "serviceAccount:$SA" --role roles/viewer
gcloud storage buckets add-iam-policy-binding "gs://$BUCKET" \
  --member "serviceAccount:$SA" --role roles/storage.objectViewer
gcloud storage buckets add-iam-policy-binding "gs://$BUCKET" \
  --member "serviceAccount:$SA" --role roles/storage.objectUser \
  --condition 'expression=resource.name.endsWith(".tflock"),title=tflock-only'

objectUser carries create, get, list, delete and update on objects and no setIamPolicy; the condition keeps it off the .tfstate objects, and conditions on a bucket require uniform bucket-level access, which the bucket has on. The AWS answer is Put and Delete on *.tflock; the GCP answer is a conditional binding. Same shape, different syntax, and the same second exit, -lock=false, with the same cost. Whoever impersonates the account needs roles/iam.serviceAccountTokenCreator on it; if the backend itself is to impersonate, GOOGLE_BACKEND_IMPERSONATE_SERVICE_ACCOUNT is the switch.

None
What the translation changes and what it does not: on the left, a panel headed "Identical on both clouds" with six rows (numbered stacks with one state each; main.tf holding only terraform and provider; one object variable per domain; this, main or the role as labels; outputs in snake_case with splat; the gates); on the right, a table with AWS and GCP columns and eight rows: state backend, provider identity, global metadata, region in names, the plan-only guardrail with its write allowed only on *.tflock, network, cluster and data.

The size of the result says where the work went. AWS is 30 files, 950 lines of HCL and 32 resource blocks; GCP is 27 files, 823 lines and 18 resources, with 17 outputs on each side. The resource count halves for structural reasons, since S3 needs four resources for what a GCS bucket does with arguments and AWS networking needs subnets, an internet gateway, route tables and associations that GCP expresses as one subnet plus Cloud NAT. The outputs match one for one, which is what lets the downstream stacks keep their shape.

In short, the translation changes 24 lines of the always-on file and four lines of hook messages, then goes stack by stack with the AWS twin open beside each one. Eight of the thirteen conventions are untouched, three swap argument names, and two are different mechanisms. The lock exception, the fail-closed lookup and the workspace on remote state cross over unchanged, because they were never about AWS.

What the Two Floors Cost

/infracost:scan runs after every plan in the loop. It also runs on the code alone, before any credentials exist. Pointed at the two directories, infracost scan aws/terraform and infracost scan gcp/terraform, it prices the AWS floor in us-east-1 at about $148 a month and the GCP floor in us-central1 at about $111, on demand, at 730 hours. On AWS the EKS control plane is $73, the NAT gateway $33, the two t4g.small nodes $29 and the db.t4g.micro instance $14, with pennies of S3. On GCP the regional GKE fee is $73, the two e2-small nodes $28 and the db-f1-micro instance $9. Both figures leave out traffic, NAT data processing and egress, which Infracost prices only from a usage file, so treat them as the minimum. Both sit far under the $500 per-resource alert the cost reviewer enforces.

Read the two totals as a shape rather than a verdict. A control plane costs the same on both sides. The AWS NAT gateway alone costs about what the two GKE nodes cost, while Cloud NAT is billed per VM and per gigabyte processed, a couple of dollars a month for two nodes, which is why the scan shows it only once a usage file supplies traffic. The scan also returns policy findings next to the prices, and those go to the reviewers first. On AWS the useful ones are an S3 bucket policy that denies non-TLS requests, lifecycle rules for non-current versions and incomplete multipart uploads, CloudWatch log exports on RDS and copying tags to snapshots; on GCP, the equivalent lifecycle rules on both buckets. The security reviewer takes the first, the cost reviewer the lifecycle ones, and what lands in a stack goes through the same loop as everything else.

Taking the Setup to Your Own Estate, or to a Third Cloud

Starting from zero on your own account is the AWS section with your values in the bootstrap block. Replace the account id, bucket names and role ARNs in CLAUDE.md and in each stack's variables.tf defaults, keep the numbered layout, and let /tf-scaffold-stack create the next stack. The setup carries the conventions; the stacks in the repository are an example of them, and you can delete all four and keep the setup.

Taking it to a cloud the repository does not cover follows the GCP section. Four things change. The backend block, with whatever locking that backend does natively. The provider identity, as a workspace map with a lookup that falls back to the read-only identity. The global metadata, tags or labels, set once in the provider. And the plan-only identity, which is always the same design: read everything, write only the lock marker, and deny by default the rest. Everything else in the setup stays as it is. Rewrite those lines in CLAUDE.md and mark them, swap the Context7 library id and the secret patterns in the Stop hook, and hand the agent the finished stacks of the cloud you already run as the spec for the new one.

Run It Yourself

Everything is in the companion repository, sized for one person, one AWS account and one GCP project.

Step one costs nothing and needs no credentials. Clone the repository and run the static gate, which checks formatting, initialises and validates all eight stacks without a backend, and runs shellcheck on the hooks and jq on every JSON file:

git clone https://github.com/filipemotta/terraform-ai-aws-gcp && cd terraform-ai-aws-gcp
bash validate.sh

It needs Terraform 1.10 or later and ends with 0 failure(s).

Step two is the MCP setup from the first section, then claude mcp list.

Step three is where the setup becomes yours. Open Claude Code with aws/ (or gcp/) as the project root, so that its CLAUDE.md, .claude/ and .mcp.json apply. Replace the placeholders in the bootstrap block and in each stack's variables.tf defaults (account 123456789012 or project atlas-demo-123, the two identities, the bucket names), and ask for /tf-scaffold-stack on a stack of your own, 04-observability say. Watch the loop run with the hook doing the validating and the guard doing the refusing, and run /infracost:scan to price what you have.

Step four needs credentials and the two identities the provider blocks assume, terraform-apply and terraform-plan-readonly on AWS (the policies are in aws/iam/), the terraform-apply@ and terraform-plan@ service accounts on GCP (the commands are in gcp/iam/plan-only-role.md). The bootstrap stack is applied once without its backend and then migrated into the bucket it created:

cd aws/terraform/00-remote-backend        # or gcp/terraform/00-remote-backend
# first apply only: comment out the backend block (see the stack README)
terraform workspace new sandbox
terraform init
terraform plan -out tfplan               # the agent stops here
terraform apply tfplan                   # a human runs this
# then uncomment the backend block and: terraform init -migrate-state

Then 01-networking, 02-eks or 02-gke, and 03-data, in order, each with terraform workspace select sandbox && terraform init && terraform plan -out tfplan from the agent and terraform apply tfplan from you.

Step five proves the guardrails. Switch a stack to the production workspace and ask the agent for an apply, to watch the hook refuse it; then run the apply yourself from another shell and watch the cloud refuse it too. That second refusal is the one that matters.

Step six is the teardown. While the floor exists it costs what the scan said, so remove it when you are done, in reverse order, with the state migrated back to local before the last destroy:

cd aws/terraform/03-data && terraform workspace select sandbox && terraform destroy
cd ../02-eks && terraform destroy
cd ../01-networking && terraform destroy
cd ../00-remote-backend && terraform init -migrate-state   # move state back to local first
# (comment out the backend block, answer yes), then: terraform destroy

deletion_protection is off and skip_final_snapshot is on for the sandbox database objects on purpose. The hooks block terraform destroy without -target for the agent, so those commands are yours to type.

Run the loop once on a stack of your own and the shape of the thing stays with you. The agent does not get better at Terraform from one day to the next. The repository does. A CLAUDE.md short enough to stay sharp, skills and subagents that own the conventions, a hook that reads stdin and says no, a role that says no again at the API, and provider documentation fetched for the version you pinned instead of the version the model remembers. Change the cloud and two dozen lines of the always-on file move.

The conventions are the prompt. The cloud is a parameter.

Filipe Motta is a Senior DevOps Engineer with 20 plus years of experience, including infrastructure-as-code at scale on AWS and GCP. He has earned the CKA, CKAD, CKS, AWS DevOps Professional, and GCP Professional DevOps Engineer certifications.