DevOps

Fix: Terraform state lock stuck after crash

Terraform's Error acquiring the state lock blocks every apply. Here's why it happens, how to force-unlock safely, and how to stop it recurring.

Terraform refuses to run and prints:

Error: Error acquiring the state lock

Lock Info:
  ID:        c9a8f2d1-7b3e-4a1c-9f2e-8d5b3a1c9f2e
  Path:      terraform.tfstate
  Operation: OperationTypeApply
  Who:       runner@ci-node-4
  Created:   2026-09-22 08:14:32 UTC

The state file is locked and every subsequent plan or apply fails until the lock is released. Here’s the diagnosis path, the safe fix, and prevention.

Why the lock exists

Terraform locks the state file at the start of a mutating operation (apply, destroy, import, state mv) so two processes can’t corrupt it. When the operation ends cleanly the lock is released. When it doesn’t — process killed, network dropped, CI runner terminated — the lock stays behind.

Storage backends that support locking:

  • S3 with native lockfile (use_lockfile = true, Terraform 1.11+) — the current recommended approach, no DynamoDB table needed
  • S3 + DynamoDB (via dynamodb_table in the backend block) — deprecated by HashiCorp, will be removed in a future minor version
  • Azure Blob (native lease)
  • Google Cloud Storage (native lock)
  • HashiCorp Consul, Postgres, Terraform Cloud (built-in)
  • Local backend (.terraform.tfstate.lock.info file)

Each backend stores the lock differently, but the symptom is identical.

Diagnose before you unlock

Never blind-unlock. First check whether another operation is genuinely running.

S3 + DynamoDB backend:

aws dynamodb get-item \
  --table-name terraform-locks \
  --key '{"LockID":{"S":"my-bucket/path/to/state-md5"}}'

Fields to inspect:

  • Who — user or CI runner that held the lock
  • Created — how long ago
  • Operation — apply, plan, or destroy
  • Info — free-form message some workflows set

Local backend:

cat .terraform.tfstate.lock.info

Any backend — check running processes:

ps aux | grep terraform

If the process is still running, wait. The lock is legitimate. Killing another engineer’s apply mid-flight corrupts state.

Cause 1 — CI job was terminated mid-apply

Most common cause. GitHub Actions job canceled, GitLab runner OOMed, Jenkins node rebooted. The Terraform process died before releasing the lock.

Fix — force unlock with the ID from the error message:

terraform force-unlock c9a8f2d1-7b3e-4a1c-9f2e-8d5b3a1c9f2e

Confirm at the prompt (yes). The lock is removed instantly.

Sanity check afterward:

terraform plan

If plan runs cleanly the state was not corrupted. If plan errors with state snapshot was created by Terraform vX.Y or shows massive drift, the state may need manual repair — jump to Cause 4.

Cause 2 — Another engineer really is applying

Two engineers hit apply simultaneously. The second sees the lock error. This is the lock working correctly.

Fix: message the person in Who. Wait for them to finish. Do not force-unlock.

Prevention section below covers how to route all applies through a single CI queue so this never happens.

Cause 3 — DynamoDB entry orphaned but no active process

Terraform crashed so hard it couldn’t clean up, but the DynamoDB row is stuck.

Fix — delete the row directly:

aws dynamodb delete-item \
  --table-name terraform-locks \
  --key '{"LockID":{"S":"my-bucket/path/to/state-md5"}}'

This is equivalent to force-unlock but works when Terraform itself can’t reach the backend (e.g., you changed AWS credentials and the current session isn’t the one that took the lock).

Cause 4 — State file corruption after force-unlock

Force-unlocking is safe if no writer was active. If a writer was mid-write when it died, the state may be truncated.

Symptoms after unlock:

  • plan shows every resource as needing recreation
  • Errors like Failed to load state: state snapshot was created by Terraform v1.5.0
  • Resource IDs missing from state that exist in the cloud

Fix: restore from S3 versioning if enabled.

aws s3api list-object-versions \
  --bucket my-terraform-state \
  --prefix path/to/terraform.tfstate

aws s3api get-object \
  --bucket my-terraform-state \
  --key path/to/terraform.tfstate \
  --version-id <last-good-version-id> \
  terraform.tfstate.recovered

# Inspect, then upload as current
aws s3 cp terraform.tfstate.recovered \
  s3://my-terraform-state/path/to/terraform.tfstate

Then terraform state pull to verify it loads cleanly.

Cause 5 — Wrong backend selected

Ran terraform apply from the wrong directory or on the wrong workspace. The lock is real but for a different state you don’t care about.

Fix: check which state you’re actually locking:

terraform state pull | head -20

If it’s not the state you intended, cd to the correct directory or run terraform workspace select <name> and retry.

Cause 6 — Consul or Postgres lock held by dead session

For non-S3 backends the lock lives in a database session that may not clean up on connection drop.

Postgres backend:

SELECT pid, state, query FROM pg_locks
JOIN pg_stat_activity USING (pid)
WHERE relation = (SELECT oid FROM pg_class WHERE relname = 'terraform_lock');

Kill the offending PID with SELECT pg_terminate_backend(<pid>); if the session is stale.

Consul backend:

consul kv get terraform/lock
consul kv delete terraform/lock

Prevention — route every apply through one queue

The right fix is to make concurrent applies impossible. Options in order of investment:

1. Terraform Cloud / Terraform Enterprise (zero-code): built-in workspace queue. Every apply serializes automatically.

2. Atlantis (self-hosted GitOps): applies triggered by PR comments, one at a time per workspace. Free.

3. Single CI job with concurrency lock: in GitHub Actions:

concurrency:
  group: terraform-prod
  cancel-in-progress: false

This makes a second apply wait for the first to finish rather than fail.

4. State versioning + backups:

resource "aws_s3_bucket_versioning" "state" {
  bucket = aws_s3_bucket.state.id
  versioning_configuration {
    status = "Enabled"
  }
}

Any force-unlock that goes wrong can be rolled back to the previous state object.

The universal recovery flow

Every stuck-lock incident, in order:

# 1. Read the lock info from the error
# 2. Verify no live process holds it
aws dynamodb get-item --table-name terraform-locks --key '{"LockID":{"S":"..."}}'
ps aux | grep terraform

# 3. Message the "Who" — wait if legit
# 4. If orphaned, force-unlock
terraform force-unlock <ID>

# 5. Sanity check
terraform plan

# 6. If plan is wrong, restore state from S3 versioning

Ninety percent of stuck locks resolve at step 4.

When force-unlock is dangerous

Never force-unlock when:

  • The Who field shows a live human currently online
  • The Created timestamp is under 10 minutes for a large apply (some deploys legitimately take 30+ minutes)
  • A CI job is still running in the Actions/Pipelines UI
  • You do not have S3 versioning enabled and can’t restore if state breaks

If any of these are true, wait or ask before unlocking.

Bottom line

Error acquiring the state lock means Terraform is refusing to let two writers collide. In most cases the writer is already dead and terraform force-unlock <ID> clears it in one command. Enable S3 versioning so any mistake is reversible, and route applies through a single queue (Terraform Cloud, Atlantis, or CI concurrency groups) so the lock never legitimately blocks anyone in the first place.

Skip the boilerplate

DevOps YAML Pack

40+ production-ready configs — Kubernetes, Docker, Terraform, Ansible, Helm, GitHub Actions. Every file commented. Copy, edit two lines, ship. MIT license, no watermarks.

Get the pack — ₹499 →