A deploy freeze policy you can copy

When to freeze deploys, who can lift a freeze, how hotfixes get through, and how to make the pipeline enforce it. With a one-page policy template.

· deploys, freezes, release-management

What a freeze is, and what it isn’t

A deploy freeze is a period during which normal releases to an environment are blocked by default. That is the whole definition. It is not a period during which nothing can ship. The difference matters, because a freeze with no hotfix path gets bypassed within a week: someone has a real fix, there is no sanctioned way through, so they deploy anyway and the freeze becomes a suggestion.

A freeze that works has three parts: a clear start and end, a default of “no”, and a short, recorded path for the exceptions. If you only build the first two, you have built the kind that gets ignored.

When a small team actually needs one

Most teams of two to fifty people need two or three freezes a year, not a permanent change-advisory process. The occasions are predictable:

  • Holidays and on-call gaps. The last two weeks of December, the week half the team is at a conference, any stretch where the person who can fix a bad deploy is not around to do it.
  • End of quarter or billing cycle. If your product runs invoices, renewals or reports on a date, the day before is a bad day to change the code that does it.
  • During an incident. While you are working out what broke, a new release makes the question harder. A freeze here is short, often hours, and lifted the moment the incident is closed.
  • Around a migration. Data migrations, infrastructure moves, a database version upgrade. The migration itself ships; other changes wait until it has settled.
  • The week before a launch. Not because the code is finished, but because the people who would notice a regression are busy with the launch.

If you find yourself freezing most weeks, the problem is the deploy process, not the calendar. Fix the pipeline and keep freezes for the occasions above.

The decisions a policy has to make

A policy is a short list of decisions made in advance so that nobody has to make them under pressure. Each one below comes with a recommended default. Change the default if you have a reason; keep the heading either way.

Scope

Which environments and which services the freeze covers. Default: production only, for the services that face customers or move money. Staging stays open; you want to keep testing during a freeze, and a frozen staging environment just moves the backlog. If you have one monolith, the scope is the monolith. If you have twenty services, name the ones that matter and leave the rest open.

Who can declare a freeze

One named role, not “the team”. Default: the tech lead, or whoever is on call at the time. The person declaring sets the end date and writes the one-line reason. Declaring a freeze should take two minutes and should not need a meeting.

Who can lift a freeze

Also one named role, and it should be the same person or their stated backup. Default: whoever declared it, or the tech lead. Lifting early is fine and common; an incident freeze is lifted when the incident closes. What you want to avoid is a freeze that nobody feels they have the standing to end, so it drifts on past its purpose.

Hotfix path

What qualifies, who approves, and how it is recorded. Default:

  • What qualifies: a fix for something currently broken for customers, a fix for a problem with the thing the freeze is protecting (the migration, the invoice run), or a dependency update that addresses a published vulnerability. New features do not qualify, however small.
  • Who approves: one approver, who is not the person who wrote the fix. The approver is whoever can lift the freeze, or anyone they have named as a backup.
  • How it is recorded: the deploy is tagged as a freeze override, with the reason and the approver’s name. The tag is what you look at afterwards.

“Anyone can propose, one person approves, the override is tagged” is the shape. It is deliberately light. If the hotfix path takes longer than the fix, people route around it.

Duration

A fixed end date and time, set when the freeze is declared. Default: the shortest period that covers the reason. An end-of-year freeze ends on the first working day of January. An incident freeze ends when the incident is closed, with a hard stop of 24 hours after which someone has to actively extend it.

Open-ended freezes do not end. They fade. Three weeks later someone asks in a channel whether the freeze is still on and nobody is sure. The end date is the one field that is not optional.

Communication

Where the freeze is announced, and what the pipeline says when it rejects a deploy. Default: one message in the channel where deploys are already discussed, stating the scope, the end date, the reason and who can approve a hotfix. Pin it. The same message goes out when the freeze is lifted.

The pipeline rejection is the more important of the two. When a deploy is blocked, the message the developer sees should say that production is frozen, until when, why, and who to ask for an override. A bare “deploy failed” sends them to the logs and then to the channel to ask what is going on. Put the answer in the failure.

Enforcement

A freeze that lives in a Slack message is not a freeze. It is a request. People miss the message, or read it and forget, or merge on a Friday without thinking. The freeze has to live where the deploy happens: in the pipeline, as a step that runs before anything reaches production and fails when a freeze is active.

The pattern is small. Keep the freeze state somewhere the pipeline can read: a status endpoint, a flag file in a bucket, a variable in the CI system. The step before the production deploy checks it and exits non-zero with a clear message if the freeze is on. An override is a second variable or input that the approver sets for one run, which the step logs and lets through.

As a GitHub Actions job, generically:

# Runs before the production deploy job. Fails with a readable message
# while a freeze is active, unless this run carries an approved override.
check-freeze:
  runs-on: ubuntu-latest
  steps:
    - name: Check deploy freeze
      env:
        FREEZE_URL: ${{ vars.FREEZE_STATUS_URL }}
        OVERRIDE_REASON: ${{ inputs.freeze_override_reason }}
        OVERRIDE_APPROVER: ${{ inputs.freeze_override_approver }}
      run: |
        status=$(curl -fsS "$FREEZE_URL")
        active=$(echo "$status" | jq -r '.active')
        if [ "$active" != "true" ]; then
          echo "No freeze on production. Continuing."
          exit 0
        fi
        until=$(echo "$status" | jq -r '.until')
        reason=$(echo "$status" | jq -r '.reason')
        if [ -n "$OVERRIDE_REASON" ] && [ -n "$OVERRIDE_APPROVER" ]; then
          echo "Freeze override: $OVERRIDE_REASON (approved by $OVERRIDE_APPROVER)"
          exit 0
        fi
        echo "::error::Production is frozen until $until: $reason."
        echo "::error::Hotfix? Re-run with freeze_override_reason and freeze_override_approver set."
        exit 1

deploy-production:
  needs: check-freeze
  # ...

The same shape in GitLab CI, reading a flag file instead of an endpoint:

check-freeze:
  stage: pre-deploy
  script:
    - |
      if [ -f freeze/production.json ]; then
        until=$(jq -r '.until' freeze/production.json)
        reason=$(jq -r '.reason' freeze/production.json)
        if [ -n "$FREEZE_OVERRIDE_REASON" ] && [ -n "$FREEZE_OVERRIDE_APPROVER" ]; then
          echo "Freeze override: $FREEZE_OVERRIDE_REASON (approved by $FREEZE_OVERRIDE_APPROVER)"
          exit 0
        fi
        echo "Production is frozen until $until: $reason."
        echo "Hotfix? Set FREEZE_OVERRIDE_REASON and FREEZE_OVERRIDE_APPROVER on the pipeline."
        exit 1
      fi
      echo "No freeze on production."

Two details worth keeping. The override needs both a reason and an approver, so a single variable can’t be left set from last time. And the check fails closed when the status can’t be read: in the GitHub example curl -fsS exits with code 22 on a 4xx or 5xx response and with another non-zero code when it can’t connect, so a broken status endpoint blocks deploys rather than silently allowing them. Decide whether that is what you want. For a small team it usually is, because a status endpoint that is down is a thing you want to find out about.

Where the freeze state lives is up to you. A JSON file in a repository with a protected branch is enough to start, and lets the person lifting the freeze do it with a commit.

The record

Each release attempted during a freeze window, whether it went through as an override or was rejected, should be listed somewhere afterwards. Not for blame. The list is what tells you whether the freeze was the right length and the right scope.

If the list shows six overrides in two weeks, the scope was too wide or the freeze too long. If it shows nothing, the freeze may have been unnecessary, or it worked, and the pipeline logs will tell you which. If it shows three rejections of the same change, someone needed a hotfix path and didn’t know it existed.

The pipeline already has this information. Overrides are tagged, rejections are failed jobs. The work is to pull them into one place and look at them once, in the week after the freeze ends. That is what makes the next freeze shorter.

Template: the one-page policy

Copy this into your team’s handbook. The defaults are filled in; change the names and anything in angle brackets.

# Deploy freeze policy

## Scope

Production only. Covers <list the services, or "all services">.
Staging and preview environments stay open during a freeze.

## Who declares

<Tech lead name> or the on-call engineer. Declaring a freeze means:
setting the end date and time, writing a one-line reason, and posting
the announcement (below). No meeting is required.

## Who lifts

Whoever declared it, or <tech lead name>. A freeze can be lifted early
at any time. Lifting means posting the announcement and clearing the
freeze status the pipeline reads.

## Hotfix path

Qualifies: a fix for something currently broken for customers, a fix
for the thing the freeze is protecting, or a dependency update for a
published vulnerability. New features do not qualify.

Anyone can propose a hotfix. One approver, who did not write the fix:
<tech lead name> or <backup name>. The deploy is tagged as a freeze
override with the reason and the approver's name.

## Default duration

A fixed end date and time, set when the freeze is declared.
Incident freezes end when the incident is closed, with a hard stop of
24 hours unless actively extended.

Standing freezes this year:

- <YYYY-MM-DD> to <YYYY-MM-DD>: end of year
- <YYYY-MM-DD> to <YYYY-MM-DD>: <reason>

## Announcement

Posted in <#channel> when declared and when lifted, and pinned while
active. States: scope, end date and time, reason, who approves hotfixes.

## Enforcement

The <pipeline name> pipeline checks freeze status before any production
deploy and fails with the end date, the reason and the hotfix
instructions. Overrides require both a reason and an approver name.

## Record

Within a week of a freeze ending, <owner> lists the releases attempted
during the window (overrides and rejections) in <place>, and notes
whether the scope or duration should change next time.

_Last reviewed: <YYYY-MM-DD>_

Where to start

Schedule your next freeze now. End of year is the obvious one, and the date is not going to move. Put it in the policy, set up the pipeline check, and then run the whole thing once at a quiet moment before you need it: declare a one-hour freeze on a weekday afternoon, try a deploy and read the failure message, push a hotfix through the override, lift the freeze, and look at the record. The first run will show you which field in the template you got wrong. Better to find that out in an hour you chose than a week you didn’t.