When an update fails
What a failed install looks like, what state your deployment is left in, and what is kept for inspection.
When an install does not complete, the Updates page shows The install did not complete, followed by the failing step in parentheses and the error message, with a Dismiss button:
The install did not complete (health_check_blue): ...
Dismiss it when you have noted the details. You can then check for updates and try again.
What state you are left in#
It depends on how far the install got.
Before the switchover#
These steps run while the current version keeps serving. If any of them fails, your running deployment is untouched:
| Failing step | Meaning |
|---|---|
begin |
Another install is still active for the company, or the run could not be recorded. |
backup |
The database dump failed, was not readable, or took longer than an hour. Nothing changed. |
pull |
An image could not be pulled by its digest. Check Registry credentials for updates. |
start_blue |
The new app container could not be created. |
health_check_blue |
The new app did not become healthy within 120 seconds, or exited or kept restarting. A common cause is a release that needs a stack change you have not applied: see What an update changes. |
launch_cutover |
The helper container could not be started. |
On a failed health_check_blue:
- The new (
-blue)appcontainer is left running, not removed, so you can read its log withsudo docker logs <container name>. The container name is<app container>-blue; find it withsudo docker ps -a. - The images that were pulled stay on the box.
- The current
appkeeps serving. The newapphas already applied its database migrations, so the database may be ahead of the version still serving.
Nothing cleans these up for you. Remove the blue container yourself when you are done inspecting it. A later install replaces a leftover blue container on its own.
During the switchover#
If one of the switchover steps fails (stop_old_app through health_check_web), the
containers can be mid-replacement. If replacing web fails after the old one was stopped, the
helper tries to put the previous web image back, and the message includes whether that
succeeded. Check the state of your containers:
sudo docker compose -p wi-prod ps -a
The helper also records a failure if it receives a stop signal part-way through.
What is kept#
- The database dump
pre-update-<timestamp>.dumpin/data/backupsinside theappcontainer, which is thewi-prod-app-datavolume. It is never deleted automatically. - The previous images, tagged
pre-update-<timestamp>. - The blue container after a failed health check.
- The helper container that ran the switchover, named
wi-prod-update-cutover-<first 8 characters of the run id>. It is not removed, and its log (sudo docker logs <helper name>) holds the detail behind a failed switchover. Find it withsudo docker ps -a.
The Updates page does not display the dump's path. The path follows the pattern above, and the newest file in that directory is the one the latest install made. To list it:
sudo docker compose -p wi-prod exec app ls -l /data/backups
Restoring#
The dump is a standard PostgreSQL custom-format file, verified readable with pg_restore when
it was written. The product does not include a restore command. Restoring a dump is a
database-administration task that you carry out with PostgreSQL's own tools against the
postgres container. Keep copies of the dump you need outside the volume, along with your
.env. See Backups and recovery on AWS
for the volume-level protection of your deployment.
Other things that look like failures#
- Check for updates fails immediately. See the error table in Checking for and installing updates.
- The page shows "Waiting for the app to come back online…" during the switchover. This is
normal. It clears when the new
appanswers. - The page shows a red error with a Dismiss button after the connection dropped for a while. Dismiss it and reopen the page. The run's outcome is recorded on the server.