Skyline

Postmortem: A supervisor Package Upgrade Left Two Horizons Running

· 20 min read · Boring Observability

Verified against Laravel 12.68 · Horizon 5.x · supervisor 4.2.5

Nobody deployed anything. A routine supervisor package upgrade restarted supervisord, and when it came back a complete set of Horizon workers was still on the box, draining the same queues, with no supervisor above them and no process manager that knew they existed. Nothing errored and no alert fired.

Two defaults produced it, neither wrong on its own: Horizon's fast_termination, which lets the master report a shutdown it has not finished, and the KillMode=process that Debian and Ubuntu ship in the supervisor systemd unit, which means nobody above it checks. We kept fast_termination on. The fix is on the systemd side, and the reasoning for that is in its own section below.

Key takeaways#

  • Upgrading the supervisor package restarts it. The Debian postinst runs deb-systemd-invoke restart supervisor.service, a full unit stop and start. With unattended-upgrades on, that happens on a day nobody shipped.
  • fast_termination => true made the master exit in 2.0 seconds with jobs 8 seconds into a 45-second run. supervisord logged stopped: horizon (exit status 0) and exited. The workers were still running.
  • KillMode=process meant systemd never cleaned them up. For a moment the box had no master, no supervisord, and live workers owned by PID 1. Then a new supervisord started a second full tree.
  • Our safeguards were pointed at a different failure. stopwaitsecs=3600 protects a master that needs more time than it is given; ours took less time than it should have. horizon:purge would have caught it and was not scheduled.
  • The fix is one drop-in file and a schedule entry. KillMode=mixed with a TimeoutStopSec above your worst-case drain makes the unit stop wait for the workers and then kill them, and horizon:purge every minute reconciles whatever escapes. fast_termination stays on.

What happened#

The application runs Horizon under supervisord on a Debian host, in an ordinary arrangement: one [program:horizon] block, autorestart=true, a generous stopwaitsecs, and a handful of supervisors with a 60-second worker timeout.

The supervisor package was upgraded, and its postinst restarted the service. Running systemctl restart supervisor by hand does the same thing.

When it came back up, the queues were being worked by two populations of processes: the fresh tree supervisord had just started, and the previous generation's workers, still finishing the jobs they had been holding when the upgrade landed.

Impact#

Nothing failed. Every layer reported success: apt exited zero, systemd marked the unit active, supervisord logged stopped then spawned, and Horizon's dashboard showed a healthy master with the expected supervisors. The exposure is entirely in the overlap.

  • Worker concurrency doubles for the duration. Whatever your maxProcesses figures were tuned against (CPU count, connection pool ceiling, memory headroom) is briefly served by twice as many processes.
  • Rate limits and third-party integrations see double the traffic. A queue deliberately pinned to a single worker so an upstream API is never called concurrently is, for that window, called concurrently. Limiter state keyed per worker does not help, because every worker exists twice.
  • The orphans are uncommandable. Horizon routes pause, continue, terminate and stop-job through the owning supervisor's Redis command list. The orphans name a supervisor that no longer exists, so nobody consumes those commands. A stop-job issued from the dashboard against one of these workers sits in Redis forever while the job runs to completion, and the UI reports success.
  • They run the previous release's code. No deploy was involved here, so this was not a factor for us. The same mechanism during a deploy leaves old code executing against a freshly migrated schema.

In the normal case the overlap is self-limiting. The orphans were signalled, so each finishes its current job and exits, and the window is your longest in-flight job. In the failure mode described in the second branch below, it is not self-limiting at all.

Root cause#

Two defaults, set by two different projects. The project boundary is most of why nobody had looked at them together.

What fast_termination skips#

The whole feature is one method. When a supervisor is told to terminate it signals its workers, then decides whether to stay and watch:

// Laravel\Horizon\Supervisor::terminate()
$this->processPools->each(function ($pool) {
    $pool->processes()->each(function ($process) {
        $process->terminate();          // SIGTERM to each worker
    });
});

if ($this->shouldWait()) {
    while ($this->processPools->map->runningProcesses()->collapse()->count()) {
        sleep(1);                       // ...the drain loop
    }
}

$this->exit($status);
protected function shouldWait()
{
    return ! config('horizon.fast_termination') ||
        app(CacheFactory::class)->get('horizon:terminate:wait');
}

With the flag on, the drain loop is skipped and the supervisor exits immediately. It does not kill the workers or shorten their jobs; it stops waiting for them, so the exit reports a shutdown that is still in progress.

That is the flag's documented purpose, and on the path it was written for it is fine. A deploy calls horizon:terminate, supervisord respawns a fresh master straight away, and the previous generation's workers drain out alongside it under the same supervisord, which can still see them. Nothing above supervisord participates. The systemd path is different, and that difference is this incident.

Why the shipped unit narrows the kill to one process#

# /usr/lib/systemd/system/supervisor.service  (Debian/Ubuntu, supervisor 4.2.5)
[Service]
ExecStart=/usr/bin/supervisord -n -c /etc/supervisor/supervisord.conf
ExecStop=/usr/bin/supervisorctl $OPTIONS shutdown
ExecReload=/usr/bin/supervisorctl -c /etc/supervisor/supervisord.conf $OPTIONS reload
KillMode=process
Restart=on-failure
RestartSec=50s

KillMode decides which processes systemd signals when a unit stops. The default, control-group, signals every process in the unit's control group, which for this unit would include every Horizon process. KillMode=process narrows that to the main process only: supervisord itself.

For supervisord that is a defensible choice. It is a process manager with its own opinions about how and in what order its children should stop, and having systemd fire SIGKILL into the control group underneath it would defeat every stopwaitsecs in the config. The setting exists so supervisord's shutdown logic stays authoritative.

It is only safe if supervisord's shutdown logic is correct. KillMode=process is systemd delegating cleanup to supervisord. fast_termination is Horizon telling supervisord that cleanup finished when it did not. With both set, systemd does not check for leftover processes, and supervisord has been told there are none.

The worse branch: when ExecStop doesn't complete#

Everything above assumes ExecStop ran to completion. It does not always. If supervisorctl shutdown fails, because the socket is gone or the Python modules were mid-upgrade, or if the stop simply outlives TimeoutStopSec (systemd's default is 90 seconds, and the shipped unit does not override it), systemd escalates to SIGKILL against the main process. With KillMode=process, that SIGKILL goes to supervisord and to nothing else, so supervisord dies without stopping any of its children.

When the new supervisord starts its own tree, the host has two complete Horizon installations, both registered in Redis and both processing the same queues. The old master has PPID 1 and nothing will ever stop it: it was never asked to terminate, and its supervisors and workers carry on as normal. It will still be there tomorrow, and every later supervisor restart adds another one.

This is the branch that produces permanently surviving workers, and turning fast_termination off makes it more likely. With fast_termination off, the drain runs inside ExecStop, and ExecStop is bounded by TimeoutStopSec. A 90-second default against a drain bounded by a 60-second worker timeout plus wind-down is a thin margin, and it shrinks every time someone raises a job's timeout. Turning the flag off without touching the unit trades a self-limiting failure for a permanent one.

Why none of our safeguards fired#

stopwaitsecs=3600 never engaged#

This is the setting everyone reaches for, and it was already set to an hour. It did nothing, because stopwaitsecs is how long supervisord waits before resorting to SIGKILL, and supervisord was never waiting. The master exited voluntarily after two seconds. A generous stopwaitsecs protects against a master that needs more time than it is given. Ours needed less time than it should have taken.

horizon:purge was not scheduled#

Horizon ships a reconciler for this class of problem and we were not running it. That is the single largest process failure in this incident, and it gets its own section below, because using it correctly requires understanding what it can and cannot see.

Remediation#

One drop-in file, one schedule entry, and one change we decided against.

The systemd drop-in#

Do not edit /usr/lib/systemd/system/supervisor.service. The next package upgrade will overwrite it, and a package upgrade is the event we are hardening against. Ship a drop-in:

# /etc/systemd/system/supervisor.service.d/override.conf

[Service]
# The shipped unit sets KillMode=process, so systemd signals only supervisord
# and leaves every Horizon process in the control group running. "mixed" keeps
# the graceful part: SIGTERM still goes to supervisord alone, so its own
# stopwaitsecs logic stays authoritative. What changes is the end of the
# sequence. The unit is not stopped until its control group is empty, and the
# final SIGKILL goes to everything still in it.
KillMode=mixed

# The ceiling on the whole stop, including the wait for the control group to
# drain. Default is 90s. A Horizon drain is bounded by the worker timeout, so
# this must exceed it with room to spare; below it, jobs are SIGKILLed
# mid-flight.
TimeoutStopSec=300
$ systemctl daemon-reload
$ systemctl show supervisor -p KillMode -p TimeoutStopUSec
KillMode=mixed
TimeoutStopUSec=5min

KillMode=mixed is the right setting rather than control-group because it preserves the reason the shipped unit chose process. The initial SIGTERM still goes to supervisord alone, so supervisord remains in charge of the graceful path and every stopwaitsecs is still honoured. Only the last-resort SIGKILL, the point at which graceful shutdown has already failed and there is nothing left to protect, widens to the whole control group.

That second half is what bounds the first branch without touching fast_termination. supervisord still exits two seconds in, and the orphaned workers are still running their jobs with no parent. But they are in the unit's control group, so the stop does not complete and the replacement supervisord does not start until they are gone. The overlap that made this an incident closes on its own, and a worker that would have outlived everything gets killed at TimeoutStopSec instead of surviving indefinitely.

The same mechanism closes the second branch. A supervisord that is SIGKILLed at the timeout no longer takes its tree with it into PID 1's care, because the SIGKILL is delivered to the control group rather than to one process. There is no path left where a fully registered duplicate Horizon survives the unit going away.

What this costs

systemctl restart supervisor now blocks for the drain instead of returning in a couple of seconds, and so does the package upgrade that started all this: apt sits on the postinst until the last worker exits. Set TimeoutStopSec above your worst-case job or the SIGKILL lands mid-flight, and note that the ceiling applies to the whole unit, so anything else supervisord manages is now subject to it too.

Why fast_termination stays on#

The obvious fix is 'fast_termination' => false, and it does close the first branch. supervisord waits out the drain instead of being told it has finished, stopwaitsecs=3600 finally has something to protect, and the worker count never exceeds what was provisioned. We considered it and kept the flag on.

The first reason is the one in the second branch above. A drain that runs inside ExecStop is a drain that can outlive TimeoutStopSec, and the failure at that ceiling is the permanent one. Turning the flag off is only safe once the unit is fixed, and once the unit is fixed the flag no longer affects whether workers survive.

The second is that the flag is useful on the path it was written for. Deploys call horizon:terminate and never involve systemd; the old workers drain under the supervisord that is already running, which can see them and command them. Making every deploy wait out the longest in-flight job to fix a systemd-layer bug puts the cost on the wrong path.

Where old code against a new schema is the concern, the lever is the deploy verb rather than the config flag. php artisan horizon:terminate --wait sets the horizon:terminate:wait cache key that shouldWait() reads, forcing a full drain on that one call while leaving the flag on everywhere else.

Schedule the backstop#

// routes/console.php
Schedule::command('horizon:purge')->everyMinute();

Covered in full below.

The backstop, in detail#

horizon:purge is Horizon's reconciler for orphaned processes, and we should have been running it. It is also narrower than the name suggests. What it catches maps exactly onto the first branch of this incident and not at all onto the second, which is why it is a backstop and not the fix.

How it decides what is an orphan#

The whole judgement is two set operations in ProcessInspector:

// Laravel\Horizon\ProcessInspector

public function current()
{
    return array_diff(
        $this->exec->run('pgrep -f [h]orizon'),        // every Horizon-looking pid
        $this->exec->run('pgrep -f horizon:purge')     // ...except this command itself
    );
}

public function monitoring()
{
    return collect(app(SupervisorRepository::class)->all())
        ->pluck('pid')                                             // registered supervisors
        ->pipe(function ($processes) {
            $processes->each(function ($process) use (&$processes) {
                $processes = $processes->merge($this->exec->run("pgrep -P {$process}"));
            });                                                    // ...and their children
            return $processes;
        })
        ->merge(Arr::pluck(app(MasterSupervisorRepository::class)->all(), 'pid'))
        ->all();                                                   // ...and registered masters
}

public function orphaned()
{
    return array_diff($this->current(), $this->monitoring());
}

In words: everything on the box that looks like Horizon, minus everything the Redis registry can account for. A process is an orphan if no registered supervisor claims it as a child and it is not a registered master. That is the signature of the first branch, workers whose supervisor de-registered and exited, which is why purge would have caught it.

It works in two phases#

// Laravel\Horizon\Console\PurgeCommand

public function purge($master, $signal = SIGTERM)
{
    $this->recordOrphans($master, $signal);          // phase 1

    $expired = $this->processes->orphanedFor(
        $master, $this->supervisors->longestActiveTimeout()
    );

    collect($expired)->each(function ($processId) use ($master, $signal) {
        $this->components->task("Process: $processId", function () use ($processId, $signal) {
            exec("kill -s {$signal} {$processId}");   // phase 2
        });

        $this->processes->forgetOrphans($master, [$processId]);
    });
}

Phase 1 records the current orphan set and signals it immediately. The recording goes into a Redis hash keyed by master name, using HSETNX so the first time a pid was seen as orphaned is preserved across sweeps, and HDELing any pid that is no longer orphaned:

// Laravel\Horizon\Repositories\RedisProcessRepository::orphaned()
$shouldRemove = array_diff($this->connection()->hkeys($key = "{$master}:orphans"), $processIds);

if (! empty($shouldRemove)) {
    $this->connection()->hdel($key, ...$shouldRemove);
}

$this->pipeline(function ($pipe) use ($key, $time, $processIds) {
    foreach ($processIds as $processId) {
        $pipe->hsetnx($key, $processId, $time);      // first observation wins
    }
});

Phase 2 is the escalation. It asks which pids have been on that list longer than longestActiveTimeout(), the maximum timeout across your registered supervisors, and signals those again. The logic is that a worker given a polite SIGTERM should have finished its job and exited within its own timeout; one that has not is stuck rather than busy.

The default signal is SIGTERM for both phases#

--signal defaults to SIGTERM and applies to both phases, so phase 2 sends the same signal as phase 1, and a process that ignored the first SIGTERM will generally ignore the second. A worker blocked in a syscall is the case phase 2 exists for, and a second SIGTERM cannot reach it. PHP dispatches asynchronous signals between VM instructions, so a process parked in a blocking socket read never gets to run its handler. We measured one of these surviving 139 seconds against a 60-second job timeout, ignoring both SIGTERM and the worker's own SIGALRM.

If you want phase 2 to kill the process, say so, and understand that you are killing a job mid-flight:

Schedule::command('horizon:purge --signal=SIGKILL')->everyMinute();

We run the SIGTERM default and alert on the phase-2 output instead, on the grounds that a process needing SIGKILL is one we want to look at rather than discard without a record.

What it cannot do#

It is a no-op when no master is registered. The command iterates registered master names and purges per-master:

foreach ($masters->names() as $master) {
    if (Str::startsWith($master, MasterSupervisor::basename())) {
        $this->purge($master, $signal);
    }
}

No registered master means the loop body never runs. That rules out the placement people reach for first, a pre-start deploy hook along the lines of "clean up before we bring Horizon back", which is the one moment it does nothing. It has to run alongside a live Horizon, from the scheduler, which in turn means your cron entry for schedule:run has to be working. If the scheduler stops running, the backstop stops with it and nothing reports it.

It cannot see a duplicate tree. In the second branch, the old master and its supervisors are all still registered in Redis, so monitoring() accounts for every one of their workers and orphaned() comes back empty. Those processes are not orphans by this definition; they belong to a Horizon that should not exist. Purge is scoped to processes whose owner has disappeared, not to owners that are illegitimate.

The orphan clock resets whenever Horizon restarts. The hash is keyed {master-name}:orphans, and the master name carries a random token. A pid recorded under web-1-Vf68:orphans is invisible to a sweep running under web-1-YzD4. Phase 2's "has been orphaned for longer than the worker timeout" is therefore measured from the current master's lifetime rather than the pid's, so a restart storm can keep resetting the escalation clock on a process that never dies.

pgrep -f [h]orizon matches on the whole command line. Anything with the string "horizon" in its arguments is a candidate for being signalled: a tail -f horizon.log, a script named horizon-deploy.sh, an editor with the config file open, a grep in someone's shell history loop. Only horizon:purge itself is explicitly excluded. On a shared box, or one where operators habitually tail the Horizon log, that is worth a moment's thought before scheduling it with --signal=SIGKILL.

Where it sits in the defence#

Failure Fixed by Caught by horizon:purge?
Workers orphaned by a fast master exit KillMode=mixed plus a real TimeoutStopSec Yes (their supervisor is gone from the registry)
Whole tree survives a supervisord SIGKILL KillMode=mixed No (the tree is fully registered)
Worker wedged in a blocking syscall Job-level timeouts on I/O; the control-group SIGKILL as a floor Only with --signal=SIGKILL
Duplicate masters from a shell-wrapped command Unwrap the command line No (both masters are registered)

horizon:purge covers one of the four, and it happens to be the one we hit first. Schedule it, but do not let it stand in for the unit fix: two of these rows have no backstop at all.

Detection#

This ran undetected because no individual signal was abnormal. What we watch now:

# Count the tree. Compare against what your provisioning actually declares.
pgrep -af 'artisan horizon' | wc -l

# The permanent-leak signature: more than one master token.
redis-cli --scan --pattern 'horizon:master:*'

# The orphan signature: a worker whose parent is init.
ps -eo pid,ppid,etime,cmd | awk '$2 == 1' | grep 'horizon:work'

# Which supervisor does each worker think it belongs to...
pgrep -af 'horizon:work' | grep -o 'supervisor=[^ ]*' | sort | uniq -c

# ...and do those supervisors still exist?
redis-cli --scan --pattern 'horizon:supervisor:*'

Two mismatches are worth alerting on. A --supervisor= value on a running worker with no matching horizon:supervisor:* key is a temporary orphan. More than one distinct master token under horizon:master:* is the permanent leak, and it will not resolve on its own. Both are cheap to check on an interval and neither requires anything to have failed first.

What we took from it#

Package upgrades are deploys. Nothing in our deployment pipeline ran, nobody was on call for it, and the change was to a package we think of as inert infrastructure. Any package whose postinst restarts a service is a production change, and unattended-upgrades makes it an unattended one.

The obvious fix was in the wrong layer. fast_termination is the setting that turns up in every search result for this symptom, and turning it off does stop the workers being orphaned. It also moves the drain inside ExecStop, under a 90-second ceiling nobody had looked at, which is how the recoverable branch becomes the permanent one.

Check what the layer above yours believes. systemd trusted supervisord to clean up its children, supervisord trusted the master's exit to mean the program had stopped, and the master had been configured not to wait long enough to know. Each hand-off was reasonable. Nobody had written down what the layer above was assuming.

Reproduce before you fix. Our first hypothesis was the well-known shell-wrapped command problem, which we did not have. Building a container that reproduced the real timings took under an hour, and it is where the second, permanent branch turned up. We would not have found that by reasoning about it, and the config change we had queued would not have closed it.

If you want the mechanics of Horizon's shutdown cascade in more depth, the troubleshooting guide covers the process tree and the signals that move through it. If your queues are backing up rather than duplicating, that is usually a balancing problem instead. And Skyline's per-queue pausing and dashboard operations exist partly so that routine interventions do not require restarting the process tree.

Frequently asked questions

Why does restarting supervisord leave Laravel Horizon workers running?

supervisord tracks only the master php artisan horizon process; supervisors and workers are Horizon's own children and invisible to it. With fast_termination enabled the master exits within about two seconds of being signalled, while its workers are still finishing jobs, so supervisord considers the program stopped and exits too. On Debian and Ubuntu the supervisor systemd unit sets KillMode=process, which tells systemd to kill only supervisord and leave everything else in its control group running, so nothing cleans the workers up.

Does upgrading the supervisor package restart it?

Yes. The Debian package postinst runs deb-systemd-invoke restart supervisor.service, which is a full unit stop and start, identical to running systemctl restart supervisor by hand. With unattended-upgrades enabled this happens without anyone triggering a deploy, which is why the symptom can appear on a day nothing shipped.

What is KillMode=process and why does it matter for Horizon?

KillMode controls which processes systemd signals when a unit stops. The default, control-group, signals every process in the unit control group. The supervisor package ships KillMode=process, which signals only the main process, supervisord itself. Any Horizon process that outlives supervisord is therefore left running, reparented to PID 1, and outside the control of both systemd and the new supervisord that replaces it.

How do I detect orphaned Horizon workers?

Two signatures. A worker process whose PPID is 1, carrying a --supervisor= argument that names a supervisor with no matching key in Redis, is a temporary orphan. More than one distinct master token under the master:* keys means an entire duplicate Horizon tree is registered and running, and that one will not resolve on its own. Compare pgrep -af for artisan horizon against what your provisioning actually declares.

Should I set fast_termination to false to stop Horizon workers being orphaned?

It closes one branch of the problem and widens another, so fix the systemd unit first. With fast_termination off, the drain runs inside the unit ExecStop, which is bounded by TimeoutStopSec, and the shipped supervisor unit does not override the 90-second default. If the drain outlives that ceiling, systemd SIGKILLs supervisord and KillMode=process leaves an entire registered Horizon tree running under PID 1, which never resolves on its own. A drop-in setting KillMode=mixed and a TimeoutStopSec above your worst-case job makes the unit stop wait for the workers and then kill whatever is left, which bounds the orphans whether the flag is on or off. Deploys can still force a full drain on demand with horizon:terminate --wait.

Does horizon:purge clean up orphaned workers?

It cleans up one kind. horizon:purge compares every Horizon process on the box against the pids the registered supervisors and masters account for, and signals the difference. That catches workers whose owner has disappeared. It cannot catch a duplicate Horizon tree, because that tree has its own registered master and supervisors, so its workers are properly accounted for. It is also a no-op when no master is registered, so it must run alongside a live Horizon rather than as a pre-start hook.