Skyline

The Life of a Laravel Queued Job: Every State, Every Transition

· Updated · 24 min read · Boring Observability

Verified against Laravel 13 · Horizon 5.x

A Laravel job spends almost none of its life as an object. It is an object for the microseconds between new SendInvoice($order) and dispatch(), and again for the milliseconds between a worker deserialising it and handle() returning. The rest of the time, while it waits, is reserved, is retried or fails, it is a JSON string in one of three Redis keys.

This article follows a job through all of that in order: what dispatch() writes, how a worker takes exactly one copy of it, why the attempt is spent at that moment rather than when something goes wrong, what the reservation guarantees when a worker is killed mid-job, and how delayed, retried, exhausted and rate-limited jobs, and deploys, move through the same small set of transitions. Each step is shown with the framework code that implements it.

Key takeaways#

  • One queue is four Redis keys. A ready list, a delayed sorted set, a reserved sorted set, and a notify list. Every state change is a move between them.
  • Exactly one worker gets each job because the pop is a Lua script. Redis runs it atomically: the lpop and the write to the reserved set cannot interleave with another worker.
  • The attempt is spent at pop, not at failure. The same Lua script increments attempts, so anything that causes a second pop (a release, a timeout, an expired lease) costs a try even if handle() never ran.
  • Nothing deletes a job until it succeeds. Completion is the only path that removes the reserved entry. That is why a killed worker loses no work, and why every job must be idempotent.
  • A restart waits for the running job, up to a limit. SIGTERM lets the current job finish; a job still running after the supervisor's timeout is killed, and recovered later by its expired reservation.

One queue, four Redis keys#

A single queue named emails on the Redis driver is four keys. Three hold jobs, and the fourth exists so a worker can sleep instead of polling. A job's state is whichever of these keys it is in:

Key Type Holds
queues:emails List Jobs ready to run, oldest at the head. Pushed with rpush, taken with lpop (FIFO).
queues:emails:delayed Sorted set Jobs that must not run yet. Score is the Unix timestamp they become available.
queues:emails:reserved Sorted set Jobs a worker is holding. Score is when that hold expires: now + retry_after.
queues:emails:notify List One token per available job, so a worker can block on blpop instead of polling.

There is no "running" key, no "processing" flag, and nothing that records which worker holds a job. The reserved set stores a payload and a deadline, and that explains most of the behaviour in the rest of this article.

Dispatch: what gets written#

dispatch(new SendInvoice($order)) ends in RedisQueue::push(), which builds a payload array and hands it to a two-command Lua script. Everything the worker later decides is decided from these fields, not from the job class:

// Illuminate\Queue\Queue::createObjectPayload()
[
    'uuid' => (string) Str::uuid(),
    'displayName' => 'App\Jobs\SendInvoice',
    'job' => 'Illuminate\Queue\CallQueuedHandler@call',
    'maxTries' => 3,          // $tries on the job, or #[Tries]
    'maxExceptions' => null,
    'failOnTimeout' => false,
    'backoff' => '30,120',    // $backoff, flattened to a string
    'timeout' => null,        // $timeout on the job, or #[Timeout]
    'retryUntil' => null,     // retryUntil() resolved to a timestamp, now
    'data' => ['commandName' => ..., 'command' => 'O:21:"App\Jobs\SendInvoice"...'],
    'createdAt' => 1756...,
    'id' => 'ZK3v...',        // 32 random chars, added by the Redis driver
    'attempts' => 0,          // added by the Redis driver
]

Two of those fields are fixed at dispatch. retryUntil is resolved here: your retryUntil() method is called at dispatch time and the resulting timestamp is stored, so the window starts when the job is queued, not when it first runs. And command is the serialised job object, which is why renaming or restructuring a job class breaks everything already in the queue (covered in more depth here).

The write itself is two commands:

-- Illuminate\Queue\LuaScripts::push()
-- Push the job onto the queue...
redis.call('rpush', KEYS[1], ARGV[1])
-- Push a notification onto the "notify" queue...
redis.call('rpush', KEYS[2], 1)

rpush onto the tail, and a token onto the notify list. The notify list is what block_for uses: a worker with nothing to do calls blpop on it and Redis wakes it as soon as a job arrives, instead of the worker sleeping for a fixed interval and finding the job late.

Two things that can happen before the write#

The transaction gate. If the connection has after_commit set, the job is ShouldQueueAfterCommit, or it was dispatched with ->afterCommit(), the push is registered as a callback on the current database transaction and runs on commit. If the transaction rolls back, the job is never queued. Without the gate, a worker can pop the job before the transaction commits and find no row (how dispatch after commit behaves with nested transactions and retries).

The uniqueness gate. A ShouldBeUnique job acquires its lock before the push. If the lock is held, the dispatch does nothing: no payload is written, no log line appears, and the dashboard shows nothing. This check happens at dispatch time and is a separate mechanism from the runtime locks discussed later (how ShouldBeUnique and WithoutOverlapping each fail).

Delayed jobs live in a separate sorted set#

dispatch(new SendInvoice($order))->delay(now()->addMinutes(10)) takes a different path and never touches the ready list:

-- Illuminate\Queue\LuaScripts::later()
-- Push the job onto the delayed queue...
redis.call('zadd', KEYS[1], ARGV[1], ARGV[2])

One zadd into queues:emails:delayed, scored with the timestamp it becomes available. No token goes onto the notify list, because there is nothing to wake a worker for yet.

Nothing polls that sorted set on a timer. A delayed job becomes available only when some worker pops that queue, because the first thing a pop does is sweep the delayed set for anything whose score has passed:

// Illuminate\Queue\RedisQueue
protected function migrate($queue)
{
    $this->migrateExpiredJobs($queue.':delayed', $queue);

    if (! is_null($this->retryAfter)) {
        $this->migrateExpiredJobs($queue.':reserved', $queue);
    }
}

migrateExpiredJobs() is a Lua script that does a zrangebyscore up to "now", zremrangebyranks what it found, and rpushes it onto the ready list with a notify token per job. Two consequences follow, and the rest of this article depends on both:

  • A delayed job runs after its due time, not at it. Its availability is discovered on the next pop of that queue. A queue with no workers running never migrates anything; the jobs are all still in the delayed set, and nothing moves them until a worker pops.
  • It rejoins at the back. rpush puts it at the tail of the ready list, behind everything already waiting, so a delay gives a job no priority. Putting a job at the head of the line is front-of-queue dispatching, a separate operation.

The same script, on the same schedule, recovers reserved jobs whose lease has expired. Delayed jobs and abandoned jobs return to the ready list by the same route.

The worker loop#

A worker is a while (true) loop. Stripped to the transitions that matter:

// Illuminate\Queue\Worker::daemon(), condensed
while (true) {
    if (! $this->daemonShouldRun($options, $connectionName, $queue)) {
        // maintenance mode, or paused by SIGUSR2
        [$status, $reason] = $this->pauseWorker($options, $lastRestart, $startTime);
        // ...
        continue;
    }

    $job = $this->getNextJob($this->manager->connection($connectionName), $queue);

    if ($supportsAsyncSignals) {
        $this->registerTimeoutHandler($job, $options);   // pcntl_alarm() armed here
    }

    if ($job) {
        $this->jobsProcessed++;
        $this->runJob($job, $connectionName, $options);
        // ...
    } else {
        $this->events->dispatch(new WorkerIdle($connectionName, $queue, $options));
        $this->sleep($options->sleep);
    }

    if ($supportsAsyncSignals) {
        $this->resetTimeoutHandler();                    // pcntl_alarm(0)
    }

    [$status, $reason] = $this->stopIfNecessary($options, $lastRestart, $startTime, $job);

    if (! is_null($status)) {
        return $this->stop($status, $options, $reason);
    }
}

Every iteration does four things in this order: decide whether to run at all, take a job, run it, then decide whether to keep running. The last check, stopIfNecessary(), is what lets a deploy stop workers without losing the job in hand, and we come back to it in the restart section. It runs after the job, so a worker never abandons a job in order to exit.

When multiple queues are listed, getNextJob() walks them in order and returns the first job it finds, which is why balance => false gives you strict priority and can starve the queues at the end of the list (the trade-off in full, and Skyline's weighted alternative).

The pop, and why exactly one worker wins#

Popping a job is one Lua script, and Redis executes Lua scripts atomically on its single command thread, so the read and the write cannot be separated:

-- Illuminate\Queue\LuaScripts::pop()
-- Pop the first job off of the queue...
local job = redis.call('lpop', KEYS[1])
local reserved = false

if(job ~= false) then
    -- Increment the attempt count and place job on the reserved queue...
    reserved = cjson.decode(job)
    reserved['attempts'] = reserved['attempts'] + 1
    reserved = cjson.encode(reserved)
    redis.call('zadd', KEYS[2], ARGV[1], reserved)
    redis.call('lpop', KEYS[3])
end

return {job, reserved}

This guarantee needs no lock. Twenty workers can call the script on the same queue in the same millisecond; Redis runs them one after another, and lpop removes what it returns, so the first script takes the payload and the other nineteen get false. No other command can run between the lpop and the zadd, so two workers never hold the same ready-list entry. This part of the lifecycle comes from Redis rather than Laravel, and no configuration setting can break it.

ARGV[1] is now + retry_after: the score that later decides whether this job is considered abandoned. The script also lpops one token off the notify list, so the number of notify tokens stays equal to the number of waiting jobs.

The script returns both payloads. job is the original as it sat on the ready list; reserved is the re-encoded copy with the incremented attempt count. The worker keeps both, because deleting the reservation later requires the exact, byte-for-byte string that was written into the sorted set.

The attempt is spent at pop, not at failure#

In that Lua script, reserved['attempts'] = reserved['attempts'] + 1 runs before the worker has seen the job, let alone executed it. tries does not count failures; it counts pops.

Anything that causes a job to be popped a second time costs a try, whether or not any of your code ran:

Cause of the extra pop Did handle() run? Attempt consumed
An exception, released with backoff Yes, partially Yes
RateLimited released the job No Yes
WithoutOverlapping could not get its lock No Yes
The worker was killed and the lease expired Partially, then died Yes
Manual $this->release(60) Yes, up to the release Yes

This is why the default tries => 1 and a throttling middleware are a bad pair. The first time the limiter says "not yet", the job is released, popped again, and fails with MaxAttemptsExceededException, having spent a retry budget meant for errors on rate limiting that was working as intended. The fix is to limit the job by time instead of attempts, with retryUntil().

RedisJob::attempts() reads the count off the payload it was given and adds one:

// Illuminate\Queue\Jobs\RedisJob
public function attempts()
{
    return ($this->decoded['attempts'] ?? null) + 1;
}

So on a job's first run, attempts() is 1, not 0. When a job is released, the payload written back is the reserved copy, which the Lua script already incremented, so the count carries through every release.

Reserved: a lease with a deadline, not a lock#

Once the job is in the reserved set, the entry is not owned by anybody. It carries a score, and the score is a deadline. Redis does not know which worker wrote it, whether that worker is alive, or whether it is still working.

Every pop on the queue sweeps that set (the second migrateExpiredJobs() call shown earlier) and moves anything past its deadline back onto the ready list. The sweep cannot tell these apart:

  • a job whose worker was killed by a deploy thirty seconds ago, and
  • a job whose worker is still running it.

Both are an expired score. That is why a crashed worker loses no work, and also why a job runs twice when its timeout is not comfortably below retry_after. The ordering rule and the two symptoms of breaking it are covered in a separate article on timeout and retry_after. For the lifecycle, the point is that the reservation only records a deadline, and the worker's timeout is the only thing that stops a job from outliving it.

Inside process(): four checks before your code runs#

The worker now has a job object. Before handle() is reached, it goes through Worker::process():

// Illuminate\Queue\Worker::process()
$this->raiseBeforeJobEvent($connectionName, $job);          // JobProcessing

$this->markJobAsFailedIfAlreadyExceedsMaxAttempts(
    $connectionName, $job, (int) $options->maxTries
);

if ($job->isDeleted()) {
    return $this->raiseAfterJobEvent($connectionName, $job);
}

$job->fire();

$this->raiseAfterJobEvent($connectionName, $job);           // JobProcessed

The second call fails a job that has already been popped too many times. It runs before execution because it exists for a job that keeps coming back without ever reporting a failure, typically one that timed out and was migrated back:

// Illuminate\Queue\Worker::markJobAsFailedIfAlreadyExceedsMaxAttempts()
$maxTries = ! is_null($job->maxTries()) ? $job->maxTries() : $maxTries;

$retryUntil = $job->retryUntil();

if ($retryUntil && Carbon::now()->getTimestamp() <= $retryUntil) {
    return;
}

if (! $retryUntil && ($maxTries === 0 || $job->attempts() <= $maxTries)) {
    return;
}

$this->failJob($job, $e = $this->maxAttemptsExceededException($job));

retryUntil takes precedence: if it is set and has not passed, the attempt count is not consulted, so a job can be popped a hundred times inside its window. And maxTries === 0 means unlimited, not zero.

Past the gate, $job->fire() resolves CallQueuedHandler@call, which unserialises your job object out of the payload and sends it through the middleware pipeline into handle(). If unserialising throws ModelNotFoundException because the row was deleted while the job waited, the job is either deleted (with deleteWhenMissingModels) or failed, and in both cases your code never runs.

Middleware: the ways a job leaves without running#

Job middleware sits between fire() and handle(), and the two that ship with Laravel both work by not calling $next. That is how a job can be reserved, counted, and returned to the queue without doing anything.

RateLimited#

// Illuminate\Queue\Middleware\RateLimited::handleJob()
foreach ($limits as $limit) {
    if ($this->limiter->tooManyAttempts($limit->key, $limit->maxAttempts)) {
        return $this->shouldRelease
            ? $job->release($this->releaseAfter ?: $this->getTimeUntilNextRetry($limit->key))
            : false;
    }

    $this->limiter->hit($limit->key, $limit->decaySeconds);
}

return $next($job);

Over the limit, the job is released with a delay of however long until the window reopens, plus three seconds, and the attempt has already been spent. Under the limit, the limiter is hit and the job proceeds. This middleware limits how often jobs may start, not how many run at once: sixty workers can all be inside the same API call and still be within "sixty per minute".

WithoutOverlapping#

// Illuminate\Queue\Middleware\WithoutOverlapping::handle()
$lock = Container::getInstance()->make(Cache::class)->lock(
    $this->getLockKey($job), $this->expiresAfter
);

if ($lock->get()) {
    try {
        $next($job);
    } finally {
        $lock->release();
    }
} elseif (! is_null($this->releaseAfter)) {
    $job->release($this->releaseAfter);
}

This one is a lock, taken for the duration of handle() and released in a finally. Three details decide whether it behaves the way you expect:

  • The default release delay is 0. A blocked job goes straight back and is very likely popped again immediately by the same worker, burning attempts in a tight loop. Set a releaseAfter.
  • The key includes the job class unless you call ->shared(). Two different job classes with the same $key do not exclude each other by default.
  • Without expireAfter, the lock has no TTL. A worker killed inside handle() never reaches the finally, and the lock stays in the cache until someone deletes it by hand.

The third detail can turn a nine-second outage into a nine-hour one, and deploys are what usually trigger it (see the restart section below). Rate limiting and concurrency in Laravel queues covers the alternatives: funnels, and single-slot supervisors that need no locks.

Completion: the only thing that deletes a job#

handle() returns. CallQueuedHandler::call() then, in order, releases the ShouldBeUnique lock, dispatches the next job in the chain and records the batch success. Only after that does it delete the job:

// Illuminate\Queue\CallQueuedHandler::call()
if (! $job->isDeletedOrReleased()) {
    $job->delete();
}

On Redis, that is a single zrem of the exact reserved payload:

// Illuminate\Queue\RedisQueue
public function deleteReserved($queue, $job)
{
    $this->getConnection()->zrem($this->getQueueRedisKey($queue).':reserved', $job->getReservedJob());
}

The job is removed from Redis after your code succeeds, never before. Anything that stops the worker between the pop and this zrem leaves the reservation in place, so the work is recoverable, and it may also be done twice. No queue driver offers a third option.

The corollary

Because zrem matches on the exact payload string, a reservation that was already migrated away deletes nothing. A worker finishing a job whose lease expired mid-run completes without an error and removes zero entries, while a second copy runs elsewhere.

Failure: release, backoff, and the delayed set again#

handle() throws. Worker::handleJobException() runs, and it asks two questions in order: should this job be failed now, and if not, how long until it is released back.

// Illuminate\Queue\Worker::handleJobException(), condensed
if (! $job->hasFailed()) {
    $this->markJobAsFailedIfWillExceedMaxAttempts($connectionName, $job, (int) $options->maxTries, $e);
    $this->markJobAsFailedIfWillExceedMaxExceptions($connectionName, $job, $e);
    $this->markJobAsFailedIfItShouldntBeRetried($connectionName, $job, $e);
}

// ...

if (! $job->isDeleted() && ! $job->isReleased() && ! $job->hasFailed()) {
    $backoff = $this->calculateBackoff($job, $options);

    $job->release($backoff);

    $this->events->dispatch(new JobReleasedAfterException($connectionName, $job, $backoff, $e));
}

throw $e;

If the job passes all three failure checks, it is released with a backoff. The backoff is resolved per attempt, from the job's backoff() or the worker's --backoff, indexed by the attempt just used:

// Illuminate\Queue\Worker::calculateBackoff()
return (int) ($backoff[$job->attempts() - 1] ?? last($backoff));

So public $backoff = [30, 120, 600]; means 30 seconds after the first failure, 120 after the second, and 600 after the third and every one after it, because last() is the fallback once the array runs out. A single integer is an array of one, so the delay never changes.

The release is also one Lua script, and it writes to the delayed set:

-- Illuminate\Queue\LuaScripts::release()
-- Remove the job from the current queue...
redis.call('zrem', KEYS[2], ARGV[1])

-- Add the job onto the "delayed" queue...
redis.call('zadd', KEYS[1], ARGV[2], ARGV[1])

A released job goes to the delayed set, not the ready list, even when the delay is zero. It becomes runnable again only when a later pop migrates it, and it rejoins at the tail of the ready list, behind everything queued in the meantime. A retry goes through the lifecycle again from the delayed set, with the attempt count carried in the payload.

When the retries run out#

There are two routes to a failed job, and they leave different-looking failures in the dashboard.

The normal route is markJobAsFailedIfWillExceedMaxAttempts(), which fails the job when the attempt that just threw was the last one allowed:

// Illuminate\Queue\Worker::markJobAsFailedIfWillExceedMaxAttempts()
if ($job->retryUntil() && $job->retryUntil() <= Carbon::now()->getTimestamp()) {
    $this->failJob($job, $e);
}

if (! $job->retryUntil() && $maxTries > 0 && $job->attempts() >= $maxTries) {
    $this->failJob($job, $e);
}

With tries => 3, the exception thrown on the third attempt fails the job immediately instead of releasing it for a fourth pop that would only be rejected. You get exactly three executions, and the failure is recorded at the third error with the real exception attached.

The other route is the pre-execution gate from earlier, markJobAsFailedIfAlreadyExceedsMaxAttempts(), which fires when a job comes back without anyone having recorded a failure: a killed worker, or an expired lease. That one fails with MaxAttemptsExceededException and a stack trace that points at the queue rather than your bug, because the code that would have thrown was killed before it could.

Either way, failJob() calls $job->fail($e), which:

  1. marks the job failed and deletes the reservation, so it can never be popped again;
  2. calls CallQueuedHandler::failed(), which releases the ShouldBeUnique lock (a leaked unique lock would block every future dispatch, and a ShouldBeUniqueUntilProcessing job that never started is exactly the case it skips), records the failure against the job's batch and chain, and then calls your job's own failed() method if it has one;
  3. fires JobFailed.

The third step is the one that persists the failure. The row in failed_jobs is not written by fail(); it is written by a listener that the worker command registers:

// Illuminate\Queue\Console\WorkCommand::listenForEvents()
$this->laravel['events']->listen(JobFailed::class, function ($event) {
    $this->writeOutput($event->job, 'failed', $event->exception);

    $this->logFailedJob($event);
});

horizon:work extends that command, so it inherits the listener. If you fail a job outside a worker process (in a test, in Tinker, or from custom code calling $job->fail()), the event fires with nobody listening and no row is written.

Retrying from the dashboard does not put that job back in the queue. It pushes a new job, with a new id, carrying the original payload with its attempt count reset. Skyline's lifecycle log records both ids on one line so the retry can be traced back to the failure it came from. queue:retry also resets the attempt count, but it pushes the stored payload as it is, so on Redis the retried job keeps its original id.

The timeout: what stops a hung job#

A job that hangs, on an HTTP call with no timeout or in an infinite loop, would otherwise occupy a worker process forever. The guard is a POSIX alarm, armed on each iteration of the worker loop:

// Illuminate\Queue\Worker::registerTimeoutHandler()
pcntl_signal(SIGALRM, function () use ($job, $options) {
    if ($job) {
        $this->markJobAsFailedIfWillExceedMaxAttempts(
            $job->getConnectionName(), $job, (int) $options->maxTries, $e = $this->timeoutExceededException($job)
        );

        $this->markJobAsFailedIfWillExceedMaxExceptions($job->getConnectionName(), $job, $e);
        $this->markJobAsFailedIfItShouldFailOnTimeout($job->getConnectionName(), $job, $e);

        $this->events->dispatch(new JobTimedOut($job->getConnectionName(), $job));
    }

    $this->kill(static::$timedOutExitCode ?? static::EXIT_ERROR, $options, WorkerStopReason::TimedOut);
}, true);

pcntl_alarm(
    max($this->timeoutForJob($job, $options), 0)
);

The handler ends in kill. A timeout does not release the job, roll anything back or let the loop continue: it terminates the worker process from inside the signal handler. The reserved entry stays where it was, and recovery is left to the lease:

  • Attempts left: the reservation sits in the sorted set until its score passes, then the next pop on that queue migrates it back and it runs again from the top.
  • Attempts exhausted: the handler's own markJobAsFailedIfWillExceedMaxAttempts() call fails it first, which deletes the reservation, so there is nothing left to migrate and the job stops for good.
  • The worker: is gone. Horizon's supervisor notices the dead process on its next monitor pass and starts a replacement, so without instrumentation a timeout leaves no visible sign.

The guard does not run in three cases: --timeout=0 cancels the alarm instead of firing it immediately; a $timeout property or #[Timeout] on the job class overrides the supervisor's value entirely; and without the pcntl extension the block is skipped, since it sits behind a supportsAsyncSignals() check.

Restarts, deploys, and why jobs survive them#

Jobs survive a restart because workers hold no state worth losing. Queued, delayed and reserved jobs are all in Redis. A worker process holds at most one job's in-flight work, so restart handling is about that one job.

The graceful path#

php artisan horizon:terminate sends SIGTERM to the master, which terminates its supervisors, which signal their workers. Inside the worker, the signal interrupts nothing:

// Illuminate\Queue\Worker::listenForSignals()
foreach ([SIGQUIT, SIGTERM, SIGINT] as $signal) {
    pcntl_signal($signal, function (int $signal) use ($connectionName, $queue, $options) {
        $this->shouldQuit = true;

        $this->events->dispatch(new WorkerInterrupted($signal, $connectionName, $queue, $options));

        $this->notifyJobOfSignal($signal);
    });
}

It sets a flag. The current job runs to completion and is deleted normally, and then stopIfNecessary(), at the bottom of the loop, sees shouldQuit and exits with status 0. Nothing is released or retried, and no attempt is spent. A job dispatched a millisecond before the deploy is still in Redis when the new workers come up.

horizon:terminate also writes the illuminate:queue:restart cache key, which plain queue:work workers check at the end of every loop iteration and use to exit at the same point.

The ungraceful path, and the number that controls it#

The graceful path has a time limit. A supervisor that has told a worker to stop keeps it in a terminating list, and stops it once that limit passes:

// Laravel\Horizon\ProcessPool::stopTerminatingProcessesThatAreHanging()
foreach ($this->terminatingProcesses as $process) {
    $timeout = $this->options->timeout;

    if ($process['terminatedAt']->addSeconds((int) $timeout)->lte(CarbonImmutable::now())) {
        $process['process']->stop();
    }
}

A job still running timeout seconds after the deploy started is killed. To the queue this is the same as a crash: the reserved entry survives, its lease expires retry_after seconds after the original pop, and the next worker to sweep the queue migrates it back and runs it again, with one attempt already used and any side effects from the partial first run still committed.

That path also skips the finally block in WithoutOverlapping, so a lock created without expireAfter() is left behind. Deploys are the most common trigger for that bug.

fast_termination

With fast_termination enabled the master exits without waiting for its workers to finish. That keeps deploys quick, but a process supervisor above Horizon can then conclude the restart is done while a full set of workers is still draining jobs. That failure mode has its own postmortem.

Pausing is not stopping#

SIGUSR2 sets paused, which makes daemonShouldRun() return false; the worker sleeps in a loop, popping nothing, until SIGCONT clears it. Jobs accumulate on the ready list and nothing is reserved, released or failed. That makes it the lowest-risk option when a downstream dependency is unhealthy. Skyline exposes it per queue from the dashboard rather than for the whole fleet at once.

The whole lifecycle, on one page#

                          dispatch()
                              |
              +---------------+---------------+
              |                               | ->delay(...)
              v                               v
      queues:emails                  queues:emails:delayed
      LIST, ready, FIFO              ZSET, score = available at
              ^                               |
              |    migrate(), run at the top of every pop()
              +-------------------------------+
              |
              |  ONE Lua script:  lpop  +  attempts++  +  zadd
              v
      queues:emails:reserved
      ZSET, score = now + retry_after
              |
    +---------+-----------+----------------------+
    |                     |                      |
 handle() ok         exception              worker killed
    |                     |                 (timeout, deploy,
   zrem            attempts >= tries ?       OOM, crash)
  (gone)             /            \               |
                   no             yes        lease expires
                    |               |             |
             zadd delayed      failed_jobs   migrate() back
             (backoff)         + failed()    to ready, retry
                    |
              back to ready on the next migrate()

Every transition in that diagram is a Redis command, most of them inside a Lua script. None of them involves one worker communicating with another: there is no leader and no heartbeat.

What the queue guarantees#

The contract is precise, and narrower than most people assume:

  • A ready job is popped by exactly one worker. Guaranteed by Lua atomicity, unconditionally.
  • A job is never lost. Guaranteed by deleting only on success, unconditionally.
  • A job runs at least once. Guaranteed by the reservation lease and migration.
  • A job runs at most once. Not guaranteed. The lease has no way to know whether the original worker is alive, so any job outliving its lease can be running in two places.

That last point is the standard trade-off of every at-least-once queue, not a Laravel limitation. Laravel's settings (timeout, retry_after, tries, backoff) decide how often you get a duplicate run. The remedy is to make the duplicate harmless: give each job a natural idempotency key and check it before doing the irreversible thing.

Watching it happen#

Almost none of these transitions is logged by default. Laravel emits events for some of them and none for others (a delayed job becoming ready, a job dropped by WithoutOverlapping, and before Laravel 13.25 a dispatch discarded by a unique lock), and the dashboard shows states, not the moves between them.

Skyline's job lifecycle logging writes one line per transition, each tagged with the job id, so one grep shows a job's full history:

[09:14:02] queue.DEBUG: [job:8f2ac1] queued onto [emails].
[09:14:02] queue.DEBUG: [job:8f2ac1] reserved from [emails] and started processing.
[09:14:03] queue.WARNING: [job:8f2ac1] released back to [emails] (reason=exception, delay=30s)
[09:14:33] queue.INFO:  [job:8f2ac1] migrated to [emails] and is ready to run.
[09:14:33] queue.DEBUG: [job:8f2ac1] reserved from [emails] and started processing.
[09:14:35] queue.DEBUG: [job:8f2ac1] completed on [emails].

Against the diagram above, that is: pushed to the ready list, popped into reserved, released into the delayed set with a 30-second backoff, migrated back to ready when the backoff elapsed, popped again, and zremed on success. That is two attempts and five state changes under one job id. Each line's level is chosen so the channel's log level controls the volume: debug to trace a specific job, info or warning as a steady state.

Further reading#

For the individual failure modes: timeout vs retry_after for the double run, uniqueness controls for the locks, rate limiting and concurrency for the release churn, and 12 best practices for the habits that make at-least-once delivery survivable. For these transitions in your own queues, see job lifecycle logging and Skyline vs Horizon.

Frequently asked questions

How does Laravel guarantee only one worker picks up a job?

On Redis, popping a job is a single Lua script that runs lpop on the ready list, increments the payload's attempt count, and writes the result to the reserved sorted set. Redis executes Lua scripts atomically on one command thread, so no other worker can act between the lpop and the zadd, and lpop removes what it returns, so the first script to run takes the payload and every other worker gets nothing. Redis provides the guarantee, so there is no lock to configure.

When is a Laravel job's attempt count incremented?

At pop, before the worker has seen the job. The same Lua script that reserves the job runs attempts = attempts + 1, so tries counts pops rather than failures. Anything causing a second pop spends an attempt even when handle() never ran: a RateLimited or WithoutOverlapping release, a manual release, a timeout, or a reservation that expired because its worker was killed.

What happens to running jobs when Horizon restarts during a deploy?

The worker receives SIGTERM, which only sets a shouldQuit flag. The current job runs to completion, is deleted normally, and the process exits at the bottom of its loop. Nothing is released and no attempt is wasted. If the job is still running once the supervisor's timeout elapses, the process is killed instead; its reserved entry stays in Redis, its lease expires retry_after seconds after the original pop, and the next worker to sweep the queue migrates it back and runs it again from the top.

What happens when a Laravel job's retries are exhausted?

The worker fails it rather than releasing it again. On the last allowed attempt an exception triggers failJob(), which deletes the reservation so it can never be popped again, writes the job to failed_jobs with its uuid and payload, calls the job's failed() method, fires JobFailed, and releases any ShouldBeUnique lock. Retrying it from the dashboard pushes a new job with a new id and a reset attempt count.

Does a released Laravel job go back to the front of the queue?

No. Release removes the job from the reserved set and zadds it to the delayed sorted set, even when the delay is zero. It becomes runnable only when a later pop migrates it, and migration uses rpush, so it rejoins at the tail of the ready list behind everything queued in the meantime. handle() then runs again from the start.