Skyline

Alerts: know when your queues stop working

Sixteen built-in checks with severities, recovery messages and a quiet period after deploys, over mail, Slack, SMS or a webhook.

Horizon ships one notification, LongWaitDetected, for a queue whose wait crosses a threshold, and it never says when the wait ended. Everything else a queue can do to you happens in silence: workers crash-looping, a supervisor killed for memory, jobs failing in bulk, a unique lock that blocks every dispatch of a job, and the one that costs the most, Horizon not running at all.

Skyline adds an alert pipeline with sixteen built-in checks, a severity on every alert, and a recovery message when the condition clears. It runs inside your application against your own Redis, with no agent and no outside service, and it sends over the mail, Slack and SMS channels Horizon already knows plus a JSON webhook.

Turning it on#

Alerts are off until you enable them, and an alerts block with no enabled key counts as off. A config/horizon.php published before 1.5 needs no changes: the package supplies the whole block. Set the flag and at least one destination:

HORIZON_ALERTS=true
HORIZON_ALERT_SLACK="https://hooks.slack.com/services/..."
HORIZON_ALERT_MAIL="ops@example.com"
HORIZON_ALERT_SMS="+15555550100"
HORIZON_ALERT_WEBHOOK="https://events.example.com/queue-alerts"

If your application already calls Horizon::routeSlackNotificationsTo(), routeMailNotificationsTo() or routeSmsNotificationsTo() in its service provider, those destinations are used when the matching variable is empty, so a working long-wait setup starts receiving alerts without being moved. SMS goes through Laravel's Vonage notification channel, as Horizon's own SMS notification does.

Then schedule the one check that cannot run inside Horizon, wherever the rest of your scheduled commands run:

// routes/console.php
Schedule::command('horizon:check')->everyMinute();

Every other check is evaluated by a running supervisor, once a minute across the whole fleet. See Noticing that Horizon is not running for why this one is different.

What it watches#

Check Fires when Default Severity
horizon_down No master supervisor has reported in after 2 min critical
queue_wait The oldest job on a queue has waited longer than its threshold your waits thresholds (60s), after 1 min warning
queue_stalled Jobs are waiting and workers are assigned, but nothing on the queue finishes after 5 min critical
queue_not_draining The backlog kept growing across the window and the time to clear it is rising past its threshold 15 min to clear, over a 10 min window warning
queue_paused A master, supervisor or queue was paused and never resumed, including queue:pause --all after 15 min warning
failure_rate Jobs are failing in bulk on one queue, or on one job class 10 failures in 5 min critical
worker_crash_loop Workers are dying without reporting why 5 crashes in 5 min critical
memory_restarts Workers keep being recycled for their memory limit. The alert names the job classes on that supervisor's queues that held on to the most heap 10 restarts in 15 min, per supervisor warning
job_timeout One job class keeps timing out 5 timeouts in 10 min warning
limiter_drop Horizon's RateLimited or WithoutOverlapping middleware dropped jobs without running them, or ThrottlesExceptions deleted them after they threw (deleteWhen()) 10 in the past hour warning
stranded_lock A unique or overlap lock outlived its job, so every dispatch of that job is skipped after 5 min critical
misconfiguration A supervisor's timeout is not below its connection's retry_after, so long jobs run twice immediately, repeats daily warning
supervisor_out_of_memory A supervisor exceeded its memory limit immediately warning
master_out_of_memory The master process exceeded memory_limit immediately critical
process_launch_failed A worker process would not start immediately critical
reservation_expired A job was still reserved when its connection's retry_after ran out, so it has been put back on the queue and will run a second time on the next evaluation, once per queue with a count critical

Each check lives under horizon.alerts.checks and takes enabled, plus for, threshold and window where they apply. Every check also accepts severity and cooldown, which override the defaults above for that check alone.

Four checks lean on data another feature records. worker_crash_loop and memory_restarts need HORIZON_INSIGHTS=true, which is what records worker restarts and per-class heap growth in the first place (see Insights). stranded_lock and limiter_drop need lock_insights, which is on by default and feeds the Locks & Limits screen.

Counting failures per job class#

failure_rate counts per queue by default. One class failing every run on a queue shared by twenty others is diluted into nothing that way, so set 'per' => 'job' to count each job class on its own. 'mode' => 'percent' compares the share of finished jobs that failed instead of the count, and a class needs at least min_jobs finished jobs in the window before its percentage counts.

'failure_rate' => [
    'enabled' => true,
    'threshold' => 25,     // percent, because of the mode below
    'window' => 300,
    'per' => 'job',
    'mode' => 'percent',
    'min_jobs' => 20,
],

Thresholds for one queue or one job class#

One threshold rarely fits every queue. A scheduled command that drops ten thousand jobs on a bulk queue leaves a backlog that takes hours to clear, and that is normal there. Four checks take overrides beside their global threshold, and a zero turns the check off for that one entry:

Check Key Keyed by
queue_not_draining queues a queue or a pool, with or without its connection (bulk, redis:bulk)
failure_rate queues while per is queue, jobs while it is job a queue name, or a job class
job_timeout jobs a job class
limiter_drop groups a limiter or job class, optionally prefixed with its kind (rate_limiter:uploads)
'queue_not_draining' => [
    'threshold' => 900,
    'queues' => [
        'redis:bulk' => 86400,  // a day to clear is fine here
    ],
],
'limiter_drop' => [
    'threshold' => 10,
    'groups' => [
        'App\\Jobs\\SyncLatestPrices' => 0,  // drops on purpose: only the latest run matters
    ],
],

An entry for the queue itself beats one for its pool, and among the rest the lowest non-zero threshold wins. Failures are counted per queue name, so a connection in a failure_rate key applies to that name on every connection. A queue that is turned off is still sampled, so removing the override later leaves no gap in the window. horizon:alerts warns about an entry that names no running queue, or whose value is not a number.

Jobs that disappear without failing#

RateLimited with dontRelease(), WithoutOverlapping without releaseAfter(), and ThrottlesExceptions with deleteWhen() all remove a job without failing it. The job is marked completed, so failure_rate never sees it. limiter_drop counts those drops from the Locks & Limits screen's ten-minute buckets, so its hour moves in steps. Set circuit_open to a number to also hear when a ThrottlesExceptions circuit held back that many jobs in an hour; it is 0, off, by default. Like the other middleware features, it only sees Skyline's middleware, imported from Laravel\Horizon\Middleware.

Jobs that time out or run twice#

A Redis queue gives a reserved job to another worker once its connection's retry_after passes. If the first worker died, that is the recovery. If the first worker is still running the job, the job now runs twice. reservation_expired reports both, once per queue with a count, and says which one it most likely was: when the job's own timeout or the supervisor's is at or above retry_after, the alert names it. A job stopped from the dashboard is left out.

A timed out job also comes back this way, because the timeout kills its worker and leaves the job reserved. That retry is expected, so while job_timeout is enabled a timeout below retry_after is left to it rather than reported again. job_timeout counts timeouts per job class, five in ten minutes by default, and the same counts are exported to Prometheus as horizon_job_timed_out_total and horizon_queue_timed_out_total.

On Laravel 13.34 and later, reservation_expired and worker_crash_loop also point to the framework's #[CountCrashesAsExceptions] attribute when a dead worker is the likely cause. A job that runs its worker out of memory or segfaults it throws nothing, so only tries stops it, at the cost of one dead worker per retry. With the attribute and #[MaxExceptions], it fails after that many crashes.

Moving from the long wait notification#

With alerts enabled, queue_wait takes over from Horizon's LongWaitDetected notification. It reads the same horizon.waits thresholds unless you give a queue its own under alerts.checks.queue_wait.queues, keyed connection:queue, where a zero turns the check off for that queue. The difference is in how it behaves: it holds before it pages, says when the wait cleared, stays quiet after a deploy and goes over every channel including the webhook.

Horizon's own long wait notification stands down while alerts are enabled, so you are not told about one queue twice. The LongWaitDetected event still fires for anything of yours that listens to it.

Why it does not page you at 4am#

  • Conditions have to hold. Each check has a for window. A backlog that clears inside five minutes is a burst being absorbed, which is what the queue is for.
  • It says it once. A still-firing alert repeats at most every cooldown seconds, 15 minutes by default, with "Still firing" in the subject. Set it to 0 to hear about it exactly once.
  • It tells you when it stopped. Every condition sends a recovery message saying how long it lasted, so a fixed problem and a muted one don't look the same.
  • Deploys are quiet. A deploy replaces every worker and leaves a backlog behind it, so notifications are withheld for suppress_after_deploy seconds (300) afterwards. The checks keep running and their timers keep counting, so anything the deploy really broke still announces itself when the window closes.

Rate checks such as failure_rate sample the counters Horizon already maintains rather than keeping a tally of their own. An evaluator that restarts or moves to another host loses a reading, not the count, and a counter that went down is read as horizon:clear-metrics rather than as a recovery.

Sending each severity somewhere different#

Every alert has a severity: critical means work is not getting done and someone should get up, warning can wait for office hours, and info is for the record. It opens the subject line and colours the Slack message:

[critical] Acme: The "default" queue has stopped moving
[critical] Still firing: Acme: The "default" queue has stopped moving
[critical] Resolved: Acme: The "default" queue has stopped moving

By default every alert goes to every configured channel. routes narrows that per severity, so only a critical alert sends the SMS:

'alerts' => [
    'routes' => [
        'critical' => ['slack', 'sms', 'webhook'],
        'warning' => ['slack'],
        'info' => ['slack'],
    ],
],

A severity you leave out still goes everywhere. A recovery goes to the same channels as the alert it recovers from. A route that names a channel with no destination is skipped, and horizon:alerts warns about it.

Noticing that Horizon is not running#

Every other alert is raised by a supervisor that is still looping, which is exactly what is missing when Horizon is down. So horizon_down runs only from horizon:check, outside the fleet, from the scheduler you set up above. Its two-minute for window covers the gap while a deploy replaces the master supervisor.

horizon:check also works as a plain health probe for a load balancer or an uptime monitor. It exits non-zero whenever a check is unhappy, whether or not it has been unhappy long enough to notify anyone yet.

A paused master is still running, so a paused fleet is reported by queue_paused from inside the fleet, per machine and with the backlog it is holding. horizon_down reports pauses only while queue_paused is turned off.

Checking your setup#

php artisan horizon:alert:test                      # send a test alert down every channel
php artisan horizon:alert:test --severity=critical  # test one severity's route
php artisan horizon:alerts                          # what is firing right now
php artisan horizon:alerts --history                # what was sent recently
php artisan horizon:alerts --clear

horizon:alert:test ignores the enabled flag, the sustain windows, the cooldown and deploy suppression on purpose. The point is to find out that the Slack webhook was revoked before an outage does. The history keeps the last 200 delivered alerts, set by alerts.history.

Sending alerts to PagerDuty, Opsgenie or your own service#

HORIZON_ALERT_WEBHOOK posts each alert, and each recovery, as JSON to any URL:

{
  "application": "Acme",
  "environment": "production",
  "state": "firing",
  "repeat": false,
  "seconds": 420,
  "alert": {
    "key": "queue_wait:redis:default",
    "check": "queue_wait",
    "severity": "warning",
    "title": "Acme: Jobs on the \"default\" queue are waiting too long",
    "summary": "The oldest job on the \"default\" queue has been waiting 7 minutes, over its threshold of 5 minutes. 4,120 jobs are waiting, estimated to clear in 12 minutes.",
    "context": {"Queue": "redis:default", "Oldest job": "7 minutes", "Threshold": "5 minutes", "Waiting": "4,120", "Workers": 6},
    "url": "https://acme.test/horizon/queues/default"
  }
}

state is firing or resolved, and alert.key stays the same across both, so an incident service can open and close one incident per condition. Add authentication headers under horizon.alerts.channels.webhook_headers; the request times out after webhook_timeout seconds (5).

To change how alerts are worded on every channel, bind your own notification class:

$this->app->bind(
    \Laravel\Horizon\Contracts\AlertNotification::class,
    \App\Notifications\QueueAlert::class
);

Alerting on something only you know about#

Thresholds cover the general cases. "The invoices queue should never sit idle on a weekday" is a rule only your application knows, so write it as a check and register it:

use Laravel\Horizon\Alerts\Alert;
use Laravel\Horizon\Alerts\Checks\Check;

class InvoiceQueueIdleCheck extends Check
{
    public function slug()
    {
        return 'invoice_queue_idle';
    }

    public function run()
    {
        if (! $this->isIdle()) {
            return [];
        }

        return [new Alert(
            'invoice_queue_idle',
            $this->slug(),
            Alert::CRITICAL,
            'No invoices have been processed today',
            'The invoices queue has run nothing since midnight, which has never happened on a weekday.'
        )];
    }

    protected function defaultSeverity()
    {
        return Alert::CRITICAL;
    }
}

// In a service provider:
Horizon::alertCheck(InvoiceQueueIdleCheck::class);

Return every alert your check owns on every call. Anything you stop returning is treated as recovered, which is what sends the resolve message. If you cannot read what you need, throw instead of returning an empty list: the manager then leaves your previous alerts alone rather than announcing a recovery it cannot vouch for. A custom check gets the same sustain window, cooldown, deploy quiet period and severity routing as the built-in ones.

Alerts and Prometheus#

If you already run Prometheus and Alertmanager, the Prometheus endpoint exports the queue lengths, wait times, failures and worker counts that several of these checks read, and you can keep alerting there. The built-in alerts are for the teams that don't run that stack, and for the conditions a scrape can't see well: a stranded lock, a worker that would not launch, a job that ran twice, or a supervisor whose timeout will make it run twice.

Common questions

How do I get alerted when Laravel Horizon stops processing jobs?

Install Skyline, set HORIZON_ALERTS=true with a Slack, mail, SMS or webhook destination, and schedule php artisan horizon:check every minute. That command runs outside the fleet, so it notices when no master supervisor has reported in for two minutes, which a check running inside Horizon never could. The other checks, such as a stalled queue, a failure spike or crash-looping workers, run inside the supervisors on their own.

Does Skyline replace Horizon's LongWaitDetected notification?

Yes, while alerts are enabled. The queue_wait check reads the same horizon.waits thresholds, but it waits a minute before it pages, sends a recovery message when the wait clears, stays quiet for five minutes after a deploy and goes over every alert channel including the webhook. Horizon's own notification stands down so you are not told twice. The LongWaitDetected event still fires for your own listeners.

How does Skyline avoid alert fatigue?

Each condition has to hold for its for window before anyone is told, a still-firing alert repeats at most every 15 minutes, and every alert sends a recovery message saying how long it lasted. Notifications are withheld for five minutes after a deploy, while the checks keep running, so a problem the deploy really caused still reports when the window closes. Severity routing lets only critical alerts reach the SMS or the pager.

Can Skyline alert on jobs that time out or run twice?

Yes, from 1.5.1. job_timeout fires when one job class times out five times in ten minutes, and reservation_expired fires when a job was still reserved after its connection's retry_after ran out, so a second copy will run. It says whether the job's timeout or a dead worker was the cause. limiter_drop covers jobs that RateLimited, WithoutOverlapping or ThrottlesExceptions dropped without running, which are marked completed and never show up as failures.

Can Skyline send queue alerts to PagerDuty or Opsgenie?

Yes. HORIZON_ALERT_WEBHOOK posts every alert and every recovery as JSON, with a stable key per condition and a state of firing or resolved, so an incident service can open and close one incident per problem. Headers for authentication go under horizon.alerts.channels.webhook_headers.

Queue control, not just queue monitoring.

Skyline is a drop-in replacement for Laravel Horizon that lets you act on what you see — pause a queue, jump a job to the front, drain a backlog. $139 once, every app you run it on.

Buy Skyline — $139 once

Secure checkout by Anystack. 30 days to change your mind.