# Paths and table size

> How identifiers become :id, rewrite rules for the ones it misses, the allow-list that is the only hard bound, and merging paths already stored.

Source: https://www.boring-observability.dev/requizon/docs/paths
Section: Configuration — Requizon documentation (version 0.5)
Updated: 2026-10-02

---

The hourly rollup, `requizon_http_request_stats`, has one row per API, host, path and hour. APIs and hosts are few. Paths are the column that can run away: a REST API puts an identifier in almost every URL, and stored raw, every order number would be its own row in every hour it was fetched. So paths are normalised before anything is stored.

## What the normaliser does

Each segment of the path is checked in turn. A segment becomes `:id` when it is:

- all digits (`8134`)
- a UUID (`2f1c3a9e-7b0d-4c1e-9f55-0c8e6f1d2a47`)
- a ULID (`01J8Z3K9Q4M2V6X7B1C5D8E0FG`)
- a 32, 40 or 64 character hex string, the shape of an MD5, SHA-1 or SHA-256
- longer than `paths.max_segment_length` (40), which catches signed tokens and encoded filenames nobody names by hand

Only the first `paths.max_segments` segments (four) are kept. When there were more, the stored path ends in `/*`, so nobody mistakes the prefix for the endpoint that was called. The query string is never part of the path.

| Called | Stored |
| --- | --- |
| `/v1/customers/cus_Q2x8fL0aZkP3m/invoices` | `/v1/customers/cus_Q2x8fL0aZkP3m/invoices` |
| `/v1/users/8134/orders` | `/v1/users/:id/orders` |
| `/v1/users/8134/orders/99/items` | `/v1/users/:id/orders/*` |
| `/files/2f1c3a9e-7b0d-4c1e-9f55-0c8e6f1d2a47.pdf` | `/files/2f1c3a9e-7b0d-4c1e-9f55-0c8e6f1d2a47.pdf` |
| `/` | `/` |

The first and fourth rows are the heuristic missing. `cus_Q2x8fL0aZkP3m` is an identifier, but it is under 40 characters and not in any of the recognised shapes. A UUID with a file extension is no longer a bare UUID. Both would add a path per customer or per file. The rest of this page is about what to do when your API looks like that.

Paths are compared case-sensitively, on every database, because `/files/abc` and `/files/AbC` are two resources.

## Choosing max_segments

Four segments is enough for most REST APIs, and it bounds the damage when an identifier the normaliser missed sits deep in a path. Raise it when the part of the path that tells your endpoints apart comes later, as with APIs that carry a long fixed prefix:

```php
// /CBMS-external/rest/yachtReservation/v6/freeYachts
// with max_segments 4: /CBMS-external/rest/yachtReservation/v6/*
// with max_segments 6: /CBMS-external/rest/yachtReservation/v6/freeYachts
'paths' => [
    'max_segments' => 6,
    'max_segment_length' => 40,
    'max_length' => 255,
    'patterns' => [],
    'other_label' => 'other',
],
```

Changing it starts new rows on the paths page from the deploy onwards. Existing rows keep the paths they were stored with until you [merge them](#merge-paths).

## Rewriting paths the heuristic misses

When an identifier does not look like one, declare the path to record instead. `paths.rewrite` maps `Str::is()` patterns, matched against the *raw* path, to the path stored in its place:

```php
'paths' => [
    // ...
    'rewrite' => [
        '/files/*' => '/files/:id',                 // /files/4313261196505276_DSCF0122.jpg
        '/v1/customers/cus_*/invoices' => '/v1/customers/:id/invoices',
    ],
],
```

Rewrites run before the heuristic, and the first pattern that matches wins, so put the narrower ones first. A rewritten path skips `max_segments`, since you wrote it, but it still goes through `patterns` below. The list applies to every API at once. For a rule that has to look at the host or the request, use a [resolver](#resolver) instead.

## The allow-list

Everything above is a heuristic, and an API with an open-ended URL space will eventually leak through it. The one hard bound is `paths.patterns`. When the list is not empty, a normalised path must match one of its `Str::is()` patterns, or it is stored as `other`:

```php
'paths' => [
    // ...
    'patterns' => [
        '/v1/customers*',
        '/v1/payment_intents*',
        '/v1/refunds*',
        '/',
    ],
    'other_label' => 'other',
],
```

Patterns are matched against the *normalised* path, so write `/v1/users/:id` rather than a pattern for the digits. They apply to every API at once, including the root path, which is why `/` is in the list above. It is empty by default because a default list would quietly relabel every path of every application that never configured one.

`other` is a real bucket: its calls, failures and response times are all counted, just not told apart. If it grows, the requests list for that API still has every call's full path to show you what to add.

## Replacing the normaliser

For an API whose identifiers the heuristic cannot recognise, register a resolver in your `RequizonServiceProvider`. Return the path to store, or `null` to hand the call to the built-in normaliser:

```php
use BoringO11y\Requizon\Requizon;

Requizon::resolvePathUsing(function ($uri, $request) {
    if ($uri->getHost() !== 'api.stripe.com') {
        return null;
    }

    // cus_Q2x8fL0aZkP3m, pi_3NcQ... → :id, while /v1/payment_intents stays
    return preg_replace('#/[a-z]{2,5}_[A-Za-z0-9]{10,}(?=/|$)#', '/:id', $uri->getPath());
});
```

A resolved path is stored as returned, cut to `paths.max_length`. It does not go through `max_segments` or `patterns`, so a resolver is responsible for keeping its own output bounded.

## Merging paths already stored

Everything on this page changes how calls are recorded from the deploy onwards. The rows already stored keep the paths they were recorded with, so the paths page shows the old, unbounded paths next to the new bucket until they age out. `requizon:merge-paths` folds them in:

```bash
php artisan requizon:merge-paths --dry-run        # every path that would move, and where to
php artisan requizon:merge-paths                  # merge
php artisan requizon:merge-paths --api=nausys     # one API only
```

It puts every stored path through the current `paths` config, renames the detail rows and adds the old hourly buckets into the new ones, so the charts keep their history under the new path. A path that new calls still land in is left where it is. A rewrite is matched against the stored path, which is the raw path unless the normaliser had already changed it.

| Option | What it does |
| --- | --- |
| `--dry-run` | Lists every path that would be merged and the path it would go into, and changes nothing |
| `--api=` | Only merges the paths of this API |
| `--ignore-path-resolver` | Merges even though `Requizon::resolvePathUsing()` is set (see below) |
| `--force` | Runs without asking for confirmation in production |

> **Merging cannot be undone**
>
> Buckets folded together stay together. Run `--dry-run` first and read the list.

Deploy the config change and restart long-running workers (`queue:restart`, `octane:reload`) before merging. A worker still on the old config keeps recording the old paths. Running the command again afterwards folds those in too.

It is safe to run while the scheduler is running: the merge and `requizon:aggregate` take turns through a cache lock, one batch of paths at a time. With more than one server, the default cache store has to be one they all share, as it does for `withoutOverlapping()`.

A [path resolver](#resolver) cannot be applied to stored rows, because it needs the request and a stored row only has its path. Merging would re-normalise the paths it produced by the config alone, while new calls keep landing where the resolver puts them. So with a resolver registered, the command refuses to merge unless you pass `--ignore-path-resolver`. Check `--dry-run` first, and use `--api` to limit the merge to the APIs your resolver leaves alone.

## max_length

The `path` column is `VARCHAR(255)`, and `paths.max_length` cuts the stored value to fit. Leave it at 255 unless you have altered the column. A longer path could be rejected by the database, and because recording never lets an error reach your call, that row would be silently lost.


## Common questions

### Why does a Requizon path end in /*?

The path had more segments than paths.max_segments (four by default), so Requizon kept the first four and marked the rest with a trailing *. /v1/users/8134/orders/99/items is stored as /v1/users/:id/orders/*. Raise max_segments if the deeper segments are what tells your endpoints apart.

### What does the "other" path mean in Requizon?

You have set paths.patterns, and the call's normalised path matched none of them. Every unmatched path is recorded under one bucket, named by paths.other_label, instead of adding its own rows to the rollup.

### How do I stop one API's file names filling the paths page?

Add a rewrite for them, such as '/files/*' => '/files/:id' under paths.rewrite in config/requizon.php. Rewrites match the raw path, before the normaliser, so they catch identifiers that do not look like one. Then deploy, restart your workers and run php artisan requizon:merge-paths to fold the paths already stored into the new one.

### Can I undo requizon:merge-paths?

No. Hourly buckets folded together stay together, and the detail rows are renamed in place. Run it with --dry-run first: that lists every stored path that would move and the path it would go into, and changes nothing.
