Paths and table size
How identifiers become :id, rewrite rules for the ones it misses, the allow-list that is the only hard bound, and merging paths already stored.
The hourly rollup, requizon_http_request_stats, has one row per API, host, path and hour. APIs and hosts are few.
Paths are the column that can run away: a REST API puts an identifier in almost every URL, and stored raw, every
order number would be its own row in every hour it was fetched. So paths are normalised before anything is stored.
What the normaliser does#
Each segment of the path is checked in turn. A segment becomes :id when it is:
- all digits (
8134) - a UUID (
2f1c3a9e-7b0d-4c1e-9f55-0c8e6f1d2a47) - a ULID (
01J8Z3K9Q4M2V6X7B1C5D8E0FG) - a 32, 40 or 64 character hex string, the shape of an MD5, SHA-1 or SHA-256
- longer than
paths.max_segment_length(40), which catches signed tokens and encoded filenames nobody names by hand
Only the first paths.max_segments segments (four) are kept. When there were more, the stored path ends
in /*, so nobody mistakes the prefix for the endpoint that was called. The query string is never part of
the path.
| Called | Stored |
|---|---|
/v1/customers/cus_Q2x8fL0aZkP3m/invoices | /v1/customers/cus_Q2x8fL0aZkP3m/invoices |
/v1/users/8134/orders | /v1/users/:id/orders |
/v1/users/8134/orders/99/items | /v1/users/:id/orders/* |
/files/2f1c3a9e-7b0d-4c1e-9f55-0c8e6f1d2a47.pdf | /files/2f1c3a9e-7b0d-4c1e-9f55-0c8e6f1d2a47.pdf |
/ | / |
The first and fourth rows are the heuristic missing. cus_Q2x8fL0aZkP3m is an identifier, but it is under 40 characters
and not in any of the recognised shapes. A UUID with a file extension is no longer a bare UUID. Both would add a path per
customer or per file. The rest of this page is about what to do when your API looks like that.
Paths are compared case-sensitively, on every database, because /files/abc and /files/AbC
are two resources.
Choosing max_segments#
Four segments is enough for most REST APIs, and it bounds the damage when an identifier the normaliser missed sits deep in a path. Raise it when the part of the path that tells your endpoints apart comes later, as with APIs that carry a long fixed prefix:
// /CBMS-external/rest/yachtReservation/v6/freeYachts
// with max_segments 4: /CBMS-external/rest/yachtReservation/v6/*
// with max_segments 6: /CBMS-external/rest/yachtReservation/v6/freeYachts
'paths' => [
'max_segments' => 6,
'max_segment_length' => 40,
'max_length' => 255,
'patterns' => [],
'other_label' => 'other',
],
Changing it starts new rows on the paths page from the deploy onwards. Existing rows keep the paths they were stored with until you merge them.
Rewriting paths the heuristic misses#
When an identifier does not look like one, declare the path to record instead. paths.rewrite maps
Str::is() patterns, matched against the raw path, to the path stored in its place:
'paths' => [
// ...
'rewrite' => [
'/files/*' => '/files/:id', // /files/4313261196505276_DSCF0122.jpg
'/v1/customers/cus_*/invoices' => '/v1/customers/:id/invoices',
],
],
Rewrites run before the heuristic, and the first pattern that matches wins, so put the narrower ones first. A
rewritten path skips max_segments, since you wrote it, but it still goes through
patterns below. The list applies to every API at once. For a rule that has to look at the host or the
request, use a resolver instead.
The allow-list#
Everything above is a heuristic, and an API with an open-ended URL space will eventually leak through it. The one
hard bound is paths.patterns. When the list is not empty, a normalised path must match one of its
Str::is() patterns, or it is stored as other:
'paths' => [
// ...
'patterns' => [
'/v1/customers*',
'/v1/payment_intents*',
'/v1/refunds*',
'/',
],
'other_label' => 'other',
],
Patterns are matched against the normalised path, so write /v1/users/:id rather than a pattern
for the digits. They apply to every API at once, including the root path, which is why / is in the list
above. It is empty by default because a default list would quietly relabel every path of every application that
never configured one.
other is a real bucket: its calls, failures and response times are all counted, just not told apart. If
it grows, the requests list for that API still has every call's full path to show you what to add.
Replacing the normaliser#
For an API whose identifiers the heuristic cannot recognise, register a resolver in your
RequizonServiceProvider. Return the path to store, or null to hand the call to the built-in
normaliser:
use BoringO11y\Requizon\Requizon;
Requizon::resolvePathUsing(function ($uri, $request) {
if ($uri->getHost() !== 'api.stripe.com') {
return null;
}
// cus_Q2x8fL0aZkP3m, pi_3NcQ... → :id, while /v1/payment_intents stays
return preg_replace('#/[a-z]{2,5}_[A-Za-z0-9]{10,}(?=/|$)#', '/:id', $uri->getPath());
});
A resolved path is stored as returned, cut to paths.max_length. It does not go through
max_segments or patterns, so a resolver is responsible for keeping its own output bounded.
Merging paths already stored#
Everything on this page changes how calls are recorded from the deploy onwards. The rows already stored keep the
paths they were recorded with, so the paths page shows the old, unbounded paths next to the new bucket until they
age out. requizon:merge-paths folds them in:
php artisan requizon:merge-paths --dry-run # every path that would move, and where to
php artisan requizon:merge-paths # merge
php artisan requizon:merge-paths --api=nausys # one API only
It puts every stored path through the current paths config, renames the detail rows and adds the old
hourly buckets into the new ones, so the charts keep their history under the new path. A path that new calls still
land in is left where it is. A rewrite is matched against the stored path, which is the raw path unless the
normaliser had already changed it.
| Option | What it does |
|---|---|
--dry-run | Lists every path that would be merged and the path it would go into, and changes nothing |
--api= | Only merges the paths of this API |
--ignore-path-resolver | Merges even though Requizon::resolvePathUsing() is set (see below) |
--force | Runs without asking for confirmation in production |
Buckets folded together stay together. Run --dry-run first and read the list.
Deploy the config change and restart long-running workers (queue:restart, octane:reload)
before merging. A worker still on the old config keeps recording the old paths. Running the command again afterwards
folds those in too.
It is safe to run while the scheduler is running: the merge and requizon:aggregate take turns through a
cache lock, one batch of paths at a time. With more than one server, the default cache store has to be one they all
share, as it does for withoutOverlapping().
A path resolver cannot be applied to stored rows, because it needs the request and a stored
row only has its path. Merging would re-normalise the paths it produced by the config alone, while new calls keep
landing where the resolver puts them. So with a resolver registered, the command refuses to merge unless you pass
--ignore-path-resolver. Check --dry-run first, and use --api to limit the merge
to the APIs your resolver leaves alone.
max_length#
The path column is VARCHAR(255), and paths.max_length cuts the stored value to
fit. Leave it at 255 unless you have altered the column. A longer path could be rejected by the database, and
because recording never lets an error reach your call, that row would be silently lost.