rl-nginx defines RateLimitly credentials, resource buckets, optional latency
guards, and the nginx locations that require decisions. Configuration values
are part of the enforcement boundary: an attacker-controlled identity can
split traffic across unlimited buckets, while an unbounded label can create
excessive telemetry cardinality or an oversized request.
This guide explains how nginx directives construct those operations. For the
underlying meanings of resource requests, latency reports, API-key quotas, and
client policy, follow the versioned rl-c-client links in the relevant
sections or start with its
Operation Model.
The tenant and API key below are deliberately non-working placeholders. Replace
both before running nginx -t.
http {
resolver 127.0.0.53 valid=30s ipv6=off;
resolver_timeout 2s;
ratelimitly_dns_srv tenant.example.invalid;
ratelimitly_auth_key rl-aes1REPLACE_WITH_YOUR_KEY;
ratelimitly_policy standard unit=50ms;
ratelimitly_fail close;
ratelimitly_zone api_per_ip
"bucket=v1|scope=api|ip=$remote_addr"
rate=100r/s;
server {
listen 8080;
location /api/ {
ratelimitly_label "scope=api";
ratelimitly zone=api_per_ip;
proxy_pass http://127.0.0.1:9000;
}
}
}
The fixed scope value gives this policy its own namespace. $remote_addr is
the only variable component and is last, so it cannot change the meaning of a
later field. If nginx’s real-IP module rewrites $remote_addr, trust forwarded
addresses only from explicitly configured proxy networks.
See examples/security-conscious.conf
for bounded route/method maps and a dynamic rate that cannot take arbitrary
client-supplied text.
RateLimitly is the final admission point before nginx content processing. nginx
first resolves its access policy, including allow/deny, auth_basic,
auth_request, and the configured satisfy all|any behavior. Pre-content
routing such as try_files also runs before the RateLimitly request.
A request rejected or finalized before that point does not reach RateLimitly
and does not consume a RateLimitly resource. A valid RateLimitly allow means
every requested resource, if any, has been consumed; the module then advances
directly to the selected content handler or upstream. A later disconnect, content error, or
upstream failure does not refund that consumption. ratelimitly_fail open is
different: it advances without a valid decision and therefore without a
guarantee that RateLimitly recorded consumption.
Internal redirects do not create a second admission. nginx can internally redirect while selecting an index, rendering an error page, or executing another content handler, but those steps remain part of the same main HTTP request. The module preserves and reuses the original completed allow, deny, or failure policy outcome across that routing. nginx subrequests are not independently rate limited by this module.
nginx complex values make many variables available, but availability does not make a value trustworthy or suitable for a bucket key.
| Variable source | Security property | Guidance |
|---|---|---|
Fixed configuration and outputs of finite map tables |
Operator-controlled and bounded | Prefer these for policy names, routes, methods, rates, services, and labels. |
$remote_addr |
Derived from the peer connection | Suitable for per-source limits unless a misconfigured real-IP/proxy-protocol path lets clients spoof it. |
$server_name |
Selected from nginx configuration | Prefer it over $host when a configured virtual-server identity is required. |
$request_method and $uri |
Parsed/normalized by nginx but selected by the client | Map them to a small fixed vocabulary before using them in buckets, services, or labels. |
$host, $arg_*, $http_*, and $cookie_* |
Directly or effectively client-controlled | Never use them as authenticated identity. Do not place raw values in bucket, service, rate, threshold, or label templates. |
$remote_user or an authentication-derived variable |
Trust depends on who populates it and when | Use only if authentication completes before rl-nginx evaluates the request, the value is canonical and bounded, and missing/invalid identity is rejected rather than mapped to a shared privileged bucket. |
Headers inserted by an upstream proxy are still untrusted unless the network
path is restricted and the proxy removes any client-provided copy. An nginx
map can bound values, but it cannot turn a client-selected plan or user ID
into authenticated identity.
The v0.5.0 C-client interface hashes the exact rendered bytes and byte length,
so an embedded NUL in $binary_remote_addr or another binary value is not
truncated. Binary identity is still easy to compose ambiguously with textual
delimiters; prefer $remote_addr for readable policies, or place a bounded
binary component in a structurally unambiguous position.
The rendered bucket string, effective window, and effective rate are hashed locally to a 128-bit resource ID. Hashing does not correct an ambiguous or attacker-controlled input. Two definitions produce the same bucket only when all three inputs match; raw user input can still create a practically unlimited set of distinct buckets.
The module passes the complete rendered value as one opaque byte string; it does not escape fields inside that value. Template authors therefore own both cardinality and structural uniqueness. A template-schema, rate, or window change produces a new resource ID and starts new bucket state, so version and roll out such a change as an identity migration.
The exact, cross-client identity contract is defined by the C client’s
Content-defined IDs.
rl-nginx owns only the rendered name and effective nginx rate/window values
passed to those helpers.
Avoid patterns such as:
# Unsafe: the user can choose the identity and create unlimited buckets.
ratelimitly_zone bad_user "bucket=user:$arg_user" rate=100r/s;
# Unsafe: URI cardinality is unbounded and ':' can occur inside values.
ratelimitly_zone bad_path
"bucket=host:$host:method:$request_method:path:$uri"
rate=100r/s;
Prefer these rules:
v1|scope=api.For example:
map $uri $rl_route_class {
default other;
~^/api/orders(?:/|$) orders;
~^/api/search(?:/|$) search;
}
map $request_method $rl_method_class {
default other;
GET read;
POST write;
}
ratelimitly_zone api_per_ip
"bucket=v1|scope=api|route=$rl_route_class|method=$rl_method_class|ip=$remote_addr"
rate=100r/s;
The map outputs have a fixed alphabet and cardinality. $remote_addr is last,
so IPv6 colons cannot be confused with another field boundary, and none of the
components can contain |.
ratelimitly_dns_srv [tenant-domain];
ratelimitly_auth_key <rl-cookie...|rl-aes...>;
ratelimitly_dns_srv is OPTIONAL. When omitted, it defaults to
c-${api-key-id}.p0.ratelimitly.com derived from the key ID in
ratelimitly_auth_key. It specifies the static DNS name used to discover:
_ratelimitly._udp.<tenant-domain>
The module has no direct server-address directive. DNS resolution for SRV and
address answers uses ratelimitly_dns_resolver (or ratelimitly_resolver),
nginx’s HTTP-scope resolver, or defaults to system DNS (/etc/resolv.conf).
When configured in HTTP scope, RateLimitly captures that resolver for its
worker-local client and deliberately ignores server/location overrides for
RateLimitly discovery. Configure a resolver that is trusted and reachable from nginx
workers. A compromised or unreliable resolver can
redirect decision traffic or turn enforcement into an outage. Set an explicit
resolver_timeout, restrict resolver and UDP egress according to the
deployment network policy, and monitor DNS failure/recovery as described in
Operations.
System-DNS derivation reads at most the first 16 nameserver entries of
/etc/resolv.conf, including IPv6 nameservers such as fd7a:115c:a1e0::53.
An entry nginx cannot use — a zone-scoped address such as fe80::1%eth0, or
any IPv6 address on an nginx built without IPv6 support — is skipped with a
warning while loading configuration, and 127.0.0.1 is used when no usable
entry remains. Declare an explicit resolver or ratelimitly_dns_resolver
whenever the deployment must not depend on that derivation.
ratelimitly_auth_key is an API-key credential. Format 1 embeds a format
version, auth mode, key ID, 32-byte secret, and six packed quotas. The locked
C client intentionally rejects legacy unversioned credentials, unknown
versions, and malformed quota words during configuration loading. Keep the
real value out of the repository and copyable examples. A deployment secret
manager can render an include file readable only by the nginx master identity,
for example:
# Main nginx configuration:
include /etc/nginx/ratelimitly/api-key.conf;
Protect the included file according to the nginx master process and deployment
model, rotate an exposed key through the RateLimitly control plane, and redact
it from support bundles. nginx -T prints included configuration; never share
its unredacted output.
The encoded fields, client-side checks, and server-enforced quota boundaries are documented in the C client’s Credentials section. DNS target naming and refresh behavior are client-owned; see DNS Refresh.
The client also rejects a complete resource request before DNS or UDP when a
rendered zone’s rate window exceeds rate_window_size_ms_max. Static and
variable-driven rate= values follow the same request-time check. This is a
configuration or dependency failure governed by ratelimitly_fail; it is not
a valid RateLimitly rejection.
ratelimitly_policy standard unit=20ms;
ratelimitly_fail open;
ratelimitly_fail close;
To select the one-transmission alternative, replace the policy line with:
ratelimitly_policy single_round unit=20ms;
ratelimitly_policy controls how the C client transmits a logical resource
request and selects a response. unit is its base scheduling unit U, not a
total timeout. It must resolve to 1..4294967295ms; zero is rejected. nginx
duration units w, d, h, m, s, and ms are accepted, and a unitless
value means seconds. Write unit=1ms when one millisecond is intended.
standard is the default. It has one initial transmission round, one replay
round, and one final receive-only unit, so its maximum admission interval and
wire deduplication TTL are 3 * U. With the default unit=20ms, that horizon
is 60ms. single_round performs no replay or completion delivery and has a
one-unit horizon. Either policy can complete earlier when its
response-selection rule is satisfied.
The names describe mechanics, not reliability guarantees. More transmissions can improve delivery opportunities and server convergence, but they also add traffic and—if server deduplication is degraded—conditional duplicate- consumption exposure.
Advanced deployments can define every C-client policy field explicitly:
ratelimitly_policy custom
unit=20ms
replays=1
replay_gap=fixed:1
final_wait_units=1
completion_delivery=on;
replays retransmit the same logical request identity within the derived
deduplication window; they are not new requests. replays accepts 0..65535
(the locked client’s R_CLIENT_HA_MAX_REPLAY_COUNT). Schedule and final-wait
values are multiples of unit. The accepted schedules are:
fixed:<units>
linear:<initial-units>:<step-units>:<maximum-units>
exponential:<initial-units>:<factor>:<maximum-units>
What the preference fields do. Within transmission round k, the client
withholds selection for the first oldest_preference(k) units, so that a valid
response from the oldest discovered server can win over a faster response
from a newer one; once those units elapse it accepts the best response it
already holds. oldest_preference=fixed:0 therefore selects the first valid
response regardless of server age, and raising the value trades decision
latency for a higher chance of converging on the oldest server’s view.
final_oldest_preference_units applies the same rule inside the final
receive-only interval. Each round’s preference must not exceed that round’s
replay_gap, and final_oldest_preference_units must not exceed
final_wait_units.
For replay rounds 0..N, the horizon is
U * (sum(replay_gap(k)) + final_wait_units). nginx validates the complete
policy and rejects an enabled configuration when that horizon exceeds the API
key’s dedup_ttl_ms_max.
Where dedup_ttl_ms_max comes from. It is a per-credential quota encoded
in the ratelimitly_auth_key value and issued by the RateLimitly control
plane; it is not an nginx setting and cannot be raised from configuration. When
the derived horizon exceeds it, nginx -t prints the effective ceiling:
ratelimitly_policy horizon is invalid or exceeds the API-key dedup_ttl_ms_max of <N>ms
So for a credential whose quota is N, the largest usable unit is
floor(N / 3) under standard (three units) and N under single_round
(one unit). A custom policy’s ceiling follows its own horizon formula above.
Because the quota travels with the credential, the same configuration can load
against one API key and be rejected against another.
See the normative Configuration DSL for every constraint and the authoritative C-client Resource-Request HA Policy for response selection, replay, final-phase, completion-delivery, and deduplication semantics.
The default is ratelimitly_fail open, but production configurations should
set the policy explicitly:
open continues normal nginx processing when DNS, UDP, timeout, protocol, or
rendered-value errors prevent a valid decision. It preserves availability
but can bypass rate-limit enforcement during an outage.close returns 429 Too Many Requests for those errors. It preserves the
enforcement boundary but can deny legitimate traffic when RateLimitly or its
dependencies are unavailable.Choose per deployment risk, capacity, and rollback plan. Do not use fail-open on a critical location merely as a substitute for monitoring. Do not use fail-close without validating that the application can tolerate an enforcement dependency outage. Keep health and recovery endpoints deliberately outside the protected location when they must remain available.
Internal nginx failures such as request-pool allocation or event-registration
failure can still return 500; the failure policy does not replace every nginx
error path. See Operations for the executable behavior matrix.
Dynamic rates and thresholds make this decision especially important: if raw client input can render an invalid value, an attacker may deliberately trigger either bypass under fail-open or denial under fail-close. Use finite maps whose outputs are all valid.
ratelimitly_zone <name> "bucket=<template>" rate=<rate>;
A zone defines one RateLimitly resource. bucket and rate are rendered per
request using nginx complex values.
Rendered buckets must contain 1..1024 bytes. Static oversized values fail
nginx -t; empty or oversized dynamic values follow ratelimitly_fail without
reaching the C client. If quoting is needed, quote the whole named argument as
"bucket=value". Do not write bucket="value": nginx makes those inner quotes
literal bucket bytes, and the module rejects that form.
Use a finite map for a dynamic rate:
map "$rl_route_class:$rl_method_class" $rl_api_rate {
default 10r/s;
orders:read 100r/s;
orders:write 20r/s;
search:read 30r/s;
}
ratelimitly_zone api_per_ip
"bucket=v1|scope=api|route=$rl_route_class|method=$rl_method_class|ip=$remote_addr"
rate=$rl_api_rate;
Every mapped rate must be valid. Never let a request argument, cookie, or header select a privileged service plan unless an authenticated authorization layer has already converted it to a bounded server-controlled value.
Rendered rates contain no spaces. Valid examples are 10r/s, 600r/m,
100r/2s, and 500r/1h.
Both rate and period milliseconds are unsigned 32-bit wire fields. Rate must be
1..4294967295. Period conversion must not exceed 4294967295ms; the largest
accepted whole-unit periods are 4294967s, 71582m, and 1193h. Decimal
overflow and the first larger value in each unit are rejected. A static invalid
value makes nginx -t fail. An invalid dynamically rendered value follows the
configured failure policy.
ratelimitly_group api_all zone=per_ip zone=per_account;
A group expands to multiple zones. Use it when one request must satisfy every listed resource limit:
location /api/ {
ratelimitly group=api_all;
proxy_pass http://127.0.0.1:9000;
}
Define the latency history once, then choose independently whether a request uses that history as an admission guard, contributes one measured sample, or does both:
ratelimitly_tracker api_latency_tracker
"service=v1|service=public-api"
ttl=30s
max_samples=128
buffer_size=32
min_sample_threshold=8;
ratelimitly_guard api_latency
tracker=api_latency_tracker
threshold=100ms;
The tracker owns every state-defining field. A guard only names that tracker and supplies an admission threshold. This permits several guards with different thresholds to evaluate the same history without defining duplicate trackers.
Prefer a fixed service name per application or a finite route-to-service map.
Do not use raw $host, $uri, request arguments, headers, cookies, or user
IDs; doing so creates attacker-controlled service cardinality and fragments
latency history.
Attach the guard when admission depends on the history:
location /api/ {
ratelimitly zone=api_per_ip guard=api_latency;
proxy_pass http://127.0.0.1:9000;
}
This does not report latency. Reporting is an explicit, independent choice:
location /api/ {
ratelimitly zone=api_per_ip guard=api_latency;
ratelimitly_report api_latency_tracker;
proxy_pass http://127.0.0.1:9000;
}
The module then measures from nginx request start to log phase and makes one best-effort report attempt after completed work. The final application status does not suppress the sample. A valid RateLimitly denial, fail-close result, internal nginx failure, or client abort does suppress it. Completed fail-open work is reported because the report describes serving the HTTP request, not proof of resource consumption. Delivery failure never changes the response.
A report does not require a guard or a Rate Request:
location /observe-only/ {
ratelimitly_report api_latency_tracker;
proxy_pass http://127.0.0.1:9000;
}
On a cold worker, the first best-effort report can be dropped while DNS
membership is still resolving; nginx does not delay the response to await
discovery. A parent report setting is inherited. Use
ratelimitly_report off; in a child location that must not contribute a
sample.
If access depends only on service health and consumes no rate-limited resource, use a guard-only Rate Request:
location /health-sensitive/ {
ratelimitly guard=api_latency;
proxy_pass http://127.0.0.1:9000;
}
A passing guard admits the request, a failing guard returns 429, and a client
or transport failure follows ratelimitly_fail. It still sends no latency
report unless ratelimitly_report is also present.
Tracker TTL and the three sample-count fields must fit an unsigned 32-bit wire
field. TTL must be positive, and a unitless duration means seconds.
max_samples must be nonzero. When buffer_size is omitted, nginx uses the
configured API key’s latency_buffer_size_max quota. An explicit value must be
nonzero; nginx does not compare it with the credential quota at nginx -t.
Because effective buffer_size contributes to the latency-tracker ID, rotating
to a credential with a different quota re-identifies a tracker that relies on
the fallback. Set buffer_size explicitly when tracker identity must survive
key rotation, and review operations first.
Rendered service keys must contain 1..1024 bytes. Static oversized values
fail nginx -t; empty or oversized dynamic values make that guard follow the
failure policy and make that report ineligible. Quote the complete
"service=value" argument when needed, never only the value. Tracker tuning
fields are static and validated while loading configuration.
A static invalid guard threshold makes nginx -t fail. A dynamic threshold is
validated per request and follows the failure policy when invalid. Use a finite
map; never render a threshold directly from client input.
min_sample_threshold=0 disables only the insertion-rate sufficiency gate. A
retained, non-expired sample is still required before minimum latency is
available. A positive value requires the estimated insertion rate to reach
that threshold. The default is 8.
The C client defines tracker fields and operation independence in Latency Guards and Independent Reports.
ratelimitly_label "route=$rl_route_class|method=$rl_method_class";
Labels are transmitted for observability rather than hashed into a resource
ID. A rendered label may contain at most 256 bytes; static oversized values
fail nginx -t, dynamic oversized values follow the failure policy, and an
empty label is omitted. Keep values low-cardinality and non-sensitive. Do not include raw
paths, query arguments, headers, cookies, API keys, session tokens, user IDs,
email addresses, or source IPs. Finite map outputs and fixed policy names are
appropriate label values.
ratelimitly_bind 192.0.2.10;
ratelimitly_debug on;
ratelimitly_bind selects the local source IP for the module’s UDP socket. It
does not configure a RateLimitly server address; servers are discovered through
DNS SRV records. Normally omit it and let the kernel choose. If policy routing
or a multi-homed host requires it, use an address owned by the host and include
that path in network and reload testing. Invalid address syntax fails
nginx -t; a syntactically valid address unavailable to a worker follows the
configured failure policy and bounded initialization-retry backoff.
Debug mode writes rn: events including DNS targets, decision status, and
hashed identifiers to the nginx error log. It does not make the error log safe
for unrestricted access. Enable it only for a bounded diagnostic window,
protect and rotate the logs, and plan for volume on hot paths.
Tenant, credential, request-policy, failure, bind, debug, zone, group, tracker,
and guard definitions belong at http scope. Protected server or location
blocks reference zones, groups, and guards with ratelimitly. They opt into one
post-response sample independently with ratelimitly_report <tracker>, or
suppress an inherited report with ratelimitly_report off. A report-only
location is observed but not protected by RateLimitly admission.
See the DSL reference for complete syntax and Operations for rollout, monitoring, outage, and recovery guidance.