Docs contents

HTTP load tests

We measured throughput, latency, memory, and connection recovery using Nagi's HTTP sample. The initial tests ran on October 1, 2026, using Nagi 0.1.2 at commit c602abb. The implementation with a request-wait deadline is compared separately in the follow-up tests.

The endpoint that receives a 4 KiB string and returns JSON handled 15,000 requests per second for 30 minutes. All 27,000,270 requests returned HTTP 200. However, connection counts and memory increased, and file descriptors remained after the load stopped. Throughput alone does not establish stability.

All traffic stayed on the same machine. These results do not establish internet throughput, DDoS resistance, or stability over 24 hours or longer.

On this page

Reading the numbers

  • Requested rate: the number of requests per second requested from the load generator.
  • Dispatched rate: the rate the client actually sent. It can fall below the requested rate when the client is overloaded.
  • Successful rate: successful requests per second, including the final response wait.
  • p99: 99% of dispatched requests finished within this time. It excludes waiting for arrivals that the client could not send.
  • RSS: the server's resident memory. File descriptors (FDs) are OS resources used for connections and other handles.

Increasing the load

Workload Requested/s Dispatched/s Successful/s p99 Errors (3 trials)
HTTP /health 10,000 10,000 10,000 2.03ms 0
HTTP /health 30,000 29,960 29,874 190.36ms 0
HTTP /health 60,000 31,990 31,739 264.83ms 0
HTTP /health 100,000 32,501 32,266 301.05ms 0
4KiB JSON 5,000 5,000 4,999 0.83ms 0
4KiB JSON 10,000 10,000 10,000 2.12ms 0
4KiB JSON 15,000 15,000 15,000 5.24ms 0
4KiB JSON 20,000 19,997 19,993 144.35ms 0
4KiB JSON 30,000 27,471 27,216 310.65ms 0
4KiB JSON 50,000 27,700 27,483 261.35ms 0
SQLite SELECT 5,000 5,000 4,999 0.64ms 0
SQLite SELECT 10,000 10,000 10,000 1.46ms 0
SQLite SELECT 20,000 19,999 19,998 6.06ms 0
SQLite SELECT 40,000 32,174 31,939 222.11ms 0
SQLite INSERT 1,000 1,000 1,000 0.39ms 0
SQLite INSERT 5,000 4,999 4,999 0.83ms 0
SQLite INSERT 10,000 10,000 10,000 1.64ms 0
SQLite INSERT 20,000 20,000 19,998 32.96ms 0

Each row is the median of three 8-second trials. The p99 is the median of each trial's p99, not a pooled percentile. Vegeta's percentiles are approximate aggregates. SQLite used an in-memory database. Reads selected the same row; writes kept inserting rows.

Requested and successful rates, with p99 latency

Open the full-size SVG

When the load generator exhausted its CPU capacity, a higher requested rate did not produce more traffic. The first consecutive sweep also exhausted client source ports. We retained that initial sweep, and the main table uses a repeat with a distinct loopback source address for each trial. All 12,177 failures in the initial sweep were client-side bind: address already in use errors.

Fixed-concurrency saturation test

We also increased connection counts using wrk, which places less overhead on the client. This client sends another request after receiving a response.

Workload Connections Responses/s p99 CPU (% of one core) Errors
HTTP /health 128 77,689 2.73ms 99.7% 0
HTTP /health 512 70,682 69.85ms 99.2% 0
HTTP /health 2,048 66,528 375.86ms 97.6% 0
4KiB JSON 128 39,022 4.75ms 96.9% 0
4KiB JSON 512 39,322 70.93ms 99.1% 0
4KiB JSON 2,048 38,344 386.19ms 96.5% 0

Values are medians of three 8-second trials. This closed-loop test excludes the latency of arrivals waiting to be sent. Treat it separately from the paced load test above.

Receiving traffic for 30 minutes

Condition 30 minutes, unrestricted client connections 10 minutes, client capped at 128 connections
Requested rate 15,000/s 15,000/s
Requests 27,000,270 9,000,090
Non-200 responses or transport errors 0 0
Overall p99 34.11 ms 11.07 ms
Maximum sampled server RSS 180.74 MiB 8.29 MiB
Maximum sampled server FDs 6,315 138
Observation during 30-second recovery 1,049 FDs remained FDs returned to the baseline of 10

Throughput, resident memory, and file descriptors during endurance tests

Open the full-size SVG

Memory increased at the same times as connection counts. The control lasted only 10 minutes, so it does not establish that a fixed connection count stays stable for 30 minutes or longer. We did not capture the types of the remaining FDs in the 30-minute run, and have not established whether connection teardown is faulty. RSS remaining high also does not, by itself, prove a memory leak.

In an additional 2-minute trial requesting 30,000/s, all 3,501,137 requests returned HTTP 200 and FDs reached 6,719. FDs returned to 10 after the load stopped and remained stable during 120 seconds of recovery. That trial records FD types and TCP states, but did not reproduce the remaining FDs from the 30-minute run. Their cause remains unresolved.

After the load stopped, we called /health every 0.5 seconds for 30 seconds and confirmed responses. The probe itself can briefly use one FD. The last sample in the 128-connection run shows 11 FDs, while other samples and independent TCP observations confirm recovery to 10.

128 is the measurement client's connection count. This PR does not add a 128-connection limit to Nagi. A small global server limit would restrict concurrency for many users or long requests.

Concurrent connection limits

Requested connections Initial HTTP success Held connections reverified RSS FDs after disconnect
1,000 1,000 1,000 28.49MiB 10
10,000 10,000 10,000 246.18MiB 10
15,000 15,000 15,000 366.97MiB 10
16,000 16,000 16,000 391.07MiB 10
16,400 16,374 16,374 400.16MiB 10
18,000 16,374 16,374 400.12MiB 10

After the initial response, we sent another HTTP request over every held connection. A successful TCP handshake alone does not count as a working connection.

When we requested 18,000 connections, 16,503 completed a TCP handshake, but only 16,374 returned the initial HTTP response and passed the held-connection verification.

Held connections and server memory

Open the full-size SVG

The OS FD limit was 16,384. The server also needs FDs for resources other than connections, so it runs out before this many clients can be accepted. This is an OS-dependent limit, not a fixed Nagi connection limit.

Idle, incomplete, and abruptly closed connections

Condition Observation over 120 seconds
Connect without sending a request All 100 remained open
Stop while sending HTTP headers All 100 remained open
Stop while sending the HTTP body All 100 received 408 and closed
Leave keep-alive idle after a response All 100 remained open
All clients disconnect normally FDs returned to 10 during 5-second recovery
200 abrupt disconnects (RST) after partial headers FDs returned to 10 during 5-second recovery

We observed 100 connections in each condition concurrently for 120 seconds. This is a small local behavior test, not a distributed attack or an attack on a public service.

Follow-up tests with the wait deadline

On October 1, 2026 (UTC), we compared Nagi 0.1.4 at 1c013868 with an implementation that adds a ten-second HTTP request-wait deadline. It uses unmodified Hyper 1.11.1. The results below come from these follow-up tests.

Connection recovery during regular traffic

We created 100 connections each for silence, partial headers, partial bodies, and unused keep-alive. Partial bodies received 408 under the existing two-second limit, and the remaining waiting connections also closed. FDs returned from 410 to 10. Connections were created sequentially before observation, so roughly ten seconds from the start of observation is not an exact lifetime for every connection. FDs also returned to 10 after 200 abrupt disconnects (RST).

In a separate test, we created 350 connections each for silence, partial headers, and unused keep-alive, totaling 1,050, while requesting 5,000 /health responses per second for 25 seconds.

Item Result
Requests actually sent 124,977
Non-200 responses or transport errors 0
Regular traffic p99 7.01ms
Observation when all waiting connections had closed 11.28 seconds after load began
Server FDs Initially 10, sampled maximum 1,092, 42 after recovery during load, finally 10
Server RSS Initially 4.34MiB, sampled maximum 20.08MiB, finally 17.73MiB

The 42 FDs after recovery include 32 connections carrying regular traffic. RSS did not return to its initial value. This short test does not establish a memory leak or long-term stability. In the first attempt, the generator sent 124,997 requests and all returned 200, but the harness failed because it required exactly 125,000 arrivals. We corrected the assertion to check actual dispatched counts and repeated the test shown above. The first response aggregate is also retained.

Worker-count correction, October 2, 2026: The recovery test above set TOKIO_WORKER_THREADS=1, but Nagi reads NAGI_THREADS. Its default is four. The original test did not record the inherited setting, so its actual worker count is unknown. The smaller lifecycle test and throughput comparisons explicitly set NAGI_THREADS=1 through their shared helper. The original results and harness remain unchanged.

We repeated the load with the same modified 0.1.4 binary, recording NAGI_THREADS=1 and one actual Tokio worker thread. The test again used 1,050 waiting connections and 5,000 regular requests per second for 25 seconds.

Item Trial recording server TCP states Trial recording both peers' TCP states
Requests actually sent 124,997 125,000
Non-200 responses or transport errors 0 0
Regular traffic p99 0.621ms 0.610ms
Observation when all waiting connections had closed 12.37 seconds after load began 11.26 seconds after load began
Server FDs after load 10 10

Two preceding attempts returned successful responses for all regular requests but failed validation because some clients did not observe connection closure. In one attempt, 25 silent clients had not observed closure even though server FDs had returned to 10. The cause is unresolved. The table shows the two later trials after adding reads to observe closure and TCP-state capture; it does not establish that every attempt observed all connections closing. TCP sampling and observation periods also differ, so these p99 values should not be used as a performance comparison with the earlier result.

At the end of both successful trials, 1,050 server-side connections were still in FIN_WAIT2, awaiting their peers' close, despite FD recovery. Reclaiming FDs does not mean every kernel TCP state has disappeared. Correction and repeat records include the attempts that failed validation.

Automated tests with a shorter deadline check that one-byte header progress does not reset the deadline, keep-alive can be reused within it, and long responses, slow response readers, streams, and upgraded WebSockets are not cut off by the request-wait deadline. These are not tests of large WebSocket fan-out or public-network capacity.

Effect on regular traffic

We alternated the implementations on the same host: before/after, after/before, then before/after. Each condition used 128 client connections and three five-second trials. The table shows the median throughput and median trial p99.

Workload Responses/s before Responses/s after Throughput change p99 (before → after) Errors
HTTP /health 78,132 70,105 −10.3% 2.412 → 2.277ms 0
4KiB JSON 40,018 38,207 −4.5% 4.540 → 5.090ms 0

The implementation that reclaims waiting connections has lower saturation throughput under these conditions. We have not profiled the cause. An earlier comparison that measured the baseline trials together also showed reductions of about 7.2% and 5.6%, respectively. The table uses the alternating-order repeat; raw data from both comparisons is retained. These brief trials on one host do not establish the same difference for every environment.

For the throughput comparisons and smaller lifecycle test, the server used one logical CPU and one worker; the load generator used two different logical CPUs. The assignments were server 0, clients 1,2, with HTTP/1.1 without TLS on loopback. OS FD limits were not raised. The 128 connections are a client-side test condition.

Raw follow-up data and script snapshots record hashes of the binaries, implementation, and harnesses. The original environments' source_commit identifies the checked-out base commit; also consult provenance.json to identify the modified implementation.

Reduce JSON response allocations

In 0.1.6, JSON success and error responses use a static Content-Type value. Comparing small JSON, 4KiB JSON, and 400/404/503 response construction in one process shows one fewer allocation and 16 fewer allocated bytes per response. Status, headers, and body bytes match. Network operations and body consumption are outside this allocation measurement.

Both versions keep the current ten-second connection wait deadline. Five alternating baseline/candidate pairs per endpoint use 128 client connections for ten seconds, one server CPU with NAGI_THREADS=1, and two separate load-generator CPUs.

Operation Baseline responses/s Candidate responses/s Change in medians p99 (before → after) Errors
/health 71,069 70,186 −1.2% 2.721 → 2.859ms 0
Small JSON 65,438 66,191 +1.2% 2.947 → 3.127ms 0
4KiB JSON 38,632 38,844 +0.5% 3.911 → 4.233ms 0

Paired differences varied in both directions, including the unchanged /health endpoint. These trials do not establish a throughput improvement. The directly measured benefit is fewer allocations during response construction. All 30 trials and reproduction steps are retained.

Connection management and DoS protection

The current serve has a 1 MiB HTTP body limit and a 2-second handler timeout. It binds only to 127.0.0.1 by default. These controls do not bound every connection's lifetime or the traffic reaching a public deployment.

The implementation adds no fixed connection cap. Silent connections and incomplete header sends now have a deadline.

The initial measurements at c602abb predate connection deadlines. The server now uses a shared ten-second deadline for initial silence, incomplete headers, and the wait from a completed response until the next complete headers. Keeping Hyper unmodified means unused keep-alive does not have a separate 60-second deadline. The deadline is configurable through an environment variable and excludes active responses, streams, and upgraded WebSockets. See HTTP for configuration.

Proposed control Intended effect Decision still needed
Investigate and fix connection teardown Release resources after disconnects Reproduce and identify the remaining FDs from the 30-minute run
Header and HTTP wait deadline Bound silent, partial-header, and unused connections Shared ten-second deadline implemented; verify deployment settings and reconnection costs
Body receive and work admission controls Bound stalled uploads and queued work Interaction with existing timeouts and queue capacity
Fronting proxy or CDN Filter excessive traffic to public services Deployment, traffic and connection budgets, and source policies

Deployment limits are needed separately from this wait deadline. If a DDoS attack fills the network link, controls inside Nagi cannot resolve that saturation; upstream protection is also needed.

Environment and reproduction

Item Condition
Nagi / Rust 0.1.2 / 1.98.1
OS Linux x86_64
CPU AMD EPYC 7763 64-Core Processor
CPU allocation cgroup quota of 4 CPUs; 5 logical CPUs available
Memory limit cgroup 16 GiB
Server / generator CPU 0 / CPUs 2,3
Server FD limit Soft and hard limits both 16,384
HTTP / SQLite HTTP/1.1 without TLS / in-memory database
Build release, opt-level=3, LTO disabled

The server used one logical CPU and one executor worker. The generator used two different logical CPUs. SQLite has a dedicated database worker. OS FD limits were not raised. Every short trial started a fresh server and database, with warmup outside the measured interval.

Use Linux/cgroup v2, Python 3.12, Rust, taskset, and Vegeta v12.13.0. The wrk test also needs wrk 4.2.0. To reproduce the same language implementation, build c602abb and copy this PR's measurement scripts into that checkout.

cargo build --locked --release -p nagic
NAGI_ROOT="$PWD" NAGI_NATIVE_TARGET_DIR="$PWD/native-target" \
  ./target/release/nagic build examples/crud.nagi --out build/crud

python scripts/http_capacity.py --phase limits --duration 8 --repeats 3 \
  --vegeta /path/to/vegeta --out build/http-limits
python scripts/http_capacity.py --phase soak --soak-seconds 1800 --soak-rate 15000 \
  --vegeta /path/to/vegeta --out build/http-soak
python scripts/http_saturation.py --wrk /path/to/wrk --out build/http-saturation
python scripts/tcp_capacity.py --out build/tcp-capacity
python scripts/http_connection_lifecycle.py --out build/connection-lifecycle

Default CPU IDs are 0 for the server and 2,3 for the client. Use --server-cpus and --client-cpus to choose disjoint available IDs if necessary. The target is fixed to 127.0.0.1:8080. Existing results are never overwritten.

Raw data, conditions, and measured script snapshots are included. Regenerate the figures while retaining the measurements. Plotting needs matplotlib; measurement does not.

python scripts/http_capacity_report.py \
  --data benchmarks/results/http-capacity-2026-10-01 \
  --assets website/assets/http-capacity

TLS, a real network link, sustained writes to persistent SQLite, WebSocket fan-out, and continuous operation over 24 hours or longer remain unmeasured. The previous measurements used different clients and CPU allocations, so these numbers are not a direct before-and-after comparison.

Nagi 0.1 documentationPage source