Latest public benchmark

Serious load, measured in public.

A single small VM pushed through heavy HTTPS concurrency with CPU, memory, latency, clean-response behavior, and raw rows published for inspection.

TLDR: Tako stayed clean through c20000 on a small exe.dev VM while doing app-aware routing, limiter accounting, forwarding headers, and instance selection.

largest run
c20000
non-200
0
client errors
0
Benchmark VMexe.dev, Tokyo

Load generator, proxy, and app all share this host.

CPU2 vCPU

AMD EPYC 9554P on KVM.

Memory7.8 GiB RAM

No swap configured.

OSUbuntu 24.04.4

Linux 6.12.90.

Tako at c500012.5k clean 200 RPS

0 non-200, 0 client errors.

Tako at c200007.3k clean 200 RPS

Stable at the largest tested HTTP concurrency.

Clean through c20000all 200 responses

0 client errors in every heavy Tako row.

Featuresclean through c8000

Channels and workflows both stay clean through c8000.

Proxy comparison

Raw HTTPS proxy path#

Same route, same self-signed TLS certificate, same upstream application, same benchmark VM. The heavy rows show saturation behavior on one small VM; Memory uses a focused five-proxy rerun with process PSS sampling.

throughput

HTTP 200 RPS by concurrency

Tako stays clean at high concurrency and beats Caddy and Envoy across the heavy rows. nginx and HAProxy show the static-proxy ceiling for this VM.

05k10k15k20kc1kc2.5kc5kc7.5kc10kc15kc20knginx c1000: 21knginx c2500: 19knginx c5000: 18knginx c7500: 16knginx c10000: 15knginx c15000: 11knginx c20000: 11kHAProxy c1000: 21kHAProxy c2500: 18kHAProxy c5000: 17kHAProxy c7500: 16kHAProxy c10000: 15kHAProxy c15000: 13kHAProxy c20000: 11kTako c1000: 15kTako c2500: 14kTako c5000: 13kTako c7500: 12kTako c10000: 10kTako c15000: 8.6kTako c20000: 7.3kEnvoy c1000: 12kEnvoy c2500: 12kEnvoy c5000: 4.7kEnvoy c7500: 4.3kEnvoy c10000: 3.7kEnvoy c15000: 3.8kEnvoy c20000: 828Caddy c1000: 6.6kCaddy c2500: 5.9kCaddy c5000: 5.2kCaddy c7500: 4.8kCaddy c10000: 1.7kCaddy c15000: 1.7kCaddy c20000: 1.3k
  • nginx
  • HAProxy
  • Tako
  • Envoy
  • Caddy

tail latency

p99 latency by concurrency

Tako completes every high-load row cleanly, with tail latency published beside RPS so the tradeoff stays visible. nginx is the tightest p99 reference in this run.

0ms5s10s15s20s25sc1kc2.5kc5kc7.5kc10kc15kc20knginx c1000: 100msnginx c2500: 315msnginx c5000: 1.1snginx c7500: 1.9snginx c10000: 1.2snginx c15000: 6.1snginx c20000: 3.8sHAProxy c1000: 103msHAProxy c2500: 278msHAProxy c5000: 1.5sHAProxy c7500: 3.5sHAProxy c10000: 6.4sHAProxy c15000: 13sHAProxy c20000: 16sTako c1000: 154msTako c2500: 527msTako c5000: 2.4sTako c7500: 5.1sTako c10000: 7sTako c15000: 12sTako c20000: 16sEnvoy c1000: 145msEnvoy c2500: 374msEnvoy c5000: 3.1sEnvoy c7500: 4.8sEnvoy c10000: 6.7sEnvoy c15000: 14sEnvoy c20000: 27sCaddy c1000: 260msCaddy c2500: 2.3sCaddy c5000: 5.2sCaddy c7500: 8.9sCaddy c10000: 20sCaddy c15000: 24sCaddy c20000: 26s
  • nginx
  • HAProxy
  • Tako
  • Envoy
  • Caddy

errors

Clean-run behavior by concurrency

The line combines non-200 responses and client-side errors, so lower is better. Tako remains at 0% through c20000 on this run.

0%10%20%30%40%c1kc2.5kc5kc7.5kc10kc15kc20knginx c1000: 0%nginx c2500: 0%nginx c5000: 0%nginx c7500: 0%nginx c10000: 0%nginx c15000: 0.15%nginx c20000: 0%HAProxy c1000: 0%HAProxy c2500: 0%HAProxy c5000: 0%HAProxy c7500: 0%HAProxy c10000: 0%HAProxy c15000: 0%HAProxy c20000: 0%Tako c1000: 0%Tako c2500: 0%Tako c5000: 0%Tako c7500: 0%Tako c10000: 0%Tako c15000: 0%Tako c20000: 0%Envoy c1000: 0%Envoy c2500: 0%Envoy c5000: 0.28%Envoy c7500: 1.06%Envoy c10000: 3.55%Envoy c15000: 39.19%Envoy c20000: 41.3%Caddy c1000: 0%Caddy c2500: 0%Caddy c5000: 0.14%Caddy c7500: 0.39%Caddy c10000: 0.06%Caddy c15000: 4.75%Caddy c20000: 7.51%
  • nginx
  • HAProxy
  • Tako
  • Envoy
  • Caddy

memory

Memory by concurrency

Memory is measured with process PSS, which avoids counting shared pages twice. At c20000, Caddy shows lower Memory than Tako, but Caddy also times out part of the load; compare the Memory line with the 200 labels below.

0MiB450MiB900MiB1.3GiB1.8GiBc5kc10kc20knginx c5000: 159MiBnginx c10000: 159MiBnginx c20000: 451MiBHAProxy c5000: 248MiBHAProxy c10000: 406MiBHAProxy c20000: 624MiBTako c5000: 511MiBTako c10000: 911MiBTako c20000: 1.7GiBEnvoy c5000: 323MiBEnvoy c10000: 554MiBEnvoy c20000: 1004MiBCaddy c5000: 621MiBCaddy c10000: 1.2GiBCaddy c20000: 1.5GiB
  • nginx
  • HAProxy
  • Tako
  • Envoy
  • Caddy
Heavy rows; secondary labels show the percentage of requests that returned 200
proxyc5000c10000c20000c20000 p99notes
nginx17.7k15.3k11.0k3.8sStatic-proxy RPS reference
HAProxy17.1k14.8k11.2k15.7sHigh RPS, wider p99
Tako12.5k10.4k7.3k15.5sClean through c20000
Envoy4.7k99.72% 2003.7k96.45% 2000.8k58.70% 20026.6s58.70% 20058.70% 200 at c20000
Caddy5.2k99.86% 2001.7k99.94% 2001.3k92.49% 20026.4s92.49% 20092.49% 200 at c20000

Caddy c20000 is not equal capacity. Its Memory is lower than Tako's, but Caddy only returned 94.14% 200 while Tako returned 100% 200.

Memory detail from the focused PSS rerun
proxyc5000 Memoryc10000 Memoryc20000 Memoryc20000 200
nginx159 MiB159 MiB451 MiB99.43% 200
HAProxy248 MiB406 MiB624 MiB100% 200
Tako511 MiB911 MiB1.7 GiB100% 200
Envoy323 MiB554 MiB1.0 GiB51.41% 200
Caddy621 MiB1.2 GiB1.5 GiB94.14% 200

Channels and workflows

Built-in feature paths, measured separately#

These rows exercise more than the proxy. The app uses the JavaScript SDK, publishes durable channel messages, and enqueues workflows with persisted steps while everything shares the same 2 vCPU budget.

built-in features

Channels and workflows 200 RPS

Both feature paths stay clean through c8000 on the same 2 vCPU VM, while still using the SDK, SQLite-backed persistence, and the proxy path.

02k4k6k8kc500c1kc2kc4kc8kChannel publish c500: 7.8kChannel publish c1000: 6.6kChannel publish c2000: 6.4kChannel publish c4000: 5.8kChannel publish c8000: 4.6kWorkflow enqueue c500: 5.6kWorkflow enqueue c1000: 5.2kWorkflow enqueue c2000: 5kWorkflow enqueue c4000: 4.7kWorkflow enqueue c8000: 4k
  • Channel publish
  • Workflow enqueue

feature tail latency

Channels and workflows p99 latency

Workflow enqueue persists steps, so it naturally carries more work than channel publish. Both paths stay clean through c8000 in this single-instance run.

0ms3s6s9sc500c1kc2kc4kc8kChannel publish c500: 94msChannel publish c1000: 213msChannel publish c2000: 896msChannel publish c4000: 3.3sChannel publish c8000: 6.4sWorkflow enqueue c500: 126msWorkflow enqueue c1000: 243msWorkflow enqueue c2000: 1.3sWorkflow enqueue c4000: 3.5sWorkflow enqueue c8000: 7.7s
  • Channel publish
  • Workflow enqueue

What it means

App-aware routing under load.#

Tako stayed clean while doing product work a static reverse proxy does not need to do: route lookup, source IP derivation, per-client limiter accounting, app and instance selection, in-flight accounting, upstream peer construction, and forwarding header normalization.

The report still keeps static proxy references, p99, Memory, and clean-run percentages in view because those are the tuning levers. Future runs can isolate larger-VM behavior, external same-region load generation, and narrower Pingora session and upstream-proxy costs under 10k to 20k live TLS connections.

Method

Same conditions, public raw data.#

The public report intentionally omits hostnames, public IPs, private network addresses, peer names, and user identifiers.

HTTP path

Load generator, proxy, and app all run on the VM. The route is bench.test:18443, resolved to loopback, with Host and SNI set to bench.test.

Proxy matrix

Tako HTTP matrix tako-server 0.0.0-09b3dc6, feature rerun tako-server 0.0.0-958986f, nginx 1.24.0, HAProxy 2.8.16, Envoy 1.38.0, Caddy 2.11.3 with rate limiting.

Timing

10 second warmup, 30 second measurement window, HTTP/1.1 over TLS, 60 second request timeout, metrics sampled from /proc once per second.