High concurrency, large messages (P99 tail latency): gRPC (186ms) << TCP (1,049ms) < REST (1,202ms)
The protocol ranked best on average is the worst at the tail. Most teams choose protocols using microbenchmarks at low concurrency — exactly the regime where this ranking is least informative for production behavior.
Communication protocols are usually chosen by convention, not measurement — REST because it's familiar, gRPC because it's "fast," raw TCP because it's "simple." None of those instincts predict what actually happens under load. This page traces, down to the kernel, exactly where and why each protocol's performance characteristic breaks.
0x01
the_verdict — read this, skip the rest if you want
Use TCP when:
You control both endpoints
Payloads are small and uniform (<2KB)
Concurrency is bounded (<50)
Minimum median latency is critical
Tail latency controlled via semaphore
→ HFT order entry, internal hot paths, game servers, custom control planes
Use REST when:
Simplicity and debuggability matter
Concurrency is low (<50)
Broad client compatibility needed
2× TCP overhead is acceptable
Message size stays below 2KB
→ Public APIs, low-traffic internal services, human-readable debugging
Use gRPC when:
P99 tail latency SLAs exist
Concurrency is high (>50)
Payloads vary in size (2KB–32KB sweet spot)
Streaming RPCs needed
Internal microservice mesh
→ Latency-sensitive distributed systems, ML inference serving, high-concurrency backends
one-line protocol selection rule
If your system has tail latency SLAs and sees burst concurrency above 50 clients: use gRPC.
If you own both endpoints and concurrency is controlled: TCP + semaphore flow control.
REST is the right default only when neither of the above applies.
systems design principle
Average latency is what your system does on a good day. P99 latency is what your system does to your users on a normal day. P99 at c=500 is what your system does during a traffic spike — the moment that matters most.
TCP's 121× P99/P50 ratio means the slowest 1% of requests are 121× worse than the median. For a service handling 10,000 req/s, that's 100 requests per second experiencing 1-second+ latency. Optimize for tail, not mean.
Everything below this point is the evidence. Three layers, each deeper than the last: the numbers, the kernel-level mechanism behind them, and the edge cases that almost broke the methodology.
0x02
layer_1 — the_numbers
what happened, measured 9.9M times
┌─────────────────────────────────────────────────────────────────────┐
│ Platform : Ubuntu 24.04 LTS | Kernel : 6.8.0-101-generic │
│ CPU : Hybrid P/E-core | P-cores: 4500 MHz max │
│ Language : Go 1.25 | Serialization: protobuf (all 3) │
└─────────────────────────────────────────────────────────────────────┘
Message Sizes → 64B 256B 1KB 4KB 64KB (+ 13-point dense sweep)
Concurrency → 1 10 50 100 500 (concurrent clients)
Protocols → TCP (raw) REST (HTTP/1.1) gRPC (HTTP/2)
75 configurations × (concurrency × 1,000 requests) = 9,915,000 samples
CPU Pinning → Server pinned to CPU1 (P-core, 4500MHz) via taskset
Client allowed on CPUs 0,3,4,5,6,7 (P-cores only)
Priority → All three protocols use identical protobuf payloads
and the same Go runtime — overhead differences are
attributable to protocol mechanics, not encoding.
Baseline latency vs message size (concurrency = 1)
Size
TCP (µs)
REST (µs)
gRPC (µs)
gRPC/TCP
64B
18.8
35.8
47.9
2.5×
256B
17.3
34.2
45.0
2.6×
1KB
18.5
42.3
46.3
2.5×
4KB ⚠
21.2
55.4
50.4
2.4×
64KB
56.6
199.1
216.3
3.8×
TCP is consistently ~2.5× faster than REST/gRPC at small sizes, widening to 3.8× at 64KB. At 4KB, REST (55.4µs) exceeds gRPC (50.4µs) — the first sign of the crossover investigated in Layer 3.
Latency vs concurrency (64B messages) — linear scaling at small messages
Concurrency
TCP (µs)
REST (µs)
gRPC (µs)
REST/TCP
1
18.8
35.8
47.9
1.9×
10
94.7
145.4
194.6
1.5×
50
486.7
777.7
889.2
1.6×
100
936.8
1,631.0
1,848.3
1.7×
500
4,996.8
9,023.5
10,830.5
1.8×
All three scale roughly linearly at small messages. The REST/TCP ratio stays stable (1.5–1.9×) — REST overhead is largely fixed per request here. Tail behavior at large messages tells a completely different story.
Throughput at concurrency = 1, 64B messages
TCP achieves ~49,000 req/s, versus 24,000 for REST and 21,000 for gRPC. At higher concurrency all three converge as the bottleneck shifts from protocol overhead to server CPU capacity.
P99/P50 tail amplification — the core result, full table at 64KB
Protocol
Conc.
P50 (µs)
P99 (µs)
P99/P50
TCP
1
41
128
3.1×
TCP
10
479
897
1.9×
TCP
50
1,753
42,010
24×
TCP
100
3,489
140,215
40×
TCP
500
8,915
1,083,462
121×
REST
1
159
395
2.5×
REST
10
1,487
12,105
8.1×
REST
50
5,877
167,801
29×
REST
100
8,108
391,420
48×
REST
500
16,887
1,177,731
70×
gRPC
1
195
422
2.2×
gRPC
10
1,775
3,066
1.7×
gRPC
50
10,424
18,781
1.8×
gRPC
100
21,276
33,193
1.6×
gRPC
500
95,729
186,705
2.0×
TCP and REST exhibit a "knee" between c=10 and c=50 where tail ratio jumps an order of magnitude. gRPC stays flat (1.6–2.2×) at every concurrency level — HTTP/2's WINDOW_UPDATE flow control prevents unbounded kernel queue buildup, regardless of client count.
0x03
layer_2 — the_mechanism
why it happened, traced to specific kernel calls and functions
TCP: 80% sys time — kernel-dominated at every concurrency level, constant regardless of load.
REST/gRPC shift toward user space as concurrency rises: REST 46%→65% user, gRPC 44%→74% user (see chart below). At scale, the bottleneck in REST and gRPC isn't the kernel network stack — it's their own abstractions.
Chart — user vs system time at low and high concurrency
User vs system time (%), c=1 vs c=500. TCP's split never moves. REST and gRPC both shift sharply toward user-space at scale.
Function-level CPU breakdown — perf report, 4KB, c=1
Attached to running client processes (120K–137K samples). Both protocols spend ~77–79% of time blocked on network I/O. Of the active CPU samples:
Category
REST %
gRPC %
Protocol layer (framing/parsing)
4.02
3.56
Memory allocation (mallocgc)
4.91
3.29
Garbage collection
3.71
0.71
Memory copy (memmove)
1.32
2.20
Go scheduler
5.11
8.31
Kernel sync (futex, psi)
4.27
4.73
I/O wait (network)
~77%
~79%
REST GC cost — 5× higher
HTTP/1.1 parsing allocates many short-lived objects per request — header maps, string slices, body wrappers. gRPC's buffer pool reuses memory instead.
gRPC scheduler cost — 63% higher
HTTP/2 keeps dedicated goroutines (loopyWriter, frame reader, keepalive pinger), all requiring scheduling. runtime.procyield appears only in gRPC — confirms mutex spin-waiting.
Flame graphs — call stack depth and structure
Generated via perf record at 1KHz + call graph capture. Wider = more CPU time. File complexity alone reflects the story: TCP 149KB, REST 241KB, gRPC 332KB of flame graph data.
TCP — shallow, kernel-dominated
2–3 user-space frames before kernel. Zero visible futex calls anywhere — no synchronization overhead at all.
REST — deep HTTP/1.1 pipeline
6 user-space frames: Client → Transport → persistConn → bufio → net.Conn → syscall. Visible futex_wait from connection-pool locking.
gRPC — distributed, most complex
8+ user-space frames spread across HTTP/2 transport, protobuf serialization, stream management, flow control, keepalive — each a separate tower. tieredBufferPool.Get visible as a distinct allocation pattern vs REST's scattered mallocgc.
Each futex call is a mutex lock acquisition or release — pure synchronization overhead with no application work. sched_yield = 0 for TCP proves zero lock contention.
REST context switches grow super-linearly (HTTP/1.1's one-goroutine-per-connection model). At c=50, gRPC (18,765) nearly matches TCP (16,657) — HTTP/2 multiplexing pays off here.
Concurrency
TCP (K)
REST (K)
gRPC (K)
gRPC/TCP
1
113
236
273
2.4×
10
254
443
492
1.9×
50
588
1,593
3,938
6.7×
100
1,095
6,479
15,758
14.4×
500
12,088
174,887
331,541
27.4×
gRPC wins on context switches at high concurrency but loses badly on cache misses — 27× more than TCP at c=500. HTTP/2 per-stream state (flow control windows, HPACK tables) for 500 concurrent streams exceeds L1/L2 capacity, causing constant cache thrashing that partially offsets the multiplexing benefit.
0x04
layer_3 — the_edge_cases
where the model almost broke, and how each anomaly was run down
Three anomalies surfaced during the experiments. Each could have been a methodology bug. Each was instead traced to a specific, confirmed mechanism — with alternative hypotheses explicitly ruled out.
Edge case 1 — page fault step function at exactly 32KB
Observed: gRPC page faults jumped discontinuously at 32KB — a step, not a gradual rise.
Page faults per 10K requests vs message size (c=1). REST climbs steadily from ~2KB; gRPC stays flat then spikes sharply at exactly 32KB.
confirmed root cause — Go allocator slab/heap boundary
Go runtime constant: _MaxSmallSize = 32768 (runtime/malloc.go).
Below 32KB → thread-local size-class pool (mcache) → slab page already mapped → no fault.
At or above 32KB → large-object allocator (mheap) → fresh heap page requested from OS → guaranteed fault per allocation.
hypothesis ruled out #1 — HTTP/2 frame splitting
Tested 8KB, 16KB, and 32KB explicit frame buffer sizes. All three show identical step functions at exactly 32KB, with jump factors of 3.2×, 3.5×, and 3.1× respectively. Frame size has zero effect on the boundary location — HTTP/2 framing ruled out.
hypothesis ruled out #2 — GC pressure
GOGC=off (GC fully disabled): absolute fault counts explode (~48–58K, since pages are never reclaimed), but the step at exactly 32KB persists in both conditions — confirming the cause is the allocator path change, not GC behavior.
Size
Normal GC
GOGC=off
Step preserved?
29KB
3,112
48,093
—
30KB
2,845
48,121
—
31KB
3,024
48,061
—
32KB
8,574
48,854
Yes
33KB
8,938
58,801
Yes
control test — the 64KB boundary, to prove 32KB is specific
If page faults simply rise with any large message, the 64KB boundary should show a similar step. It doesn't — page faults grow linearly through 64KB (9,618 → 15,590 across 62–66KB, no discontinuity), confirming 32KB specifically is the Go allocator threshold, not a general large-message effect.
Size
Normal GC
GOGC=off
62KB
9,618
88,823
63KB
12,020
88,804
64KB
14,693
88,991
65KB
15,405
98,999
66KB
15,590
99,003
Edge case 2 — page fault three-zone behavior across message sizes
Observed: REST vs gRPC page fault ordering is non-monotonic — three behaviorally distinct zones.
Zone 1: <2KB gRPC overhead > REST
Zone 2: 2–31KB REST overhead > gRPC
Zone 3: ≥32KB gRPC overhead rises sharply
Size
REST PF
gRPC PF
REST/gRPC
Zone
512B
3,573
3,615
1.0×
Zone 1
1KB
3,360
3,363
1.0×
Zone 1
2KB
4,688
3,674
1.3×
Zone 2
4KB
7,178
3,706
1.9×
Zone 2
8KB
8,513
5,362
1.6×
Zone 2
16KB
17,185
5,889
2.9×
Zone 2
32KB
23,736
19,326
1.2×
Zone 3 onset
64KB
33,125
25,035
1.3×
Zone 3
continuous vs fragmented memory — the architectural root
HTTP/1.1 allocates the response body as one fresh contiguous buffer per request. HTTP/2 pre-allocates fixed-size frame buffers via tieredBufferPool and reuses them. Below 32KB, gRPC's reuse wins. At 32KB, both protocols cross Go's large-object threshold and the pre-allocation advantage disappears — gRPC's extra per-stream state then means it hits the threshold more often, narrowing the Zone 2 gap.
Edge case 3 — REST latency exceeds gRPC at 4KB (the Zone 1→2 crossover, confirmed in latency too)
At 4KB, c=1: REST (55.4µs) > gRPC (50.4µs) — reversing the 1KB ordering. Reproduced across multiple runs, confirmed by hardware counters showing REST consuming 15% more cycles and 18% more cache references at 4KB — the inverse of the 1KB relationship.
Crossover 1 (~1–2KB): GRADUAL — gap closes over 4 data points. Matches the architectural cause (HTTP/2 setup cost gradually overtaken by REST's buffer growth).
Crossover 2 (24KB→32KB): SHARP — an 82µs swing in one step. Matches the page fault step function exactly. Hard boundary, not architecture.
0x05
the_fix — semaphore flow control for TCP
TCP's 121× tail ratio was diagnosed as a missing application-level flow control problem, not a fundamental TCP limitation. Five lines of Go test that hypothesis directly.
root cause
At c=500 with 64KB messages, all 500 goroutines call write immediately. The kernel send buffer saturates; late arrivals queue inside the kernel, growing unboundedly. The last goroutine waits for all 499 others to drain.
// Original — all 500 goroutines blast simultaneously
go func() { sendRecv(conn, data) }() // × 500
// Optimized — semaphore limits max in-flight to 50
sem := make(chan struct{}, maxInflight) // maxInflight = 50
go func() {
sem <- struct{}{} // acquire — blocks when 50 already in-flight
sendRecv(conn, data)
<-sem // release
}()
The semaphore limits the sender side but provides no receiver-driven backpressure. The 50 active goroutines still write without acknowledgment that the server has consumed the data. gRPC's WINDOW_UPDATE is receiver-driven — a full TCP solution requires an equivalent, left as future work.
page fault reduction, explained
Before: 500 goroutines simultaneously allocate 64KB buffers → 500 large-object allocations → 1.5M page faults. After: at most 50 hold buffers at once → 47× fewer faults. Confirms the original explosion was caused by simultaneous large-object allocation, not the protocol itself.
the tradeoff
Context switches rise 1.7× — goroutines now block/unblock on the semaphore channel, generating extra scheduling events. Queue discipline simply moves from the kernel (unbounded, unmeasurable) to the application layer (bounded, measurable, controllable). Throughput drops only 7% in exchange for a 16.8× P99 improvement — a clearly worthwhile trade for any tail-sensitive system.
0x06
limits — what this doesn't prove
Methodology validity and acknowledged limitations
Go runtime inseparability: GC and goroutine scheduler costs can't be fully isolated from protocol costs without a bare-metal C++ baseline. Partially mitigated by function-level perf attribution.
Single-node only: network latency would dominate µs-level differences in a distributed deployment. Relative ratios, tail patterns, and sync costs remain applicable; tail effects would likely be amplified by network variability.
Symmetric echo design: send-path and receive-path overhead are combined per measurement. Asymmetric designs (small ACK, or large server-initiated response) would isolate each direction — future work.
Client-side profiling only: all perf measurements were collected on the client process. Server-side profiling would reveal server-side overhead separately — future work.
Go allocator threshold: definitive confirmation requires recompiling Go runtime with a modified _MaxSmallSize. Two indirect experiments (frame-size sweep + GOGC=off + the 64KB control) provide strong but indirect evidence.
Warmup: one warmup request per client. gRPC's HPACK table init and HTTP/2 stream 0 setup may not be fully amortized in a single warmup.
Why this method is still valid despite the limits
Protocol overhead is measured by holding payload size and serialization constant across all three implementations — same Go runtime, same protobuf payloads. Observed differences in CPU cycles, cache references, and page faults are therefore attributable to protocol-level mechanisms, not encoding cost.
direct validation
Function-level perf data confirms protobuf deserialization is visible in both REST's and gRPC's call graphs at equivalent depth — present in both, costing both, proving the serialization-control assumption holds.
0x07
references
B. Raghavan et al., "Network stack overhead analysis," Cornell University, Tech. Rep., 2023.
R. T. Fielding, "Architectural styles and the design of network-based software architectures," Ph.D. dissertation, UC Irvine, 2000.
This is the same instinct I want to bring to order entry systems eventually — don't trust the average, find the mechanism behind the tail. The 9.9M samples here are a proxy for the kind of obsession that matters more once the unit is microseconds instead of milliseconds.