Protocol Performance Analysis

TCP vs REST vs gRPC — kernel-level characterization on a single-node Linux system
Lokesh Aravapalli (IMT2022577) · Dharoori Srinivas Acharya (IMT2022066) · IIIT Bangalore, SSP Course Project
SYSTEMS_INFRA KERNEL_PROFILING HW_COUNTERS Go 1.25 perf 6.8.12 FlameGraph protobuf LINUX_KERNEL_6.8.0 Ubuntu_24.04 9.9M_SAMPLES 75_EXPERIMENTS
0x00

the_claim

10² 10³ 10⁴ 10⁵ 10⁶ c=1c=10c=50c=100c=500 TCP — fastest at c=1, worst at c=500 REST — never wins gRPC — worst at c=1, best at c=500

P99 latency (µs, log scale) vs concurrency, 64KB messages. The ranking inverts completely between the left edge and the right edge of this chart.

the protocol inversion effect
Low concurrency, small messages (avg latency):
  TCP (18µs) < REST (35µs) < gRPC (47µs)

High concurrency, large messages (P99 tail latency):
  gRPC (186ms) << TCP (1,049ms) < REST (1,202ms)

The protocol ranked best on average is the worst at the tail. Most teams choose protocols using microbenchmarks at low concurrency — exactly the regime where this ranking is least informative for production behavior.

Communication protocols are usually chosen by convention, not measurement — REST because it's familiar, gRPC because it's "fast," raw TCP because it's "simple." None of those instincts predict what actually happens under load. This page traces, down to the kernel, exactly where and why each protocol's performance characteristic breaks.

0x01

the_verdict — read this, skip the rest if you want

Use TCP when:
→ HFT order entry, internal hot paths, game servers, custom control planes
Use REST when:
→ Public APIs, low-traffic internal services, human-readable debugging
Use gRPC when:
→ Latency-sensitive distributed systems, ML inference serving, high-concurrency backends
one-line protocol selection rule
If your system has tail latency SLAs and sees burst concurrency above 50 clients: use gRPC.
If you own both endpoints and concurrency is controlled: TCP + semaphore flow control.
REST is the right default only when neither of the above applies.
systems design principle
Average latency is what your system does on a good day. P99 latency is what your system does to your users on a normal day. P99 at c=500 is what your system does during a traffic spike — the moment that matters most.

TCP's 121× P99/P50 ratio means the slowest 1% of requests are 121× worse than the median. For a service handling 10,000 req/s, that's 100 requests per second experiencing 1-second+ latency. Optimize for tail, not mean.

Everything below this point is the evidence. Three layers, each deeper than the last: the numbers, the kernel-level mechanism behind them, and the edge cases that almost broke the methodology.

0x02

layer_1 — the_numbers

what happened, measured 9.9M times
┌─────────────────────────────────────────────────────────────────────┐ │ Platform : Ubuntu 24.04 LTS | Kernel : 6.8.0-101-generic │ │ CPU : Hybrid P/E-core | P-cores: 4500 MHz max │ │ Language : Go 1.25 | Serialization: protobuf (all 3) │ └─────────────────────────────────────────────────────────────────────┘ Message Sizes → 64B 256B 1KB 4KB 64KB (+ 13-point dense sweep) Concurrency → 1 10 50 100 500 (concurrent clients) Protocols → TCP (raw) REST (HTTP/1.1) gRPC (HTTP/2) 75 configurations × (concurrency × 1,000 requests) = 9,915,000 samples CPU Pinning → Server pinned to CPU1 (P-core, 4500MHz) via taskset Client allowed on CPUs 0,3,4,5,6,7 (P-cores only) Priority → All three protocols use identical protobuf payloads and the same Go runtime — overhead differences are attributable to protocol mechanics, not encoding.
Baseline latency vs message size (concurrency = 1)
SizeTCP (µs)REST (µs)gRPC (µs)gRPC/TCP
64B18.835.847.92.5×
256B17.334.245.02.6×
1KB18.542.346.32.5×
4KB ⚠21.255.450.42.4×
64KB56.6199.1216.33.8×

TCP is consistently ~2.5× faster than REST/gRPC at small sizes, widening to 3.8× at 64KB. At 4KB, REST (55.4µs) exceeds gRPC (50.4µs) — the first sign of the crossover investigated in Layer 3.

Latency vs concurrency (64B messages) — linear scaling at small messages
ConcurrencyTCP (µs)REST (µs)gRPC (µs)REST/TCP
118.835.847.91.9×
1094.7145.4194.61.5×
50486.7777.7889.21.6×
100936.81,631.01,848.31.7×
5004,996.89,023.510,830.51.8×

All three scale roughly linearly at small messages. The REST/TCP ratio stays stable (1.5–1.9×) — REST overhead is largely fixed per request here. Tail behavior at large messages tells a completely different story.

Throughput at concurrency = 1, 64B messages

TCP achieves ~49,000 req/s, versus 24,000 for REST and 21,000 for gRPC. At higher concurrency all three converge as the bottleneck shifts from protocol overhead to server CPU capacity.

P99/P50 tail amplification — the core result, full table at 64KB
ProtocolConc.P50 (µs)P99 (µs)P99/P50
TCP1411283.1×
TCP104798971.9×
TCP501,75342,01024×
TCP1003,489140,21540×
TCP5008,9151,083,462121×
REST11593952.5×
REST101,48712,1058.1×
REST505,877167,80129×
REST1008,108391,42048×
REST50016,8871,177,73170×
gRPC11954222.2×
gRPC101,7753,0661.7×
gRPC5010,42418,7811.8×
gRPC10021,27633,1931.6×
gRPC50095,729186,7052.0×

TCP and REST exhibit a "knee" between c=10 and c=50 where tail ratio jumps an order of magnitude. gRPC stays flat (1.6–2.2×) at every concurrency level — HTTP/2's WINDOW_UPDATE flow control prevents unbounded kernel queue buildup, regardless of client count.

0x03

layer_2 — the_mechanism

why it happened, traced to specific kernel calls and functions
Hardware performance counters — 64B, c=1, 10K requests
CounterTCPRESTgRPC
CPU Cycles567M1,350M1,539M
Instructions751M1,549M1,674M
IPC1.321.141.08
Cache References2.4M14.5M18.6M
Cache Misses222K456K548K
Branch Misses2.2M3.8M4.6M
Context Switches14,70842,62945,193
User Time33ms (20%)187ms (46%)245ms (53%)
Sys Time131ms (80%)215ms (54%)210ms (47%)
where the work happens
TCP: 80% sys time — kernel-dominated at every concurrency level, constant regardless of load.
REST/gRPC shift toward user space as concurrency rises: REST 46%→65% user, gRPC 44%→74% user (see chart below). At scale, the bottleneck in REST and gRPC isn't the kernel network stack — it's their own abstractions.
Chart — user vs system time at low and high concurrency
0 20 40 60 80 TCP REST gRPC user (solid=c1, light=c500) sys (solid=c1, light=c500)

User vs system time (%), c=1 vs c=500. TCP's split never moves. REST and gRPC both shift sharply toward user-space at scale.

Function-level CPU breakdown — perf report, 4KB, c=1

Attached to running client processes (120K–137K samples). Both protocols spend ~77–79% of time blocked on network I/O. Of the active CPU samples:

CategoryREST %gRPC %
Protocol layer (framing/parsing)4.023.56
Memory allocation (mallocgc)4.913.29
Garbage collection3.710.71
Memory copy (memmove)1.322.20
Go scheduler5.118.31
Kernel sync (futex, psi)4.274.73
I/O wait (network)~77%~79%
REST GC cost — 5× higher
HTTP/1.1 parsing allocates many short-lived objects per request — header maps, string slices, body wrappers. gRPC's buffer pool reuses memory instead.
gRPC scheduler cost — 63% higher
HTTP/2 keeps dedicated goroutines (loopyWriter, frame reader, keepalive pinger), all requiring scheduling. runtime.procyield appears only in gRPC — confirms mutex spin-waiting.
Flame graphs — call stack depth and structure

Generated via perf record at 1KHz + call graph capture. Wider = more CPU time. File complexity alone reflects the story: TCP 149KB, REST 241KB, gRPC 332KB of flame graph data.

TCP — shallow, kernel-dominated
2–3 user-space frames before kernel. Zero visible futex calls anywhere — no synchronization overhead at all.
REST — deep HTTP/1.1 pipeline
6 user-space frames: Client → Transport → persistConn → bufio → net.Conn → syscall. Visible futex_wait from connection-pool locking.
gRPC — distributed, most complex
8+ user-space frames spread across HTTP/2 transport, protobuf serialization, stream management, flow control, keepalive — each a separate tower. tieredBufferPool.Get visible as a distinct allocation pattern vs REST's scattered mallocgc.
Syscall analysis — perf trace, 64B, c=1, 10K requests
SyscallTCPRESTgRPCSignificance
futex1,33548,25655,33236–41× higher in REST/gRPC
read29,64420,30526,384TCP does most raw I/O directly
write19,83610,38114,137REST/gRPC buffer and batch writes
epoll_pwait19,49918,89932,667gRPC monitors more I/O event sources
nanosleep3,8778,8126,932lock-contention sleep-and-retry
sched_yield03151TCP never contends on any lock
futex calls per request — the dominant overhead metric
TCP: 0.13/request  |  REST: 4.8/request  |  gRPC: 5.5/request

Each futex call is a mutex lock acquisition or release — pure synchronization overhead with no application work. sched_yield = 0 for TCP proves zero lock contention.
Concurrency scaling — context switches, cache misses, user/sys split
ConcurrencyTCPRESTgRPCREST/TCP
11,4394,4375,1273.1×
102,67614,89717,6495.6×
5016,65772,12518,7654.3×
10038,281149,32931,9663.9×
500122,112914,456165,8407.5×

REST context switches grow super-linearly (HTTP/1.1's one-goroutine-per-connection model). At c=50, gRPC (18,765) nearly matches TCP (16,657) — HTTP/2 multiplexing pays off here.

ConcurrencyTCP (K)REST (K)gRPC (K)gRPC/TCP
11132362732.4×
102544434921.9×
505881,5933,9386.7×
1001,0956,47915,75814.4×
50012,088174,887331,54127.4×

gRPC wins on context switches at high concurrency but loses badly on cache misses — 27× more than TCP at c=500. HTTP/2 per-stream state (flow control windows, HPACK tables) for 500 concurrent streams exceeds L1/L2 capacity, causing constant cache thrashing that partially offsets the multiplexing benefit.

0x04

layer_3 — the_edge_cases

where the model almost broke, and how each anomaly was run down

Three anomalies surfaced during the experiments. Each could have been a methodology bug. Each was instead traced to a specific, confirmed mechanism — with alternative hypotheses explicitly ruled out.

Edge case 1 — page fault step function at exactly 32KB

Observed: gRPC page faults jumped discontinuously at 32KB — a step, not a gradual rise.

0 10K 20K 30K 512B1K2K4K8K16K32K64K 32KB boundary REST gRPC

Page faults per 10K requests vs message size (c=1). REST climbs steadily from ~2KB; gRPC stays flat then spikes sharply at exactly 32KB.

At 31KB — three gRPC frame sizes
8KB frame: ~4,117 PF · 16KB frame: ~4,106 PF · 32KB frame: ~3,394 PF

→ low and stable across all variants
At 32KB — three gRPC frame sizes
8KB frame: ~16,828 PF · 16KB frame: ~14,043 PF · 32KB frame: ~16,081 PF

→ 4× simultaneous jump, independent of frame size
confirmed root cause — Go allocator slab/heap boundary
Go runtime constant: _MaxSmallSize = 32768 (runtime/malloc.go).

Below 32KB → thread-local size-class pool (mcache) → slab page already mapped → no fault.
At or above 32KB → large-object allocator (mheap) → fresh heap page requested from OS → guaranteed fault per allocation.
hypothesis ruled out #1 — HTTP/2 frame splitting
Tested 8KB, 16KB, and 32KB explicit frame buffer sizes. All three show identical step functions at exactly 32KB, with jump factors of 3.2×, 3.5×, and 3.1× respectively. Frame size has zero effect on the boundary location — HTTP/2 framing ruled out.
hypothesis ruled out #2 — GC pressure
GOGC=off (GC fully disabled): absolute fault counts explode (~48–58K, since pages are never reclaimed), but the step at exactly 32KB persists in both conditions — confirming the cause is the allocator path change, not GC behavior.
SizeNormal GCGOGC=offStep preserved?
29KB3,11248,093—
30KB2,84548,121—
31KB3,02448,061—
32KB8,57448,854Yes
33KB8,93858,801Yes
control test — the 64KB boundary, to prove 32KB is specific
If page faults simply rise with any large message, the 64KB boundary should show a similar step. It doesn't — page faults grow linearly through 64KB (9,618 → 15,590 across 62–66KB, no discontinuity), confirming 32KB specifically is the Go allocator threshold, not a general large-message effect.
SizeNormal GCGOGC=off
62KB9,61888,823
63KB12,02088,804
64KB14,69388,991
65KB15,40598,999
66KB15,59099,003
Edge case 2 — page fault three-zone behavior across message sizes

Observed: REST vs gRPC page fault ordering is non-monotonic — three behaviorally distinct zones.

Zone 1: <2KB
gRPC overhead > REST
Zone 2: 2–31KB
REST overhead > gRPC
Zone 3: ≥32KB
gRPC overhead rises sharply
SizeREST PFgRPC PFREST/gRPCZone
512B3,5733,6151.0×Zone 1
1KB3,3603,3631.0×Zone 1
2KB4,6883,6741.3×Zone 2
4KB7,1783,7061.9×Zone 2
8KB8,5135,3621.6×Zone 2
16KB17,1855,8892.9×Zone 2
32KB23,73619,3261.2×Zone 3 onset
64KB33,12525,0351.3×Zone 3
continuous vs fragmented memory — the architectural root
HTTP/1.1 allocates the response body as one fresh contiguous buffer per request. HTTP/2 pre-allocates fixed-size frame buffers via tieredBufferPool and reuses them. Below 32KB, gRPC's reuse wins. At 32KB, both protocols cross Go's large-object threshold and the pre-allocation advantage disappears — gRPC's extra per-stream state then means it hits the threshold more often, narrowing the Zone 2 gap.
Edge case 3 — REST latency exceeds gRPC at 4KB (the Zone 1→2 crossover, confirmed in latency too)

At 4KB, c=1: REST (55.4µs) > gRPC (50.4µs) — reversing the 1KB ordering. Reproduced across multiple runs, confirmed by hardware counters showing REST consuming 15% more cycles and 18% more cache references at 4KB — the inverse of the 1KB relationship.

At 1KB (gRPC more expensive)
REST 42.3µs · gRPC 46.3µs
REST PF 3,360 · gRPC PF 3,363
gRPC's fixed HTTP/2 setup cost dominates.
At 4KB (REST more expensive)
REST 55.4µs · gRPC 50.4µs
REST PF 7,178 · gRPC PF 3,706
REST's buffer reallocation overtakes gRPC's setup cost.
Dense latency sweep — 13 message sizes confirm both crossovers directly
SizeTCP (µs)REST (µs)gRPC (µs)Note
512B193744gRPC 7µs slower than REST
1KB ←183942gap closing — only 3µs
2KB194344nearly tied
4KB ⚠215347REST clearly slower — confirmed
8KB216351REST 23% slower than gRPC
16KB268562peak gRPC advantage
24KB3210270gRPC still 32µs faster
32KB ←32121171sharp flip — gRPC 50µs SLOWER
64KB48180199REST 19µs faster than gRPC
two crossovers, two distinct signatures
Crossover 1 (~1–2KB): GRADUAL — gap closes over 4 data points. Matches the architectural cause (HTTP/2 setup cost gradually overtaken by REST's buffer growth).

Crossover 2 (24KB→32KB): SHARP — an 82µs swing in one step. Matches the page fault step function exactly. Hard boundary, not architecture.
0x05

the_fix — semaphore flow control for TCP

TCP's 121× tail ratio was diagnosed as a missing application-level flow control problem, not a fundamental TCP limitation. Five lines of Go test that hypothesis directly.

root cause
At c=500 with 64KB messages, all 500 goroutines call write immediately. The kernel send buffer saturates; late arrivals queue inside the kernel, growing unboundedly. The last goroutine waits for all 499 others to drain.
// Original — all 500 goroutines blast simultaneously go func() { sendRecv(conn, data) }() // × 500 // Optimized — semaphore limits max in-flight to 50 sem := make(chan struct{}, maxInflight) // maxInflight = 50 go func() { sem <- struct{}{} // acquire — blocks when 50 already in-flight sendRecv(conn, data) <-sem // release }()
Results — before vs after (64KB, c=500)
Before (raw TCP, no flow control)
44,735µsavg latency 121×P99/P50 tail ratio 1,083,462µsP99 latency 1,557,815page faults
After (semaphore, maxInflight=50)
5,132µsavg latency — 8.7× better 18×P99/P50 tail ratio — 6.7× better 64,388µsP99 latency — 16.8× better 32,961page faults — 47× fewer
MetricBeforeAfterImprovement
Avg latency44,735µs5,132µs8.7× faster
P508,915µs3,556µs2.5× faster
P991,083,462µs64,388µs16.8× faster
P99/P50 ratio121×18×6.7× better
Page faults1,557,81532,96147× fewer
Context switches998,6901,743,6851.7× more
Throughput10,466 req/s9,703 req/s7% lower
why 18× residual, not 2× like gRPC
The semaphore limits the sender side but provides no receiver-driven backpressure. The 50 active goroutines still write without acknowledgment that the server has consumed the data. gRPC's WINDOW_UPDATE is receiver-driven — a full TCP solution requires an equivalent, left as future work.
page fault reduction, explained
Before: 500 goroutines simultaneously allocate 64KB buffers → 500 large-object allocations → 1.5M page faults. After: at most 50 hold buffers at once → 47× fewer faults. Confirms the original explosion was caused by simultaneous large-object allocation, not the protocol itself.
the tradeoff
Context switches rise 1.7× — goroutines now block/unblock on the semaphore channel, generating extra scheduling events. Queue discipline simply moves from the kernel (unbounded, unmeasurable) to the application layer (bounded, measurable, controllable). Throughput drops only 7% in exchange for a 16.8× P99 improvement — a clearly worthwhile trade for any tail-sensitive system.
0x06

limits — what this doesn't prove

Methodology validity and acknowledged limitations
Why this method is still valid despite the limits

Protocol overhead is measured by holding payload size and serialization constant across all three implementations — same Go runtime, same protobuf payloads. Observed differences in CPU cycles, cache references, and page faults are therefore attributable to protocol-level mechanisms, not encoding cost.

direct validation
Function-level perf data confirms protobuf deserialization is visible in both REST's and gRPC's call graphs at equivalent depth — present in both, costing both, proving the serialization-control assumption holds.
0x07

references

  1. B. Raghavan et al., "Network stack overhead analysis," Cornell University, Tech. Rep., 2023.
  2. R. T. Fielding, "Architectural styles and the design of network-based software architectures," Ph.D. dissertation, UC Irvine, 2000.
  3. gRPC Authors, "grpc documentation," grpc.io/docs, 2024.
  4. B. Gregg, "The flame graph," Communications of the ACM, vol. 59, no. 6, pp. 48–57, 2016.
  5. Go Team, "net/http package documentation," pkg.go.dev/net/http, 2024.
  6. Perf Wiki, "Linux perf tools documentation," perf.wiki.kernel.org, 2024.
why I actually care about this
This is the same instinct I want to bring to order entry systems eventually — don't trust the average, find the mechanism behind the tail. The 9.9M samples here are a proxy for the kind of obsession that matters more once the unit is microseconds instead of milliseconds.