eBPF service mesh: kill sidecars, reclaim your nodes
SEO Meta Description: Replace Envoy sidecars with eBPF XDP and TC hooks. Learn BPF map design, per-core overhead math, and kTLS offload to boost Kubernetes node density.
TL;DR
Envoy sidecars are eating your node capacity. A production-grade eBPF-native L7 proxy using XDP ingress hooks and TC egress hooks can eliminate per-pod sidecar overhead entirely — freeing 0.2–0.5 vCPU and 60–120 MB RAM per pod. The math compounds fast. Wire it up right and your ops team won’t hate you for it.
The sidecar tax is real
The sidecar model made sense in 2017. In 2026, running 300 pods per node, it’s a liability.
Most teams measure sidecar overhead in isolation. One Envoy proxy looks cheap. Multiply across node density and the picture changes:
| Metric | Envoy Sidecar (per pod) | eBPF Native | Delta |
|---|---|---|---|
| Memory overhead | ~80 MB | ~2 MB (BPF maps) | −97% |
| Steady-state CPU | 0.1–0.3 cores | 0.01–0.05 cores | −85% |
| p99 add-on latency | 3–8 ms | 0.1–0.4 ms | −94% |
| TLS termination path | Userspace (Envoy) | kTLS kernel offload | kernel vs. user |
At 200 pods per node, the sidecar tax burns 20–60 cores on pure proxy work. That’s capacity you’re paying for and not shipping features with.
The architecture: XDP + TC hooks + BPF maps
Layer 1 — XDP ingress at wire speed
XDP (eXpress Data Path) runs your BPF program before a socket buffer (sk_buff) is even allocated. This is the fast path for connection dispatch and early drop decisions. At this layer you parse Ethernet/IP/TCP headers, look up connection state in a BPF map, and either pass, drop, or redirect.
Note: Pseudocode for illustration — helper functions and error paths omitted for clarity.
SEC("xdp")
int xdp_l4_dispatch(struct xdp_md *ctx) {
struct conn_key key = parse_five_tuple(ctx);
struct conn_state *state = bpf_map_lookup_elem(&conn_track_map, &key);
if (!state) return XDP_PASS; // new connection, let TC handle it
return xdp_redirect_map(&backend_map, state->backend_idx);
}
XDP can’t do L7 parsing alone — it lacks access to reassembled TCP streams. That’s where TC takes over.
Layer 2 — TC hooks for L7 inspection
TC (Traffic Control) cls_bpf hooks attach at tc ingress and tc egress on the veth pair of each pod’s network namespace. Here you have full sk_buff access: HTTP/2 headers, gRPC metadata, TLS SNI. The constraint to plan around: the BPF verifier enforces a 1M instruction limit and 512-byte stack. Stateful L7 parsing requires spilling state into per-CPU maps between tail calls.
Note: Pseudocode for illustration — helper functions and error paths omitted for clarity.
SEC("tc")
int tc_l7_inspect(struct __sk_buff *skb) {
struct l7_ctx *ctx = bpf_map_lookup_elem(&per_cpu_ctx, &zero);
if (!ctx) return TC_ACT_OK;
parse_http2_frame(skb, ctx);
return route_to_backend(ctx);
}
Layer 3 — BPF map design for connection tracking
This is where most eBPF service mesh POCs fall apart. Your map design determines both performance and correctness.
BPF_MAP_TYPE_LRU_PERCPU_HASHfor connection state: per-CPU eliminates lock contention, LRU evicts stale entries automaticallyBPF_MAP_TYPE_ARRAY_OF_MAPSfor backend pools: allows atomic backend rotation without lockingBPF_MAP_TYPE_RINGBUFfor telemetry export: replaces perf events with lower overhead
Per-CPU hash maps on a 64-core node with 10K active connections show lookup times of 50–200 ns versus 800–2000 ns for shared hash maps under contention. That gap matters at scale.
kTLS: removing the last userspace bottleneck
Without kTLS, even a sidecar-free eBPF mesh hits a wall: TLS termination forces a context switch into a userspace process. kTLS moves symmetric crypto (AES-GCM via AES-NI) into the kernel’s TLS record layer. Combined with the SO_KTLS socket option, your BPF program can inspect plaintext after kernel decrypt — no userspace roundtrip.
Cipher suite compatibility matters. TLS 1.2 with AES-GCM offloads cleanly; CBC suites don’t. TLS 1.3 is strongly preferred but not strictly required. Audit your negotiated cipher suites before assuming offload is active — a misconfigured client advertising only CBC suites will silently fall back to userspace TLS and negate the gains. Kernel 5.19+ is required for the full kTLS send/receive path.
Per-core overhead math: how many pods can you actually run?
On a 32-core node at 10 Gbps line rate, XDP handles ~15 Mpps per core in busy-poll mode. Realistic L7 eBPF path: 400K–800K RPS per core including map lookups and tail calls.
At 500 RPS/pod average: one core services ~800 pods’ worth of proxy work. Compare that to Envoy sidecars at 0.2 cores each — 200 pods burns 40 cores on proxying alone.
eBPF proxy overhead for 200 pods: 0.25 cores. Sidecar overhead: 40 cores. That’s your node density win.
Operational reality
The hardest part isn’t the BPF code — it’s the tooling. You need BTF-enabled kernels, libbpf 1.0+, and CO-RE (Compile Once, Run Everywhere) to ship across kernel versions without recompiling. Plan for this from day one.
Security gap: mTLS responsibility shifts to you
Critical: Eliminating userspace sidecars means eliminating the component that historically owned mTLS certificate chain validation. Envoy verifies peer certificates as part of its connection handling. Move to an eBPF-native path and that verification doesn’t happen automatically — nothing in the kernel validates your workload identity certificates unless you explicitly implement it.
Your BPF-layer architecture must account for:
- SPIFFE/SVID validation — certificate fetching from the trust domain and validation against the workload identity store must be re-implemented, typically via a privileged userspace agent that populates a BPF map with allowed peer identities
- Certificate rotation — BPF maps holding allowed peer certificates or derived session keys must update atomically on rotation without dropping live connections
- Revocation — without a sidecar handling CRL or OCSP checks, you need an explicit revocation propagation path into your BPF policy maps
Teams migrating from Istio or Linkerd have shipped with mutual auth silently disabled at the workload level because they assumed the eBPF layer inherited this behavior. It doesn’t. Treat mTLS validation as a first-class design requirement, not an afterthought.
Conclusion: 4 actionable takeaways
-
Measure before you migrate. Capture baseline per-pod CPU and memory for your Envoy sidecars. You need the before/after numbers to justify the kernel expertise investment to leadership.
-
Start with TC hooks, not XDP. XDP’s constraints are non-trivial. Get L7 parsing working in TC first, then push hot paths to XDP once your BPF map schema is stable.
-
Design BPF maps for per-CPU access from the start. Retrofitting per-CPU maps into a shared-map design mid-migration is painful. Get the schema right in week one.
-
Explicitly re-implement mTLS validation. Don’t assume the kernel inherits certificate chain verification from your old sidecar. Design a privileged userspace agent to populate peer-identity BPF maps and treat certificate rotation and revocation as hard requirements before going to production.
Tags: kubernetes microservices backend devops cloud