~/blog/kubernetes $ cat kubernetes-network-stack-2026.md
title:

Our Kubernetes Network Stack in 2026: Still Calico, BGP and BFD

date:
categories: kubernetes
reading: 10 min

Every year or so I sit down and ask the same question: is it time to change the network stack in our Kubernetes clusters? In 2026 the answer is once again no. This post is my attempt to write down why, mostly so I don’t have to re-discover the same dead ends next year.

What we run today

At work, we have Kubernetes clusters with a mix of bare-metal machines and VMs. To handle networking for them, we run Calico with the BIRD backend and plain BGP. No overlay, no IPIP, no VXLAN. Pod and service CIDRs are routed in the fabric like any other network, so every pod has a real, reachable address and traffic leaves the node without SNAT. Each node peers with a pair of route reflectors, the node-to-node mesh is off, and nodes announce their pod block plus the aggregated service CIDR. Felix runs the iptables dataplane (with the nft backend, as on any modern distro), and kube-proxy is in nftables mode.

This design has served us well. Without encapsulation there is no overhead on the data path, so pod-to-pod traffic runs at the speed of the NIC, and MTU is whatever the fabric gives you. Because nothing is NATed, the source IP you see on a firewall or in an application log is the real pod IP, which makes network policies and audits a lot simpler. Pod routes are visible from the network side too, so when we work with the network team on an incident we look at the same routing table. Announcing the service CIDR means ClusterIPs are reachable from outside the cluster without a load balancer in front. And the same setup works on a bare-metal host and in a VM, there is no “VM flavour” of the network to maintain.

Our clusters live in isolated data centers, so they are reasonably well protected from the outside already. We still don’t trust that alone and use Calico network policies heavily, both NetworkPolicy and GlobalNetworkPolicy. That got complicated enough that I wrote a Calico policy visualiser earlier this year, which some of you may have already seen or used. Keeping the policy model we have built around is one more reason not to switch CNI lightly.

It’s a boring setup, and I mean that as a compliment. When something breaks, birdcl show protocols, ip route get and nft list table ip kube-proxy tell you almost everything. There is nothing exotic in the data path, so standard Linux tools like tcpdump, conntrack and ip route show you the truth. The whole team knows how to read it.

Why Calico

Two reasons, and both are about BGP.

The first is that Calico’s BGP support is mature. Route reflectors, iBGP and eBGP, per-node ASNs, custom listen ports, service CIDR advertisement, all of it has worked for years. We don’t need anything fancy here, we need it to be predictable.

The second is BFD. With BGP alone, when a node goes down for any reason (a kernel panic, a dead hypervisor host, a pulled cable), the route reflectors keep believing it is there until the hold timer expires. For all that time the fabric keeps sending traffic for that node’s pod block, and its share of the service CIDR, into a black hole. Kubernetes itself notices nothing for a while either, so from the user’s side it is simply 90+ seconds of timeouts. That is not acceptable for us. With BFD the failure is detected in well under a second, the routes are withdrawn, and the rest of the cluster carries on as if nothing happened.

Calico open source does not support BFD. The feature request has been open since 2021, Calico Enterprise has it, and the latest open-source release (3.33.0, October 2026) still ships without it. The good news is that BIRD itself supports BFD perfectly well, and Calico generates its BIRD config from templates. A small customization of those templates is enough to enable BFD on the sessions towards the route reflectors. It’s not an officially supported path, and we have to re-check it on every Calico upgrade, but the change is tiny and has been stable for us.

So we have BGP plus BFD with a well-understood CNI. Which is exactly what makes the alternatives a hard sell.

Why not Calico VPP

Calico VPP is the option I wanted to like the most. It is GA since Calico 3.27, it keeps the Calico API and policy model, and it promises a lot of performance. Every time I look at it, though, I end up with the same list.

The first item is BGP itself. The VPP dataplane doesn’t use BIRD, it runs its own agent based on GoBGP, and that agent has no BFD. Our BIRD template patch has nowhere to live. This one may change, though. GoBGP merged built-in BFD in May 2026 and shipped it in v4.6.0. But the VPP dataplane still pins GoBGP v3.35 from spring 2025, and v3 to v4 is a major version jump, so nothing has reached Calico VPP yet. Until it does, moving to VPP means giving up the one thing we modified Calico for in the first place.

Then there is the NIC. VPP takes over the physical interface, and the host is left with a virtual tap device behind it. All host traffic (SSH, kubelet, kube-vip) now passes through VPP before it reaches the kernel. That is an extra hop and a new place to break for everything running with hostNetwork. We rely on kube-vip in BGP mode to announce the API server VIP, and I could not find anything documented about how it behaves behind VPP. “Probably fine” is not a good answer for the thing kubectl depends on.

VPP also replaces kube-proxy with its own service implementation. Probably an upgrade in raw numbers, but the documented gaps are real (no session affinity, some HostEndpoint options ignored). More importantly, all our runbooks are built around iptables, nft, conntrack and tcpdump on the host. With VPP the dataplane is simply not visible there anymore, you debug with vppctl. Everyone on call has to learn a new toolbox. Add the node prerequisites on top (hugepages, a VPP driver matching each NIC type across VMs and bare metal), and the node image changes as well.

And there is release lag. The VPP dataplane release for Calico 3.32 came out in May 2026. Calico 3.33 is out, and there is no matching VPP release at the time of writing. Being pinned to N-1 on the CNI is not a blocker, but it is one more thing to track.

To be clear, the performance gain is real. In our own iperf3 tests between pods on VM nodes, plain Calico averaged 6.04 Gbits/sec and Calico VPP 8.48 Gbits/sec, roughly 40% more throughput on the same hardware. Keep in mind this was on VMs, where the hypervisor already spends CPU on networking before VPP gets the packet, so bare-metal numbers could look different. Still, 40% is not something I can wave away. That is why we keep coming back to it, and we will evaluate it again. For now, though, the balance is against it: it is not stable enough for us, it has no BFD, and the community around it is tiny compared to plain Calico or Cilium, which matters when you hit a bug at 3 a.m. None of this is impossible to solve. Together it turns “swap the dataplane” into a multi-quarter project, with BFD lost at the end of it.

Why not Cilium

Cilium comes up in every discussion about Kubernetes networking, and for good reason. I know where the community is: Cilium is far more popular these days, most new clusters I hear about start with it, and Calico is treated as the older choice. I don’t disagree with the popularity, but Calico still holds its ground for what we need. For a BGP-based, no-overlay setup like ours Cilium has the same problem as VPP, in a different shape.

Open-source Cilium does not support BFD. The documentation says so directly, and the tracking issue has been open since late 2022. The building blocks are there now: Cilium already runs on a GoBGP v4 release that includes the new built-in BFD, and in August 2026 a community pull request appeared that wires it into the BGP peer config. It is still a draft, and the maintainers asked for a formal design proposal before going further, so it is work in progress rather than something I can plan an upgrade around. Isovalent Enterprise has had BFD since their 1.16 release, which tells you where the feature sits commercially. When BFD ships in open-source Cilium we will evaluate it seriously, and I mean that: it would remove our main objection.

The second issue is how much the BGP side of Cilium has moved. The MetalLB-based BGP was deprecated in 1.14 and removed in 1.17. The first BGP Control Plane API (CiliumBGPPeeringPolicy) was deprecated in 1.18 and removed in 1.19. The v2 CRDs went from v2alpha1 to v2, and the alpha versions are being dropped as well. If we had migrated to Cilium two or three years ago, we would have rewritten our BGP configuration at least twice by now, with each rewrite landing in the middle of a cluster upgrade. The current API looks stable, but I would rather watch it stay stable for a while than find out the hard way.

The third one is troubleshooting. eBPF is the reason Cilium is fast, and I don’t dispute the numbers. But it also moves the dataplane out of the places where Linux tools look. With kube-proxy replacement there are no Service rules in iptables or nft at all, the docs even tell you to check that they are absent. Connection tracking lives in BPF maps, so conntrack -L is empty of pod flows. With BPF host routing packets skip netfilter and most of the host stack, and anything handled at the XDP layer never reaches tcpdump. What you get instead is cilium-dbg monitor, cilium-dbg bpf ct list and friends, pwru for the hard cases, and Hubble, which is basically the one tool that shows you what happened to a packet and why it was dropped. Hubble Relay and UI are off by default, so you also have to deploy and run them. These tools are good. They are also a completely different skill set, and every runbook we have for “pod cannot reach X” would need to be rewritten from scratch. That cost is the same one I listed for VPP, just with better documentation.

Plus, migrating CNI on a running cluster is painful regardless of the destination. We would need a very good reason, and “the same BGP without BFD” is the opposite of one.

Conclusion

For 2026 we stay on Calico with BIRD, BGP to the route reflectors, and BFD enabled through a small template change. The alternatives are not bad. Calico VPP is GA and Cilium’s BGP Control Plane is solid. They just don’t give us anything we need badly enough to lose BFD and rebuild half of our operations around it.

Things I’m watching for next year: native BFD in open-source Calico, the Cilium BFD pull request, and whether the VPP dataplane moves to GoBGP v4 and picks up BFD with it. If any of those land, this post gets a sequel, and possibly a migration plan.

~/blog
$