Skip to content
Aiman Ismail

Why new Kubernetes connections disappeared between two nodes

This is a note — quick thoughts, possibly AI-assisted. Not a fully fleshed article.

kubernetesciliumtailscalevxlannetworkingtroubleshootingnixos

A new connection from one Kubernetes Pod to another node kept timing out, even though older connections between the same nodes still worked. The bug turned out to be two networking tools accidentally using the same internal label for different purposes.

The short version:

  • Cilium wrapped the Pod's packet so it could travel to the other node.
  • Cilium attached a numeric label to that wrapped packet.
  • Tailscale mistook part of that label for one of its own labels.
  • Linux therefore sent the packet toward the normal LAN router instead of through Tailscale.
  • A narrow routing rule made Cilium's wrapped packets take the Tailscale path before the conflicting rule could see them.

That fixed the missing new connections, but a later fragment-drop alert exposed a second interaction between the same layers: Tailscale had learned Cilium's router addresses as possible peer endpoints. Its own UDP transport then entered Cilium's VXLAN overlay, which was itself carried over Tailscale. A second narrow rule prevented that recursive path without changing the already-correct MTUs.

A small networking glossary

  • Pod: a running Kubernetes workload. Each Pod gets its own cluster-internal IP address.
  • Node: the machine that runs Pods. This cluster had two nodes.
  • Cilium: the component that gives Pods network connectivity and applies network policy.
  • VXLAN: a way to carry a Pod packet inside another packet. Think of putting a letter inside a courier envelope. The inner letter is the Pod traffic; the outer envelope is addressed between nodes.
  • Tailscale: the encrypted private network connecting the two machines.
  • tailscale0: the Linux network interface used to send packets into Tailscale.
  • TCP: the protocol applications commonly use for reliable connections. A new TCP connection starts with a SYN packet.
  • UDP: a simpler transport used here for the outer VXLAN envelope. Cilium sends these envelopes to UDP port 8472.
  • DNS: the system that turns a service name into an IP address.
  • ClusterIP Service: a stable virtual Kubernetes IP that forwards connections to one or more Pods.
  • Route: Linux's answer to “which interface should carry this packet?”
  • Policy routing: extra routing rules that can choose a route using more than the destination address.
  • Packet mark: an internal number attached to a packet while it is inside the Linux kernel. It is metadata, not bytes sent to the remote application. Routing rules can inspect it.
  • MTU: the largest packet an interface or route can carry without further splitting.

Two more terms appear in networking tools:

  • The overlay is the virtual Pod network created by Cilium.
  • The underlay is the real path carrying that virtual network. Here, the underlay was Tailscale.

How a cross-node Pod packet should travel

Suppose Pod A runs on node 1 and Pod B runs on node 2:

A packet travelling from Pod A to Pod B on another node An animated packet moves from Pod A through Cilium, Linux routing, the encrypted Tailscale tunnel, and Cilium on the second node before reaching Pod B. Pod A creates a packet for Pod B Cilium on node 1 wraps it in a VXLAN/UDP envelope Linux routing chooses the outgoing interface Tailscale encrypted tunnel carries it securely to node 2 Cilium on node 2 opens the VXLAN envelope Pod B receives the original packet packet repeats
The application sees one Pod-to-Pod connection. Cilium and Tailscale handle the extra transport layers underneath it.

Cilium uses UDP destination port 8472 for the outer VXLAN envelope. The application does not know about VXLAN or Tailscale; it only sees the original inner packet.

What “wrapping” adds to the packet

Each outer layer supplies information needed by the layer below it. The animation highlights the layers from the original application data outward:

Layers added around application data for cross-node delivery Nested boxes show application data inside TCP, the inner Pod IP packet, inner Ethernet, VXLAN, UDP, and the outer node IP packet. The layers highlight from inside to outside. Outer node IP which node? UDP port 8472 VXLAN this carries a virtual network Inner Ethernet Inner Pod IP TCP application data sender: add layers from inside → outside · receiver: remove them outside → inside
The original data remains at the center. Cross-node delivery adds addressing and transport information around it.

What actually happened

The journey broke at Linux routing, after Cilium had created the outer envelope:

The wrong route before the fix and the correct route after it A VXLAN packet reaches Linux routing. Before the fix, an orange packet follows the physical LAN route and is lost. After the fix, a green packet follows Tailscale table 52 to tailscale0 and reaches node 2. Cilium creates VXLAN packet packet mark: 0x1d080400 Linux policy routing reads the packet mark AFTER FIX · CORRECT Tailscale table 52 tailscale0 → node 2 BEFORE FIX · WRONG Normal routing table physical LAN gateway → lost corrected packet misrouted packet
The packet contents did not change. The fix changed which routing rule handled the VXLAN envelope first.

This explains the confusing evidence: Cilium said it had sent the packet “to overlay,” but no matching packet appeared on tailscale0. Cilium had completed its part; Linux chose the wrong exit afterward.

Why Linux chose the wrong route

Cilium and Tailscale both use packet marks. This is normally fine, but packet marks are one shared 32-bit number. The two tools must avoid assigning conflicting meaning to the same bits.

Cilium put two things in the mark on this VXLAN packet:

  • the Pod's Cilium security identity;
  • a value meaning “this is overlay traffic.”

The failing Cilium identity was decimal 7432, or hexadecimal 0x1d08. The final mark was:

Cilium identity:       0x1d08
Outer VXLAN mark:      0x1d080400

Tailscale had a policy rule that looked only at part of the mark:

fwmark 0x80000/0xff0000 lookup main

Read this as: “If these selected bits equal 0x80000, use the normal Linux routing table.” The mask after / selects which bits matter. Apply that mask to Cilium's mark:

0x1d080400 & 0x00ff0000 = 0x00080000

It accidentally matched. Tailscale saw the 08 inside Cilium's identity and treated the packet as traffic that should bypass Tailscale's routing table.

The important point is not the hexadecimal arithmetic. It is this:

Cilium attached a label for one reason. Tailscale read part of the same label and gave it a different meaning.

Cilium assigns identities dynamically. This was not permanently tied to one application: any endpoint identity whose low byte was 0x08 could hit the same failure.

Packet size constraints

The three MTU values were intentionally different:

remote Pod route  1180
Cilium links      1230
tailscale0        1280

Why leave space? Wrapping the inner Pod packet adds an outer IP header, UDP header, VXLAN header, and Ethernet header. The smaller inner limit leaves room for that envelope. Tailscale's 1280-byte interface limit then accounts for its own encrypted transport.

The routing fix preserved all three values. This mattered because reducing MTU had solved an earlier packet-fragmentation problem, but it could not solve a packet being sent through the wrong interface.

Symptoms

  • The CloudNativePG database controller could not establish a new connection to a PostgreSQL instance on the other node.
  • Moving the controller out of the Pod network and onto the node's own network (hostNetwork) restored service, but only as a temporary recovery.
  • Traffic sent by the node itself, rather than by a Pod, worked across nodes.
  • Some already-established Pod connections worked.
  • Cilium reported traffic at to-overlay.
  • No corresponding UDP/8472 packet appeared on tailscale0.
  • Lowering MTU had previously fixed fragment drops, but did not fix this new-flow failure.

In plain language, that mix of results said:

  • the other machine and the Tailscale tunnel were reachable;
  • Cilium decided to wrap the packet for cross-node delivery;
  • something went wrong between wrapping it and sending it into Tailscale;
  • Linux's connection tracking could preserve an older working path, so testing only an existing connection was misleading.

Reproduce without the affected application

The first useful step was to remove CloudNativePG from the experiment.

Create ordinary Pods pinned to opposite nodes, run a simple TCP listener on one, and start a fresh connection from the other. Delete and recreate the source Pod between attempts so the test does not reuse Cilium endpoint state or Linux's memory of an older connection.

Test all four combinations:

  1. core Pod → worker Pod IP;
  2. worker Pod → core Pod IP;
  3. core Pod → worker-backed ClusterIP Service;
  4. worker Pod → core-backed ClusterIP Service.

Also run a control Pod with a different Cilium security identity. In this incident:

  • the identity matching the failed application reproduced the failure;
  • a fresh control identity succeeded;
  • therefore the application and general cross-node MTU were not the discriminating variables.

Trace from endpoint to wire

The goal was to follow one new TCP connection attempt through each layer. A TCP connection begins with a packet carrying the SYN flag, so that first SYN is a useful marker.

Use Cilium's monitor to confirm that the Pod packet reached the overlay path. At the same time, use tcpdump to watch for the outer VXLAN envelope:

# Cilium's logical datapath decision
kubectl -n kube-system exec ds/cilium -c cilium-agent -- \
  cilium-dbg monitor --type trace --type drop

# Expected encrypted underlay emission
sudo tcpdump -ni tailscale0 'udp port 8472'

# Check whether it escaped on the physical interface instead
sudo tcpdump -ni <physical-interface> 'udp port 8472'

Seeing to-overlay means Cilium decided to wrap the packet. It does not mean Linux sent the new outer packet through tailscale0.

The final command asks Linux: “Where would you send a UDP/8472 packet to this Tailscale peer if it had the exact mark Cilium produced?”

ip -4 rule show
ip route show table 52
ip route get <peer-tailscale-ip> \
  ipproto udp dport 8472 mark 0x1d080400

Before the fix, the exact marked lookup selected each node's physical LAN default route rather than Tailscale table 52.

The durable fix

Add a higher-priority rule that is specific to Cilium VXLAN and the exact peer:

ip -4 rule add priority 5205 \
  to <peer-tailscale-ip>/32 \
  ipproto udp dport 8472 \
  lookup 52

Read the rule one line at a time:

  • priority 5205: check this rule before Tailscale's conflicting rule at 5210. Linux checks lower numbers first.
  • to the peer /32: match only the other node's exact Tailscale IP address. /32 means one IPv4 address.
  • UDP destination port 8472: match only Cilium's VXLAN envelopes.
  • lookup 52: use the routing table Tailscale maintains, which sends the packet through tailscale0.

The result is: “Before considering the more general Tailscale mark rule, send Cilium's VXLAN packets for this node through Tailscale.”

The rule is deliberately narrow:

  • one peer /32, not the whole tailnet;
  • UDP only;
  • destination port 8472 only;
  • no mutation of Cilium or Tailscale marks;
  • table 52 remains owned and populated by Tailscale.

Running the ip rule add command manually would fix the current machine, but the rule would disappear after rebuilding or replacing the host. The lasting implementation therefore put the same rule in the NixOS host configuration:

{ peerAddress }:
{ pkgs, ... }:
{
  systemd.services.cilium-tailscale-routing = {
    description = "Route Cilium VXLAN packets through the Tailscale underlay";
    wantedBy = [
      "multi-user.target"
      "tailscaled.service"
    ];
    partOf = [ "tailscaled.service" ];
    after = [ "tailscaled.service" ];
    before = [ "k3s.service" ];
    path = [ pkgs.iproute2 ];
    serviceConfig = {
      Type = "oneshot";
      RemainAfterExit = true;
    };
    script = ''
      while ip -4 rule del priority 5205 \
        to ${peerAddress}/32 ipproto udp dport 8472 lookup 52 \
        2>/dev/null; do :; done
      ip -4 rule add priority 5205 \
        to ${peerAddress}/32 ipproto udp dport 8472 lookup 52
    '';
    preStop = ''
      while ip -4 rule del priority 5205 \
        to ${peerAddress}/32 ipproto udp dport 8472 lookup 52 \
        2>/dev/null; do :; done
    '';
  };
}

You do not need to understand every Nix expression to understand the behavior:

  • remove any old copy of this exact rule;
  • add one fresh copy;
  • start it during normal system activation;
  • remove and reinstall it around a Tailscale restart;
  • install it before Kubernetes starts.

The delete loop makes repeated activation safe and removes duplicates left by an interrupted attempt. PartOf=tailscaled.service plus the wantedBy relationship reinstalls the rule when Tailscale restarts.

A lifecycle bug caught during rollout

The first version said only “start this when Tailscale starts.” During nixos-rebuild test, Tailscale was already running, so systemd did not start it again. The new routing service remained inactive and no rule went live.

Adding multi-user.target also said “start this as part of the normal running system.” That made configuration activation start the rule immediately, while retaining the Tailscale restart coupling. The automated check now verifies both the rule text and when the service starts.

This is why an automated configuration test should check when a service starts, not just the command it will eventually run.

Verification

Do not qualify this fix with one curl. Check new state, established state, encapsulation, and the applications that exposed the problem.

Routing and MTU

On both nodes:

ip -o link show tailscale0      # mtu 1280
ip -o link show cilium_host     # mtu 1230
ip -o link show cilium_vxlan    # mtu 1230
ip route show                   # remote Pod CIDR: mtu 1180
ip rule show                    # priority 5205 present

ip route get <peer-tailscale-ip> \
  ipproto udp dport 8472 mark 0x1d080400
# must resolve through tailscale0 and table 52

New Pod eth0 interfaces also reported MTU 1230.

Dataplane

  • Newly recreated ordinary Pods established direct TCP both ways.
  • ClusterIP Services worked both ways in repeated attempts.
  • Three-second TCP transfers remained active and transferred about 30–35 MB each way.
  • TLS 1.3 handshakes and multiple asymmetric application records passed both ways.
  • DNS resolved both test Services from both nodes.
  • Large ping packets passed 5/5 both ways. ping -s 1152 creates an 1180-byte packet after adding the 28 bytes of IPv4 and ICMP headers, so this tested the route's size limit rather than only tiny packets.
  • tcpdump on both tailscale0 interfaces showed VXLAN packets moving in both directions, including large inner TCP segments, with zero kernel capture drops.
  • Bounded Cilium drop monitoring showed no logical-fragment drops.
  • Both Cilium agents reported cluster health 2/2, including host and endpoint ICMP/HTTP.

Application recovery

Only after the ordinary-Pod tests passed:

  • Flux removed CloudNativePG's temporary settings that bypassed the normal Pod network.
  • The controller returned to an ordinary Pod IP on the core node, Ready with zero restarts.
  • The PostgreSQL cluster reported healthy and continuous archiving remained healthy.
  • pg_isready, pg_is_in_recovery() = false, and a simple SQL expression passed.

Temporary listeners, capture Pods, identity probes, and the test namespace were removed after qualification.

A second failure: Tailscale tried to travel through itself

Several hours after the fwmark fix, CiliumOverlayFragmentDrops fired again. It was tempting to conclude that the earlier MTU correction had regressed. The live state contradicted that explanation:

tailscale0         1280
Cilium links       1230
Pod endpoint links 1230
remote Pod routes  1180
configuration drift   0

The priority-5205 VXLAN rule was active on both nodes, marked UDP/8472 lookups still selected Tailscale table 52, and packet captures showed outer VXLAN packets no larger than 1230 bytes with 1180-byte inner packets. There was no outer IPv4 fragmentation.

The alert itself was also narrower than “the network MTU is wrong.” It meant that Cilium received a non-first IPv4 fragment but could not find the first fragment's cached transport ports. Only core-01 ingress was affected. The five-minute rate was bursty and peaked at 10.43 drops per second; sandbox-01 had no matching series.

The fragment map identified the traffic class

Cilium keeps a fragment map so later fragments can reuse information learned from the first fragment. On core-01, 7,935 of 8,192 entries were occupied. Of those, 4,448 had exactly this tuple:

10.42.1.36:41641 -> 10.42.0.91:41641 UDP

Those IPs were not application Pods. They were the Cilium router addresses on the two nodes. UDP port 41641 is Tailscale's transport port.

The same tuple appeared inside VXLAN captures. Tailscale endpoint discovery had learned the Cilium router addresses and was attempting to send its own encrypted transport between them. Linux routed those addresses through Cilium, producing this loop:

Tailscale UDP/41641
  -> Cilium Pod-CIDR route
  -> VXLAN UDP/8472
  -> tailscale0
  -> Tailscale carrying its own transport

The intended stack was “Pod traffic inside Cilium VXLAN inside Tailscale.” The accidental stack was “Tailscale transport inside Cilium VXLAN inside Tailscale.” This recursive traffic heavily occupied the fragment map and coincided with the fragment-cache misses on core-01.

The evidence did not distinguish every individual miss caused by fragment reordering or loss from one caused by LRU eviction in the nearly full map. It did establish the offending traffic class and the recursive route. Flushing the map would only erase evidence and provide temporary relief, so it was left to age naturally.

Block only the recursive local transport

The durable fix added a rule immediately before the existing VXLAN rule:

ip -4 rule add priority 5204 \
  iif lo \
  to 10.42.0.0/16 \
  ipproto udp sport 41641 dport 41641 \
  prohibit

Each selector matters:

  • priority 5204: evaluate it before the priority-5205 VXLAN exception;
  • iif lo: match traffic originating from the host, not packets forwarded from Pods;
  • Pod CIDR destination: reject only attempts to use a Cilium address as the Tailscale underlay endpoint;
  • UDP source and destination port 41641: match Tailscale transport, not ordinary application UDP;
  • prohibit: make this invalid endpoint candidate fail instead of entering Cilium recursively.

The existing priority-5205 rule remained responsible for legitimate Cilium UDP/8472 envelopes to the peer's Tailscale /32. Forwarded Pod traffic also remained routable because it does not match iif lo.

The NixOS service now owns both rules and reinstalls them together around Tailscale restarts:

ip -4 rule add priority 5204 iif lo \
  to 10.42.0.0/16 \
  ipproto udp sport 41641 dport 41641 prohibit

ip -4 rule add priority 5205 \
  to ${peerAddress}/32 \
  ipproto udp dport 8472 lookup 52

This change was merged in PR #451.

Qualifying the recursion fix

The rule was activated one node at a time. Qualification checked both what should fail and what must keep working:

  • a matching locally originated UDP/41641 route lookup returned prohibit;
  • the marked UDP/8472 lookup still selected tailscale0 and table 52;
  • forwarded Pod traffic still resolved through the Cilium route;
  • both nodes and both Cilium health endpoints remained reachable;
  • the 1280/1230/1180 MTU chain and configuration-drift value remained unchanged;
  • a bounded Cilium drop monitor was clean;
  • 15 consecutive metric samples over 7½ minutes were zero, the five-minute drop rate reached zero, and the alert became inactive.

No fragment map was flushed. The dominant recursive entries on core-01 aged down naturally while producing no new drops. A later independent check again found both nodes Ready, Cilium health 2/2, direct node-to-node Tailscale transport over native IPv6, and no renewed fragment-drop rate.

Takeaways

  • Internal packet labels are shared. Two independent networking systems can accidentally assign different meanings to the same bits.
  • An underlay must not select an overlay address as its own transport path. Otherwise the tunnel can recursively carry itself.
  • An MTU alert is a starting point, not a diagnosis. Confirm live MTUs, configuration drift, direction, rate, routes, and the actual traffic tuple before changing packet sizes.
  • Inspect state tables, not only packet captures. The fragment map revealed the dominant UDP/41641 router-to-router tuple that short monitor windows missed.
  • Follow the whole packet journey. to-overlay means Cilium decided to wrap a packet; it does not prove Linux sent the wrapper through the intended interface.
  • Ask Linux about the real packet. A basic route lookup can look correct while the packet's mark triggers another policy rule. Include its destination, protocol, port, and mark in the test lookup.
  • Reproduce with an ordinary Pod. This separated the infrastructure defect from CloudNativePG immediately.
  • Test fresh and established connections. Linux's memory of existing connections can make a broken new-connection path look healthy.
  • Keep the exception narrow. Match the exact peers, protocol, and port instead of overriding Tailscale policy broadly.
  • Preserve declarative ownership. NixOS owns the host rule; Flux owns removal of the Kubernetes workaround; neither depends on a live-only patch.
  • Test when a systemd service starts. Correct command text is not enough if the service never runs during configuration activation.