
LoRaWAN over WireGuard looked like a weekend job
Somewhere out there, one of our Edgepilot LoRaWAN® gateways sits behind a mobile carrier’s CGNAT, quietly having its public IP and port rotated out from under it. Its job is to forward LoRa packets over UDP, using the Semtech GWMP protocol, to ChirpStack’s gateway bridge. That worked fine, until the carrier changed the gateway’s IP or port and ChirpStack quietly stopped listening to it.
Not crashed. Not erroring. Just gone silent, while the gateway kept sending packets into a void and calling it a job well done… “Classic UDP, am I right?” ba dum tss
The shape of the fix was obvious from the start: the gateway needed a stable IP and port. So we contacted the SIM carrier and asked about static IP addresses, or whatever solution they offered for keeping the gateway’s endpoint stable.
While we waited for an answer, I put in a stopgap. The gateway checks its own public IP and restarts its UDP forwarder when it notices a change. That papered over the IP part. It did nothing for port rotation (not reliably, anyway), which happens more often and just as silently.
Then the carrier came back with an insane quote and a document describing what was, essentially, just a VPN.
DING 💡
It clicked: a VPN, of course! Fine, we’ll run it ourselves, because apparently I love turning carrier invoices into infrastructure.
WireGuard, specifically, because its endpoint roaming and PersistentKeepalive mean a gateway can vanish and reappear on a totally different public IP and port while keeping the same VPN-internal address. The carrier can rotate whatever it likes; the tunnel just heals itself. We already had a separate OpenVPN tunnel for fleet management that I’d always wanted to replace with WireGuard anyway, so this was the perfect moment.
Here’s the part that should have been boring: pipe the gateway’s UDP through WireGuard, land it on a pod in our Kubernetes cluster, and forward it to ChirpStack’s gateway bridge Service. It’s just UDP forwarding. How hard can that be? subtle foreshadowing
Two ways to do this badly
Quick note: if your WireGuard server happens to live on the same host as your LNS or gateway bridge, none of what follows applies to you. Just skip ahead a few paragraphs to where it works. It's genuinely that easy when tunnel and destination share a network. Ours didn't, because ours was Kubernetes with Cilium, and that's where things got interesting.
First attempt: run a WireGuard sidecar inside the gateway bridge pod itself. It worked great, technically. In our setup, it also meant one WireGuard instance per LoRa region (EU868, US915, AU915, etc), each needing its own keys, its own peers, its own everything. Three VPN servers doing the job of one, purely because I’d wired the tunnel directly into the thing it was feeding. Some would call it “region-based sharding for scalability.” I called it quits after testing just one region.
Second attempt: a small UDP relay that would open a fresh local socket per gateway and forward manually. More flexible, in theory. In practice it meant tracking per-gateway session state by hand, and I had no good sense yet of how much of that Kubernetes would just hand me for free if I found the right lever. I shelved it before writing more than a sketch.
What actually worked was almost insultingly simple: one standalone WireGuard pod, with iptables DNAT rules forwarding the tunnel’s fixed internal address to the real ChirpStack Service ClusterIP. Gateway sends to 10.9.0.1:1700 (EU868 gateway bridge), iptables rewrites that to the ClusterIP of ChirpStack’s eu868-gateway-bridge Service. Done.
Except it didn’t work.
The pod could see the destination and still couldn’t reach it
Packets arrived at the WireGuard pod. iptables rewrote them correctly. And then they vanished, same as the gateway’s original CGNAT problem, just one hop later and inside our own cluster this time.
Nerd talk activated
We run Cilium with kube-proxy replacement, which handles Kubernetes Service translation in eBPF. By default, that happens at the socket layer: when an application calls connect() or sendmsg(), Cilium checks the destination against its Service table right there and picks a backend. Works beautifully for traffic an application inside the pod actually creates. Our packets weren’t created by the WireGuard pod, though. They were merely passing through it on their way from wg0 to eth0, the way a router forwards packets it didn’t originate. No application socket ever touched them, so there was nothing for that hook to rewrite.
Hubble showed our forwarded packets labeled world, as if they were leaving the cluster instead of heading to a Service two namespaces over. Cilium wasn’t broken. It was doing exactly what it was configured to do, for a use case that didn’t fit the usual model: a pod acting as a router instead of an application.
The fix is one Helm value: socketLB.hostNamespaceOnly: true. It bypasses that socket-level rewrite inside pod network namespaces and falls back to Cilium’s per-packet lookup at the veth instead, the same tc/eBPF layer that handles everything else. That’s where our forwarded traffic could finally be recognized and translated. One line, and packets that had been silently dying started arriving at ChirpStack within a second of the change.
A one-line fix that touches how Service traffic from every pod in the cluster gets load-balanced doesn’t get rolled out on faith, however convincing the diff looks. So I rolled it out as a CiliumNodeConfig override on exactly one worker node first and confirmed the gateway’s real LoRa traffic showed up in ChirpStack. Only then did it go cluster-wide with a plain helm upgrade. Even then, one node at a time: control-plane nodes first since nothing user-facing runs there, then each worker, checking pod health after every single one before moving to the next.
What it actually looks like
Once it worked, the shape of it turned out to be pretty small:
One pod as the hub. The gateway’s LoRa traffic and an admin’s SSH/web-UI access both arrive as ordinary WireGuard peers, and the pod’s iptables rules sort out where each one is actually headed.
The tunnel subnet is 10.9.0.0/16, deliberately far bigger than the fleet will ever need, so it’s not something to think about again for a long time.
The takeaway
Forwarding UDP through a tunnel is genuinely trivial, right up until the tunnel and the thing it feeds live in different networks with their own opinions about what counts as “local” traffic. The moment a packet stops being something your own pod sent and becomes something merely passing through it, half your networking stack’s usual assumptions quietly stop applying. Ours was hiding inside Cilium’s socket load balancer. Yours may be hiding somewhere else, but it’ll be there. At least now you know.