Context
My homelab runs on a single physical server: an old i7 with 32 GB of RAM.
On top of it, I built a Kubernetes cluster with microk8s. It scales easily, the community is huge, and there is a tool for everything: logs, security, backups. It can also run in hybrid mode, with my local node and some nodes in the cloud if I ever need them.
I didn’t start with a few containers and migrate later. I went straight to Kubernetes.
Today, that’s 28 services in 16 namespaces. On one node.
⚠️ One warning before going further ⚠️. This setup is a proof of concept, a lab I built to learn network management and threat detection on Kubernetes. Don’t run it as is in production! If you have public and private services, put them in separate clusters.
Ansible
VM creation is still manual. Everything after that is Ansible, so I can rebuild the cluster on any server with exactly the same configuration.
I wrote two playbooks:
- the first one provisions and secures one or more microk8s VMs
- the second one syncs the Kubernetes resources with the cluster, so my repo is the single source of truth for what is installed
That’s 384 tasks and 88 secrets, applied the same way every single time. No manual install step, so no manual install mistake.
It’s also what let me rebuild the cluster around a hundred times to test every access and security rule. The single node is where I am today, not the architecture. Going multi-node, changing servers, migrating, recovering from a disaster or failing over part of it to the cloud all start with the same command.
Two IPs, two Traefiks
Most of my services are private, a few are public. I didn’t split them by cluster, I split them by IP (for the moment).
MetalLB has two pools with one address each. .210 for public, .220 for private. Both with autoAssign: false.
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
name: public-pool
namespace: metallb-system
spec:
addresses:
- 192.168.1.210/32
autoAssign: false # nobody takes this IP without asking for it
---
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
name: private-pool
namespace: metallb-system
spec:
addresses:
- 192.168.1.220/32
autoAssign: false
Most people see autoAssign: false as a way to avoid exposing things by accident. For me it’s collision protection. With one address per pool, it’s first come, first served. I learned it the hard way: a service I spun up without thinking took the IP, Traefik didn’t have it anymore, and everything behind Traefik went down. Now a Service gets an address only if it asks for it.
The router only knows the public address. Port forwarding goes to .210, and nothing ever goes to .220.
Traefik has 6 entrypoints, including web-public (8080) / web-private (8081) pair. Each IP has its own Traefik Service, listening only on its own entrypoints.
In the end, I have 2/3 private routes for 1/3 public ones. Most of my cluster is never supposed to see the internet.
externalTrafficPolicy: Local
By default, kube-proxy SNATs the source IP when it forwards a packet. Traefik would then see a node IP instead of the Cloudflare edge IP. And CrowdSec, which bans by IP, would end up banning my own node.
externalTrafficPolicy: Local on the public Service keeps the original source IP. Then Traefik trusts the Cloudflare ranges in forwardedHeaders.trustedIPs, and reads the real visitor IP from X-Forwarded-For. Those ranges are refreshed at deploy time from cloudflare.com/ips-v4 and /ips-v6.
Here is the flow:
| From | To | Expected |
|---|---|---|
| Internet (through cloudflare proxy) | public (.210) | real visitor IP from X-Forwarded-For, passes |
| LAN | public (.210) | rejected by cloudflare-ips |
| Tailscale | private (.220) | passes, same as from the LAN |
| — | — | ban decision triggered on purpose, then verified |
cloudflare-ips rejecting a LAN request means the real source IP makes it all the way to Traefik.
The trade-off with Local: packets landing on a node without a Traefik pod are dropped. I have one node. For once, my hardware constraint makes my life easier.
Private on purpose, and poorer for it
Private routes have no certificate and no HTTPS redirect. Public side has let’s encrypt certificats.
For DNS, CoreDNS holds a wildcard record (* IN A) pointing to the private address. A new private service is one IngressRoute, and that’s it. CoreDNS shares .220 with the private Traefik thanks to allow-shared-ip, then traefik redirect to the correct service based on domain used (reverse proxying).
Then there is the Tailscale trade-off. I want the tailnet. But when I’m at home, I also want to reach my private services without launching anything. So the whitelist-local middleware is broad, it accepts the whole local range.
Let’s be honest, it’s not strong protection. An IoT device or a guest on the same Wi-Fi gets through. The real lock on private isn’t the middleware, it’s “can you route to that IP at all?”. From the internet, you can’t. The middleware is the second barrier, not the first.
NetworkPolicy helpers
130 helper calls, 10 helpers, 28 services.
Without them, it’s around 80 lines of identical YAML per service. Roughly 1,700 lines to edit in lockstep every time I change a rule.
The one I like most is block-local-access. Egress is open to 0.0.0.0/0, except the LAN range. So a public compromised pod can still reach the internet (because in anyway it needs to have it to run in normal situation), but it can’t scan my local network.
# what block-local-access renders in each namespace
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: block-local-access
spec:
podSelector: {}
policyTypes:
- Egress
egress:
- to:
- ipBlock:
cidr: 0.0.0.0/0
except:
- 192.168.1.0/24
Two bouncers
Some services don’t speak HTTP, and that breaks the setup.
It goes through an IngressRouteTCP with HostSNI(*) on a dedicated entrypoint with a specific port. No public-security chain on it: an ipAllowList or an HTTP bouncer has nothing to grip on a protocol without SNI. It isn’t in my DDNS scope either, cloudflare.records only covers HTTP services. Cloudflare doesn’t proxy arbitrary TCP anyway, that’s their paid Spectrum offer.
So the port is exposed directly. The only protection left is the firewall bouncer.
That’s why I run two bouncers, each with its own API key:
- L7: the Traefik plugin, on HTTP traffic
- L3/L4: a DaemonSet with
hostNetwork: true,NET_ADMIN/NET_RAWcapabilities, writing nftables rules withdeny_action: DROPon the input and forward hooks
Both only consume one decision type: ban.
CrowdSec
The Traefik plugin runs in crowdsecMode: stream and pulls decisions every 300 seconds. I raised httpTimeoutSeconds to 30: the LAPI pod is sometimes slow to answer.
In the public-security chain, cloudflare-ips runs before crowdsec-bouncer. So the L7 bouncer has less to process by only sees traffic that already passed the allowlist.
Logs
The only logs that carry a decision are the Traefik access logs. They’re enabled with all fields kept, because that’s what CrowdSec needs to decide who bans.
The rest is classic monitoring: ServiceMonitors, PrometheusRules for alerts, and Uptime Kuma monitors registered automatically through a CRD: when I add a service, its monitor comes with it.
The limits I’m keeping
I don’t call debts, because I chose them. CPU is overcommitted 5.6×, RAM 2.1×, on a single node. It’s a bet that peaks won’t happen at the same time. So far, it holds.
VM creation is still manual : I accept it because this single node is a starting point, the hundred rebuilds are the proof it can move when it has to.
Evolutions
This setup works. But I know it’s not the best answer to the security problem. I built it this way to learn Kubernetes networking, logging and security tooling, and for that, it did the job.
The next step is to drop the single cluster for two separate ones: a public one and a private one. Most of the complexity in this article comes from having both worlds on the same node:
- Two MetalLB pools, and the IP collision risk that comes with them
- Whitelist-local and cloudflare-ips living side by side in the same Traefik
- The 18 namespaces I maintain by hand in a traefik_backend_namespaces list
- A good part of the 130 NetworkPolicy calls
With two clusters, a compromised public service can’t even see the private one. The boundary isn’t a rule anymore, it’s the infrastructure. Just some firewal rules. That’s A LOT less to maintain, and a lot fewer ways to get a rule wrong.