Users expect every service to be secure by default. You never know who might be listening, even on your own network. That’s why HTTPS has become standard in any modern setup. This article covers how I set it up on a Kubernetes cluster with Cloudflare, Traefik and Let’s Encrypt’s DNS-01 challenge.
One of my constraints was to be able to reinstall the whole server very quickly. And that’s exactly what got me a Let’s Encrypt 429, rate-limited. So if you don’t want to wait 24 hours for nothing like I did, read this carefully.
The setup
Everything runs on a single-node Kubernetes cluster with very limited RAM, so any memory saving was welcome.
Cloudflare sits in front of it. It hides part of my server’s network information, keeps all my DNS configuration in one place, gives me access to a CDN and blocks some malicious traffic.
Behind it, Traefik acts as the reverse proxy. It routes requests to the right service based on the domain name, and only accepts connections coming from Cloudflare’s public IPs.
Cloudflare’s proxy already handles HTTPS on the public side. But I also wanted the link between Cloudflare and my cluster to be encrypted, so the whole path is a two-hop HTTPS chain. The certificate used on that second hop is issued by Let’s Encrypt and stored in the cluster by cert-manager.
Finally, the whole system is described in Ansible playbooks, so it can be rebuilt anywhere in a few minutes. That keeps me agile. I can migrate from one provider to another quickly, and spin up test environments on demand for functional or security testing. It’s also the reason I hit involuntary the Let’s Encrypt daily rate limit in the first place.
Challenges: HTTP-01 vs DNS-01
Let’s Encrypt uses a challenge system to make sure the person asking for a certificate really owns the domain. Two kinds of challenges are available, and in both cases the renewal is handled automatically.
HTTP-01 — the requester must host a proof file that Let’s Encrypt fetches over port 80, at /.well-known/acme-challenge/. I didn’t go with it, for three reasons tied to my environment.
First, each domain needs its own proof. That would have meant running a small container per domain (a sidecar), which is excessive given the memory I have available. The alternative — writing custom Traefik routing rules to expose the challenge path for every domain — was doable, but would have taken too long, and the other method suited me better.
Second, Cloudflare sits in front of my infrastructure in strict mode: traffic to the origin is forced over HTTPS, which prevents Let’s Encrypt from retrieving the challenge file served over plain HTTP on port 80.
Finally, Let’s Encrypt rate limits apply per IP address. With several domains validated from a single public IP, you hit them quickly, especially with renewals or repeated failures.
DNS-01 — this challenge proves ownership with a TXT record in the DNS zone. That’s the one I chose, because it has no hosting requirement at all: nothing to expose, nothing to route. All I needed was a Cloudflare API token allowed to edit my DNS zones.
The flow is simple. Each domain asks Let’s Encrypt for a certificate. cert-manager writes the proof in a TXT record through the Cloudflare API. Let’s Encrypt reads it and signs the certificate. cert-manager then stores it in Kubernetes as a TLS secret. It is also egress-only, so a certificate can exist before the service is even exposed.
How it is wired in the cluster
The Cloudflare token comes from a vaulted variable, and is pushed as a secret in the cert-manager namespace. Two ClusterIssuers use it, staging and production. They are identical apart from the ACME directory URL. Each one has a single DNS-01 solver and no selector, so they cover every domain.
Every app that needs HTTPS then declares its own Certificate, with issuerRef: letsencrypt-prod and an explicit secretName. The Traefik IngressRoute references that secret by hand. There is no cluster-issuer annotation anywhere, which is why I end up with one Certificate per namespace.
Two details cost me some time. First, before requesting a certificate, I delete the leftover _acme-challenge TXT records when the TLS secret is missing locally. An interrupted run leaves orphan records, and those confuse the next validation.
Second, cert-manager polls the TXT record itself on public resolvers before telling Let’s Encrypt to verify. So it needs a dedicated network policy allowing UDP/TCP 53 to the internet. My default DNS policy only allows kube-dns.
A certificate that no one sees?
At the end of the day, you could tell me that this article is about a certificate nobody ever sees. Well, it’s true. But the backbone of this article is the end-to-end encrypted tunnel. Not just a certificate that signals some security concern. We want the full communication encrypted.
The visitor uses Cloudflare’s proxy certificate for the first part of the path. Then the proxy forwards the request to our reverse proxy, over the connection secured by the Let’s Encrypt certificate. From there, the request is routed to our services.
In the Cloudflare configuration, I set the encryption mode to “Full (strict)”. This forces the proxy to use a valid HTTPS connection, and it returns an error otherwise.
Rebuilding cost
As I said earlier, I built my system so I could reinstall it as many times as I want, for testing and cloud migration reasons.
The problem is that Let’s Encrypt limits the number of certificates it signs per day on its production environment.
So don’t get caught like I did, twice. Point the issuer at the staging environment when you’re testing non-production stuff. Otherwise their rate limiter blocks you for a while.
Bonus: get the real user’s IP through Cloudflare
I use CrowdSec to detect and block malicious activity. It reads Traefik’s logs and extracts client actions and IPs. For that to work, the IP in the logs has to be the visitor’s real one.
By default, Traefik logs the peer IP of the TCP connection it accepted (ClientHost). In our architecture, that is the IP of the Cloudflare proxy which routed the request. So every visitor looks like a Cloudflare edge IP. CrowdSec can neither group actions per client, nor ban one without banning everyone.
To fix that issue, you just have to:
- Fetch Cloudflare’s IP ranges at deploy time into a
cloudflare_ip_rangesvar. - Trust them on your public entrypoints, so Traefik reads the real visitor IP from
X-Forwarded-Forinstead of the peer address:
websecure-public:
port: 8443
exposedPort: 443
protocol: TCP
# Ensure real client IP is extracted behind Cloudflare
forwardedHeaders:
insecure: false
trustedIPs: "{{ cloudflare_ip_ranges }}"
expose:
public: true # expose HTTPS on public service
- Also apply
externalTrafficPolicy: Localon Traefik’s public Service, to disable the default multi-node routing. It’s a simple shortcut. It makes sure the client IP isn’t replaced by an internal Kubernetes address when the request is forwarded to another node than the one that received it. We have a basic cluster here, so we don’t need complex routing or scaling capabilities.
service:
enabled: false # Disable default service since we use additionalServices
additionalServices:
public: # Public service for external access
enabled: true
single: true
spec:
type: LoadBalancer
externalTrafficPolicy: Local
The same ranges are passed to the CrowdSec bouncer plugin (forwardedHeadersTrustedIPs), so its own decisions apply to the real client IP too.