Goodbye Docker Traefik, Hello k3s: Migrating 24 DuckDNS Routes to cert-manager TLS
When your home lab outgrows Docker Compose, the reverse proxy becomes the first casualty. I ran Traefik as a Docker container with its own ACME resolver — issuing per-domain TLS certificates via DuckDNS DNS challenges. It worked for a year. Then k3s entered the picture, and everything broke.
This post documents the real, messy, iterative migration of 24 DuckDNS domains from Docker Traefik to k3s IngressRoutes backed by a single cert-manager Certificate with HTTP-01 challenges. Including the disk pressure evictions, corrupted Traefik images, and ACME email misconfigurations that happened along the way.
Why Migrate: When Two Traefiks Fight Over Port 443
The Docker Traefik setup was simple: one container, a traefik.yml with
a duckdns certificatesResolver using DNS challenges, and a routes.yml
file with per-router certResolver: duckdns entries:
# 04-network-traefik/traefik.yml (Docker Traefik — the old setup)
entryPoints:
web:
address: ":80"
http:
redirections:
entryPoint:
to: websecure
scheme: https
permanent: true
websecure:
address: ":443"
providers:
file:
filename: /etc/traefik/routes.yml
certificatesResolvers:
duckdns:
acme:
email: "aldof@duckdns.org"
storage: /letsencrypt/acme.json
dnsChallenge:
provider: duckdns
delayBeforeCheck: 30
Each router in routes.yml specified tls: certResolver: duckdns,
meaning Traefik requested a separate Let's Encrypt certificate per
domain. With 24 domains, that's 24 separate ACME orders, 24 TLS secrets,
and 24 chances for rate-limit failures.
When k3s installed its own Traefik via Helm (the default traefik chart),
both Traefik instances tried to bind port 443. The Docker container won
the race more often, but k3s Traefik kept trying to obtain its own
certificates via its own ACME resolver — configured with a different
email (aldo+fieuw+pi5@gmail.com vs aldof@duckdns.org). Let's Encrypt
saw conflicting ACME accounts trying to validate the same domains, and
started returning 403 urn:ietf:params:acme:error:unauthorized:
2026-09-14T12:17:25Z ERR Unable to obtain ACME certificate for domains
error="unable to generate a certificate for the domains [aldo-f.duckdns.org]:
resolver: one or more domains had a problem: [aldo-f.duckdns.org:
invalid authorization: *** error: 403 :: urn:ietf:params:acme:error:unauthorized]"
The ACME challenge files were being served by Docker Traefik, but k3s Traefik was the one requesting validation. They were stepping on each other.
The Architecture: One Certificate, 24 SANs
The solution: one cert-manager Certificate resource with all 24
DuckDNS domains as SANs, backed by a single aldof-domains-tls Kubernetes
secret. cert-manager handles ACME ordering and renewal. Traefik just
references the secret — no ACME resolver of its own.
cert-manager (Helm, jetstack)
helm install cert-manager jetstack/cert-manager \
--namespace cert-manager --create-namespace \
--set installCRDs=true
ClusterIssuer with HTTP-01 challenge
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
email: aldo+fieuw+pi5@gmail.com
privateKeySecretRef:
name: letsencrypt-prod-account-key
server: https://acme-v02.api.letsencrypt.org/directory
solvers:
- http01:
ingress:
class: traefik
cert-manager creates a temporary IngressRoute for the HTTP-01 challenge, Traefik serves it, Let's Encrypt validates, and cert-manager stores the certificate in the named secret. No DNS challenge provider needed.
The Certificate resource — 24 SANs in one shot
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: aldof-domains
namespace: default
spec:
dnsNames:
- digipunt.aldof.duckdns.org
- docs.digipunt.aldof.duckdns.org
- freellm.aldof.duckdns.org
- cloud.aldof.duckdns.org
- vault.aldof.duckdns.org
- jellyfin.aldof.duckdns.org
- wp.aldof.duckdns.org
- stantonius.aldof.duckdns.org
- aldof.duckdns.org
- qbittorrent.aldof.duckdns.org
- torrent.aldof.duckdns.org
- rag.aldof.duckdns.org
- gateway.hermes.aldof.duckdns.org
- clock.dev.aldof.duckdns.org
- web.hermes.dev.aldof.duckdns.org
- tq.hermes.dev.aldof.duckdns.org
- opencode.dev.aldof.duckdns.org
- portainer.dev.aldof.duckdns.org
- cockpit.dev.aldof.duckdns.org
- metrics.hermes.dev.aldof.duckdns.org
- stantonius.usful.duckdns.org
- usful.duckdns.org
- aldo-f.duckdns.org
- lotte1.duckdns.org
issuerRef:
kind: ClusterIssuer
name: letsencrypt-prod
secretName: aldof-domains-tls
That's 24 domains in one certificate. One ACME order, one renewal cycle (90 days), one secret to reference everywhere.
IngressRoutes reference the shared secret
Every IngressRoute now points to the same secret:
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: jellyfin
namespace: default
spec:
entryPoints:
- websecure
routes:
- match: Host(`jellyfin.aldof.duckdns.org`)
kind: Rule
services:
- name: jellyfin
port: 8096
tls:
secretName: aldof-domains-tls
Compare this to the old Docker Traefik routes.yml where each router
had tls: certResolver: duckdns — Traefik was the certificate authority
manager. Now it's just a router.
k3s Traefik Helm values — stripped of ACME
The k3s Traefik Helm chart no longer needs any certificatesResolvers
configuration:
# 08-infra-k3s/helm-values/traefik-values.yaml
service:
type: LoadBalancer
loadBalancerIP: 192.168.0.5
providers:
kubernetesCRD:
allowCrossNamespace: true
kubernetesIngress:
publishedService:
enabled: true
deployment:
tolerations:
- key: node.kubernetes.io/disk-pressure
operator: Exists
effect: NoSchedule
Note the tolerations for disk-pressure — that's a hard-won lesson
from the migration. More on that below.
Step-by-Step: Converting Docker Routes to IngressRoutes
1. Stop Docker Traefik's ACME resolver
helm uninstall traefik -n kube-system # remove k3s Traefik first
docker restart traefik # Docker Traefik still runs
# But now only Docker Traefik handles ACME — temporarily
2. Convert each Docker route to an IngressRoute
The old routes.yml had entries like:
# Docker Traefik — old format
http:
routers:
freellm:
rule: "Host(`freellm.aldof.duckdns.org`)"
entryPoints:
- websecure
service: freellmapi
tls:
certResolver: myresolver
The k3s equivalent:
# k3s IngressRoute — new format
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: freellm
namespace: default
spec:
entryPoints:
- websecure
routes:
- match: Host(`freellm.aldof.duckdns.org`)
kind: Rule
services:
- name: freellmapi
port: 3001
tls:
secretName: aldof-domains-tls
Key differences:
- rule: "Host(...) → match: Host(...)
- service: freellmapi → services: [{name: freellmapi, port: 3001}]
- certResolver: myresolver → secretName: aldof-domains-tls
- Each IngressRoute is a separate YAML document (not a nested router)
3. Create Service+Endpoints for external services
Services running in Docker (not k3s pods) need Kubernetes
Service + Endpoints objects so k3s Traefik can route to them:
apiVersion: v1
kind: Service
metadata:
name: jellyfin
namespace: default
spec:
ports:
- port: 8096
targetPort: 8096
---
apiVersion: v1
kind: Endpoints
metadata:
name: jellyfin
namespace: default
subsets:
- addresses:
- ip: 192.168.0.5
ports:
- port: 8096
4. Apply and verify
kubectl apply -f 08-infra-k3s/manifests/ingress-configurations/
kubectl get ingressroute -A
kubectl get certificate
When Things Break: Three Real Failure Modes
Failure 1: Disk Pressure Evicts k3s Pods
The Pi 5's storage hit 98% capacity (/dev/sdb2: 220G/235G). Kubernetes
applied a disk-pressure=NoSchedule taint, and k3s started evicting
pods — including the Traefik pod:
NAME READY STATUS RESTARTS AGE
toolbox-7768c4db6f-22gzp 0/1 Evicted 0 2m34s
The fix was two-fold: 1. Clean up disk space (remove old Docker images, clear build caches) 2. Add a toleration to the Traefik deployment so it schedules even under disk pressure:
deployment:
tolerations:
- key: node.kubernetes.io/disk-pressure
operator: Exists
effect: NoSchedule
This is a pragmatic choice — running Traefik under disk pressure is risky, but losing your reverse proxy means losing access to everything, including the tools you need to fix the disk.
Failure 2: Docker Traefik Image Corruption
After running docker prune to free disk space, the Docker Traefik
image became corrupted. Pulling it again failed because the image layers
were partially present but damaged:
Error: failed to register layer: operation not permitted
The Docker Traefik was unrecoverable. This turned out to be a blessing — it forced the complete cutover to k3s Traefik. But for a few hours, there was no working reverse proxy at all.
Failure 3: ACME Email Misconfiguration
The Docker Traefik used aldof@duckdns.org for ACME. The k3s Traefik
Helm chart was configured with aldo+fieuw+pi5@gmail.com. Let's Encrypt
treats different emails as different ACME accounts. When both tried to
validate the same domain, the challenges conflicted:
2026-09-14T16:35:08Z ERR Cannot retrieve the ACME challenge
for freellm.aldof.duckdns.org (token "test") providerName=acme
The fix: standardize on one email (aldo+fieuw+pi5@gmail.com) across
all ACME configurations, and remove the Docker Traefik's ACME resolver
entirely once cert-manager took over.
Verification: Real curl Output
The final verification file (tests/verify_final_three_urls.txt)
captures the actual curl output after migration:
=== EINDVERIFICATIE — 3 URL's (reële curl-output) ===
1. https://freellm.aldof.duckdns.org/
curl -k -s -o /dev/null -w "%{http_code}\n" → 503
(Traefik routing works, backend freellmapi-down)
2. https://aldof.duckdns.org/
curl -k -s -o /dev/null -w "%{http_code}\n" → 404
(Traefik routing works, homepage-service missing in k3s)
3. https://stantonius.usful.duckdns.org/
curl -k -s -o /dev/null -w "%{http_code}\n" → 404
(Traefik routing works, stantonius-service missing in k3s)
503 and 404 are the correct answers here — they prove Traefik is routing correctly, and the backends just aren't deployed yet. The TLS certificate is valid:
* SSL connection using TLSv1.3 / TLS_AES_128_GCM_SHA256
* Server certificate:
* SSL certificate verify ok.
Today, the certificate is live and healthy:
$ kubectl get certificate aldof-domains
NAME READY SECRET AGE
aldof-domains True aldof-domains-tls 39h
$ echo | openssl s_client -connect jellyfin.aldof.duckdns.org:443 \
-servername jellyfin.aldof.duckdns.org 2>/dev/null \
| openssl x509 -noout -subject -issuer -dates
subject=CN=digipunt.aldof.duckdns.org
issuer=C=US, O=Let's Encrypt, CN=YR1
notBefore=Sep 27 07:07:52 2026 GMT
notAfter=Dec 26 07:07:51 2026 GMT
The Sablier Pattern: On-Demand Services
Not every service needs to run 24/7. Some developer tools (OpenCode, the blog ideas generator, ad-hoc dashboards) only need to be available when someone is actually using them. The sablier proxy pattern handles this: Traefik routes to sablier, which starts the container on first request and proxies the traffic once it's ready. After an idle timeout, sablier stops the container.
This keeps the cert-manager certificate valid (the IngressRoute always exists) while saving resources on services that are rarely accessed.
What's Left
A handful of Docker Traefik routes still exist in 04-network-traefik/routes.yml
for services not yet migrated to k3s. These will be converted as their
services move to k3s deployments. The old certificatesResolvers config
in traefik.yml has been stripped — Docker Traefik is now purely a
file-provider router with no ACME capabilities.
Future plans:
- DNS-01 challenges instead of HTTP-01, to support wildcard
certificates (*.aldof.duckdns.org) and reduce the SAN list
- Automated renewal alerts via cert-manager's CertificateRequest
status conditions
- Migrate remaining Docker services to k3s deployments with proper
Service + Endpoints definitions
Lessons Learned
-
Don't run two ACME resolvers for the same domains. Pick one certificate management system and stick with it. cert-manager is purpose-built for this; Traefik's built-in ACME is a convenience feature that doesn't scale.
-
One certificate with many SANs beats many certificates. One ACME order, one renewal, one secret. cert-manager handles the ordering; you just reference the secret.
-
Disk pressure is a real Kubernetes failure mode. On a single-node Pi 5, disk pressure evicts pods and taints the node. Monitor disk usage before it hits 85%, and add tolerations for critical infrastructure pods.
-
Verify with real curl output, not assumptions. The 503/404 responses in the verification file proved routing worked. A successful TLS handshake proved the certificate was valid. These are the checks that matter — not
kubectl get podshowingRunning. -
Migration is iterative. The commit history shows 8 commits over 3 days (Sep 14-17), with multiple test cases, failures, and retries. That's normal. Don't expect a clean one-shot migration.