A cluster you can rebuild
Kubernetes, Traefik, Let's Encrypt, object storage and one web server, arranged so a new site takes seconds and the whole thing can be rebuilt from a private repo
I have a lot of small things. Static sites, documentation, a few PWAs I wrote on a Sunday and will never promote, a game server, a compiler someone might poke at once a month. Each one is tiny. Together they are enough work that if deploying takes an afternoon, I stop making them.
A full Kubernetes distribution is overkill for this and I use one anyway, because a lightweight single-binary distribution like k3s gives you the whole object model on one machine without the operational weight. Namespaces, routes, policies, rollouts. The same words I would use at work, on a box I can reboot without asking anyone.
So the whole setup below is arranged around one requirement: putting a new thing online should take seconds, and losing the machine should cost a day, not a year.
What follows is the shape of it, not my configuration. Nothing here is a copy of what is running, and I would not publish that anyway.
One server for every site
The piece people over-think first is the web server. You do not need one per site.
Run a single web server and let it pick the document root from the hostname. In Caddy that is roughly:
root * /srv/{host}
try_files {path} {path}/ /index.html
file_server
Now a folder named after a domain is a website. Copy files in, the site exists. Delete the folder, it is gone. nginx does the same thing with a map or a wildcard server block if you prefer nginx.
A new site is a directory with the domain as its name. That is the whole deployment step.
For a PWA this is even better, because a PWA is a pile of static files and a service worker. The try_files line above makes client-side routing work, which is the only thing most single page apps need from a server.
One warning from experience: that same fallback turns every missing file into your homepage with a 200 status. Search engines index the duplicates happily. If you care, carve out a real 404 for the paths that should fail.
Traefik at the door, Let’s Encrypt configured once
Traefik sits in front and decides which hostnames are allowed to reach anything at all. This is the part that surprised me the first time: the web server can be perfectly willing to serve a new domain while the cluster still returns a flat 404, because nothing has told the ingress controller that the hostname exists.
The point of the setup is the certificate resolver. You configure ACME once, as a resolver, and then every route refers to it by name. A new domain then needs:
- a DNS record pointing at the cluster
- a route that matches
Host(...)and names the resolver
No certificate object, no renewal job, no reminder in your calendar for eighty-nine days from now. The resolver obtains and renews everything it is asked for.
Two practical notes. Point the ACME storage at something persistent, or you will re-issue certificates every restart and meet the rate limit in an afternoon. And keep an HTTP route alongside the HTTPS one so the redirect works and the challenge can complete.
Cloudflare in front, and the rule everyone skips
Put a proxy in front of the cluster. I use Cloudflare, mostly for three things: the origin address stops being public, a volumetric attack hits somebody else’s capacity first, and static assets get cached at the edge so a small machine serves a surprising amount of traffic.
Then comes the part that is actually the point, and the part people leave undone for years:
A proxy in front of an origin that still answers the whole internet is decoration. The origin has to stop taking calls from anyone else.
Lock the firewall so the only source allowed to reach your HTTP ports is the proxy’s published address ranges. Everything else is dropped. Without that rule, the moment somebody learns the address, whether from an old DNS record, a certificate log, a misconfigured subdomain or a leaky email header, they talk to your origin directly and every protection you configured is simply skipped.
Three practical notes, because this is the part that bites.
The address ranges change. The provider publishes them and they are not static, so refresh them automatically on a schedule rather than pasting a list once and forgetting which year that was. A site that goes dark after a quiet range update is a miserable thing to debug.
Addresses alone are a weak identity. Anyone can rent space inside those ranges. If it matters, add authenticated origin pulls as well, a client certificate the proxy presents and your origin requires, so the rule becomes “requests from the proxy, proving it is the proxy”, not just “requests from that neighbourhood”.
And keep a way back in that does not depend on the rule being right. A firewall change that locks out your own access is the single most common way to turn a small evening into a long one.
The certificate story survives all of this. The resolver still obtains certificates for the origin, and encryption runs on both legs: browser to proxy, proxy to origin. If challenge traffic ever becomes awkward, switch the resolver to the DNS challenge, which never needs an inbound connection at all and is the better choice for wildcards anyway.
Storage you can find
The orthodox answer is a PersistentVolumeClaim and a storage class, and then your data lives somewhere under a generated directory name with a UUID in it, and finding it means querying the API.
For a single-node or small cluster serving static files, I would rather know where things are. A directory tree on the host, mounted into the pod, named after what it holds. Sites under one parent, databases under another, backups under a third.
If I cannot tar a directory and know I have the data, I do not consider it stored.
Be honest about the trade-offs, because they are real:
- it pins the workload to the node that has the files
- nothing is replicated and nothing is backed up unless you do it
- the container runs as some user id, and the directory has to permit it, which is the single most common reason a pod comes up and then cannot write anything
If you want to stay closer to the orthodoxy and still keep the data findable, there is a middle option and it is the one I would recommend to most people: a path-based provisioner. The default in several lightweight distributions writes every volume into one predictable directory on the host. You get real PersistentVolumeClaims, so the manifests look normal and nothing is unusual to a future reader, and you also get a parent directory you can ls, tar and copy to another machine.
Use it for the things that genuinely want a volume, like a database, a queue or an upload directory; keep the plain mounted tree for the things that are just files.
The actual rule is not “avoid PVCs”. It is that during an incident, at an hour when you are not at your best, you should be able to find your data without reading documentation.
Storage with a user interface
Sooner or later you want to look at the files without opening a terminal, and so does anyone you share the machine with.
There are two levels of this and they solve different problems.
The simple one is a file browser. A small container that mounts the directory tree and serves a web interface over it: upload, download, rename, preview. For a web root, for backups, for dropping a build someone asked for, that is the entire requirement. It takes one deployment and one route, and it is the fastest way to make a directory feel like a service.
The serious one is an S3-compatible object store. MinIO is the usual choice and it ships a console, which is the part that matters here: buckets, objects, access keys and policies in a browser, and an API every language already has a client for. Reach for this when applications are the ones reading and writing, as with uploads, artefacts and backups from other services; then you stop writing path-handling and permission code in every app.
Either way, back it with the same findable directory, give each application its own key and its own bucket, and scope those keys so one application cannot read another’s data. A storage UI is also an administrative surface, so it gets the same treatment as every other one, which is the next section.
A private registry, fed by GitHub Actions
Static files you copy. Anything with a runtime gets built into an image, and that image should live somewhere private.
Two honest options. Run your own registry in the cluster, behind Traefik and behind authentication. Or use the registry attached to your source host; a private package registry on GitHub works fine and is one less service to babysit. I have done both. Running your own is more satisfying and more work.
The pipeline is the same either way:
- a workflow builds the image on push to the main branch or on a tag
- it logs in with a scoped token stored as a repository secret
- it pushes the image tagged with the commit
- the cluster pulls it with an image pull secret in that namespace
Tag images with the commit, not with latest. latest means every rollout is a mystery and every rollback is archaeology.
Zero trust, for one person
Zero trust is a serious discipline and I am one person with a server. But the principles scale down, and they are mostly about refusing a single assumption: that being inside the network means you are trusted.
What that looks like in practice, kept general:
- nothing is reachable because of where it sits. Every service authenticates every caller, including services talking to each other.
- the default is deny. Start with a policy that blocks pod-to-pod traffic and open only the specific paths that are needed.
- least privilege, per workload. A namespace per project rather than one big shared one, its own service account, its own credentials, no cluster-wide permissions handed out because it was quicker. Namespaces cost nothing and they make the blast radius a thing you can point at.
- secrets never live in the repo as plain text. Encrypt them so the repository is useless to whoever reads it.
- administrative interfaces are not published. Consoles, dashboards and the API server do not get a public hostname just because that is convenient.
- assume one container is already compromised and ask what it can reach. That question has improved my setup more than any checklist.
Convenience and trust are the same decision wearing different clothes.
One more, less glamorous: you cannot reason about a cluster you cannot see. A metrics stack with a few alerts is not a security control, but “that pod has been restarting for three days” and “disk is at ninety percent” are how most incidents announce themselves before they become incidents.
I will not pretend the result is secure. It is a lot better than the default, and the default is what most small clusters actually run.
Everything in a private repo
This is the piece I would keep if I had to throw away all the rest.
Every manifest lives in one private repository. Namespaces, deployments, routes, policies, the storage layout, the resolver. If something exists in the cluster and is not described there, it is a bug, and one day it will be the reason a rebuild does not come back up.
What does not go in: secrets in plain text, data, and anything that describes the host specifically enough to be useful to a stranger.
The test is simple and worth running for real. Take a fresh machine, install the distribution, apply the repository, restore the data directory. If the sites come back, you have infrastructure. If you find yourself remembering a command you once ran by hand at two in the morning, you have a pet, and pets die.
What it adds up to
Point a domain, add a route, drop a folder. The certificate appears on its own. The site is live before I have finished deciding whether it was a good idea.
That last part is the entire point. The setup is not impressive, and it is not meant to be. It exists so the cost of trying something is low enough that I keep trying things.
Comments
Threads live in GitHub Discussions. Nothing from GitHub is requested until you press the button.