Rootless Podman Quadlets
Two years ago, I reinstalled my server on which runs a few services like emails, blogs, Nextcloud… So far I always used Debian with, as much as possible, provided packages. But this time I wanted to go the container way.
Docker kind of pioneered the container world, but it usually runs as root, which is not ideal for security reasons. Podman, on the other hand, is designed to run containers as unprivileged users, making it a more secure option.
For the containers themselves, quite often the service running in the container runs as root. That’s not great either. But some people/organizations offer non-root container images, which run as unprivileged users, making them more secure as well.
Anyway, here are a few of the issues I encountered when trying to run rootless containers with Podman. It’s a bit of an outdated post, so most of it would not apply to a new installation, but I didn’t manage to write it earlier.
Podman was young
So I started that installation around August 2024, on a Debian 12 system.
And the version of podman provided by Debian 12 was getting old.
At first, I thought I would use the podman’s equivalent of docker compose, where you have a file describing the different containers with their configuration.
To make it more modular, I wanted to use the include directive, which allows linking multiple compose files.
But that required a newer version of podman-compose.
So I tried to install a newer version via a dedicated Python virtual environment. But that was not realistic from a maintainability point of view. And as I was reading about podman, I saw several times this thing called Quadlet and everyone saying that it was “the future”.
But for that… Debian 12 was a no-go. So I upgraded to Debian 13 with Podman 5. Debian 13 was about 1 year before being released. So… running a development version of an operating system in production is not what I would call a good move 😉. But I took the risk (also by the time my installation would be ready, I would be closer to Debian 13 release date).
Quadlets
Docker Compose organizes your containers, defines the volumes, the networks, the services that are running, services dependencies on other services, etc. When you are running Linux, chances are that you have systemd running. Systemd organizes the services running on your computer, the dependencies between them, and the order in which they are started.
Do you see a kind of duplication between Docker Compose and Systemd?
So I guess at some point some people thought that it would be a good idea to use Systemd to organize the containers and get rid of Compose or similar tools.
So systemd units replace Docker Compose files.
So I have a podman user, but the thing is, by default, systemd services linked to a user start and stop when the user logs in and out.
Which is not the expected behavior of a server.
First, make sure to have the podman socket enabled (if some of your containers (like Traefik) need to access it): systemctl --user enable podman.socket
Then, to have the user’s systemd daemon start on boot and not stop when the user logs out, lingering needs to be enabled: sudo loginctl enable-linger podman (where podman is the unprivileged user)
Then, to make your services start on boot, in their unit file, put:
[Install]
WantedBy=default.targetFor example, here is what this blog’s nginx quadlet looks like:
[Unit]
Description=Blog desgrange.net
[Install]
WantedBy=default.target
[Service]
Restart=always
# Extend Timeout to allow time to pull the image.
TimeoutStartSec=900
[Container]
ContainerName=blog
Image=docker.io/library/nginx:latest
AutoUpdate=registry
NoNewPrivileges=true
UIDMap=0:1100001:65536
GIDMap=0:1100001:65536
Network=proxy.network
ExposeHostPort=8080
Environment=TZ=UTC
Volume=blog_config.volume:/etc/nginx/conf.d
Volume=/path/to/blog:/app
Label=traefik.enable=true
Label=traefik.docker.network=proxy
Label=traefik.http.routers.blog.rule=Host(`blog.desgrange.net`)
Label=traefik.http.routers.blog.entrypoints=websecure
Label=traefik.http.routers.blog.tls=true
Label=traefik.http.routers.blog.tls.certresolver=letsencrypt-xxx
Label=traefik.http.routers.blog.middlewares=blog
Label=traefik.http.middlewares.blog.headers.stsseconds=63072000
Label=traefik.http.middlewares.blog.headers.stsincludesubdomains=true
Label=traefik.http.middlewares.blog.headers.stspreload=true
Label=traefik.http.services.blog.loadbalancer.server.port=8080Disk full
I use Ansible to do the installation, so it’s repeatable. And I tested everything on a virtual machine first. That VM had at least 20 GB of disk space. And I was quite surprised that I had to deal with disk full issues after installing a few small containers.
By default, podman was using VFS, which is not the most efficient storage driver. A container image is usually built in several steps, each step is called a layer. And basically, VFS stores each layer in a separate directory, and each layer contains the whole content of the image filesystem so far. So a 7-layer image, stored that way, may take ~7 times the size of the image filesystem.
Here the solution was to switch to the OverlayFS storage driver. OverlayFS stores only the differences between layers, which is way more efficient than VFS. Also, if you have several images using the same base image, OverlayFS will store the base image only once.
So I had to install the userland version of OverlayFS: sudo apt install fuse-overlayfs
[storage]
driver="overlay"
[storage.options]
mount_program="/usr/bin/fuse-overlayfs"Note that was not necessary once I switched to Debian 13 + Podman 5, as it comes with OverlayFS by default.
Inter-containers networking
The architecture was that if a few containers have to work together, let’s say a WordPress with a MariaDB and a Redis, they would be in the same network. And only WordPress would also be in the proxy network to receive requests from the outside.
But somehow, the containers could not reliably communicate with each other.
Turns out Podman was delegating network management (bridge, NAT, DNS…) to an old CNI plugin, and it seemed it was known to not be done in a very reliable way.
The solution was simple, switch to podman’s integrated network backend “Netavark” and the userland tooling: sudo apt install aardvark-dns iptables passt netavark slirp4netns
[network]
network_backend = "netavark"Note that this configuration is not necessary with Podman 5 if the required packages are installed.
Also note that changing the storage driver or the network backend requires podman system reset, which deletes all images and containers.
That’s why I was doing all my tests in a VM.
Rootless low-port binding
On Linux, ports below 1024 are reserved for the root user.
So my regular podman user can’t bind it by default, and running web servers, e-mail servers, etc. requires low ports.
There are a few solutions for this problem:
- run a privileged proxy… but I already have a proxy in my containers, so that would be a bit too redundant, increase complexity and attack surface;
- add the
CAP_NET_BIND_SERVICEcapability to the container… but it didn’t work. AFAIR the process opening the port was not the container itself; in fact, I use systemd socket units, so the process to grant the capability to was probablyrootlessportorpasta. Also, I think on each update of the binary (after anapt update), I would have to apply the capability again; - tell Linux that there are no privileged ports (or specify the lowest port you intend to use):The main drawback of this solution is that any application can bind to low ports, but it can be mitigated with a firewall opening only the known ports used by the containers. That’s the solution I chose./etc/sysctl.d/unprivileged_port_start.conf
net.ipv4.ip_unprivileged_port_start = 0
Getting real IP addresses
One important thing in the containers is to have the “real” IP address of the machine connecting to the server. That’s useful for multiple reasons:
- debugging/troubleshooting (so you see if your requests are logged at the right place);
- IP filtering/firewalling (restricting a service for local network only, for instance);
- security (detect and filter attacking IPs with tools like fail2ban or CrowdSec);
- …
But here I have Internet → Router → Server → Podman network → Proxy container → App container. So of course at first I saw only local IP addresses in the logs 🙁.
What ended up working for me is to use systemd sockets; they open ports directly on the host so they can directly see the real IP addresses:
[Socket]
ListenStream=0.0.0.0:443
FileDescriptorName=websecure
Service=traefik.service
[Install]
WantedBy=sockets.target default.targetAnd use them directly by the proxy server (in this case Traefik):
[Service]
Sockets=https.socket
[Container]
ExposeHostPort=443Namespaces
To add an extra layer, in the defense-in-depth spirit, I wanted to isolate the UIDs and GIDs of the containers. The idea is that a user with UID 33 in container A and a user with UID 33 in container B, by default, map to the same subordinate UID in the host. So technically, if an attacker gains access to container A and escapes from it, they can access some of the resources of container B.
To prevent that, I first started with UserNS=auto, but that gives a random subordinate UID mapping on each start of the container.
So persistent files would be readable the first time, but not after a reboot.
I tried various settings for UserNS but didn’t achieve what I wanted.
AFAIR, UserNS is the new way to declare namespaces, the old way (but still supported) being to declare a UIDMap and a GIDMap.
But in the end, I managed to get what I wanted with that “old” way, not without any issues, like podman’s issue 22803.
I went the brute way.
First, I assigned a huge range of subordinate UIDs/GIDs to the non-privileged user: usermod --add-subuids 1000000-9999999 --add-subgids 1000000-9999999 podman
Then I gave a full non-overlapping range of UIDs/GIDs to each container.
By default, a container needs 65 536 UIDs/GIDs, so that’s what I gave them, and a “few” empty ones between each container for simplification (counting in binary in base 10 is not easy):
| Starting host UID/GID | Ending host UID/GID | |
|---|---|---|
| Container A | 1000000 | 1065535 |
| Container B | 1100000 | 1165535 |
| Container C | 1200000 | 1265535 |
That works quite well. The main issue being when sharing resources between containers, I have discovered options in ACLs that I didn’t know existed (and that I have already forgotten).
Non-root images provider
I started this by using Bitnami’s non-root images. They use a mini-Debian base image, ensure the images run with an unprivileged user, and share similar config patterns. And they are free and open source. So a great choice…
Bitnami is sponsored by Bitrock, which was acquired by VMware in 2019. So far so good, nothing changed. Then, in 2023, Broadcom acquired VMware. To put it simply, investors love Broadcom; their clients don’t.
Broadcom realized they were giving something for free, how outrageous! So, outside a (very expensive) paying plan, they first limited the Bitnami images to the “latest” tag. Then at some point, they stopped updating the images at all. And then moved all the “latest” images to a different repository. So lots of people still using those outdated images had their system break, the images not being where they were at all.