Rootless Podman Quadlets

Two years ago, I reinstalled my server on which runs a few services like emails, blogs, Nextcloud… So far I always used Debian with, as much as possible, provided packages. But this time I wanted to go the container way.

Docker kind of pioneered the container world, but it usually runs as root, which is not ideal for security reasons. Podman, on the other hand, is designed to run containers as unprivileged users, making it a more secure option.

For the containers themselves, quite often the service running in the container runs as root. That’s not great either. But some people/organizations offer non-root container images, which run as unprivileged users, making them more secure as well.

Anyway, here are a few of the issues I encountered when trying to run rootless containers with Podman. It’s a bit of an outdated post, so most of it would not apply to a new installation, but I didn’t manage to write it earlier.

Podman was young

So I started that installation around August 2024, on a Debian 12 system. And the version of podman provided by Debian 12 was getting old. At first, I thought I would use the podman’s equivalent of docker compose, where you have a file describing the different containers with their configuration. To make it more modular, I wanted to use the include directive, which allows linking multiple compose files. But that required a newer version of podman-compose.

So I tried to install a newer version via a dedicated Python virtual environment. But that was not realistic from a maintainability point of view. And as I was reading about podman, I saw several times this thing called Quadlet and everyone saying that it was “the future”.

But for that… Debian 12 was a no-go. So I upgraded to Debian 13 with Podman 5. Debian 13 was about 1 year before being released. So… running a development version of an operating system in production is not what I would call a good move 😉. But I took the risk (also by the time my installation would be ready, I would be closer to Debian 13 release date).

Quadlets

Docker Compose organizes your containers, defines the volumes, the networks, the services that are running, services dependencies on other services, etc. When you are running Linux, chances are that you have systemd running. Systemd organizes the services running on your computer, the dependencies between them, and the order in which they are started.

Do you see a kind of duplication between Docker Compose and Systemd?

So I guess at some point some people thought that it would be a good idea to use Systemd to organize the containers and get rid of Compose or similar tools.

So systemd units replace Docker Compose files.

So I have a podman user, but the thing is, by default, systemd services linked to a user start and stop when the user logs in and out. Which is not the expected behavior of a server. First, make sure to have the podman socket enabled (if some of your containers (like Traefik) need to access it): systemctl --user enable podman.socket Then, to have the user’s systemd daemon start on boot and not stop when the user logs out, lingering needs to be enabled: sudo loginctl enable-linger podman (where podman is the unprivileged user)

Then, to make your services start on boot, in their unit file, put:

[Install]
WantedBy=default.target

For example, here is what this blog’s nginx quadlet looks like:

[Unit]
Description=Blog desgrange.net

[Install]
WantedBy=default.target

[Service]
Restart=always
# Extend Timeout to allow time to pull the image.
TimeoutStartSec=900

[Container]
ContainerName=blog
Image=docker.io/library/nginx:latest
AutoUpdate=registry
NoNewPrivileges=true
UIDMap=0:1100001:65536
GIDMap=0:1100001:65536

Network=proxy.network
ExposeHostPort=8080

Environment=TZ=UTC

Volume=blog_config.volume:/etc/nginx/conf.d
Volume=/path/to/blog:/app

Label=traefik.enable=true
Label=traefik.docker.network=proxy
Label=traefik.http.routers.blog.rule=Host(`blog.desgrange.net`)
Label=traefik.http.routers.blog.entrypoints=websecure
Label=traefik.http.routers.blog.tls=true
Label=traefik.http.routers.blog.tls.certresolver=letsencrypt-xxx
Label=traefik.http.routers.blog.middlewares=blog
Label=traefik.http.middlewares.blog.headers.stsseconds=63072000
Label=traefik.http.middlewares.blog.headers.stsincludesubdomains=true
Label=traefik.http.middlewares.blog.headers.stspreload=true
Label=traefik.http.services.blog.loadbalancer.server.port=8080

Disk full

I use Ansible to do the installation, so it’s repeatable. And I tested everything on a virtual machine first. That VM had at least 20 GB of disk space. And I was quite surprised that I had to deal with disk full issues after installing a few small containers.

By default, podman was using VFS, which is not the most efficient storage driver. A container image is usually built in several steps, each step is called a layer. And basically, VFS stores each layer in a separate directory, and each layer contains the whole content of the image filesystem so far. So a 7-layer image, stored that way, may take ~7 times the size of the image filesystem.

Here the solution was to switch to the OverlayFS storage driver. OverlayFS stores only the differences between layers, which is way more efficient than VFS. Also, if you have several images using the same base image, OverlayFS will store the base image only once.

So I had to install the userland version of OverlayFS: sudo apt install fuse-overlayfs

~/.config/containers/storage.conf
[storage]
driver="overlay"
[storage.options]
mount_program="/usr/bin/fuse-overlayfs"

Note that was not necessary once I switched to Debian 13 + Podman 5, as it comes with OverlayFS by default.

Inter-containers networking

The architecture was that if a few containers have to work together, let’s say a WordPress with a MariaDB and a Redis, they would be in the same network. And only WordPress would also be in the proxy network to receive requests from the outside.

But somehow, the containers could not reliably communicate with each other.

Turns out Podman was delegating network management (bridge, NAT, DNS…) to an old CNI plugin, and it seemed it was known to not be done in a very reliable way. The solution was simple, switch to podman’s integrated network backend “Netavark” and the userland tooling: sudo apt install aardvark-dns iptables passt netavark slirp4netns

~/.config/containers/containers.conf
[network]
network_backend = "netavark"

Note that this configuration is not necessary with Podman 5 if the required packages are installed.

Also note that changing the storage driver or the network backend requires podman system reset, which deletes all images and containers. That’s why I was doing all my tests in a VM.

Rootless low-port binding

On Linux, ports below 1024 are reserved for the root user. So my regular podman user can’t bind it by default, and running web servers, e-mail servers, etc. requires low ports.

There are a few solutions for this problem:

Getting real IP addresses

One important thing in the containers is to have the “real” IP address of the machine connecting to the server. That’s useful for multiple reasons:

But here I have Internet → Router → Server → Podman network → Proxy container → App container. So of course at first I saw only local IP addresses in the logs 🙁.

What ended up working for me is to use systemd sockets; they open ports directly on the host so they can directly see the real IP addresses:

~/.config/systemd/user/https.socket
[Socket]
ListenStream=0.0.0.0:443
FileDescriptorName=websecure
Service=traefik.service

[Install]
WantedBy=sockets.target default.target

And use them directly by the proxy server (in this case Traefik):

~/.config/containers/systemd/traefik.container
[Service]
Sockets=https.socket

[Container]
ExposeHostPort=443

Namespaces

To add an extra layer, in the defense-in-depth spirit, I wanted to isolate the UIDs and GIDs of the containers. The idea is that a user with UID 33 in container A and a user with UID 33 in container B, by default, map to the same subordinate UID in the host. So technically, if an attacker gains access to container A and escapes from it, they can access some of the resources of container B.

To prevent that, I first started with UserNS=auto, but that gives a random subordinate UID mapping on each start of the container. So persistent files would be readable the first time, but not after a reboot.

I tried various settings for UserNS but didn’t achieve what I wanted. AFAIR, UserNS is the new way to declare namespaces, the old way (but still supported) being to declare a UIDMap and a GIDMap. But in the end, I managed to get what I wanted with that “old” way, not without any issues, like podman’s issue 22803.

I went the brute way. First, I assigned a huge range of subordinate UIDs/GIDs to the non-privileged user: usermod --add-subuids 1000000-9999999 --add-subgids 1000000-9999999 podman Then I gave a full non-overlapping range of UIDs/GIDs to each container. By default, a container needs 65 536 UIDs/GIDs, so that’s what I gave them, and a “few” empty ones between each container for simplification (counting in binary in base 10 is not easy):

Starting host UID/GIDEnding host UID/GID
Container A10000001065535
Container B11000001165535
Container C12000001265535

That works quite well. The main issue being when sharing resources between containers, I have discovered options in ACLs that I didn’t know existed (and that I have already forgotten).

Non-root images provider

I started this by using Bitnami’s non-root images. They use a mini-Debian base image, ensure the images run with an unprivileged user, and share similar config patterns. And they are free and open source. So a great choice…

Bitnami is sponsored by Bitrock, which was acquired by VMware in 2019. So far so good, nothing changed. Then, in 2023, Broadcom acquired VMware. To put it simply, investors love Broadcom; their clients don’t.

Broadcom realized they were giving something for free, how outrageous! So, outside a (very expensive) paying plan, they first limited the Bitnami images to the “latest” tag. Then at some point, they stopped updating the images at all. And then moved all the “latest” images to a different repository. So lots of people still using those outdated images had their system break, the images not being where they were at all.

Comments Add one by emailing me.