Runner: containers on the dind daemon can't resolve DNS, so uncached ci builds fail #22

Closed
opened 2026-09-30 14:45:36 +00:00 by piscis · 0 comments
Owner

What's wrong

On the docker runner, containers on the dind daemon's default bridge can't resolve names. That includes BuildKit RUN steps. So every ci build without a warm cache fails in the Dockerfile's RUN apk add --no-cache git:

#9 5.153 WARNING: fetching https://dl-cdn.alpinelinux.org/alpine/v3.24/main/x86_64/APKINDEX.tar.gz: DNS: transient error (try again later)
#9 10.16 ERROR: unable to select packages:
#9 10.16   git (no such package):

Runs that failed this way: 10, 11, 14, 15, 16 (main), 17 (main). Runs 12 and 13 were green only because every build step came from cache.

Cause

A probe on the runner, run 18 (throwaway dns-probe.yml, not merged), showed this:

Where /etc/resolv.conf nslookup dl-cdn.alpinelinux.org
Job container (docker:cli) 127.0.0.11 works
docker run on the dind daemon (default bridge) 8.8.8.8, 8.8.4.4 ("Used default nameservers") times out
same, asking 8.8.8.8 or 1.1.1.1 directly times out
docker run --network host on the dind daemon 127.0.0.11 works
docker buildx build RUN step 8.8.8.8, 8.8.4.4 + IPv6 Google Network unreachable after 10s
docker buildx build --network host RUN step 127.0.0.11 works

The dind container sits on a user-defined network, so its own resolver is Docker's embedded DNS at 127.0.0.11. dockerd inside dind can't hand that address to containers on its bridge, so it falls back to Google DNS. From the runner host, DNS queries to public resolvers don't get through: they time out, or at best work only some of the time. HTTPS to 1.1.1.1 by IP did get an answer, so it's DNS specifically, not all egress.

This is not a collision between concurrent runs. Runs 10 and 11 failed together because BuildKit ran their identical apk add step once and streamed that one failure to both jobs, which is why the two logs are identical down to the millisecond. Run 16 ran alone and failed the same way. Image tags, container and network names, the builder and ports are all unique per run or unused.

Fix (runner, needs admin)

Give the dind daemon a resolver its bridge containers can reach. For example, start dockerd with --dns <ip> (or "dns": [...] in its daemon.json), pointing at the runner host's resolver or at any resolver the host firewall lets through. Alternatively, allow outbound DNS to the fallback resolvers.

Workaround in this repo

ci builds with --network host (BUILD_NETWORK=host in the Makefile), so RUN steps use the dind container's resolver. Once the runner is fixed, drop BUILD_NETWORK=host from .forgejo/workflows/ci.yml.

Don't let the workaround become permanent. With --network host, RUN steps share the dind container's network namespace, so they can reach anything listening there, including the daemon's own port. The job already has full daemon access, so for this repo's own PRs that adds little. It is still one more path to the shared daemon.

Acceptance criteria

  • docker run --rm alpine nslookup dl-cdn.alpinelinux.org on the runner's dind daemon resolves
  • ci is green with BUILD_NETWORK=host removed and the build cache cold
## What's wrong On the `docker` runner, containers on the dind daemon's default bridge can't resolve names. That includes BuildKit `RUN` steps. So every `ci` build without a warm cache fails in the Dockerfile's `RUN apk add --no-cache git`: ```text #9 5.153 WARNING: fetching https://dl-cdn.alpinelinux.org/alpine/v3.24/main/x86_64/APKINDEX.tar.gz: DNS: transient error (try again later) #9 10.16 ERROR: unable to select packages: #9 10.16 git (no such package): ``` Runs that failed this way: [10](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/10), [11](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/11), [14](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/14), [15](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/15), [16](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/16) (main), [17](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/17) (main). Runs [12](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/12) and [13](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/13) were green only because every build step came from cache. ## Cause A probe on the runner, [run 18](https://code.vicoli.de/vicoli-oss/docker-forgejo-mcp/actions/runs/18) (throwaway `dns-probe.yml`, not merged), showed this: | Where | `/etc/resolv.conf` | `nslookup dl-cdn.alpinelinux.org` | |---|---|---| | Job container (`docker:cli`) | `127.0.0.11` | works | | `docker run` on the dind daemon (default bridge) | `8.8.8.8`, `8.8.4.4` ("Used default nameservers") | times out | | same, asking `8.8.8.8` or `1.1.1.1` directly | | times out | | `docker run --network host` on the dind daemon | `127.0.0.11` | works | | `docker buildx build` `RUN` step | `8.8.8.8`, `8.8.4.4` + IPv6 Google | `Network unreachable` after 10s | | `docker buildx build --network host` `RUN` step | `127.0.0.11` | works | The dind container sits on a user-defined network, so its own resolver is Docker's embedded DNS at `127.0.0.11`. dockerd inside dind can't hand that address to containers on its bridge, so it falls back to Google DNS. From the runner host, DNS queries to public resolvers don't get through: they time out, or at best work only some of the time. HTTPS to `1.1.1.1` by IP did get an answer, so it's DNS specifically, not all egress. This is not a collision between concurrent runs. Runs 10 and 11 failed together because BuildKit ran their identical `apk add` step once and streamed that one failure to both jobs, which is why the two logs are identical down to the millisecond. Run 16 ran alone and failed the same way. Image tags, container and network names, the builder and ports are all unique per run or unused. ## Fix (runner, needs admin) Give the dind daemon a resolver its bridge containers can reach. For example, start dockerd with `--dns <ip>` (or `"dns": [...]` in its `daemon.json`), pointing at the runner host's resolver or at any resolver the host firewall lets through. Alternatively, allow outbound DNS to the fallback resolvers. ## Workaround in this repo `ci` builds with `--network host` (`BUILD_NETWORK=host` in the Makefile), so `RUN` steps use the dind container's resolver. Once the runner is fixed, drop `BUILD_NETWORK=host` from `.forgejo/workflows/ci.yml`. Don't let the workaround become permanent. With `--network host`, `RUN` steps share the dind container's network namespace, so they can reach anything listening there, including the daemon's own port. The job already has full daemon access, so for this repo's own PRs that adds little. It is still one more path to the shared daemon. ## Acceptance criteria - [ ] `docker run --rm alpine nslookup dl-cdn.alpinelinux.org` on the runner's dind daemon resolves - [ ] `ci` is green with `BUILD_NETWORK=host` removed and the build cache cold
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
vicoli-oss/docker-forgejo-mcp#22
No description provided.