Runner: containers on the dind daemon can't resolve DNS, so uncached ci builds fail #22
Labels
No labels
bug
enhancement
needs-info
needs-triage
ready-for-agent
ready-for-human
wayfinder:grilling
wayfinder:map
wayfinder:prototype
wayfinder:research
wayfinder:task
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
vicoli-oss/docker-forgejo-mcp#22
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What's wrong
On the
dockerrunner, containers on the dind daemon's default bridge can't resolve names. That includes BuildKitRUNsteps. So everycibuild without a warm cache fails in the Dockerfile'sRUN apk add --no-cache git:Runs that failed this way: 10, 11, 14, 15, 16 (main), 17 (main). Runs 12 and 13 were green only because every build step came from cache.
Cause
A probe on the runner, run 18 (throwaway
dns-probe.yml, not merged), showed this:/etc/resolv.confnslookup dl-cdn.alpinelinux.orgdocker:cli)127.0.0.11docker runon the dind daemon (default bridge)8.8.8.8,8.8.4.4("Used default nameservers")8.8.8.8or1.1.1.1directlydocker run --network hoston the dind daemon127.0.0.11docker buildx buildRUNstep8.8.8.8,8.8.4.4+ IPv6 GoogleNetwork unreachableafter 10sdocker buildx build --network hostRUNstep127.0.0.11The dind container sits on a user-defined network, so its own resolver is Docker's embedded DNS at
127.0.0.11. dockerd inside dind can't hand that address to containers on its bridge, so it falls back to Google DNS. From the runner host, DNS queries to public resolvers don't get through: they time out, or at best work only some of the time. HTTPS to1.1.1.1by IP did get an answer, so it's DNS specifically, not all egress.This is not a collision between concurrent runs. Runs 10 and 11 failed together because BuildKit ran their identical
apk addstep once and streamed that one failure to both jobs, which is why the two logs are identical down to the millisecond. Run 16 ran alone and failed the same way. Image tags, container and network names, the builder and ports are all unique per run or unused.Fix (runner, needs admin)
Give the dind daemon a resolver its bridge containers can reach. For example, start dockerd with
--dns <ip>(or"dns": [...]in itsdaemon.json), pointing at the runner host's resolver or at any resolver the host firewall lets through. Alternatively, allow outbound DNS to the fallback resolvers.Workaround in this repo
cibuilds with--network host(BUILD_NETWORK=hostin the Makefile), soRUNsteps use the dind container's resolver. Once the runner is fixed, dropBUILD_NETWORK=hostfrom.forgejo/workflows/ci.yml.Don't let the workaround become permanent. With
--network host,RUNsteps share the dind container's network namespace, so they can reach anything listening there, including the daemon's own port. The job already has full daemon access, so for this repo's own PRs that adds little. It is still one more path to the shared daemon.Acceptance criteria
docker run --rm alpine nslookup dl-cdn.alpinelinux.orgon the runner's dind daemon resolvesciis green withBUILD_NETWORK=hostremoved and the build cache cold