Moving House: Migrating 29 Repositories from Codeberg to Codefloe in a Day

Why Vidocq moved to codefloe.com, and the complete field guide -- what the Forgejo migrator copies, what it silently doesn't, and the seven traps we hit so you don't have to.

The code flow from Codeberg to Codefloe

Vidocq just moved house for the second time in four months. In May our self-hosted Forgejo died with the machine it lived on, and Codeberg gave the project a home within days — 29 repositories, CI, mirrors, the lot. Today the whole ecosystem lives on Codefloe, a managed Forgejo instance, and this post is both the story of that day and the field guide we wish we had found beforehand.

Why we moved

Codeberg has taken the position that projects built with AI assistance do not belong on their platform. Vidocq is openly and transparently an AI-assisted project — it says so in its provenance policy, in its commit trailers, and in the very first post of this blog. So the two were no longer a fit, and we want to be clear about the spirit of the move: we understand and we respect that choice. Codeberg is a community-funded association; deciding what it hosts is exactly its prerogative. They sheltered this project when our own forge burned down, and we left the way you leave a friend’s house — grateful, and tidying up behind us: every repository was archived, not deleted, so history, links and old action references keep resolving.

Codefloe, on the other hand, told us explicitly that AI-assisted projects are welcome. And because a 29-repository ecosystem with TCK suites is not a light CI tenant, we run our own runners there — four Docker-in-Docker stacks on our infrastructure — rather than saturating the shared ones.

The shape of the migration

We did a full dry run ten days earlier: migrate all 29 repositories with Forgejo’s built-in migrator, look at what survived, throw it away. That dry run produced the script we used on the real day (POST /repos/migrate with service: forgejo — not gitea, which 500s — issues, pull requests, labels and topics included, plus a source-vs-target verification of branches, tags, issues, PRs and releases).

The real day then ran in four phases:

  1. Clean the house first. Merge every PR of ours that was ready (eight of them, including one whose CI turned out to depend on a snapshot artifact that had never been deployed — more below), push every local branch that existed nowhere else, snapshot work-in-progress. Migrating a forge is the worst possible moment to discover that a branch only existed on one laptop.

  2. Infrastructure before repositories. Runners registered on the target (org-level registration tokens: GET /orgs/<org>/actions/runners/registration-token), then the ~35 organization secrets. A migrated repository whose first CI run fails for a missing secret teaches you nothing.

  3. Repo by repo, infra repos first. Our composite-actions repo (ci), governance, the dependency graph, the sites — then the flagship (vauban) with a full smoke test (throwaway PR, real build, real governance gate), then everything else in batches. Each repository ended its runbook archived on Codeberg, only after verifying that its head commit existed on Codefloe.

  4. The sweeps. Workflows (uses: URLs), then POM <scm> blocks, then a git grep codeberg.org origin/main across all 29 repositories until it returned zero.

What the Forgejo migrator does not copy

The migrator is genuinely good: code, branches, tags, issues, pull requests, labels, topics all arrive. Here is the checklist of what does not travel, all of which we had to recreate by API:

  • has_actions — migrated repositories arrive with Actions disabled even when the source had them enabled;

  • branch protections — and when you recreate them, usernames that have no account on the target make the call fail, so filter the whitelists;

  • organization teams and their members;

  • Actions secrets (by design — values are never readable);

  • push mirrors (our GitHub mirrors had to be recreated, and only for repositories that actually exist on the GitHub side);

  • deploy webhooks, SSH keys, and of course the runners themselves.

One more quirk: the migrator sometimes resurrects branches of merged PRs that had been deleted on the source. Three of our repositories arrived with one extra branch each; harmless, but it makes a naive branch-count verification cry wolf.

The seven traps

This is the part we wish someone had written before us.

1. Your CI actions hardcode the forge. Our composite actions called https://codeberg.org/... in a dozen places — API calls, raw fetches, clones. On the new forge those calls either 404 or, worse, authenticate with the new forge’s token against the old forge and 401. The fix is to derive every URL from $GITHUB_SERVER_URL, which the runner always provides. Do this before migrating: it makes the actions forge-agnostic forever.

2. Shared runners will happily steal your labels. Codefloe’s shared runners serve ubuntu-latest and docker. Our image-building jobs (runs-on: docker) landed on a shared runner that had no Docker-in-Docker sidecar and died with Docker daemon unreachable. The fix: give your own runners an exclusive label (vidocq-dind) and target that.

3. Docker-in-Docker inherits the host’s DNS. On one of our runner hosts the dind daemon inherited a LAN-only resolver, and job containers could not resolve public hostnames — while Maven downloads worked fine on the other host. An explicit dockerd --dns 1.1.1.1 ended the mystery.

4. A cancelled run leaves a lying commit status. Force-push twice in a row and the cancelled run writes a failure status on the sha that the newer run does not immediately overwrite. Our watcher declared a smoke test dead while it was green and still running. Trust the jobs of the run, not the commit status, when force-pushes are involved.

5. The silent no-op sweep. Our per-repository sweep did git checkout main && sed …​ && git push origin main. On one repository the checkout failed — uncommitted work from a live agent session — so the sed ran on the feature branch, the commit landed there, and git push origin main reported success because local main was simply up to date. The sweep looked green and had done nothing. Two lessons: verify sweeps with git grep against origin/main, not with your script’s output; and use git worktree for any repository whose working tree you don’t fully control.

6. CDNs forget, archives don’t. A workflow pinned Maven 4.0.0-rc-5 from dlcdn.apache.org, which purges superseded releases. It had been silently un-buildable for weeks. archive.apache.org keeps everything; add it as a fallback.

7. Migration is a flashlight. Two of the three CI failures we chased on the new forge were pre-existing bugs the old forge had been hiding in plain sight — a javadoc jar failing on a module that no longer has public classes, red on main for a month. Re-running every pipeline on fresh infrastructure is the best audit you will ever get for free.

Repointing every clone

The forge moves; your checkouts don’t. If you have Vidocq clones (or are migrating your own org), three commands cover it. Once per machine, trust the new host and make sure your SSH key is registered on it:

ssh-keyscan codefloe.com >> ~/.ssh/known_hosts

Then either repoint each clone explicitly:

git remote set-url origin git@codefloe.com:Vidocq/<repo>.git

or, for a whole workspace of clones, let a loop do it:

for d in */; do
  old=$(git -C "$d" remote get-url origin 2>/dev/null) || continue
  git -C "$d" remote set-url origin "${old/codeberg.org/codefloe.com}"
done

On Windows the same two steps in PowerShell (the bundled OpenSSH client provides ssh-keyscan):

ssh-keyscan codefloe.com | Add-Content "$env:USERPROFILE\.ssh\known_hosts"

Get-ChildItem -Directory | ForEach-Object {
  $old = git -C $_.FullName remote get-url origin 2>$null
  if ($old) {
    git -C $_.FullName remote set-url origin ($old -replace 'codeberg\.org', 'codefloe.com')
  }
}

There is also a zero-touch option: a scoped insteadOf rule rewrites the old URLs on the fly for every clone on the machine, without editing any of them — scoped to the organizations so the rest of your Codeberg usage is untouched:

git config --global url."git@codefloe.com:Vidocq/".insteadOf "git@codeberg.org:Vidocq/"
git config --global url."https://codefloe.com/Vidocq/".insteadOf "https://codeberg.org/Vidocq/"

And for people who arrive at the old address with no idea a move happened, every archived repository on Codeberg now carries a "moved to codefloe.com" pointer in its description.

Did it actually work?

Three proofs, because a green job proves less than you think:

  • the smoke PR on vauban went through the full gate — governance checks, 20-module build, dependency-graph impact analysis — on Codefloe runners;

  • maven-metadata.xml on the Central Portal snapshots repository showed fresh timestamps minutes after the sweep pushes for every family — and since Codeberg was already archived, only the new CI could have produced them;

  • git push against the archived source answers Forgejo: Repo: Vidocq/vauban is archived. — read-only enforced by the server, not by good intentions.

On method: a human, an AI, and 29 repositories

True to how Vidocq is built, this migration was pair work between a human and an AI assistant. The AI wrote and ran the API scripts, watched the CI runs, diffed the logs, and drafted the fixes; the human made every decision that mattered — what to merge, what to archive, which secrets to hand over and when — and supplied the credentials the machine had no business finding on its own. Every commit of the day carries its provenance trailer, as our AI policy requires. The traps above were found the honest way: by hitting them, together, in one afternoon.

Thank you, Codeberg — hello, Codefloe

Four months ago Codeberg was the safety net that kept this project alive; today it hosts our archived history, and we mean the word gratitude. The project now lives at codefloe.com/Vidocq and codefloe.com/VidocqTools — same code, same philosophy, same transparency. Issues, PRs and contributions are welcome there. Bring your own opinions; we already bring our own runners.