Skip to content

fix(ci): repair never-green Coq kernel build gate (opam root + t27c path) - #2321

Merged
gHashTag merged 1 commit into
masterfrom
fix/coq-kernel-opam-root
Aug 21, 2026
Merged

gHashTag merged 1 commit into
masterfrom
fix/coq-kernel-opam-root

Conversation

@gHashTag

Copy link
Copy Markdown
Owner

Closes #2320

Which build this is

Two workflows in this repo emit a check named build. cli-tri.yml's was repaired earlier today; this is the other one — .github/workflows/coq-kernel.yml, workflow Coq kernel, job build, step Install Flocq (opam). The check-runs API returns only the newest run per check name, so one has been masking the other.

Never green

All 12 runs on master concluded failure, from run #3 (2026-04-06) to run #173 (2026-08-21). The gate was added broken; it is not a regression. The failing step moved once, which disguises that:

run date died at rest of job
#3 2026-04-06 step 3, actions/checkout@v4 steps 4-8 skipped
#173 2026-08-21 step 4, Install Flocq (opam) steps 5-9 skipped

Steps 5 through 9 have never executed once.

Cause 1 — opam root is not root's

The job runs coqorg/coq:8.19-ocaml-4.14-flambda with options: --user root. That image does ship an initialised opam root, but it belongs to the image's coq user at /home/coq/.opam. Under --user root, HOME=/root, so bare opam finds no .opam and exits 50 with Opam has not been initialised.

Two downstream steps, Build coq/ (T27 + Flocq) and coqchk PhiFloat (consistency), both eval $(opam env) and depend on that same root and switch.

Fix — a job-level OPAMROOT: /home/coq/.opam.

Deliberately not opam init --disable-sandboxing -y: a fresh root has no Coq in it, so opam install coq-flocq would rebuild the entire compiler, and the coqc/coqchk already on PATH would still resolve to the image's switch. Reusing the image's root also inherits its already sandbox-disabled config, which is the thing --disable-sandboxing exists to arrange.

No official setup action (ocaml/setup-ocaml, coq-community/docker-coq-action) appears anywhere in this repository, so there was no demonstrably-working in-repo configuration to copy.

Cause 2 — wrong binary path, latent behind cause 1

Since steps 5-9 never ran, the next failure was never observed. Build t27c and validate phi f64 parameters did:

cd bootstrap && cargo build --release
./target/release/t27c validate-phi

The root Cargo.toml is a workspace whose members include bootstrap, so cargo writes the binary to the workspace target directory at the repo root. Line 2 resolves against bootstrap/, i.e. bootstrap/target/release/t27c — a path that cannot exist. That is an exit 127 which reads like a compiler failure.

Fixed to cargo build --release -p t27c from the repo root, then ./target/release/t27c validate-phi.

The step is not weakened

No || true, no continue-on-error, no dropped dependency. coq-flocq is still installed, and it has to be: coq/Kernel/PhiFloat.v carries From Flocq Require Import IEEE754.Binary. Two diagnostic lines (opam var root, opam switch list) and set -eux were added so a residual failure names itself rather than exiting 50 mutely.

Verification

I could not run this locally — opam and the Coq container are not available to me in this environment, and I did not test it. CI is the verification. Because the workflow's own paths: filter includes .github/workflows/coq-kernel.yml, this PR does trigger the gate, so the check runs rather than being silently absent.

Known sibling, not touched

.github/workflows/coq-proofs.yml (job compile-proofs, step Install Coq Interval) has the identical --user root opam defect. Its check name is compile-proofs so it masks nothing, and its paths: filter (proofs/trinity/**.v) means a change there could not be verified by this PR's run.

The 'build' check owned by .github/workflows/coq-kernel.yml has failed on
every master run since it was added (run #3, 2026-04-06 .. run #173).
Steps 5-9 of the job have never executed once. This is a different check
from cli-tri.yml's 'build', which shares the name and masks it in the
check-runs API.

Cause 1: the job runs coqorg/coq:8.19 with --user root, but the image's
initialised opam root belongs to the coq user at /home/coq/.opam. As root
HOME=/root has no .opam, so opam exited 50. Set job-level OPAMROOT rather
than running opam init, which would rebuild Coq into an empty root while
coqc/coqchk on PATH still came from the image switch.

Cause 2, latent behind cause 1: 'cd bootstrap && cargo build --release'
then './target/release/t27c' resolves to bootstrap/target/release/t27c,
but bootstrap is a member of the root cargo workspace, so the binary
lands in the repo-root target/. Build with -p t27c from the root.

The step is not weakened: coq-flocq is still installed and no failure is
suppressed. coq/Kernel/PhiFloat.v has a hard Flocq import.

Closes #2320
@github-actions

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

@github-actions

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-08-21 09:08:31 UTC

Summary

Status Count
Total Open PRs 3
PRs with Failing Checks 1
PRs with All Checks Green 2
READY 1
FAILING 1
PENDING 0

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=65f033d04125 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Coq kernel: build check has never been green — opam root unreachable as root, plus wrong t27c path

1 participant