feat(hosting): meet DigitalOcean's build standard, and write the catalog copy (#2281) - #2680
Conversation
…log copy (#2281) DigitalOcean granted Vendor Portal access and named `digitalocean/marketplace-partners` as the process of record. Auditing the bundle against that repo and against `droplet-1-clicks@master` found four gaps. Three are fixed here; the fourth needs a measurement. **Security updates were never installed.** Their `99-img-check.sh` scores a pending security update as a hard `[FAIL]`, not a warning, and any FAIL exits non-zero — which fails the build at the last provisioner, furthest from the cause. Nothing in the bundle ran an upgrade, so every green build so far was luck about what the base image carried that day, or about whether Ubuntu's own boot-time timers happened to land first. Neither is ours to own. Adds the `full-upgrade` from their reference `marketplace-image.json`, before the first install so the packages we add are patched too. **`X-DO-MARKETPLACE` was missing** from the Caddyfile first boot writes. Rule 10 of their build standard specifies it on the reverse-proxy block and their own catalog apps ship it. It breaks nothing at runtime, which is why it survived review — it is visible only to a Marketplace reviewer. **There was no `listing.md`.** Rule 11 asks for catalog copy in-repo so the page and the image are versioned together. Ours lived nowhere. It is also where the >=8 GB sizing guidance and the "DigitalOcean does not build or support Trinity" line required by the vendor terms actually get written down. **Submission is now automatable.** `manifest` + `shell-local` post-processors read the snapshot id from the build's own record and PATCH it to the vendor API with the real payload (`imageId` required, plus `reasonForUpdate`, `osVersion`, `softwareIncluded[]`). A `post-processors` CHAIN rather than sibling blocks, because shell-local reads what manifest writes and siblings run in parallel. `mp-submit.sh` is a no-op unless `TRINITY_DO_APP_ID` is set, so a plain `packer build` still just builds, and it names the 400 that means "a previous submission is still in review" — which otherwise reads as a bad token. The runbook gains a **submission gate**: nothing goes to DigitalOcean until another team member has QA'd a droplet from that snapshot. A green `img_check` says the image is acceptable, not that Trinity works on it. **Still open: `build_size`.** It is `s-2vcpu-4gb` (80 GB), and the comment above it claimed the opposite of the value. The build droplet's CPU and RAM never reach the snapshot — Trinity does not run during the build — but the disk size does, and DigitalOcean does not allow shrinking it. So 80 GB excludes every plan with a smaller boot disk, and on the v5 line disk is decoupled from memory: `s5-8vcpu-16gb-30gb` is 16 GiB of RAM on a 30 GiB disk, exactly the plan a Trinity operator should pick, and this snapshot cannot deploy to it. The comment now states the real criterion — the smallest boot disk the baked images fit on — and the value waits on measuring them. Guards are mutation-checked: reverting any one fix turns its test red. The no-op default is executed rather than pattern-matched, per #2522. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB
Measured the snapshot the build actually produces:
Name Min Disk Size Size
trinity-v0-9-5-rc2-20260903 80 9.01 GiB
The content is 9 GiB. The 80 GB floor came entirely from the build droplet's
own disk, and DigitalOcean does not let a snapshot deploy to anything smaller —
"you must select a disk size equal to or larger than the Droplet used to create
the snapshot." So every customer was being pushed onto an 80 GB boot disk, and
paying for it, to hold 9 GiB. On the v5 plans, where the boot disk is sized
independently of RAM, that is the difference between choosing a disk and being
told one.
s-1vcpu-2gb (50 GB) rather than DigitalOcean's recommended $6 s-1vcpu-1gb
(25 GB): the 9.01 GiB is their COMPRESSED stored size, while Docker's overlay2
tree on the build droplet is uncompressed — actual build-time use is nearer
18-20 GB before Ubuntu and the apt caches. Too tight against 25 GB to commit
without a build proving it, and a build droplet that exhausts its disk fails
only after pulling every image. 50 GB is also what three of DigitalOcean's own
catalog apps build on (openclaw, jellyfin, craftcms).
25 GB is probably reachable and would match their guidance exactly; one
experimental build settles it, and it is not worth blocking the listing on.
Marked with a ponytail: note.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB
…asure it (#2281) A droplet booted from the 2026-09-10 snapshot reports: /dev/vda1 77G 8.8G 68G 12% / 8.8 GB, with everything installed and Trinity running. The comment justifying s-1vcpu-2gb assumed Docker's uncompressed overlay2 tree would reach 18-20 GB and called 25 GB too tight. That was a guess dressed as a reason, and it was wrong by more than a factor of two. Value unchanged — 50 GB still works and is what three of DigitalOcean's own catalog apps build on — but the stated reason is now the measurement rather than the guess, and the ponytail note names what is actually left: 25 GB is settled on disk, and the only open question is whether 1 GB of RAM survives `apt full-upgrade` and `docker pull`. Their own exa-24-04 builds on that size. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB
…olliding (#2281) Two builds of v0.9.5-rc2, identical except for build_size, settle what was guesswork: s-1vcpu-2gb (50 GB) 14m38s img_check 8/0 snapshot min disk 50 s-1vcpu-1gb (25 GB) 11m53s img_check 8/0 snapshot min disk 25 Both open questions about the $6 droplet are now measured. Disk: a droplet from the snapshot reports 8.8 GB used with everything installed and Trinity running, against 25 GB available. RAM: a full build on 1 GB completed with no OOM and no disk exhaustion, through `apt full-upgrade` and the agent base image pull, the two places 1 GB would have bitten. It is also faster than the size it replaces, and DigitalOcean's own exa-24-04 builds here. That is AC 2 of #2281 met literally rather than in substance, and it takes the floor a customer must deploy onto from 80 GB down to 25 GB. DigitalOcean does not allow a snapshot to deploy to a smaller disk than it was built on, so this value is the listing's real hardware constraint: on the v5 line, where boot disk is sized independently of RAM, it decides whether an operator picks their disk or is handed one. Also: `snapshot_name` gains minute resolution. Those two builds produced two snapshots named `trinity-v0-9-5-rc2-20260910`, distinguishable only by ID and min disk size, while the Vendor Portal's "select a system image" step picks by what it shows you. Submitting the wrong image costs a review cycle and fails silently, since both are valid Trinity snapshots. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB
|
merge-train 2026-09-10: on this train. No commits pushed to your branch — it was already current with What I changed and why. The body carried a section headed "Still open: The coverage answer, proven rather than asserted.
Your mutation claim does hold at the text level — I reproduced all four, and each reverts exactly one guard to red. The guards discriminate spelling, and only spelling. The one genuine execution is And nothing else runs it either: The two riskiest changes are the two with zero execution anywhere — Smaller findings, none fixed by me:
Outward-facing check: clean. No credential, internal URL, IP, or real email enters any file that ships in the image or the listing. Correctly left with no closing keyword — "Related to #2281. Does not close it" is right, and #2281 already carries One process note that isn't about your PR: the merge-train lane classifier triggers deep review on |
…ers (#2281) Found by booting a droplet from the 25 GB snapshot and looking at the Console: it renders `trinity-25gb-verify login:` and nothing else. `passwd -S root` reports `L` — locked. DigitalOcean's create page requires either a password or an SSH key, and the choice decides which route to the MOTD exists: password auth root has a password, Droplet -> Console works, no SSH needed SSH key auth root is LOCKED, the Console cannot accept a login at all; the customer must ssh root@<ip> `listing.md` and this README both described the Console as *the* way to get the admin password, unconditionally. For every customer who picks key auth — which DigitalOcean's own create page nudges toward — that instruction leads to a prompt they cannot satisfy, for the one credential without which the listing is unusable. Not fixable by setting a root password: `img_check.sh` scores "User root has no password set" as a PASS condition, so an image that sets one fails review. Both routes reach the MOTD, so documenting both is the fix. The MOTD itself is correct and was never the problem — it renders the password, the URL, the HTTPS status and the support line as designed. Also adds the `cat /etc/trinity/admin-credentials` fallback for re-reading it later. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB
#2680 landed two build-standard changes on the files this branch restructured: - 01-provision.sh: the pre-install `apt-get full-upgrade` (DO's img-check fails on pending security updates) is kept, now between `apt-get update` and the checkout's own package install. - firstboot.sh: the body is this branch's `start.sh --provision --site-only` call; the `X-DO-MARKETPLACE` header #2680 added to the Caddyfile moves to `provision_site` in start.sh, where the Caddyfile is written now, keyed on the do-marketplace provenance since the same path serves the doc install. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015owqMKD5QDjzZrUTF2Joht
…ddyfile is written now #2680's guard pinned the header inside firstboot.sh's Caddyfile heredoc; #2380 moved that heredoc into `start.sh --provision --site-only`, so the merge left the test asserting a file that no longer writes a Caddyfile. Point it at start.sh and also pin that the header is keyed on the do-marketplace provenance, since the same function serves the doc install. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015owqMKD5QDjzZrUTF2Joht
…er, and a domain that gets a certificate (#2380) (#2707) * feat(deploy): one installer provisions the machine, three callers use it (#2380) Provisioning a bare cloud VM for Trinity existed in three copies: the Packer bakery, the Packer first-boot script, and a hand-written script pasted into the DigitalOcean deploy doc. They had already drifted — one carried a DOCKER-USER DROP list missing 8081, so the login page answered plain HTTP past the certificate and past the http->https redirect, on the image whose headline design note is that everything reaches users through Caddy on 80/443. `start.sh --provision --cloud <name>` is now the only implementation, in two phases because a snapshot-based image splits them: `--machine-only` (packages, pinned Caddy with the IP-certificate floor asserted, ufw, the firewall unit) is bakeable; `--site-only` (the instance's own IP, the .env keys, the Caddyfile, the certificate) is per-instance. The bakery calls the first, first boot calls the second and continues into the install, and a doc-driven install calls both. firstboot.sh drops from 230 lines to 118 and 01-provision.sh from 176 to 99. It is off by default and refuses unless it is root on Linux with a cloud metadata service answering, so it stays inert on a developer laptop. An ADMIN_PASSWORD in the environment is now written through to .env, so an install script does not have to hand-roll its own .env writer. The firewall rule is inverted rather than enumerated. Everything entering a container from off-box is dropped whatever the port, with two RETURNs ahead of it: replies to connections a container opened (without which agents lose outbound internet), and traffic from Docker's own bridges. Naming what is inside rather than which interface is outside also covers DigitalOcean's private eth1. There is no list left to drift, so the test that policed the list is replaced by one that pins the ordering and the absence of a fourth copy. Also widens the first-run hardening guide to doc-driven DigitalOcean installs, which #2380's author amended their own acceptance criterion to require. The installer records `do-script` — honest, because it refuses to run at all unless DO's metadata service answers — and `hardening_guide_eligible` is a separate gate from `marketplace_install` rather than a widening of it, so the guide reaches these installs without anyone claiming a vendor listing they never came from. Deliberately still provenance and never TLS state: the managed fleet runs plain HTTP behind a tunnel with no domain and a 100.x address, so a posture-based gate would fire on every paying client forever. The https-ip copy gains the renewal caveat, hedged the way that block's own acceptance criterion requires: a ~6-day certificate renewed only while the machine runs is a property of the profile such an install is known to use, so it is stated as what happens after a long shutdown, never as a claim about this instance's live certificate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * feat(deploy): a prompting DigitalOcean installer, not a file to edit (#2380) The deploy doc's flow was eight manual steps in DigitalOcean's web console, one of which is "paste this script into a textarea after editing three lines". Seven of those eight can be a script, and the one that cannot — open the URL and sign in — is the human's job anyway. A console click-path also cannot satisfy this work's own acceptance criterion that every command in the doc has been executed verbatim at least once: nobody can execute "tick Advanced Options", so the most error-prone step is the one that ships permanently untested. It prompts rather than shipping a file to edit. "Download this and change four lines" fails the audience the doc exists for, and it puts two secrets in a file on disk; `read -rs` keeps both off the terminal, out of shell history, and out of every file but the droplet's own user-data. Checks come before questions, so a missing doctl is never discovered after collecting two credentials. The password is asked twice because it is the credential for a server that does not exist yet — a typo is not recoverable by retrying, it is a rebuild — and the paid resource is confirmed before it is created. Prompts still read from a pipe, so the whole flow stays drivable in a test: printf 'pw\npw\ntoken\n\n\ny\n' | bash scripts/deploy/trinity-do-create.sh The droplet-side half is unchanged and stays one copy: clone, then `start.sh --provision --cloud digitalocean --hosted --unattended`, the same installer the Packer image runs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * fix(deploy): containers cannot read the instance metadata service (#2380) A cloud instance serves its own user-data verbatim from link-local, for the life of the machine, to anything that can reach it — and on a script-installed droplet that user-data carries the Trinity admin password and the operator's Claude subscription token. Agent containers could reach it. An agent running untrusted model-authored code is precisely the case that must not be able to read the credential that owns the box. Dropped outbound from containers as the RFC 3927 range rather than the well-known 169.254.169.254 host: the property being blocked is "link-local, host-adjacent, not routable", and every cloud publishes its metadata endpoint somewhere in that range, so the single address would have been a DigitalOcean- shaped fix to a general problem. Position is load-bearing and is what the test pins. The DROP sits after the conntrack RETURN (so only the opening packet of a metadata connection is ever evaluated, which is enough — it never establishes) and BEFORE the two bridge RETURNs. Behind them it would be decoration: container-originated traffic RETURNs out of the chain before reaching it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * feat(deploy): make "add a domain" actually work, not just be advertised (#2380) The hardening card's step one told the operator to point DNS at the server and set the Public URL in Settings, then said "whatever terminates TLS in front of it picks up the name". Nothing does. `public_chat_url` is a display and webhook-base setting, and the card's own source comment says so — no code reads it and reconfigures a proxy, a listener or a certificate. On every install this card is shown to, Caddy is holding a single `https://<IP>` site block written at provision time. So the sequence the card produced was: operator sets the URL, install_tls_posture flips to `https-domain` by string-parsing that setting, the card treats step one as complete and advances to the tunnel step, and the domain serves a certificate error because the web server still answers only to the IP. The card retired on a state it had helped break, and the instance was worse off than before it gave the advice. `scripts/deploy/set-domain.sh` does the job the card was describing. It refuses unless DNS already points at this machine — that check is first, before anything is touched, because issuance validates over the name and a stale A record is what turns a working instance into a broken one. Then it rewrites the Caddyfile for the name (no `shortlived` profile: a real name earns an ordinary ~90-day certificate, which is the upgrade being made), validates, reloads, waits for the certificate against the real hostname with the system trust store, and only then tells Trinity its new address. Every failure after the rewrite restores the previous config and reloads, so a bad run ends where it started. The bare IP is kept as a redirect rather than dropped, with its short-lived certificate, so links handed out before the domain existed keep working — a redirect that cannot complete a handshake is not a redirect. It runs on the server because it has to: Trinity is in a container with no host privileges, so it can neither write /etc/caddy/Caddyfile nor reload Caddy. The card now says that, names the command, and keeps every honest claim the previous copy made — it only moves the actor from "whatever is out there, somehow" to something the operator can actually run. That constraint is the same one that already keeps the tunnel step to prose, so the card is now consistent about it. `set_env_key` moves to scripts/deploy/env-file.sh, sourced by both writers: a .env writer that disagrees with itself corrupts credentials silently, and this branch exists because two copies of provisioning logic drifted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * fix(onboarding): the card's one action is the one that works (#2380) Walking the journey a real operator takes found the previous fix had covered the wrong half. The card said the right thing in its disclosure and still pointed its primary button at the trap. The operator who HAS a domain — exactly who the button was for — clicked "Add a domain", landed on Settings → General → Public URL (a bare input, a Save button, no guard, placeholder `https://your-domain.com`), typed their domain and saved. That sets the name Trinity hands out and reconfigures nothing: Caddy still held only the `https://<IP>` site written at provision time, so the domain served a certificate error while the old IP link kept working and the instance looked fine. `install_tls_posture` then flipped to `https-domain` by string-parsing their own input, the card advanced to the tunnel stage, and the address section carrying the command that would have fixed it — `v-if` on that stage — disappeared. The UI congratulated them on a step it had just helped them break, and removed the way back. So the face carries the command now, not a link to the field. There is no navigation button because there is nowhere useful to navigate: `set-domain.sh` sets the Public URL itself, and setting it by hand is not enough on its own. `select-all` on the block so one click takes the whole line. The disclosure keeps the reasoning, including the sentence saying explicitly that the settings field alone does not do it. The three tests that pinned the old shape are updated rather than deleted, each carrying why the assertion inverted — "links to Settings → General" was a true description of a defect. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * feat(deploy): adding a domain is a settings field, not a shell command (#2380) The last two attempts at this both ended with an operator needing a root shell on the droplet, which a non-engineer following a deploy guide does not have. The constraint that forced it was real — Trinity is containerised and can neither rewrite /etc/caddy/Caddyfile nor reload Caddy — but it was the wrong thing to work around. Inverting the direction removes it entirely: Caddy asks Trinity whether a hostname is allowed, Trinity answers from what an admin saved, and no privilege moves into the container. The provisioned Caddyfile gains a catch-all site with on-demand TLS and an `ask` gate at /api/public/tls-allowed. Save a domain in Settings, point its A record here, and the first visit obtains the certificate. That is the whole step, which is what this card claimed from the beginning. The gate carries the entire security model, so it allows exactly one name: the host of the saved public_chat_url, parsed with urlparse and compared for equality. Substring matching would hand `evil-example.com` a certificate request from a server configured for `example.com`. It fails closed on an unset URL and on a failed settings read — `on_demand` without a working gate turns the instance into a certificate requester for anyone who points DNS at its address, until the ACME account is rate-limited and the operator's own renewals start failing. Executed rather than grepped in the tests: a static check cannot tell an exact match from a substring one, and that difference is the whole of it. set-domain.sh is deleted. It was written for this job two commits ago and is now a second way to do something that works from the browser; keeping it as a "fallback" would have been keeping a path nobody tests for a case that no longer exists. Caddyfile syntax verified against the 2.11.x docs rather than memory: `interval` and `burst` are removed from on_demand_tls, `ask` receives ?domain= and authorises on 2xx. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * fix(deploy): attach the account's SSH keys to the droplet it creates (#2380) Without this the droplet has no key on it, DigitalOcean emails a root password to the account owner, and the only way to a shell is the browser console. That is fine until something needs looking at, which on a first deploy is exactly when it happens. Attaches keys that already exist on the account — the operator's own, on the operator's account, going onto the operator's server. Nothing is created, uploaded or generated, and an account with no keys behaves exactly as before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * fix(deploy): portable mktemp in the DigitalOcean installer (#2380) `mktemp -t trinity-user-data` takes a bare prefix on macOS/BSD, but GNU coreutils reads the same argument as a template and requires the trailing X's — so on Ubuntu (i.e. every Linux operator) the installer died with "mktemp: too few X's in template" immediately after collecting both secrets and confirming the create. Use a full path template instead: portable on both, and it keeps the umask-077 0600 mode the user-data file needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgRv6koYJKzor7ZHoRkGYS * test(deploy): the X-DO-MARKETPLACE guard reads start.sh, where the Caddyfile is written now #2680's guard pinned the header inside firstboot.sh's Caddyfile heredoc; #2380 moved that heredoc into `start.sh --provision --site-only`, so the merge left the test asserting a file that no longer writes a Caddyfile. Point it at start.sh and also pin that the header is keyed on the do-marketplace provenance, since the same function serves the doc install. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015owqMKD5QDjzZrUTF2Joht * test(deploy): restore the two firewall guards the port-list rewrite dropped (#2380) 6104b0e deleted test_2281_firstboot_port_exposure.py because the DOCKER-USER rule no longer enumerates ports, so the tests that policed the list had nothing left to check. Two of its five tests were never about the list, and they went with the file: - ufw and iptables-persistent are never both installed. ufw Breaks iptables-persistent, so installing it to persist the rules makes apt remove ufw, and the install dies later at `ufw --force reset`. - the firewall rules are re-applied on every boot. Without iptables-persistent, trinity-docker-firewall.service is the only thing that survives a reboot; without it every Docker-published port reopens after the first restart (#2281 review I1). Both are back in test_2380_provision_single_source.py, pointed at where the code lives now: start.sh --provision installs ufw and writes the unit as a heredoc, so the tests read that heredoc instead of a shipped unit file. Each was checked against a mutation of start.sh (iptables-persistent added, the enable line removed, After=docker.service removed) and fails on each. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * fix(deploy): snap doctl cannot read /tmp, and an apostrophe kills first boot (#2380) (#2683) * fix(deploy): snap doctl cannot read /tmp, and an apostrophe kills first boot (#2380) Two fixes to the DigitalOcean installer, found QA'ing it across macOS and Linux. Neither is visible on the machine it was written on. **A snap-installed doctl has a private /tmp.** On most Linux distributions doctl comes from snap, and a snap runs in its own mount namespace — a file written to the real /tmp is not there when doctl opens it, so it fails on a file that demonstrably exists: Error: open /tmp/trinity-user-data.XXXXXX: no such file or directory Snap's `home` interface can read non-hidden paths under $HOME, so a snap doctl gets the file there instead; everyone else keeps TMPDIR. The name stays visible because that interface denies dotfiles too, `umask 077` still makes it 0600, and the EXIT trap still removes it. Reported by dolho, who hit it on Linux. The `mktemp -t` half of his finding is already on this branch (aaa3de5). A snap doctl with no usable $HOME now fails at that point with an explanation, rather than at the droplet-create call with a path error, after both prompts. **A single quote in either secret breaks the droplet's first boot.** Both are interpolated into single-quoted assignments in the user-data, so a `'` inside one closes the string early: export ADMIN_PASSWORD='Tr0ub4dor's!Horse' bash: unexpected EOF while looking for matching `'' Not a hypothetical input — the password rules require a special character and `'` is one, so that password passes this script's own check and the backend's. The failure is expensive out of proportion to the typo: the droplet is created and billing before first boot runs, the operator watches a 15-minute progress bar end in a timeout, and the cause is invisible without reading the install log. Both secrets now go through `_shquote` (the portable `'\''` idiom). Verified on both platforms rather than reasoned about — macOS bash 3.2 and ubuntu:24.04 bash 5.2 / GNU coreutils 9.4, driving the real prompt flow with a stubbed doctl and diffing the generated user-data: - `mktemp -t trinity-user-data` → macOS OK, Linux `too few X's in template` - the full template → both OK, mode 0600 - snap-detection truth table → identical on both - password round-trip → 7/7 both platforms, including apostrophes, backticks, backslashes, tabs and non-ASCII; the unfixed script fails the apostrophe cases with a syntax error `tests/unit/test_2380_installer_portability.py` pins all three. The round-trip guard is executed, not pattern-matched — it runs the script's own `_shquote` and lets a shell parse the result back, so it catches any breakage rather than the one spelling that was wrong. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB * test(deploy): the installer's default tag must track VERSION (#2380) `trinity-do-create.sh` hardcodes the release it installs: TRINITY_IMAGE_TAG="${TRINITY_IMAGE_TAG:-v0.9.5-rc2}" That default IS what operators get — the guide tells them to run the script straight off a tag with no environment set. So on the day `v0.9.5` is cut, a script fetched from the v0.9.5 tag still clones and pulls `v0.9.5-rc2` unless someone remembers this one line. Nothing catches it today. It is not a syntax error, the rc2 images still exist and still pull, and the install SUCCEEDS — it just installs the previous release candidate. The operator cannot tell, and neither can CI: the release checklist's "VERSION match" step compares the VERSION file to the tag, not this script to either. The guard ties them together. It passes now (VERSION `0.9.5-rc2`, default `v0.9.5-rc2`) and fails the moment VERSION is bumped for the cut — which is exactly when someone needs telling. Verified by mutation: setting VERSION to `0.9.5` produces AssertionError: trinity-do-create.sh installs v0.9.5-rc2 but VERSION says 0.9.5. A second case pins the `v` prefix. One string feeds both `git clone --branch` (needs the git ref) and `docker pull` (published under both spellings), so only the `v` form works for both — the #2471 failure, where `v0.9.5-rc1` shipped without its v-prefixed image alias and left no tag value able to build the marketplace snapshot. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB * fix(deploy): refuse before the prompts, quote the tag, and test the installer by running it (#2380) Addresses review I1-I3 on #2683. I1: the snap-doctl-without-$HOME refusal ran after both secret prompts, breaking the script's own "checks first, questions second" rule. The USER_DATA_DIR decision moves into the pre-flight block beside `doctl account get`; mktemp stays where it was. I2: TRINITY_IMAGE_TAG was the third single-quoted interpolation in the user-data heredoc and the only unquoted one. It now goes through _shquote, and the static guard asserts the rule (every '${...}' in the heredoc is a _Q value) instead of two spellings of it. I3: the guards grepped for the quoted spelling and never ran the script. New tests run the real installer against a stub doctl (package and snap paths, apostrophes and shell metacharacters in the password, token and tag), then `bash -n` the captured user-data and evaluate its real export lines and subscription call in context, and assert a snap doctl without $HOME is refused before any question. Running it found a fourth defect: under `set -u`, bash < 4.4 (macOS's /bin/bash is 3.2) treats an empty "${SSH_ARGS[@]}" as unbound, so an account with no SSH keys died at "Creating the droplet...", after both secrets were typed. Expanded as ${SSH_ARGS[@]+"${SSH_ARGS[@]}"}. Each fix was checked by reverting it: the tests fail on every revert. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(deploy): describe the provisioning path, installer, firewall and TLS gate (#2380) /validate-pr on #2707 found the branch shipping new capability without the requirements and feature-flow updates CLAUDE.md requires. Written from the code at this commit, not from the PR body. Requirements (docs/memory/requirements/infrastructure.md): - Corrected: HOST-010 and the section 8.10 description, PROV-006 (flags now carry hardening_guide_eligible), PROV-009 (the second path has been a Cloudflare Tunnel since #2564, not a VPN), PROV-010 (a two-stage card gated on hardening_guide_eligible, with a domain field), PROV-011 (start.sh --provision now writes TRINITY_INSTALL_SOURCE). - Added: PROV-012 one provisioning implementation, three callers; PROV-013 the portless container firewall and the metadata-service block; PROV-014 the prompting installer and do-script provenance; PROV-015 the on-demand TLS ask gate. Feature flows: install-provenance (two-stage card, do-script, hardening_guide_eligible, the domain field and ask gate; fixes both stale claims in #2593), hosted-install (a Provision Layer section), telemetry-sharing (the IntersectionObserver gate from #2593), and their index rows. Stale sentences in agent-lifecycle.md, backend.md, api-endpoints.md and DEPLOYMENT.md corrected. Fixes #2593 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * chore(deploy): drop references to set-domain.sh, which no longer exists (#2380) 0423e6c replaced scripts/deploy/set-domain.sh with a Settings field, but two comments still named it as env-file.sh's second caller, and the provisioning test kept an unused path constant pointing at it. start.sh is now the only script that sources env-file.sh; the comments say so. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * test(deploy): pin the on-demand site block and reach the ask gate through FastAPI (#2380) Review follow-ups on #2707, all mechanical. The Caddyfile guard asserted `"on_demand" in body`, which the global `on_demand_tls` option already satisfies: deleting the catch-all `https:// { tls { on_demand } }` site — the entire "add a domain" feature — passed every test. It now extracts the heredoc and asserts that site as a block, and that the bare-IP site still asks for a short-lived certificate. Checked by deleting the block: the test fails. Nothing reached GET /api/public/tls-allowed through FastAPI. The decision (exact hostname, fail closed) is proved by exec-slicing the function, which cannot see the route's path, its query parameter, or whether it gained an auth dependency — and Caddy holds no credential, so a 401 there would refuse every certificate with the failure visible only in a live handshake. A new test mounts the real router and calls the real URL. Settings' install-source label map had no entry for `do-script`, so a script-installed droplet showed the raw value. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Oleksii Dolhov <oleksii.dolhov@gmail.com>
…ials without a terminal (trinity-enterprise#580, #581, #582) (#2715) * feat(deploy): one installer provisions the machine, three callers use it (#2380) Provisioning a bare cloud VM for Trinity existed in three copies: the Packer bakery, the Packer first-boot script, and a hand-written script pasted into the DigitalOcean deploy doc. They had already drifted — one carried a DOCKER-USER DROP list missing 8081, so the login page answered plain HTTP past the certificate and past the http->https redirect, on the image whose headline design note is that everything reaches users through Caddy on 80/443. `start.sh --provision --cloud <name>` is now the only implementation, in two phases because a snapshot-based image splits them: `--machine-only` (packages, pinned Caddy with the IP-certificate floor asserted, ufw, the firewall unit) is bakeable; `--site-only` (the instance's own IP, the .env keys, the Caddyfile, the certificate) is per-instance. The bakery calls the first, first boot calls the second and continues into the install, and a doc-driven install calls both. firstboot.sh drops from 230 lines to 118 and 01-provision.sh from 176 to 99. It is off by default and refuses unless it is root on Linux with a cloud metadata service answering, so it stays inert on a developer laptop. An ADMIN_PASSWORD in the environment is now written through to .env, so an install script does not have to hand-roll its own .env writer. The firewall rule is inverted rather than enumerated. Everything entering a container from off-box is dropped whatever the port, with two RETURNs ahead of it: replies to connections a container opened (without which agents lose outbound internet), and traffic from Docker's own bridges. Naming what is inside rather than which interface is outside also covers DigitalOcean's private eth1. There is no list left to drift, so the test that policed the list is replaced by one that pins the ordering and the absence of a fourth copy. Also widens the first-run hardening guide to doc-driven DigitalOcean installs, which #2380's author amended their own acceptance criterion to require. The installer records `do-script` — honest, because it refuses to run at all unless DO's metadata service answers — and `hardening_guide_eligible` is a separate gate from `marketplace_install` rather than a widening of it, so the guide reaches these installs without anyone claiming a vendor listing they never came from. Deliberately still provenance and never TLS state: the managed fleet runs plain HTTP behind a tunnel with no domain and a 100.x address, so a posture-based gate would fire on every paying client forever. The https-ip copy gains the renewal caveat, hedged the way that block's own acceptance criterion requires: a ~6-day certificate renewed only while the machine runs is a property of the profile such an install is known to use, so it is stated as what happens after a long shutdown, never as a claim about this instance's live certificate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * feat(deploy): a prompting DigitalOcean installer, not a file to edit (#2380) The deploy doc's flow was eight manual steps in DigitalOcean's web console, one of which is "paste this script into a textarea after editing three lines". Seven of those eight can be a script, and the one that cannot — open the URL and sign in — is the human's job anyway. A console click-path also cannot satisfy this work's own acceptance criterion that every command in the doc has been executed verbatim at least once: nobody can execute "tick Advanced Options", so the most error-prone step is the one that ships permanently untested. It prompts rather than shipping a file to edit. "Download this and change four lines" fails the audience the doc exists for, and it puts two secrets in a file on disk; `read -rs` keeps both off the terminal, out of shell history, and out of every file but the droplet's own user-data. Checks come before questions, so a missing doctl is never discovered after collecting two credentials. The password is asked twice because it is the credential for a server that does not exist yet — a typo is not recoverable by retrying, it is a rebuild — and the paid resource is confirmed before it is created. Prompts still read from a pipe, so the whole flow stays drivable in a test: printf 'pw\npw\ntoken\n\n\ny\n' | bash scripts/deploy/trinity-do-create.sh The droplet-side half is unchanged and stays one copy: clone, then `start.sh --provision --cloud digitalocean --hosted --unattended`, the same installer the Packer image runs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * fix(deploy): containers cannot read the instance metadata service (#2380) A cloud instance serves its own user-data verbatim from link-local, for the life of the machine, to anything that can reach it — and on a script-installed droplet that user-data carries the Trinity admin password and the operator's Claude subscription token. Agent containers could reach it. An agent running untrusted model-authored code is precisely the case that must not be able to read the credential that owns the box. Dropped outbound from containers as the RFC 3927 range rather than the well-known 169.254.169.254 host: the property being blocked is "link-local, host-adjacent, not routable", and every cloud publishes its metadata endpoint somewhere in that range, so the single address would have been a DigitalOcean- shaped fix to a general problem. Position is load-bearing and is what the test pins. The DROP sits after the conntrack RETURN (so only the opening packet of a metadata connection is ever evaluated, which is enough — it never establishes) and BEFORE the two bridge RETURNs. Behind them it would be decoration: container-originated traffic RETURNs out of the chain before reaching it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * feat(deploy): make "add a domain" actually work, not just be advertised (#2380) The hardening card's step one told the operator to point DNS at the server and set the Public URL in Settings, then said "whatever terminates TLS in front of it picks up the name". Nothing does. `public_chat_url` is a display and webhook-base setting, and the card's own source comment says so — no code reads it and reconfigures a proxy, a listener or a certificate. On every install this card is shown to, Caddy is holding a single `https://<IP>` site block written at provision time. So the sequence the card produced was: operator sets the URL, install_tls_posture flips to `https-domain` by string-parsing that setting, the card treats step one as complete and advances to the tunnel step, and the domain serves a certificate error because the web server still answers only to the IP. The card retired on a state it had helped break, and the instance was worse off than before it gave the advice. `scripts/deploy/set-domain.sh` does the job the card was describing. It refuses unless DNS already points at this machine — that check is first, before anything is touched, because issuance validates over the name and a stale A record is what turns a working instance into a broken one. Then it rewrites the Caddyfile for the name (no `shortlived` profile: a real name earns an ordinary ~90-day certificate, which is the upgrade being made), validates, reloads, waits for the certificate against the real hostname with the system trust store, and only then tells Trinity its new address. Every failure after the rewrite restores the previous config and reloads, so a bad run ends where it started. The bare IP is kept as a redirect rather than dropped, with its short-lived certificate, so links handed out before the domain existed keep working — a redirect that cannot complete a handshake is not a redirect. It runs on the server because it has to: Trinity is in a container with no host privileges, so it can neither write /etc/caddy/Caddyfile nor reload Caddy. The card now says that, names the command, and keeps every honest claim the previous copy made — it only moves the actor from "whatever is out there, somehow" to something the operator can actually run. That constraint is the same one that already keeps the tunnel step to prose, so the card is now consistent about it. `set_env_key` moves to scripts/deploy/env-file.sh, sourced by both writers: a .env writer that disagrees with itself corrupts credentials silently, and this branch exists because two copies of provisioning logic drifted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * fix(onboarding): the card's one action is the one that works (#2380) Walking the journey a real operator takes found the previous fix had covered the wrong half. The card said the right thing in its disclosure and still pointed its primary button at the trap. The operator who HAS a domain — exactly who the button was for — clicked "Add a domain", landed on Settings → General → Public URL (a bare input, a Save button, no guard, placeholder `https://your-domain.com`), typed their domain and saved. That sets the name Trinity hands out and reconfigures nothing: Caddy still held only the `https://<IP>` site written at provision time, so the domain served a certificate error while the old IP link kept working and the instance looked fine. `install_tls_posture` then flipped to `https-domain` by string-parsing their own input, the card advanced to the tunnel stage, and the address section carrying the command that would have fixed it — `v-if` on that stage — disappeared. The UI congratulated them on a step it had just helped them break, and removed the way back. So the face carries the command now, not a link to the field. There is no navigation button because there is nowhere useful to navigate: `set-domain.sh` sets the Public URL itself, and setting it by hand is not enough on its own. `select-all` on the block so one click takes the whole line. The disclosure keeps the reasoning, including the sentence saying explicitly that the settings field alone does not do it. The three tests that pinned the old shape are updated rather than deleted, each carrying why the assertion inverted — "links to Settings → General" was a true description of a defect. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * feat(deploy): adding a domain is a settings field, not a shell command (#2380) The last two attempts at this both ended with an operator needing a root shell on the droplet, which a non-engineer following a deploy guide does not have. The constraint that forced it was real — Trinity is containerised and can neither rewrite /etc/caddy/Caddyfile nor reload Caddy — but it was the wrong thing to work around. Inverting the direction removes it entirely: Caddy asks Trinity whether a hostname is allowed, Trinity answers from what an admin saved, and no privilege moves into the container. The provisioned Caddyfile gains a catch-all site with on-demand TLS and an `ask` gate at /api/public/tls-allowed. Save a domain in Settings, point its A record here, and the first visit obtains the certificate. That is the whole step, which is what this card claimed from the beginning. The gate carries the entire security model, so it allows exactly one name: the host of the saved public_chat_url, parsed with urlparse and compared for equality. Substring matching would hand `evil-example.com` a certificate request from a server configured for `example.com`. It fails closed on an unset URL and on a failed settings read — `on_demand` without a working gate turns the instance into a certificate requester for anyone who points DNS at its address, until the ACME account is rate-limited and the operator's own renewals start failing. Executed rather than grepped in the tests: a static check cannot tell an exact match from a substring one, and that difference is the whole of it. set-domain.sh is deleted. It was written for this job two commits ago and is now a second way to do something that works from the browser; keeping it as a "fallback" would have been keeping a path nobody tests for a case that no longer exists. Caddyfile syntax verified against the 2.11.x docs rather than memory: `interval` and `burst` are removed from on_demand_tls, `ask` receives ?domain= and authorises on 2xx. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * fix(deploy): attach the account's SSH keys to the droplet it creates (#2380) Without this the droplet has no key on it, DigitalOcean emails a root password to the account owner, and the only way to a shell is the browser console. That is fine until something needs looking at, which on a first deploy is exactly when it happens. Attaches keys that already exist on the account — the operator's own, on the operator's account, going onto the operator's server. Nothing is created, uploaded or generated, and an account with no keys behaves exactly as before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QvfFoqB2v2zPyXJXeyBGbP * fix(deploy): portable mktemp in the DigitalOcean installer (#2380) `mktemp -t trinity-user-data` takes a bare prefix on macOS/BSD, but GNU coreutils reads the same argument as a template and requires the trailing X's — so on Ubuntu (i.e. every Linux operator) the installer died with "mktemp: too few X's in template" immediately after collecting both secrets and confirming the create. Use a full path template instead: portable on both, and it keeps the umask-077 0600 mode the user-data file needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgRv6koYJKzor7ZHoRkGYS * test(deploy): the X-DO-MARKETPLACE guard reads start.sh, where the Caddyfile is written now #2680's guard pinned the header inside firstboot.sh's Caddyfile heredoc; #2380 moved that heredoc into `start.sh --provision --site-only`, so the merge left the test asserting a file that no longer writes a Caddyfile. Point it at start.sh and also pin that the header is keyed on the do-marketplace provenance, since the same function serves the doc install. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015owqMKD5QDjzZrUTF2Joht * test(deploy): restore the two firewall guards the port-list rewrite dropped (#2380) 6104b0e deleted test_2281_firstboot_port_exposure.py because the DOCKER-USER rule no longer enumerates ports, so the tests that policed the list had nothing left to check. Two of its five tests were never about the list, and they went with the file: - ufw and iptables-persistent are never both installed. ufw Breaks iptables-persistent, so installing it to persist the rules makes apt remove ufw, and the install dies later at `ufw --force reset`. - the firewall rules are re-applied on every boot. Without iptables-persistent, trinity-docker-firewall.service is the only thing that survives a reboot; without it every Docker-published port reopens after the first restart (#2281 review I1). Both are back in test_2380_provision_single_source.py, pointed at where the code lives now: start.sh --provision installs ufw and writes the unit as a heredoc, so the tests read that heredoc instead of a shipped unit file. Each was checked against a mutation of start.sh (iptables-persistent added, the enable line removed, After=docker.service removed) and fails on each. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * feat(onboarding): browser admin claim, one first-run overlay, credentials without a terminal Abilityai/trinity-enterprise#580 — marketplace admin claim A DigitalOcean 1-Click first boot no longer generates an admin password. With no user-data password, ADMIN_PASSWORD stays blank and ADMIN_PASSWORD_SOURCE= browser tells start.sh that is deliberate, so the first browser visitor creates the admin at /setup (email, password, product-updates consent). The MOTD prints the URL to claim, never a password. An operator-supplied password (cloud-init user-data, including trinity-do-create.sh) still pre-provisions the admin and keeps the wizard closed. The prod/hosted compose files move from `:?` to `?`: an unset password still refuses to render; an explicitly blank one is the claim path. The backend already handled a blank password; only docstrings change. The accepted risk (the window before the first visit) is written down in DEPLOYMENT.md -> Security Recommendations. Abilityai/trinity-enterprise#581 — one first-run overlay Replaces the dashboard card ladder (HardeningGuide, FinishSetupCard, FrontDeskPanel, OnboardingWizard) with one blocking, teleported overlay: a rail over a conditional step registry (firstRunSteps.js, pure and unit-tested), steps secure / email / claude / keys / agent / sharing, welcome and done panels, the schematic illustrations, and TrinityMark.vue shared with /setup. Completion is derived from existing state, and only skips and "Finish later" persist. Unlike the spec, `keys` and `agent` never open the overlay on their own, so an established fleet does not see it on every new browser. ActivationChecklist stays inline. /setup logs the operator in after the claim. Re-runnable from Settings -> General and ?onboarding=1. Abilityai/trinity-enterprise#582 — credentials inside first-run The Claude step (the only required step) takes a subscription token or an API key, validates it with Anthropic before saving, and names the fix on bad input. The first credential is connected to the agents that had none, so the starter fleet can run immediately. Optional GitHub, Resend and Gemini keys persist through the encrypted secret-settings path (ent#435) with admin-only endpoints, and runtime reads resolve the saved setting before the env. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * fix(onboarding): close the claim paths review found outside the marketplace Pre-landing review and a diff-scoped security audit of the previous commit. Claim path (Abilityai/trinity-enterprise#580): - docker-compose.prod.yml is back to `${ADMIN_PASSWORD:?}`. Neither start.sh nor the marketplace uses it, so relaxing it only let a source build that forgot its password boot with an open /setup. - docker-compose.hosted.yml passes ADMIN_PASSWORD_SOURCE (default `unset`) to the backend, and /api/setup/admin-password refuses a blank-password claim unless the source is `browser`. The #2381 existing-admin check still runs first; the dev compose (variable absent) keeps its wizard. - start.sh decides a blank password with env_value, so `ADMIN_PASSWORD=""`, `''` or a trailing space no longer count as set, and writes a generated password with set_env_key. First credential (Abilityai/trinity-enterprise#582): - Agents come from the DB rather than a Docker listing that swallows errors. Agents with a successful execution are skipped, so fleets that authenticate per agent keep their own key. Agents with a running execution are not restarted. The responses carry an int `connected_agents`, and seeding re-runs the idempotent connect so agents created during the save are not missed. - Gemini-runtime agents resolve the saved Gemini key before the env. - subscriptions router gains its `# mcp:` header. Overlay (Abilityai/trinity-enterprise#581): - Auto-login after /setup uses the claimed email, not a hardcoded 'admin'. - The Claude step reports the real connected count and says a re-paste replaces the saved credential. - setup_started is recorded only when the overlay opens on its own and the claude or agent step applies. - Exports orphaned by the absorbed components are deleted; docs that still described them now point at FirstRunOverlay / firstRunSteps.js. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * chore(release): 0.9.5-rc3 Release candidate cut from feature/ent580-582-first-run so the marketplace browser-claim, the first-run overlay and the in-browser credential steps (Abilityai/trinity-enterprise#580, #581, #582) can be tested on a real 1-Click snapshot before their PR opens. rc1 and rc2 were cut from feature/2380-do-provision the same way; rc2 stays immutable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * fix(deploy): snap doctl cannot read /tmp, and an apostrophe kills first boot (#2380) (#2683) * fix(deploy): snap doctl cannot read /tmp, and an apostrophe kills first boot (#2380) Two fixes to the DigitalOcean installer, found QA'ing it across macOS and Linux. Neither is visible on the machine it was written on. **A snap-installed doctl has a private /tmp.** On most Linux distributions doctl comes from snap, and a snap runs in its own mount namespace — a file written to the real /tmp is not there when doctl opens it, so it fails on a file that demonstrably exists: Error: open /tmp/trinity-user-data.XXXXXX: no such file or directory Snap's `home` interface can read non-hidden paths under $HOME, so a snap doctl gets the file there instead; everyone else keeps TMPDIR. The name stays visible because that interface denies dotfiles too, `umask 077` still makes it 0600, and the EXIT trap still removes it. Reported by dolho, who hit it on Linux. The `mktemp -t` half of his finding is already on this branch (aaa3de5). A snap doctl with no usable $HOME now fails at that point with an explanation, rather than at the droplet-create call with a path error, after both prompts. **A single quote in either secret breaks the droplet's first boot.** Both are interpolated into single-quoted assignments in the user-data, so a `'` inside one closes the string early: export ADMIN_PASSWORD='Tr0ub4dor's!Horse' bash: unexpected EOF while looking for matching `'' Not a hypothetical input — the password rules require a special character and `'` is one, so that password passes this script's own check and the backend's. The failure is expensive out of proportion to the typo: the droplet is created and billing before first boot runs, the operator watches a 15-minute progress bar end in a timeout, and the cause is invisible without reading the install log. Both secrets now go through `_shquote` (the portable `'\''` idiom). Verified on both platforms rather than reasoned about — macOS bash 3.2 and ubuntu:24.04 bash 5.2 / GNU coreutils 9.4, driving the real prompt flow with a stubbed doctl and diffing the generated user-data: - `mktemp -t trinity-user-data` → macOS OK, Linux `too few X's in template` - the full template → both OK, mode 0600 - snap-detection truth table → identical on both - password round-trip → 7/7 both platforms, including apostrophes, backticks, backslashes, tabs and non-ASCII; the unfixed script fails the apostrophe cases with a syntax error `tests/unit/test_2380_installer_portability.py` pins all three. The round-trip guard is executed, not pattern-matched — it runs the script's own `_shquote` and lets a shell parse the result back, so it catches any breakage rather than the one spelling that was wrong. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB * test(deploy): the installer's default tag must track VERSION (#2380) `trinity-do-create.sh` hardcodes the release it installs: TRINITY_IMAGE_TAG="${TRINITY_IMAGE_TAG:-v0.9.5-rc2}" That default IS what operators get — the guide tells them to run the script straight off a tag with no environment set. So on the day `v0.9.5` is cut, a script fetched from the v0.9.5 tag still clones and pulls `v0.9.5-rc2` unless someone remembers this one line. Nothing catches it today. It is not a syntax error, the rc2 images still exist and still pull, and the install SUCCEEDS — it just installs the previous release candidate. The operator cannot tell, and neither can CI: the release checklist's "VERSION match" step compares the VERSION file to the tag, not this script to either. The guard ties them together. It passes now (VERSION `0.9.5-rc2`, default `v0.9.5-rc2`) and fails the moment VERSION is bumped for the cut — which is exactly when someone needs telling. Verified by mutation: setting VERSION to `0.9.5` produces AssertionError: trinity-do-create.sh installs v0.9.5-rc2 but VERSION says 0.9.5. A second case pins the `v` prefix. One string feeds both `git clone --branch` (needs the git ref) and `docker pull` (published under both spellings), so only the `v` form works for both — the #2471 failure, where `v0.9.5-rc1` shipped without its v-prefixed image alias and left no tag value able to build the marketplace snapshot. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB * fix(deploy): refuse before the prompts, quote the tag, and test the installer by running it (#2380) Addresses review I1-I3 on #2683. I1: the snap-doctl-without-$HOME refusal ran after both secret prompts, breaking the script's own "checks first, questions second" rule. The USER_DATA_DIR decision moves into the pre-flight block beside `doctl account get`; mktemp stays where it was. I2: TRINITY_IMAGE_TAG was the third single-quoted interpolation in the user-data heredoc and the only unquoted one. It now goes through _shquote, and the static guard asserts the rule (every '${...}' in the heredoc is a _Q value) instead of two spellings of it. I3: the guards grepped for the quoted spelling and never ran the script. New tests run the real installer against a stub doctl (package and snap paths, apostrophes and shell metacharacters in the password, token and tag), then `bash -n` the captured user-data and evaluate its real export lines and subscription call in context, and assert a snap doctl without $HOME is refused before any question. Running it found a fourth defect: under `set -u`, bash < 4.4 (macOS's /bin/bash is 3.2) treats an empty "${SSH_ARGS[@]}" as unbound, so an account with no SSH keys died at "Creating the droplet...", after both secrets were typed. Expanded as ${SSH_ARGS[@]+"${SSH_ARGS[@]}"}. Each fix was checked by reverting it: the tests fail on every revert. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(deploy): describe the provisioning path, installer, firewall and TLS gate (#2380) /validate-pr on #2707 found the branch shipping new capability without the requirements and feature-flow updates CLAUDE.md requires. Written from the code at this commit, not from the PR body. Requirements (docs/memory/requirements/infrastructure.md): - Corrected: HOST-010 and the section 8.10 description, PROV-006 (flags now carry hardening_guide_eligible), PROV-009 (the second path has been a Cloudflare Tunnel since #2564, not a VPN), PROV-010 (a two-stage card gated on hardening_guide_eligible, with a domain field), PROV-011 (start.sh --provision now writes TRINITY_INSTALL_SOURCE). - Added: PROV-012 one provisioning implementation, three callers; PROV-013 the portless container firewall and the metadata-service block; PROV-014 the prompting installer and do-script provenance; PROV-015 the on-demand TLS ask gate. Feature flows: install-provenance (two-stage card, do-script, hardening_guide_eligible, the domain field and ask gate; fixes both stale claims in #2593), hosted-install (a Provision Layer section), telemetry-sharing (the IntersectionObserver gate from #2593), and their index rows. Stale sentences in agent-lifecycle.md, backend.md, api-endpoints.md and DEPLOYMENT.md corrected. Fixes #2593 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * chore(deploy): drop references to set-domain.sh, which no longer exists (#2380) 0423e6c replaced scripts/deploy/set-domain.sh with a Settings field, but two comments still named it as env-file.sh's second caller, and the provisioning test kept an unused path constant pointing at it. start.sh is now the only script that sources env-file.sh; the comments say so. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * fix(subscriptions): run the narrowed first-credential connect before the #2572 sweep The merge put dev's #2572 credential-less sweep ahead of ent#582's connect_agents_to_first_credential. The sweep has no notion of an agent that already authenticates some other way, so running first it would assign and restart agents ent#582 deliberately skips (any agent with a successful execution, and any agent mid-execution), and leave the first-run step reporting `connected_agents: 0` because nothing was left to connect. Swapped: the narrow pass runs first and reports its count, and the sweep then picks up whatever is genuinely credential-less. Also rewrites the first-run section of the setup guide, which #2746 synced while describing the card stack ent#581 retires ("at most one first-run card", "Not now snoozes for two weeks"). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PdRMazCDYVoxqLFw6uKm3a * fix(ci): linear from-address parser (CodeQL ReDoS) and stub get_gemini_api_key in telegram backfill test - platform_keys_service.from_address_domain: replace the single regex that CodeQL flagged as py/polynomial-redos with partition + single-class fullmatch; same accept/reject set, pinned by a parametrized test and a hostile-input timing check. - test_telegram_webhook_backfill: the settings_service stub lacked get_gemini_api_key, which routers/settings.py now imports (ent#582). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011y53gTcB3fTAHhJc3xxpkx --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Oleksii Dolhov <oleksii.dolhov@gmail.com>
What
DigitalOcean enabled our Vendor Portal account and pointed at
digitalocean/marketplace-partnersas the process of record. Auditing the bundleagainst that repo — and against
droplet-1-clicks@master, where their own buildstandard and catalog apps live — found four gaps. Three are fixed here; the
fourth needs a measurement only the DO account can give.
The bundle already matched DO on everything with a stated requirement: vendored
cleanup.sh+img_check.sh(pinned atb708788, which is master HEAD),per-instance boot hook,
99-MOTD, password generated at boot, Caddy withprofile shortlived. These are the gaps.1. The build installed no security updates
Their
99-img-check.sh(lines 438–454) scores a pending security update as ahard
[FAIL], not a warning:Any FAIL exits non-zero, which fails the build — at the last provisioner,
furthest from the cause, reading as a flake.
01-provision.shranapt-get updateand targeted installs, never an upgrade.So the 2026-09-03 build passed on luck: either DO's base image was current that
day, or Ubuntu's boot-time
unattended-upgradestimers landed before Packer'sSSH session. Neither is ours to own.
Adds the
full-upgradefrom their referencemarketplace-image.json, with their--force-confdef/--force-confoldflags, placed before the first install sothe packages we add are patched too — and before the Caddy repo is added, so the
caddy=2.11.4pin is untouched.2.
X-DO-MARKETPLACEwas missingRule 10 of
droplet-1-clicks/.github/agents/1-click-agent.mdspecifies thisheader on the reverse-proxy block, and their live catalog ships it — openclaw's
Caddyfile sets it to
"openclaw". Our first-boot Caddyfile had thetls/issuer acme/profile shortlivedhalf and not this line.It breaks nothing at runtime, which is exactly why it survived review: it is
visible only to a Marketplace reviewer.
3. There was no
listing.mdRule 11 asks for the catalog copy in-repo, so the page and the image are
versioned together. Ours lived nowhere —
README.mdhere is a builder documentfor us, not customer-facing copy.
listing.mdis also where two acceptance criteria on #2281 finally get writtendown: the ≥8 GB sizing recommendation, and the vendor-terms-required statement
that DigitalOcean does not build or support Trinity. Content is checked against
what the droplet actually does — the MOTD's Console-based password path, the
#cloud-config write_filesoverride, the same-disk caveat on our automaticbackups.
4. Submission is now automatable, and the runbook gained the state that blocks it
Their README documents the update path we had only half-recorded. Adds
scripts/mp-submit.shplusmanifest+shell-localpost-processors: thesnapshot id is read from the build's own record and PATCHed to
/api/v1/vendor-portal/apps/<app_id>with the real payload —imageId(required),
reasonForUpdate,osVersion,softwareIncluded[].A
post-processorschain block, not two siblingpost-processorblocks:siblings run in parallel and the submit would race the manifest file it reads.
That is a race which passes on a fast disk and fails on a release day.
mp-submit.shis a no-op unlessTRINITY_DO_APP_IDis set, so a plainpacker buildstill just builds. It names the two API states worthrecognising — notably 400, meaning the app is still
pending/in reviewandcannot be updated, which otherwise reads as a bad token on release day.
The runbook also gains a submission gate: nothing goes to DigitalOcean until
another team member has QA'd a droplet from that snapshot. A green
img_checksays the image is acceptable, not that Trinity works on it.
build_size: 80 GB boot disk → 50 GBFixed here after all, in
e7687f3c8and9cd6cdfd4— this section previouslysaid it was deliberately left alone, which is no longer true.
build_sizewass-2vcpu-4gb(80 GB disk), under a comment reading "DO's ownguidance builds on the smallest size" — the comment and the value contradicted
each other.
The build droplet's CPU and RAM never reach the snapshot: a snapshot is a disk
image, and Trinity never runs during the build (
01-provision.shinstalls andpulls;
start.shruns at first boot on the customer's droplet). The one propertythat propagates is the disk size, and DO does not allow shrinking it — "You
can increase the boot disk size when resizing, but you cannot decrease it."
So 80 GB excluded every plan with a smaller boot disk — not just the cheap tiers:
on DO's v5 line disk is decoupled from memory, and
s5-8vcpu-16gb-30gbis 8 vCPUand 16 GiB RAM on a 30 GiB disk, exactly the plan a Trinity operator should
pick, and that snapshot could not deploy to it. Measured against the snapshot the
build actually produces:
trinity-v0-9-5-rc2-20260903, min disk 80, size 9.01GiB. The 80 GB floor came entirely from the build droplet's own disk, so every
customer was pushed onto an 80 GB boot disk, and paying for it, to hold 9 GiB.
Now
s-1vcpu-2gb(50 GB). A droplet booted from the 2026-09-10 snapshot reports/dev/vda1 77G 8.8G 68G 12% /— 8.8 GB with everything installed and Trinityrunning. An earlier justification for 50 GB assumed Docker's uncompressed
overlay2 tree would reach 18-20 GB and called 25 GB too tight; that was a guess
dressed as a reason and it was wrong by more than a factor of two. The value is
unchanged by the correction — 50 GB is also what three of DigitalOcean's own
catalog apps build on (openclaw, jellyfin, craftcms) — but the stated reason is
now the measurement.
Still open, and named in the
ponytail:marker: 25 GB is settled on disk;the remaining question is whether 1 GB of RAM survives
apt full-upgradeanddocker pull, which is what standing on DO's recommendeds-1vcpu-1gbwouldrequire. Their own
exa-24-04builds on that size. One experimental buildsettles it.
Tests
tests/unit/test_2281_marketplace_build_standard.py, four guards, eachmutation-checked — reverting exactly one fix turns exactly one guard red,
verified against the committed baseline:
full-upgradepresent and positioned before the first install (an upgradeafterwards leaves our own packages unpatched and still fails
img_check)X-DO-MARKETPLACEpresent inside the Caddyfile heredoc, not merelysomewhere in the script
post-processorschain withmanifestordered first
static rule only knows the spelling it was written against): the script is run
with a clean environment in an empty directory and must exit 0
All 24 tests in the
2281suite pass.Not verified
packer buildhas not been re-run against these changes; no DO token in thissession. The
full-upgradelengthens the build.build_sizedecision will need.Related to #2281. Does not close it — the Vendor Portal submission is a
human-gated step, now explicitly behind the QA gate above.
🤖 Generated with Claude Code
https://claude.ai/code/session_011sj48trT3XmyaSus4GUtXB