Agent workflow: branch off main, rehearse on PRE, never act on PROD, tear PRE down #52

Closed
opened 2026-10-07 14:34:28 +00:00 by pit · 0 comments
Owner

Problem Statement

The Operator works this repo with Agents. AGENTS.md and docs/agents/ say how an Agent reads the issue tracker and the domain docs, but nothing says how an Agent works: which branch to start from, where to work, what it may run, and what must not outlive the work. Four rules live in the Operator's head and in the git history:

  1. always work on a fresh branch off main;
  2. always rehearse on PRE — the same code PROD runs;
  3. never run a mutating step against PROD;
  4. always destroy PRE when the testing is done.

The rule that matters most — rehearse on PRE — is not reachable from the place an Agent works. A fresh worktree carries no gitignored secrets (tofu/.env, tofu/npm/.env, ansible/.env, ansible/vault-pass) and no OpenTofu state, so make pre refuses at its first gate; and PRE's identity is fixed (vmid 142, 10.12.0.142), so two Agents rehearsing at once would collide on the Node and on the shared base image. The convention cannot be followed even by an Agent that wants to follow it.

Solution

Write the workflow down where Agents already look, and make the rules reachable.

  • A new repo-only doc, docs/agents/workflow.md ("How an Agent works in this repo"), states the four rules as prose and points at the runbook's PRE rehearsal for the procedure.
  • A new make worktree Target makes rule 1 real: it creates a fresh worktree on a new branch off origin/main, links the host-level secrets into it, gives it its own OpenTofu state copy, and gives its PRE its own identity — so a rehearsal from the worktree works, and two Agents can rehearse at once without colliding.
  • PRE's identity moves from a fixed address to DHCP plus a provider-assigned vmid; PROD stays pinned, because PROD's address is load-bearing (the Edge and the router rule depend on it) and PRE's is not.
  • The repo's terms normalise to PROD and PRE, and Rehearsal joins the glossary.

User Stories

  1. As an Agent, I want the workflow written where I already look, so that I do not have to infer it from the history or ask the Operator each time.
  2. As an Agent, I want to start from a known-fresh branch off main, so that my change never rides on a stale base.
  3. As an Agent, I want one command that creates that worktree and its branch, so that I cannot get the base or the name wrong.
  4. As an Agent, I want my worktree's secrets linked from the host, so that a rehearsal from the worktree runs instead of refusing at a missing .env.
  5. As an Agent, I want my worktree to have its own OpenTofu state, so that my rehearsal does not contend with another Agent's.
  6. As an Agent, I want my PRE to have its own identity (address and vmid), so that my rehearsal and another Agent's do not collide on the Node.
  7. As the Operator, I want Agent rehearsals to address their PRE by a DHCP-assigned address, so that a rehearsal can never squat on an address a future Guest needs.
  8. As an Agent, I want to learn my PRE's address after it is created, so that I can run the playbook against it without hardcoding an address.
  9. As an Agent, I want to know exactly which PROD Targets are forbidden and which read-only ones are allowed, so that I can inspect without stepping over the line.
  10. As an Agent, I want to know that a mutation on PROD needs the Operator's approval, so that a direct request from the Operator is the only path to a PROD change.
  11. As an Agent, I want the coverage of a rehearsal stated honestly, so that I do not claim a rehearsal proves the Edge, TLS or the WAN path.
  12. As an Agent, I want my completion bar to be "every acceptance criterion or user story in the issue exercised on PRE", so that I know when the work is done.
  13. As an Agent, I want to report each criterion's pass/fail with its command, so that the Operator sees evidence without a wall of logs.
  14. As an Agent, I want to tear PRE down before I hand the work back, so that I leave nothing behind.
  15. As the Operator, I want PRE torn down as part of the work, so that I never inherit a stray Guest.
  16. As an Agent, I want to know the rules are convention today and gates tomorrow, so that I do not rely on enforcement that does not exist yet.
  17. As an Agent, I want the CI plan recorded, so that I know I will write tests for the acceptance criteria and CI will run them.
  18. As a reader, I want one word for each environment (PROD, PRE), so that no document, issue or commit drifts to "live" or "staging".
  19. As a reader, I want "rehearsal" defined, so that the verb in the rules matches the glossary.
  20. As an Agent, I want the doc not to restate the runbook's steps, so that there is one procedure, not two.
  21. As the Operator, I want the doc repo-only, so that agent tooling never reaches the published wiki.
  22. As an Agent, I want to know the one remaining hazard of concurrent rehearsals, so that I do not assume perfect isolation.
  23. As an Agent, I want the read-only PROD Targets named, so that "never act on PROD" does not read as "never look at PROD".
  24. As the Operator, I want PROD's pinned address untouched by this change, so that the Edge and router rule keep working.
  25. As a future maintainer, I want the per-worktree state deviation from ADR 0004 recorded, so that the "one state" decision is not silently overridden.

Implementation Decisions

  • The doc is docs/agents/workflow.md, repo-only. AGENTS.md gains a ### Workflow pointer under ## Agent skills; docs/wiki-pages.yml lists the new page under exclude (an unpublished doc that is neither published nor excluded is a coverage gap in that manifest).
  • The doc's shape: the four rules as prose, each with its why and its limit; a pointer to the runbook for the procedure (no restatement); the CI note. It uses glossary vocabulary throughout (PROD, PRE, Guest, Stack, Service, Target, rehearsal, Operator, Agent).
  • Rule 1 as written: git fetch origin && git switch -c hermes/<issue>-<slug> origin/main, one branch per issue, in a worktree.
  • Rule 3 as written: forbid every mutating PROD Target (prod, guest, edge, service, snapshot, state-backup) and ssh; explicitly allow the read-only ones (guest-plan, edge-plan, service-check) and verify (which reads PROD from the WAN and writes only to a temp dir); a mutation on PROD needs the Operator's approval. The deploy model is stated plainly: merge is not deploy — after merge the Operator runs make prod by hand, for now.
  • Rule 2 as written: divergence between PROD and PRE is allowed only in inventory and inventory-scoped variables (inventories/prod|pre/hosts, group_vars/forgejo/* for PROD, host_vars/forgejo-pre.yml for PRE); no role or template may branch on environment (verified: none does today). A rehearsal covers the roles, the playbook, the templated config, the units and the package installs; it cannot cover the Edge, TLS, the WAN-published URL, the SSH stream or per-repo Actions (forgejo_actions_repos: [] on PRE).
  • Rule 4 as written: PRE must not outlive the work unit — the completion bar is met, then make pre-down, then the PR. Evidence is one line per acceptance criterion / user story, not full logs.
  • make worktree creates the worktree (git worktree add under .worktrees/), links the four gitignored secrets from the primary checkout (they are host-level, not per-branch), seeds the worktree's OpenTofu state copy, and records the worktree's PRE identity. A clean worktree missing those secrets cannot run make pre; the Target creates the links.
  • Each worktree gets its own OpenTofu state copy. This deviates from ADR 0004's "one Stack, one state": the primary state stays the source of truth for PROD, and a worktree's copy is an ephemeral rehearsal state that only ever targets PRE. Recorded as an amendment to ADR 0004 (a Consequences / Update-when line), not a new decision record.
  • PRE's identity becomes dynamic: ip_config { ipv4 { address = "dhcp" } } with the gateway omitted, and no vm_id (the provider assigns one; random_vm_ids = true in the provider block makes concurrent assignment collision-safe). PROD keeps vmid = 141 and 10.12.0.141 pinned.
  • The container resource exposes the DHCP-assigned address as a computed ipv4 map (per network device). The rehearsal reads it from the apply's state/output and writes a generated PRE inventory and host_vars for the worktree — PRE's host variables stop being a committed static file for rehearsal purposes.
  • The shared base image (proxmox_download_file.debian_13_base, a node-level template pinned by checksum) is imported into each worktree's state, or resolved by a proxmox_files data-source lookup, so no worktree re-downloads it and none collides on it.
  • Terminology: prose and comments read "PROD" (today "Prod" — 61 occurrences across 14 files). Identifiers do not move: the make prod Target, PROD_WORD := prod (the typed apply-gate word), PROD_INVENTORY, inventories/prod, the ADR 0004 slug, and lowercase prod in commands all stay as they are. The glossary's PROD entry states this so nobody "fixes" them.
  • GLOSSARY.md gains Rehearsal (definition only, no procedure) and the PROD/PRE entries carry the PROD/PRE spelling.
  • No new ADR for the workflow policy itself (it is a convention, cheap to reverse); the PRE-state change is an amendment to ADR 0004.
  • CI note: rules 2–4 are convention until CI lands (.forgejo/workflows does not exist; issue #2 covers the first workflow). The eventual gate: an Agent writes tests that verify each acceptance criterion / user story, and CI runs them automatically.

Testing Decisions

  • The seam is the rehearsal itself — the playbook run against PRE, from the worktree. It is the existing seam (make pre), used at the highest point; no new seam is introduced.
  • Good tests here assert external behaviour: the playbook's observable effect on the PRE Guest (the Service is up and configured, the templates render, the units are enabled), not the role's internals.
  • What a rehearsal covers: the roles, the playbook, the templated config, the units, the package installs, the read-only Postgres probes, and any scenario the issue names.
  • What it cannot cover: the Edge, TLS, the WAN-published URL, the SSH stream, and per-repo Actions.
  • The completion bar: every acceptance criterion or user story in the issue exercised on PRE, each reported pass/fail with the command that produced it.
  • The make worktree Target is verified by make -n worktree (the composed commands) and one real run; the DHCP/vmid change is verified by the first PRE apply showing a distinct address and vmid from PROD's.
  • Prior art: the make pre / make pre-down Targets and the runbook's PRE rehearsal section.

Out of Scope

  • The CI gates themselves (the workflow file, the test harness). Recorded as a follow-up; this spec only writes the plan down.
  • Automating the PROD deploy. It stays manual until it does not.
  • PROD's addressing. It stays pinned.
  • The behaviour of the Services themselves (Forgejo, the runner, forgejo-mcp).
  • Any environment beyond PROD and PRE.

Further Notes

  • ADR 0004 gains a line for the per-worktree state; its single-state claim is amended, not reversed, because the primary state remains the source of truth for PROD.
  • The remaining hazard of concurrent rehearsals: two worktrees are safe only if each owns a distinct PRE identity. The dynamic identity (DHCP address plus provider-assigned vmid) is what makes that true; a rehearsal that pinned a fixed address would reintroduce the collision.
  • The lesson-5-end-with-pre precedent survives only in an old PR title; the rule is written fresh here rather than cited.
## Problem Statement The Operator works this repo with Agents. `AGENTS.md` and `docs/agents/` say how an Agent reads the issue tracker and the domain docs, but nothing says how an Agent *works*: which branch to start from, where to work, what it may run, and what must not outlive the work. Four rules live in the Operator's head and in the git history: 1. always work on a fresh branch off `main`; 2. always rehearse on PRE — the same code PROD runs; 3. never run a mutating step against PROD; 4. always destroy PRE when the testing is done. The rule that matters most — rehearse on PRE — is not reachable from the place an Agent works. A fresh worktree carries no gitignored secrets (`tofu/.env`, `tofu/npm/.env`, `ansible/.env`, `ansible/vault-pass`) and no OpenTofu state, so `make pre` refuses at its first gate; and PRE's identity is fixed (vmid 142, `10.12.0.142`), so two Agents rehearsing at once would collide on the Node and on the shared base image. The convention cannot be followed even by an Agent that wants to follow it. ## Solution Write the workflow down where Agents already look, and make the rules reachable. - A new repo-only doc, `docs/agents/workflow.md` ("How an Agent works in this repo"), states the four rules as prose and points at the runbook's PRE rehearsal for the procedure. - A new `make worktree` Target makes rule 1 real: it creates a fresh worktree on a new branch off `origin/main`, links the host-level secrets into it, gives it its own OpenTofu state copy, and gives its PRE its own identity — so a rehearsal from the worktree works, and two Agents can rehearse at once without colliding. - PRE's identity moves from a fixed address to DHCP plus a provider-assigned vmid; PROD stays pinned, because PROD's address is load-bearing (the Edge and the router rule depend on it) and PRE's is not. - The repo's terms normalise to `PROD` and `PRE`, and `Rehearsal` joins the glossary. ## User Stories 1. As an Agent, I want the workflow written where I already look, so that I do not have to infer it from the history or ask the Operator each time. 2. As an Agent, I want to start from a known-fresh branch off `main`, so that my change never rides on a stale base. 3. As an Agent, I want one command that creates that worktree and its branch, so that I cannot get the base or the name wrong. 4. As an Agent, I want my worktree's secrets linked from the host, so that a rehearsal from the worktree runs instead of refusing at a missing `.env`. 5. As an Agent, I want my worktree to have its own OpenTofu state, so that my rehearsal does not contend with another Agent's. 6. As an Agent, I want my PRE to have its own identity (address and vmid), so that my rehearsal and another Agent's do not collide on the Node. 7. As the Operator, I want Agent rehearsals to address their PRE by a DHCP-assigned address, so that a rehearsal can never squat on an address a future Guest needs. 8. As an Agent, I want to learn my PRE's address after it is created, so that I can run the playbook against it without hardcoding an address. 9. As an Agent, I want to know exactly which PROD Targets are forbidden and which read-only ones are allowed, so that I can inspect without stepping over the line. 10. As an Agent, I want to know that a mutation on PROD needs the Operator's approval, so that a direct request from the Operator is the only path to a PROD change. 11. As an Agent, I want the coverage of a rehearsal stated honestly, so that I do not claim a rehearsal proves the Edge, TLS or the WAN path. 12. As an Agent, I want my completion bar to be "every acceptance criterion or user story in the issue exercised on PRE", so that I know when the work is done. 13. As an Agent, I want to report each criterion's pass/fail with its command, so that the Operator sees evidence without a wall of logs. 14. As an Agent, I want to tear PRE down before I hand the work back, so that I leave nothing behind. 15. As the Operator, I want PRE torn down as part of the work, so that I never inherit a stray Guest. 16. As an Agent, I want to know the rules are convention today and gates tomorrow, so that I do not rely on enforcement that does not exist yet. 17. As an Agent, I want the CI plan recorded, so that I know I will write tests for the acceptance criteria and CI will run them. 18. As a reader, I want one word for each environment (`PROD`, `PRE`), so that no document, issue or commit drifts to "live" or "staging". 19. As a reader, I want "rehearsal" defined, so that the verb in the rules matches the glossary. 20. As an Agent, I want the doc not to restate the runbook's steps, so that there is one procedure, not two. 21. As the Operator, I want the doc repo-only, so that agent tooling never reaches the published wiki. 22. As an Agent, I want to know the one remaining hazard of concurrent rehearsals, so that I do not assume perfect isolation. 23. As an Agent, I want the read-only PROD Targets named, so that "never act on PROD" does not read as "never look at PROD". 24. As the Operator, I want PROD's pinned address untouched by this change, so that the Edge and router rule keep working. 25. As a future maintainer, I want the per-worktree state deviation from ADR 0004 recorded, so that the "one state" decision is not silently overridden. ## Implementation Decisions - The doc is `docs/agents/workflow.md`, repo-only. `AGENTS.md` gains a `### Workflow` pointer under `## Agent skills`; `docs/wiki-pages.yml` lists the new page under `exclude` (an unpublished doc that is neither published nor excluded is a coverage gap in that manifest). - The doc's shape: the four rules as prose, each with its *why* and its *limit*; a pointer to the runbook for the procedure (no restatement); the CI note. It uses glossary vocabulary throughout (`PROD`, `PRE`, Guest, Stack, Service, Target, rehearsal, Operator, Agent). - Rule 1 as written: `git fetch origin && git switch -c hermes/<issue>-<slug> origin/main`, one branch per issue, in a worktree. - Rule 3 as written: forbid every mutating PROD Target (`prod`, `guest`, `edge`, `service`, `snapshot`, `state-backup`) and `ssh`; explicitly allow the read-only ones (`guest-plan`, `edge-plan`, `service-check`) and `verify` (which reads PROD from the WAN and writes only to a temp dir); a mutation on PROD needs the Operator's approval. The deploy model is stated plainly: merge is not deploy — after merge the Operator runs `make prod` by hand, for now. - Rule 2 as written: divergence between PROD and PRE is allowed only in inventory and inventory-scoped variables (`inventories/prod|pre/hosts`, `group_vars/forgejo/*` for PROD, `host_vars/forgejo-pre.yml` for PRE); no role or template may branch on environment (verified: none does today). A rehearsal covers the roles, the playbook, the templated config, the units and the package installs; it cannot cover the Edge, TLS, the WAN-published URL, the SSH stream or per-repo Actions (`forgejo_actions_repos: []` on PRE). - Rule 4 as written: PRE must not outlive the work unit — the completion bar is met, then `make pre-down`, then the PR. Evidence is one line per acceptance criterion / user story, not full logs. - `make worktree` creates the worktree (`git worktree add` under `.worktrees/`), links the four gitignored secrets from the primary checkout (they are host-level, not per-branch), seeds the worktree's OpenTofu state copy, and records the worktree's PRE identity. A clean worktree missing those secrets cannot run `make pre`; the Target creates the links. - Each worktree gets its own OpenTofu state copy. This deviates from ADR 0004's "one Stack, one state": the primary state stays the source of truth for PROD, and a worktree's copy is an ephemeral rehearsal state that only ever targets PRE. Recorded as an amendment to ADR 0004 (a Consequences / Update-when line), not a new decision record. - PRE's identity becomes dynamic: `ip_config { ipv4 { address = "dhcp" } }` with the gateway omitted, and no `vm_id` (the provider assigns one; `random_vm_ids = true` in the provider block makes concurrent assignment collision-safe). PROD keeps `vmid = 141` and `10.12.0.141` pinned. - The container resource exposes the DHCP-assigned address as a computed `ipv4` map (per network device). The rehearsal reads it from the apply's state/output and writes a generated PRE inventory and `host_vars` for the worktree — PRE's host variables stop being a committed static file for rehearsal purposes. - The shared base image (`proxmox_download_file.debian_13_base`, a node-level template pinned by checksum) is imported into each worktree's state, or resolved by a `proxmox_files` data-source lookup, so no worktree re-downloads it and none collides on it. - Terminology: prose and comments read "PROD" (today "Prod" — 61 occurrences across 14 files). Identifiers do not move: the `make prod` Target, `PROD_WORD := prod` (the typed apply-gate word), `PROD_INVENTORY`, `inventories/prod`, the ADR 0004 slug, and lowercase `prod` in commands all stay as they are. The glossary's PROD entry states this so nobody "fixes" them. - `GLOSSARY.md` gains `Rehearsal` (definition only, no procedure) and the PROD/PRE entries carry the `PROD`/`PRE` spelling. - No new ADR for the workflow policy itself (it is a convention, cheap to reverse); the PRE-state change is an amendment to ADR 0004. - CI note: rules 2–4 are convention until CI lands (`.forgejo/workflows` does not exist; issue #2 covers the first workflow). The eventual gate: an Agent writes tests that verify each acceptance criterion / user story, and CI runs them automatically. ## Testing Decisions - The seam is the rehearsal itself — the playbook run against PRE, from the worktree. It is the existing seam (`make pre`), used at the highest point; no new seam is introduced. - Good tests here assert external behaviour: the playbook's observable effect on the PRE Guest (the Service is up and configured, the templates render, the units are enabled), not the role's internals. - What a rehearsal covers: the roles, the playbook, the templated config, the units, the package installs, the read-only Postgres probes, and any scenario the issue names. - What it cannot cover: the Edge, TLS, the WAN-published URL, the SSH stream, and per-repo Actions. - The completion bar: every acceptance criterion or user story in the issue exercised on PRE, each reported pass/fail with the command that produced it. - The `make worktree` Target is verified by `make -n worktree` (the composed commands) and one real run; the DHCP/vmid change is verified by the first PRE apply showing a distinct address and vmid from PROD's. - Prior art: the `make pre` / `make pre-down` Targets and the runbook's PRE rehearsal section. ## Out of Scope - The CI gates themselves (the workflow file, the test harness). Recorded as a follow-up; this spec only writes the plan down. - Automating the PROD deploy. It stays manual until it does not. - PROD's addressing. It stays pinned. - The behaviour of the Services themselves (Forgejo, the runner, forgejo-mcp). - Any environment beyond PROD and PRE. ## Further Notes - ADR 0004 gains a line for the per-worktree state; its single-state claim is amended, not reversed, because the primary state remains the source of truth for PROD. - The remaining hazard of concurrent rehearsals: two worktrees are safe only if each owns a distinct PRE identity. The dynamic identity (DHCP address plus provider-assigned vmid) is what makes that true; a rehearsal that pinned a fixed address would reintroduce the collision. - The `lesson-5-end-with-pre` precedent survives only in an old PR title; the rule is written fresh here rather than cited.
pit closed this issue 2026-10-07 16:40:43 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
olympus/infra-forge#52
No description provided.