Tier-2 edge-brain (M7) — async edge LLM remediation, OFF by default #15

Merged
jmz merged 27 commits from feat/tier2-edge-brain into main 2026-09-17 18:43:14 +00:00
Owner

Tier-2 edge-brain (M7) — async, orchestrator-dispatched, OFF by default

Relocates warden's guardrailed LLM auto-remediation to a mesh node, dispatched over a new fire-and-forget NATS brain plane. Design converged over four adversarial review rounds (SAFE-TO-SPEC); built via subagent-driven development (fresh implementer per task + spec/quality review after each + a whole-branch opus review). Spec: docs/superpowers/specs/2026-09-17-tier2-edge-brain-design.md; plan: docs/superpowers/plans/2026-09-17-tier2-edge-brain.md.

26 commits · 1057 passed / 1 skipped · ruff clean.

What ships (all inert unless enabled)

  • mesh/brain.py async wire contract (payload-capped codec, brain subjects)
  • Revocable grant: Grant.jti + revoked_grants store + /invoke revocation check (before actuation)
  • Store schema v2: brain_job tracker — atomic terminal CAS + per-fingerprint lease + terminal_at per-job grace sweep; revoked_grants; agent brain/invoke_cmd
  • BrainRunner (edge AutoResolver over an in-memory shim), Runbook.tier2 + drift hash
  • Orchestrator: _dispatch_brain (persist-before-publish, concurrency cap, edge-interpreter invoke-tools), _on_brain_result (CAS + local-only fail-closed verify — the trust anchor), durable deadline sweep + cancel/revoke, Tier-0→2→1→escalate wiring, per-fingerprint lease enforced in _run_chain
  • Node wiring (announces brain/invoke_cmd; fail-closed on loopback URL / missing model cred), telemetry envelope, docs/mesh.md section + enable runbook, end-to-end integration test

Safety (verified end-to-end by the whole-branch review)

  • No false RESOLVE without the orchestrator's own local detector re-run; _verify_remote never decides a brain resolve
  • Grant revocation closes the zombie on every terminal path (bounded one-in-flight TOCTOU, documented)
  • Per-fingerprint lease blocks a 2nd dispatch / central Tier-1 while a job is live; terminal CAS exactly-once; durable sweep re-drives a crash-after-CAS
  • OFF by default (WARDEN_BRAIN_ENABLED orchestrator + WARDEN_EDGE_BRAIN node); untrusted edge telemetry never feeds durable counters/budget/drafter

Enabling (separate, gated deploy step — NOT done here)

Set the flags + a routable WARDEN_ORCHESTRATOR_URL + WARDEN_EDGE_MODEL_CRED file-secret + warden-nats ACLs for warden.brain.* / .result.* / .cancel.* + the go/no-go: confirm ≥1 real check is BOTH orchestrator-locally-verifiable AND edge-remediation-routed (else enabled-for-zero-checks). See docs/mesh.md.

Notes

  • Deferred minors (non-blocking) are listed in the build ledger; none affect the default (OFF) path.
  • Seven tasks each took one review fix round; the whole-branch review caught + fixed one load-bearing gap (the per-fingerprint lease was dead code — now enforced in _run_chain).

🤖 Generated with Claude Code

https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af

## Tier-2 edge-brain (M7) — async, orchestrator-dispatched, OFF by default Relocates warden's guardrailed LLM auto-remediation to a mesh node, dispatched over a new fire-and-forget NATS brain plane. Design converged over four adversarial review rounds (SAFE-TO-SPEC); built via subagent-driven development (fresh implementer per task + spec/quality review after each + a whole-branch opus review). Spec: `docs/superpowers/specs/2026-09-17-tier2-edge-brain-design.md`; plan: `docs/superpowers/plans/2026-09-17-tier2-edge-brain.md`. **26 commits · 1057 passed / 1 skipped · ruff clean.** ### What ships (all inert unless enabled) - `mesh/brain.py` async wire contract (payload-capped codec, brain subjects) - Revocable grant: `Grant.jti` + `revoked_grants` store + `/invoke` revocation check (before actuation) - Store schema v2: `brain_job` tracker — atomic terminal CAS + per-fingerprint lease + `terminal_at` per-job grace sweep; `revoked_grants`; agent `brain`/`invoke_cmd` - `BrainRunner` (edge `AutoResolver` over an in-memory shim), `Runbook.tier2` + drift hash - Orchestrator: `_dispatch_brain` (persist-before-publish, concurrency cap, edge-interpreter invoke-tools), `_on_brain_result` (CAS + **local-only fail-closed verify** — the trust anchor), durable deadline sweep + cancel/revoke, Tier-0→2→1→escalate wiring, per-fingerprint lease enforced in `_run_chain` - Node wiring (announces `brain`/`invoke_cmd`; fail-closed on loopback URL / missing model cred), telemetry envelope, `docs/mesh.md` section + enable runbook, end-to-end integration test ### Safety (verified end-to-end by the whole-branch review) - No false RESOLVE without the orchestrator's own **local** detector re-run; `_verify_remote` never decides a brain resolve - Grant revocation closes the zombie on every terminal path (bounded one-in-flight TOCTOU, documented) - Per-fingerprint lease blocks a 2nd dispatch / central Tier-1 while a job is live; terminal CAS exactly-once; durable sweep re-drives a crash-after-CAS - OFF by default (`WARDEN_BRAIN_ENABLED` orchestrator + `WARDEN_EDGE_BRAIN` node); untrusted edge telemetry never feeds durable counters/budget/drafter ### Enabling (separate, gated deploy step — NOT done here) Set the flags + a routable `WARDEN_ORCHESTRATOR_URL` + `WARDEN_EDGE_MODEL_CRED` file-secret + `warden-nats` ACLs for `warden.brain.*` / `.result.*` / `.cancel.*` + the **go/no-go**: confirm ≥1 real check is BOTH orchestrator-locally-verifiable AND edge-remediation-routed (else enabled-for-zero-checks). See `docs/mesh.md`. ### Notes - Deferred minors (non-blocking) are listed in the build ledger; none affect the default (OFF) path. - Seven tasks each took one review fix round; the whole-branch review caught + fixed one load-bearing gap (the per-fingerprint lease was dead code — now enforced in `_run_chain`). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
jmz added 26 commits 2026-09-17 18:33:10 +00:00
Design converged over four adversarial review rounds (SAFE-TO-SPEC);
plan reviewed twice (READY-TO-EXECUTE after Task-5 test corrections).
Grant gains a jti field (default "" to keep existing positional
constructors working) stamped by issue() with secrets.token_hex(16)
and included in the signed payload so it can't be tampered with.
verify() parses it back into the returned Grant but stays store-less
- revocation lookup belongs to a later task in invoke_api, not here.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
The plan's Global Constraints wrongly copied the SGM repos' no-attribution
rule; warden-oss convention is to keep the harness trailer.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Adds the Tier-2 edge-brain jti-revocation denylist (revoke_grant/
is_grant_revoked/sweep_revoked) and the brain_job table (methods land in
a later task), plus agent.brain/invoke_cmd columns for the upcoming
brain-capable agent registry entries. Also fixes a latent migration bug:
PRAGMA user_version was only ever stamped inside the `if ver < 1` branch,
so a DB already at v1 would never re-stamp and would re-run (and fail)
the v2 ALTERs on every boot; the stamp now happens once at the end of
_migrate after all version branches have run.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
test_migration_v1_to_v2_stamps_user_version builds a fresh DB, which
migrates 0->1->2 in a single _migrate() call regardless of where the
PRAGMA user_version stamp sits — it can't distinguish the fixed stamp
placement from the old bug (stamp only inside `if ver < 1`).

Add test_migration_from_genuine_v1_db_reaches_v2: hand-constructs a raw
sqlite file already in the genuine post-v1 state (source column present,
agent table without brain/invoke_cmd, PRAGMA user_version = 1 already
stamped) so opening it only ever takes the `if ver < 2` branch. Verified
this test fails against the old (reverted-and-restored) buggy stamp
placement and passes against the current code.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
/invoke now checks store.is_grant_revoked(grant.jti) after grant.verify
and grant_mod.allows() succeed, strictly before the target is resolved
or the Invoker is called. A revoked jti now 403s with {"error":
"grant_revoked"} and bumps warden_invoke_denied_total{reason="revoked"}
instead of proceeding to actuation.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Add brain/invoke_cmd to AgentRecord, build_announce's emitted body, and
parse_announce's clean-field output (default brain=False/invoke_cmd=None
when absent). Registry.record_to_row/load_from persist and restore both
through the Task-3 agent.brain/invoke_cmd columns, JSON-encoding
invoke_cmd.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Extend the existing parametrized rejection table with a non-list
invoke_cmd and a list with non-string elements, both expecting
Reject("bad_invoke_cmd"). Follow-up to the Task 5 review: the branch
added in parse_announce had no test exercising it.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Add an opt-in `tier2: bool = False` field to Runbook, parsed the same way as
`auto`/`verify` (unknown top-level frontmatter keys still only warn, never
crash). Add `resolve.runbook_set_hash(runbooks)`, a sha256 fingerprint over
each runbook's load-bearing fields (name, body, auto, verify, max_attempts,
allowed_tools, tier2), sorted before hashing so it is order-independent.
These will gate edge-brain dispatch to tier2-flagged runbooks and let the
edge detect drift against central's runbook set.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Review fix round 1/5: timeout_seconds is load-bearing for the drift hash —
the grant TTL is sized as max_attempts * timeout_seconds + margin, so drift
on timeout_seconds alone (with max_attempts unchanged) let the edge's resolve
loop outlive the orchestrator-sized TTL without the hash catching it, i.e.
the exact mid-run /invoke 403 this hash exists to prevent. Add it to the
hashed tuple, add a test that isolates timeout_seconds drift, and note in
the docstring that the M3 invoke: map is deliberately out of scope (it drives
the Tier-0 deterministic path, not the tier2/edge-brain Tier-1 path this hash
gates).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Review ruling: grant TTL = max_attempts × timeout_seconds + margin, so both
sizing inputs must be in the drift hash.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Node-side Tier-2 edge-brain: BrainRunner runs the existing guardrailed
AutoResolver locally against a throwaway in-memory _ShimStore (implementing
exactly the store methods attempt() calls, with attempts persisting/advancing
across strikes), refusing when the node's runbook set hash doesn't match the
job's, and mapping the resolver's outcome to a BrainResult/BrainStatus.
Inert by default: no run_agent -> UNSUPPORTED without touching the resolver.

Adds a best-effort abort flag to AutoResolver.attempt(), checked between
strikes only (not mid-agent-run — ClaudeAgentRunner blocks on subprocess.run,
so there's no live Popen handle to kill; left a TODO for that refactor). No
safety property depends on it firing promptly.

Adds FakeAgent + fake_agent_resolves/fake_agent_two_strikes to mesh_fakes.py
implementing run_agent's .with_env() contract for tests.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Fix round 1/5 (coordinator review), two Important findings, both latent:

1. attempt() returns outcome "gone" whenever mark_resolving finds the
   incident no longer open — that covers either terminal state (RESOLVED or
   ESCALATED), not specifically "resolved". Mapping it to BrainStatus.RESOLVED
   would falsely report a resolve this run never made. Unreachable today
   (_ShimStore._match can't stop matching its own fixed fingerprint mid-
   attempt) but changed to FAILED so a future shim generalization can't
   false-resolve; the orchestrator's own terminal handling should own an
   already-closed incident, not this brain result.

2. The attempts-persistence property (mark_resolving incrementing/persisting
   across strikes) was asserted by name only: attempt()'s loop is bounded by
   range(max_attempts), not by .attempts, and .attempts is never surfaced on
   BrainResult, so no existing assertion could tell a persisting counter from
   one reset to 0 every call (confirmed: monkeypatching it to always reset
   left all 9 tests green). Added a white-box test that drives _ShimStore +
   AutoResolver directly and asserts shim.incident.attempts == 2 after a
   2-strike run. Verified by mutation: reverting the increment to a reset
   makes only this new test fail.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Adds the WardenStore methods the orchestrator lifecycle relies on:
insert_brain_job, get_brain_job, cas_brain_terminal (exactly-once
winner via UPDATE...WHERE state='running' + rowcount==1),
mark_brain_finalized, fingerprint_has_live_brain (a non-terminal row
IS the per-fingerprint lease), sweep_brain_jobs (past-deadline running
+ unfinalized-terminal past grace) and count_live_brain_jobs.

Task 3's schema-v2 brain_job table predates these fields, so this
extends that still-unreleased migration in place (column-existence
checks, same idiom as the existing agent.brain/invoke_cmd columns)
rather than bumping to v3 — schema v2 hasn't shipped yet on this
branch, so folding the fields into it keeps one coherent migration.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Review round 1 findings on Task 8, both pre-release cleanups since v2
hasn't shipped yet:

- Rewrite the v2 CREATE TABLE brain_job to define the real columns
  directly (fingerprint/node/deadline/jti/state/finalized/terminal_at
  + created_at/updated_at) instead of Task 3's 5 dead placeholder
  columns (agent_id/incident_id/status/request/response) plus an
  ALTER-patch bolted on top. One clean CREATE TABLE, no dead columns;
  user_version still lands at 2 and both migration-pinning tests still
  pass unchanged.

- sweep_brain_jobs's terminal arm previously compared `now > grace`
  with no per-row timestamp — the same boolean for every unfinalized
  terminal row on a given call (reap-all-or-reap-none), not a real
  grace window. cas_brain_terminal now takes `now` and stamps a new
  terminal_at column; sweep_brain_jobs's terminal arm is now
  `terminal_at + grace < now`, a genuine per-job window. Added
  test_sweep_terminal_grace_is_per_job_not_a_global_cutoff to pin it
  (fails under the old degenerate form).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Review rulings: Task 8 owns the brain_job schema (no dead placeholder columns);
sweep grace is a per-job duration via terminal_at, and cas_brain_terminal(job_id, now)
stamps it. Propagated the cas signature to Task 11/12 references.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Add _edge_invoke_tools(rb, target, incident) so a Tier-2 brain job's
composed Bash(warden invoke ...) allowlist shells out through the
TARGET edge node's own advertised interpreter (AgentRecord.invoke_cmd),
never the orchestrator's sys.executable/_default_invoke_cmd() -- a
brain job runs on the edge, so baking in the orchestrator's own venv
path would produce commands that can't run where they're dispatched.
Falls closed (returns None) when the target has no invoke_cmd or is
None, matching compose_invoke_tools' existing decline convention so
callers escalate rather than guess a path.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Remediator._dispatch_brain(incident, evidence, rb, target) -> bool dispatches
a Tier-2 BrainJob to an edge node: it declines (no persist, no publish) over
the brain_max_concurrent cap, when the target has no invoke_cmd, or when
_edge_invoke_tools can't safely compose the allowlist. On success it issues a
per-attempt grant (TTL sized from the runbook's max_attempts*timeout_seconds,
margined and capped), commits the brain_job lease via store.insert_brain_job
BEFORE publishing the encoded job — so a crash between the two orphans a lease
for the sweep to reap rather than ever letting a job run unleased.

Remediator.__init__ gains publish (fire-and-forget, distinct from the
request/reply invoker), brain_enabled, brain_max_concurrent, env, and
runbook_hash — all defaulted so existing construction is unchanged. spine.py
wires conn.publish (only when the bus is connected), WARDEN_BRAIN_ENABLED /
WARDEN_BRAIN_MAX_CONCURRENT, settings.env, and runbook_set_hash(runbooks)
into the Remediator it builds.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Remediator._on_brain_result is the terminal handler for a Tier-2 edge-brain
job result: atomic CAS (cas_brain_terminal) so a duplicate/racing delivery
for the same job_id no-ops on the loser, then the incident-open guard
(get_open_by_fingerprint), then the trust anchor before honoring any
edge-claimed RESOLVE — _brain_final_verify never falls back to
_verify_remote when the check has no live local spec (a config reload could
have dropped it between dispatch and verify), because that path reads the
brain node's own just-published finding_state rows and would let it
self-certify. A verify failure (or any non-resolved status) escalates
instead. The edge's self-reported events land on the incident timeline under
a distinct edge_reported kind, never feeding counters/budget/the runbook
drafter. jti is always revoked and the job always finalized for the CAS
winner, freeing the concurrency slot regardless of outcome.

spine.py subscribes brain.subject_result(env) (the whole env's result
traffic — no per-node segment to glob) to this handler, tolerating a BrainAck
arriving on the same subject as a liveness no-op.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Remediator.sweep_brain(now) reaps store.sweep_brain_jobs: a still-running
lease past its own deadline is CAS-terminaled, published as a BrainCancel
on the cancel subject, and escalated to a human (never re-run as a fresh
LLM job); a terminal-but-unfinalized lease (result won the CAS but the
orchestrator crashed before finishing bookkeeping) is re-driven through
the same idempotent finalize path without a second CAS. Both branches
share _sweep_finalize (revoke_grant + incident-open-guarded escalate +
mark_brain_finalized), mirroring _on_brain_result's terminal bookkeeping.

Registered on the spine main loop next to _sandbox_reap (_brain_sweep):
its own interval guard (WARDEN_BRAIN_SWEEP_EVERY, default 30s) plus a
startup sweep, so a job whose deadline passed during downtime is reaped
on the next tick with no special-casing. Per-job grace window for the
crashed-terminal case is WARDEN_BRAIN_SWEEP_GRACE (default 2m).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Wire the Tier-2 edge-brain dispatch into _run_chain, between Tier 0's
failure and central Tier 1. Eligibility requires all of: brain_enabled,
rb.tier2, the finding's owning agent (registry.get(finding.source) --
the same resolution `recheck` uses, not derive_locus) advertising
brain=True, and the check being locally verifiable
(_local_spec_for(check_id) is not None -- the trust-anchor precondition
also used by _brain_final_verify, now factored out of it).

A successful _dispatch_brain returns immediately from the chain -- the
persisted lease means central Tier 1 must never also run concurrently
against the same incident. A clean pre-dispatch decline (e.g. over the
concurrency cap) falls through to Tier 1 exactly as before.

brain_enabled defaults to False, so default orchestrator behavior is
unchanged.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
MeshSettings gains edge_brain/edge_model_cred_path/brain_max_concurrent,
parsed from WARDEN_EDGE_BRAIN/WARDEN_EDGE_MODEL_CRED/WARDEN_BRAIN_MAX_CONCURRENT.
mesh_settings_from_env fails closed when edge_brain is on and the
orchestrator url isn't routable (loopback/localhost/unset) or no model
credential is configured.

run_role wires the BrainRunner for the node role only, when edge_brain is
enabled: loads runbooks, builds the default ClaudeAgentRunner-backed edge
run_agent (injectable), and subscribes subject_dispatch/subject_cancel,
acking then publishing a BrainResult on subject_result. A stock node
(edge_brain unset, the default) never subscribes at all -- run_node itself
is untouched, per the collaborator-building glue living in run_role.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Critical gap: run_node called _start_announcer with no brain/invoke_cmd,
so build_announce always defaulted them to False/None regardless of
settings.edge_brain. The orchestrator's dispatch eligibility
(brain_target.brain) and _dispatch_brain's invoke_cmd guard both
hard-require these announced fields, so an edge-brain-enabled node
subscribed and could run a directly-delivered job but was never actually
selected by the orchestrator -- the registry never learned it was
brain-capable or how to invoke it.

_start_announcer now takes brain/invoke_cmd kwargs (default False/None,
unchanged for every other caller) threaded into build_announce. run_node
passes brain=settings.edge_brain and, when enabled,
invoke_cmd=[sys.executable, "-m", "warden.cli", "invoke"] -- this node's
own interpreter, matching what compose_invoke_tools bakes into a
dispatched job's tool allowlist.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Verify (no leak found, no production code change needed) that Task 11's
edge_reported envelope already fully isolates a Tier-2 edge's self-reported
events/telemetry from durable counters: _record_edge_result never calls
incr_counter, and warden_agent_cost_usd_total/the token counters are written
only by AutoResolver._record_run (central Tier-1). Prove it end-to-end in
tests/test_brain_integration.py: a real orchestrator Remediator + a real
edge BrainRunner wired over a shared FakeConn, driving the actual
dispatch -> publish -> wire -> _on_brain_result path to a verified
RESOLVED, plus a forged-telemetry test asserting the cost/token counters
never move. Document the whole Tier-2 subsystem in docs/mesh.md (async
brain plane, the revocable grant, the local-only fail-closed verify,
OFF-by-default) and the enable runbook (NATS ACLs mirroring the
exec/findings split, the model-credential file-secret, the routable-URL
fail-closed requirement, and the go/no-go check), correcting two doc
sections that pre-dated this build and still described Tier-2 as
unimplemented/deferred.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
fingerprint_has_live_brain was dead code: nothing consulted it before
Tier-0 exec, Tier-2 dispatch, or the central Tier-1 branch, so a flap
(natural resolve) + re-flap of the same fingerprint while a brain job
was in flight could dispatch a second brain job or fall through to
central Tier-1 concurrently with it.

Add a guard in Remediator._run_chain, after the auto-runbook check and
before any tier actuates: a live lease for the incident's fingerprint
now short-circuits the whole chain with {"outcome": "skipped",
"reason": "brain_in_flight"}, covering Tier 0, Tier 2, and Tier 1 in
one place. The lease's owner (_on_brain_result / sweep_brain) is
unaffected since neither path goes through _run_chain.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
Remove the dead MeshSettings.brain_max_concurrent field: it parsed
WARDEN_BRAIN_MAX_CONCURRENT with a different default (1) than the live
orchestrator, which reads the same env var directly in spine.py
(default 2) and never consulted this field. Verified with a full-repo
grep that nothing reads settings.brain_max_concurrent. Removing it
collapses the dual-reader trap onto the one path actually used.

Also close four coverage gaps flagged in review of the Tier-2
edge-brain build:
- test_mesh_brain: decode_job/decode_result reject malformed/non-JSON
  bytes via Reject (already handled by _decode_dict; adds the assertion).
- test_mesh_grant: mutating jti inside a signed grant while keeping the
  original signature is caught as a bad signature, proving the HMAC
  covers jti.
- test_brain_dispatch: a publish that raises AFTER insert_brain_job
  commits still leaves the lease row live (a sweepable orphan) -
  persist-before-publish holds under a publish fault, and the
  exception propagates rather than being swallowed.
- config.py: note on _is_loopback_url that a scheme-less
  WARDEN_ORCHESTRATOR_URL (e.g. orch.prod:8899) parses as hostless and
  is therefore also treated as loopback -> fail-closed.

Full suite: 1062 passed, 1 skipped.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
jmz merged commit bfb47faef6 into main 2026-09-17 18:43:14 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
public/warden!15
No description provided.