Deterministic, plugin-based ops incident engine: detect failure-by-absence, dedupe into incidents, notify, and (opt-in) auto-resolve.
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
root bfb47faef6 chore(mesh): converge brain_max_concurrent dual-reader + close test gaps
Remove the dead MeshSettings.brain_max_concurrent field: it parsed
WARDEN_BRAIN_MAX_CONCURRENT with a different default (1) than the live
orchestrator, which reads the same env var directly in spine.py
(default 2) and never consulted this field. Verified with a full-repo
grep that nothing reads settings.brain_max_concurrent. Removing it
collapses the dual-reader trap onto the one path actually used.

Also close four coverage gaps flagged in review of the Tier-2
edge-brain build:
- test_mesh_brain: decode_job/decode_result reject malformed/non-JSON
  bytes via Reject (already handled by _decode_dict; adds the assertion).
- test_mesh_grant: mutating jti inside a signed grant while keeping the
  original signature is caught as a bad signature, proving the HMAC
  covers jti.
- test_brain_dispatch: a publish that raises AFTER insert_brain_job
  commits still leaves the lease row live (a sweepable orphan) -
  persist-before-publish holds under a publish fault, and the
  exception propagates rather than being swallowed.
- config.py: note on _is_loopback_url that a scheme-less
  WARDEN_ORCHESTRATOR_URL (e.g. orch.prod:8899) parses as hostless and
  is therefore also treated as loopback -> fail-closed.

Full suite: 1062 passed, 1 skipped.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GfHTnzfxkwRhcrjcpdP5af
2026-09-17 18:39:45 +00:00
docs feat(mesh): edge-brain telemetry envelope + docs + integration test 2026-09-17 18:09:23 +00:00
src/warden chore(mesh): converge brain_max_concurrent dual-reader + close test gaps 2026-09-17 18:39:45 +00:00
tests chore(mesh): converge brain_max_concurrent dual-reader + close test gaps 2026-09-17 18:39:45 +00:00
.gitignore mesh M1: docs (docs/mesh.md, README fleet section), reject-reason reference, uv.lock ignore, v0.13.0 2026-09-09 23:33:23 +00:00
LICENSE warden: deterministic, plugin-based ops incident engine (initial OSS release) 2026-09-08 19:05:30 +00:00
pyproject.toml docs: release 0.26.0 (roadmap sweep) + bump version pins 2026-09-16 16:43:24 +00:00
README.md docs: release 0.26.0 (roadmap sweep) + bump version pins 2026-09-16 16:43:24 +00:00

warden

A small, deterministic operational incident engine for devops. warden turns any check — an error rate, a latency SLO, a full disk, an expiring certificate, a downed endpoint, config drift, a crash-looping container, and the failure-by-absence class error-based tools miss — into deduped incidents, notifies, and (opt-in, later) auto-resolves them via runbook-driven agents behind a hard allowlist.

It is plugin-based and domain-agnostic: the core knows nothing about your systems or the kind of problem. You point it at your world with detector plugins and a declarative checks.yaml; each detector decides what "failing" means for its own check.

Why

One engine for the whole operational surface, not another single-purpose alerter. Point-tools each watch one slice (metrics, logs, uptime, certs, jobs) and route nowhere useful; warden gives every kind of check the same lifecycle — stable-fingerprint dedupe, transition-only notification, crash-resumable incidents, and a path to automated remediation. It also covers one class the others structurally can't: failure-by-absence (a run that didn't happen, a feed that stopped) throws no error, so warden treats "did the expected thing happen, and is its output current?" as a first-class check — and reports a check that can't run as loudly as one that fails, never silently OK.

Design in one breath

  • Deterministic spine, no LLM in the hot path. Routing, dedupe, and severity are plain code keyed on a stable fingerprint; untrusted evidence text never decides an action.
  • Detector → Finding → Incident. A detector is pure: check(ctx) -> [Finding]. The store diffs findings per fingerprint and emits only transitions (ok→problem opens one incident, problem→ok resolves it, a severity change re-notifies). All state is in sqlite, so the process is crash-resumable.
  • Everything is a plugin. Built-in generic detectors ship in this repo; third-party add-ons register via the warden.detectors entry-point group — no fork required.

Install

pip install warden                 # core (light: pyyaml + structlog)
pip install "warden[delta]"        # + the delta_freshness detector's deps (deltalake/pyarrow)

Quickstart

# checks.yaml — your detection surface (reviewed like code)
checks:
  nightly-etl:
    plugin: heartbeat
    trigger: {every: 5m}
    params: {key: nightly-etl, max_age: 26h}   # the job POSTs /heartbeat on success
  api-up:
    plugin: probe
    trigger: {every: 1m}
    params: {kind: http, url: https://api.example.com/health}
WARDEN_CHECKS=checks.yaml WARDEN_DB=warden.db warden
# GET :8891/health · GET :8891/incidents · GET :8891/metrics (Prometheus) · POST :8891/heartbeat {"key":"nightly-etl"}
# SIGHUP reloads checks.yaml live (a bad file is refused, keeping the running config).
# WARDEN_RESOLVE_DWELL=5m suppresses flapping (resolve only after OK holds that long).

Built-in detectors

plugin detects
heartbeat failure-by-absence — a keyed heartbeat older than max_age (or never seen)
probe an HTTP endpoint (status/body), a TCP port, or an SSH/SFTP banner not answering
promql a Prometheus expression that crosses a threshold (turns any recording rule into a warden check)
prom_scrape a metric read straight from an exporter's /metrics crosses a threshold (when the Prometheus server isn't reachable but the exporter is)
disk free space on a path below a floor (severity scales with how far below)
cert_expiry a TLS certificate within warn_days/crit_days of expiry — or already expired
systemd a systemd unit not in the active state
json_check polls a JSON health/readiness endpoint and maps each record ({name,status,detail}) to a finding — adopts a service's own checklist
dagster a Dagster sensor not RUNNING / stale last-tick, a job run FAILURE, or a run-storm (retry loop), via GraphQL
delta_freshness a Delta table stale beyond a business-day threshold (weekend/holiday-aware)

Every built-in is pure and capability-injected, so it unit-tests without touching the network or a real host. Since 0.26.0 warden also ships the agent_audit drift sweep and ingests external error/business events (error_events, WARDEN_EVENTS_INGEST, on warden.events.<env>.<source>); container-health remains a roadmap plugin. Bring your own with the warden.detectors entry point (docs/adding-a-plugin.md) — no fork required.

Notifications

Transitions fan out to every configured channel (one failing channel never silences the others): a structured log (always on), an HTTP push (WARDEN_NOTIFY_URL — ntfy/webhook), Matrix (WARDEN_MATRIX_* — posts an m.notice to a room), and a Gitea/Forgejo issue tracker as the durable case record (WARDEN_FORGEJO_URL + _TOKEN + _OWNER + _REPO): an incident opens an issue, severity changes and the resolution append comments (the timeline), and resolving closes it — idempotent per incident. Wire the channels your deployment has; secrets stay in your private overlay. (For an internal tracker behind an expired/self-signed cert, WARDEN_TLS_INSECURE_HOSTS skips TLS verification for those exact hosts only.)

Telemetry

GET /metrics is Prometheus exposition. Beyond warden's own health (open incidents, last-tick age) it exposes durable counters for the whole incident lifecycle and — the part that matters for an agentic system — per-fix token and cost accounting parsed from each resolver run's stream-json result:

  • warden_transitions_total{kind,check} — opened / resolved / severity_changed
  • warden_incidents_resolved_total{check,mode} · _escalated_total · _waiting_total
  • warden_agent_runs_total{check,mode} · warden_agent_cost_usd_total{check}
  • warden_agent_{input,output,cache_read_input,cache_creation_input}_tokens_total{check}
  • warden_agent_turns_total{check} · warden_agent_duration_ms_total{check}

Every agent run also writes a usage event onto the incident timeline (tokens, cost, turns, duration), so you can see exactly what a given fix cost. Point Prometheus at /metrics for the metrics half.

For logs, set WARDEN_OTLP_LOGS_URL to an OTLP/HTTP collector endpoint and warden ships its structured logs there (batched, best-effort, zero extra dependencies — a plain OTLP-JSON POST to /v1/logs), which forwards them to Loki. A dead collector never blocks the loop.

Configuration & secrets

warden reads its detection surface from WARDEN_CHECKS and its state db from WARDEN_DB. Keep your checks.yaml, runbooks, and secrets in your own private repo/overlay — they are deployment-specific and never belong in this open-source repo. See docs/checks.example.yaml.

Writing a plugin

A detector is ~30 lines. Ship it in this repo (generic) or as your own package via the warden.detectors entry point (proprietary/add-on). See docs/adding-a-plugin.md.

Status

Phases 1–3 are implemented and tested (0.26.0 landed the last roadmap items: error_events NATS ingest, the agent_audit drift sweep, the fleet-hub first-connect retry, and sandbox.run mounted on /invoke):

  • Detect + notify — the detectors above, a Prometheus self-exporter (/metrics), and log / webhook / Matrix channels; SIGHUP config reload and flap suppression.

  • Runbook resolution — on an incident, warden matches a runbook and composes an exact, allowlist-scoped claude -p … remediation command (see docs/runbooks.example/).

  • Auto-resolution (opt-in, WARDEN_AUTO_RESOLVE=1) — for runbooks that set auto: true, warden launches that command itself behind hard guardrails: deny-by-default allowlist, per-attempt wall-clock timeout, 2-strikes, spine-side verify (the agent never self-certifies), a full action timeline on the incident, and escalation carrying the resumable agent session id. A runbook without auto: true is only ever composed for a human. Untrusted evidence can never widen the allowlist or inject instructions.

  • Guide-the-agent (Phase 4, with auto-resolve + Matrix) — when auto-resolution can't fix an incident but has a resumable session, warden hands off to WAITING_HUMAN and listens in the room. A human's reply resumes the same agent session under the same allowlist (a chat message can't widen the agent's tools), then the spine re-verifies. Escalation carries the claude --resume <session> id.

  • Chat issue-reporting (0.19.0) — an allowlisted Matrix sender can open a tracked incident directly with @warden report .../!warden report ... on the same room the guidance loop already listens in; gated by its own allowlist + per-sender rate limit, source-scoped so it can't collide with a detector finding, and escalates to a human unless a runbook is explicitly authored for it. See docs/mesh.md.

  • Auto-draft Tier-0 runbooks (0.20.0) — the "border-collie" learning loop: after the LLM (Tier 1) resolves a novel incident with exactly one clean, reproducible mesh invoke, warden drafts a DRAFT invoke: runbook (auto: false) and attaches it to the Forgejo case; a human reviews, merges, and flips auto: true to make it a $0, no-LLM Tier-0 fix next time. Free-form/ambiguous fixes are never drafted. See docs/mesh.md.

  • NATS TLS (0.21.0) — the ops bus can run over verified TLS: an opt-in tls:// mode with mandatory server-cert + hostname verification (no way to disable it) and optional mutual TLS, configured with three file-path env vars; a missing/unreadable CA or cert/key fails loud instead of falling back to plaintext. See docs/mesh.md.

  • Fleet across envs (0.22.0) — a prod "hub" orchestrator can aggregate a read-only, view-only pane of sibling per-env orchestrators' incidents (GET /fleet, warden_fleet_incidents{env,kind}): a remote orchestrator best-effort mirrors its own incidents to the hub over a second, publish-only, TLS-required connection, and the hub's FleetView is data-only — no actuation collaborator, so a remote incident can never trigger a local (or remote) exec, route, grant, or remediation. Gated on both sides; unset ⇒ byte-for-byte the prior single-env behavior. See docs/mesh.md.

  • MCP bridge (read-only, 0.23.0) — warden mcp (the optional pip install "warden[mcp]" extra) runs a stdio MCP server so an operator's Claude Code/Desktop session can query live warden state during triage — open incidents, an incident's detail+timeline, topology, fleet — through four tools that proxy the orchestrator's own read-only HTTP endpoints; it holds no NATS credential, no DB handle, and no invoker/grant/exec collaborator, so it cannot mutate state or invoke a capability. See docs/mesh.md.

  • sandbox.run ephemeral fixers (0.24.0; mounted on /invoke in 0.26.0) — a grant-gated capability to run a bounded fix in a TTL'd, resource-capped, one-shot, hardened container (pinned-digest image allowlist only, read-only rootfs, CapDrop: ALL, no host mounts ever, no fallback to the restart-proxy) that the caller cannot widen. POST /invoke now dispatches capability: sandbox.run to it (grant-gated exactly like the mesh exec path); it stays inert without a configured spawn-proxy (WARDEN_SANDBOX_DOCKER_HOST). See docs/mesh.md.

  • Roadmap sweep (0.26.0) — the last roadmap items: the fleet-hub first-connect retry/backoff (a hub down at startup is picked up later with no orchestrator restart), the error_events external-event NATS ingest (opt-in, source-scoped, WARDEN_EVENTS_INGEST), the agent_audit pure-Python cross-source freshness/drift sweep, and mounting sandbox.run on /invoke (above).

  • Off-box fate-sharing twin (Phase 5) — the primary warden reports its own liveness (WARDEN_SELF_HEARTBEAT_URL → a twin's /heartbeat each loop). A second warden on another host watches that key with the heartbeat detector, so if the primary — or its whole host — goes dark, the twin pages. The watcher is watched; they share fate only if both hosts die at once. See docs/twin.example.yaml.

All five phases are implemented and tested (0.26.0 landed agent_audit and the error_events ingest detector; container-health remains a roadmap plugin — see docs/design.md). Auto-resolution, guidance, and the twin are opt-in and gated, so warden is safe to run in detect-and-notify mode from day one and graduate toward autonomy at your pace.

Fleet / mesh (optional, NATS; M1–M6 through 0.26.0 — the 0.21–0.26 items (NATS TLS, cross-env fleet, MCP bridge, sandbox.run + its /invoke mount, hub retry, error_events) are detailed in docs/mesh.md). warden-core stays the deterministic orchestrator: one orchestrator per environment, over a dedicated ops bus (the nats extra), never a bus that carries business traffic. Agents (WARDEN_ROLE=sensor|actuator| node) now actually run: a sensor runs detectors and publishes findings, an actuator executes typed capabilities from its manifest in a locked-down subprocess and replies, and both self-register and re-announce on a lease. The orchestrator tracks them in a topology registry, diffs an optional topology.yaml (desired state) against what actually announced, and serves it at GET /topology plus fleet metrics on /metrics (warden_agent_up/_info/_lease_age_seconds/_capability, warden_agents_registered, warden_mesh_connected). A bus outage is always one SKIPPED incident (mesh-bus), never a storm of per-agent pages. WARDEN_NATS_FINDINGS_SUBJECT (a single hardcoded subject) is deprecated in favour of the versioned per-agent subject warden.findings.<env>.<agent_id> with the {"v":1,...} envelope, which every sensor is expected to use going forward. Ingested findings are untrusted, subject to the same fenced, allowlist-bounded handling as local detectors. An operator (and, since 0.16.0, the Tier-1 LLM fallback) drives a capability through warden invoke or a grant-gated POST /invoke on the orchestrator — a per-attempt HMAC grant bound to {incident, capabilities, expiry} so the caller never needs a NATS credential. Full wire protocol, manifest shape, the docker.restart capability, and the /invoke contract are in docs/mesh.md.

Tier-0 deterministic self-heal is now live (0.15.0). A runbook can declare invoke: — a capability call with literal params, routed to a fresh actuator and confirmed by a polling verify check — that the spine tries before any LLM is ever consulted. When it resolves the incident, it's a $0, no-LLM fix; only if it can't (or isn't configured) does the existing LLM auto-resolver run, and only if neither can does the incident escalate to a human. A warden with no mesh, or a runbook with no invoke:, behaves exactly as before — this is opt-in per runbook. See Remediation tiers in docs/mesh.md for the full tier sequence, the invoke: frontmatter, and the metrics it emits.

The LLM fallback can now act through the mesh too (0.16.0). When Tier 0 fails, Tier 1's LLM is no longer stuck with a runbook author's hand-written allowed_tools — it's handed an exact, fully-enumerated warden invoke <cap> --target <id> --param k=v toolset composed from the routed actuator's own registered capabilities, plus a short-lived, per-attempt HMAC grant bound to {incident, capabilities, expiry}. The child gets that grant and the orchestrator's URL, never a NATS credential; a capability whose param space can't be safely and exactly enumerated declines the whole composition and escalates to a human rather than handing the LLM a wildcard tool. See Tier-1 composition in docs/mesh.md.

Verify can now nudge a remote sensor, and routing can be pinned ad hoc (0.17.0). A sensor/node serves POST /recheck — re-run one of its own checks and re-publish immediately — and Tier 0's polling verify can best-effort trigger it once, on its remote-observe path only, via the mesh warden.recheck capability, instead of only ever waiting out that sensor's own scheduled cadence; a trigger failure never fails or skips the observe, and with nothing wired the verify step is byte-for-byte the pre-0.17.0 observe-only behavior. Separately, warden invoke and POST /invoke can now pin a --host/--service locus directly (server-side precedence: explicit target, then host/service, then env-wide) — the same routing a runbook's invoke: scope: already used, now reachable from an ad hoc operator or Tier-1-composed call, not only a declared runbook. See Remote recheck and Locus routing in docs/mesh.md.

License

Apache-2.0.