## autonomous coding harness

chug

Given a spec and a goal, it keeps on chugging.

design rule #1 — the loop is code, not conversation. the model never decides whether to continue: the driver does.

live queue · tests · cycles → self-improving in public adversarial validation built in
$ chug run --spec SPEC.md \ --goal "Build X and make the check pass" chug <version> (<commit>) cwd=/srv/app spec=SPEC.md model=… ▪ iter 01 read SPEC.md · seeded LEDGER.md ▪ iter 07 edit_file src/api.rs · bash: cargo test ✓ ▪ iter 12 chug: budget low — 8 iteration(s) remain ✓ goal_complete accepted — check exited 0
one run, sketched — the real driver appends every event to .chug/events.jsonl

## 01 · how it works

The loop is code, not conversation

A run takes a spec file with a check: line and a goal, and drives a fixed loop until the check exits zero. Goal and progress live in files on disk, re-read every iteration — so context trimming can never kill the run.

SPEC + GOAL

SPEC.md with a check: line, plus --goal. Plain files — the driver re-reads them every iteration.

→

DRIVER LOOP

Continue/stop is the driver's call, never the model's. Anti-stall kick, stuck tripwire, budget-low warnings, same-cwd driver lock.

→

THE TOOLS

read_file, write_file, edit_file, bash, grep, tgrep, glob, list_dir, update_ledger, decision_log, delegate, web_fetch, goal_complete — sandboxed to --cwd.

→

GOAL_COMPLETE

Verified, not taken on faith: the spec's check: re-runs and the claim is accepted only on exit 0.

✗ check fails → keep chugging

The failure goes back to the model with the loop still in charge. It fixes, re-runs, re-claims. Rejected claims never end the run — the budget does, and a budget death prints a resume line and leaves the ledger intact.

LEDGER.md → external memory

Updated every iteration, injected into every turn, so transcript trimming never loses progress. A fresh run archives and reseeds it; --resume keeps it.

## 02 · proof it runs itself

Built by its own loop, measured by its own gates

chug improves chug. The loopd supervisor runs the self-improvement cycle back to back, and at every wrap the synced stats band below is regenerated from the repo's own ledgers — done rows in the queue ledger, gate counts in commit messages, the git log. Numbers live there, quoted and never rounded up; this section is the story behind them.

The tgrep arc: validation round after round

When the loop built its own token-budgeted context search, the adversarial validator refused to pass it — again and again: a merge-radius contradiction, symbols mode dropping declarations, vacuous pins, then one flaky perf pin with every mutant red. Each refusal forced a fix-up round; the final round closed it with a fresh mutant batch, all dead. A full artifact set came out of it, and the post-merge gates ran green — the counts are in the stats band.

The bug the pipeline existed to catch

Building its own hooks system, the implementer passed gates — and the validator still failed it. First refusal: PostToolUse was firing on calls that PreToolUse had vetoed, a real semantic bug no test had pinned. The fix-up swept the whole class — the sibling risk-gate-block path had the same hole — with killing tests proven red first. Second round: PASS, zero blocking findings.

Dogfood, verbatim

Its interactive chat mode: written for it, by it. Its delegate tool — a Rust module of its own, launching child chug runs in git worktrees: written by a child run, validated by another model family, then grown through wait_secs long-polling, resume, collect, and terminal-wait mode across later cycles. Its decision_log, hooks, plan mode, TUI, MCP client, Langfuse tracing: same story. And this very page — the one you are reading — was written by chug running its own loop, with this site's check gate re-running on every claim.

Honest failure books: one cycle hit its iteration ceiling mid-wrap and the wrap was lost — the next cycle reconstructed it from git history and ledgers alone. Failed validation rounds are the norm, not the exception; the tgrep arc alone logged refusal after refusal before its pass.

→ live counts: queue, tests, cycles — rewritten at every wrap

## live stats · rewritten by scripts/site-sync.sh at every cycle wrap

141/145queue items landed — done rows in TODO.md, 145 item rows total
1121tests green at the newest full-suite gate count in a commit message (7d121c3)
50cycles completed by the loopd supervisor — "cycle OK" lines in .chug/loopd/loopd.log
2026-09-29last EVALUATION.md write (git log, chug repo)

Last five items landed — item · ref · date, quoted from the chug repo's git log:

## 03 · the timeline

Day one → chug.sh

Every date and hash below is quoted straight from git history — the harness repo and this site's own repo. The loop wrote this page too, so this timeline is the loop reading its own commit log back to you.

2026-09-20

Day one — the core harness f911488

First commit at 17:42: a driver-in-code loop, LEDGER.md as external memory, and goal completion that only counts when the check exits zero. Thirty-one commits land before midnight.

2026-09-20

The TUI, eight minutes later 76ac019

A ratatui dashboard at 17:50 — live activity stream, LEDGER panel, steering input, operator abort.

2026-09-20

The risk gate db6fea6

glob, list_dir and replace_all join the tools, and laya starts classifying every bash command before it runs — destructive calls are blocked with an error the model can route around.

2026-09-20

Chat mode, dogfooded fb842f2

Interactive mode lands with the claim in the commit message itself: "dogfooded by chug itself".

2026-09-20

The meta-loop era begins 7deb7f0

META-SPEC writes down chug-orchestrating-chug; SELF-SPEC makes improvement continuous. That evening "muse implements + kimi validates" becomes the default round policy (0611c33).

2026-09-21

MCP — stdio, then streamable HTTP 3386d2b

SPEC-7 merges: config discovery, handshake, tools/list registry, seven review fixes. The streamable-HTTP/SSE transport (SPEC-9) lands the same day — remote servers appear as mcp__name__tool too.

2026-09-21

Langfuse observability 25a36c1

SPEC-8 merges: a trace per run, a generation per LLM call, spans and scores per tool call. Fire-and-forget — telemetry never changes run behavior.

2026-09-21

The doctrine becomes code 672d04a · aac3629 · 3c795b3

Mutation testing enters the validation template — every landing must now prove its tests can fail. LOOP-SPEC writes the whole evaluate→queue→implement→validate→merge→push loop into one file. glm-5-3-flash takes the implementer seat; kimi-k3 keeps the verdicts.

…and 116 earlier milestones (T1–T127)

2026-09-25

loopd goes continuous 548d494

The supervisor lands after lunch and starts its first cycle the same minute — the first line of the loopd log. Cycles run back-to-back, each wrap the next cycle's input. It is still running today.

2026-09-25

delegate — chug spawns chug 1012dca

T23: launch and status for bounded child runs in git worktrees. Later cycles grow wait_secs long-polling, zombie reaping, resume, collect, and terminal-wait mode.

2026-09-26

Plan mode 87fe53f

chug plan (T73): read-only exploration under a five-tool contract — read_file, grep, glob, list_dir — with submit_plan as the single write and exit path.

2026-09-27

Per-phase model routing c1daaca

T81: glm orchestrates routine cycles, kimi takes the judgment calls — a mechanical freshness predicate in loopd.sh, never model vibes.

2026-09-27

Hooks ccb828a

PreToolUse veto + PostToolUse advisory (T83). The first validation round FAILED the merge — it caught PostToolUse firing on calls that never executed, a real bug — and the class sweep fixed the sibling risk-gate path with it.

2026-09-27

Permissions e9afed9

The deny-list lands (T90): .chug/permissions.json, first-match-wins, fail-closed on match, first in the policy chain before hooks and the risk gate. Its FEATURES.md roadmap row flips to LANDED.

2026-09-27

The move to K7 — always-on com.tampajohn.chug-loopd.plist

loopd moves under launchd on the operator's host (K7HC2K125R): the plist wraps loopd.sh in caffeinate -dims with RunAtLoad. The improvement loop no longer lives in a terminal.

2026-09-27

chug.sh — this site e998d45

chug builds its own public site: eight commits in one evening, every section landed against the verify.sh check gate, GitHub Pages serving a CNAME. You are reading the dogfood.

2026-09-28

META-META-SPEC estimate calibration: all-in counts, test/doc density ~1.5-3x, ~400 should-split band (532c403)

2026-09-28

F10 phase 1: chug mcp-serve stdio MCP server skeleton + chug_status tool (d6264be)

2026-09-28

META-SPEC §6 validator template: write /tmp helper scripts via bash heredoc (eef7a29)

2026-09-28

chug_status weak pin: missing-.chug leg asserts its distinctive phrase (T124 M3 survivor) (976e4ae)

2026-09-28

F10 phase 2a: chug_collect MCP tool (read-only structured child result over mcp-serve) (85ca4c1)

2026-09-28

goal_complete denial bypass (codex review) (9b36a2e)

2026-09-28

SSE truncation acceptance (codex review) (70ec4a7)

2026-09-28

loopd grep spoofing (codex review) (3bc3169)

2026-09-28

Sandbox escape: file tools follow symlinks; bash has no confinement (codex review §2) (d5d9c28)

2026-09-28

driver.lock acquisition race — two starters can both own it (codex review §1) (611916e)

2026-09-28

Crash mid-tool-batch leaves unresumable transcript (codex review §1) (bf672b6)

2026-09-28

loopd launches stale binary after failed build (codex review §1) (08f0a39)

2026-09-28

mcp.json executes before permission enforcement (codex review §2) (37f8edc)

2026-09-28

max_tokens too small for thinking models (GLM truncation class) (00adb88)

2026-09-29

Editable spec check: bypasses bash deny + risk gate (codex review §2) (b5f1fa0)

2026-09-29

F10 phase 2b: chug_launch MCP write leg, flag-gated by mcp-serve --allow-launch (default-deny) (2b4490b)

2026-09-29

eval-digest reader staleness check excludes the reader's own live stream (b52e62c)

2026-09-29

Goal-gate check + driver-spawned shells must not inherit CARGO_TARGET_DIR (cross-worktree… (8cabc79)

2026-09-29

update_ledger writes through fsatomic::write_atomic (completes T136 crash-safety class) (52a0abf)

2026-09-29

F2 phase 2a: chug run --approve plan.md gate + web_fetch in plan mode (ROADMAP PULL) (55d59c3)

The quiet patch stays on the chart: zero commits between the last early-cycle landing and loopd's first start. Honest books include the blank pages.

## 04 · feature grid

What's in the box

Every capability below is quoted from the repo's README and FEATURES roadmap — including the phases still deferred with written reasons, because a harness that keeps honest books about what it hasn't done yet is the whole point.

builtin tools

read_file write_file edit_file bash grep tgrep glob list_dir update_ledger decision_log delegate web_fetch goal_complete — all paths sandboxed to the run's working directory.

delegate sub-agentslanded

Launch, observe, and collect bounded child chug runs in git worktrees: launch / status / collect, long-poll wait_secs, terminal-wait mode, and resume of an aborted child.

plan modelanded

chug plan explores the repo and drafts an implementation plan with a read-only contract: exactly five tools plus submit_plan, its only write and exit path.

hookslanded

.chug/hooks.json policy-as-config: PreToolUse hooks can veto a tool call before it executes; PostToolUse hooks advise. Fails open, per-checkout, never fires a hook from a hook.

permissionslanded

.chug/permissions.json deny rules evaluated first in the policy chain — before hooks and the risk gate; fail-closed on match, first-match-wins. The first phase is landed; ask-mode and settings.json unification are deferred phases with written reasons.

risk gate (Laya)

Every bash command is classified destructive / risky / safe by a local Laya judge server before it runs; destructive is blocked with an error the model can route around. Verdicts logged to .chug/risk_verdicts.jsonl.

MCP servers

Consumes tools from MCP servers over stdio or streamable HTTP with Claude Code-compatible config. Server tools appear as mcp__name__tool alongside the builtins; per-server fail-soft.

Langfuse observability

Optional tracing to self-hosted Langfuse v3: a trace per run, a generation per LLM call, a span per tool call, outcome and iteration scores. Fire-and-forget — off means zero cost, telemetry never changes run behavior.

TUI

--tui: live activity stream, LEDGER.md panel, and a status bar with model, iteration/budget, elapsed time, and cumulative tokens. i steers, q aborts gracefully.

web_fetch

Read-only HTTP(S) GET with hard bounds: 5 redirects, 10s connect / 30s total, output capped and HTML stripped to text, binary types refused. Bounded and audited where raw curl is neither.

tgrep

Token-budgeted ranked context search: ranked match clusters with tight context windows, quoted phrases, a symbols mode for Rust signatures. Deterministic scoring — no embeddings, no LLM.

chat

Interactive TUI sessions: @file attachments, tab autocomplete for commands and paths, slash commands, esc to interrupt. Steering notes pass into autonomous runs at iteration boundaries.

decision_loglanded

Every loop judgment — validation verdicts, routing calls, recoveries — emits a structured record with class, inputs, options, choice, and a confidence score, building the corpus for confidence-gated routing.

loop hardening

Driver lock against concurrent runs, transcript trimming that keeps the cached prefix byte-stable, jq-mineable events.jsonl, budget-low warnings, a stuck tripwire, and an anti-stall kick.

Image inputlanded

read_file on png/jpg returns image blocks (vision); chat accepts pasted/dragged screenshots. UI work needs eyes.

Session forkin-flight

Clone transcript+ledger at iteration N into a new session id; explore two approaches from one state.

Streaming UXlanded

Text deltas to sinks as they arrive (TUI live typing, headless progress); watchdog gets byte-level liveness for free.

Structured todo toollanded

todo_add/update/list driver-visible tools (statuses enforced) as an alternative to freeform LEDGER edits — the orchestration ledger becomes queryable.

Slash-command packslanded

.chug/commands/*.md repo-local commands invocable from chat (/review, /triage) and as run goals. Community-extensible without code.

chug as MCP serverlanded

chug mcp-serve: expose run/delegate/status as MCP tools so Claude Code, the bridge fleet, or another chug can drive it. The fleet primitive.

MCP resources+promptsqueued

Consume MCP resource/prompt capabilities (the wired surface is tools).

Web searchqueued

Provider-pluggable search tool complementing web_fetch.

## 05 · the loop doctrine

One cycle, start to push

evaluate→ queue→ implement · glm→ validate · kimi + mutation testing→ merge→ push

The evaluator reads the corpus — events, ledgers, specs, code — and writes the assessment plus the next queue rows, each with its own spec before any child is dispatched. Children implement in git worktrees, one at a time, under iteration and token budgets; glm-5-3-flash implements. Then validation is always kimi-k3 — a different model family, so the verdict is an independent second opinion, and it is adversarial by design: gates re-run from a clean checkout, mutants are injected to prove the tests can actually fail, weak pins get called out. A FAIL verdict is not a crisis; it is the pipeline working. Only then: merge, flip the queue row, push. The doctrine is written down, versioned, and pinned by tests — the loop edits its own rules in the open, never by vibes.

The full protocol — phases, budgets, recovery recipes, hard rules — is one file: LOOP-SPEC.md.

## 06 · get started

Clone, build, chug

Everything below is the repo's own quickstart, verbatim. Put a check: line in your spec — goal_complete is only accepted when the check exits 0.

1 · INSTALL

$ curl -fsSL https://chug.sh/install.sh | sh

Prebuilt binaries for macos-arm64, linux-x86_64 and linux-aarch64 — detects your platform, verifies the sha256, installs to ~/.local/bin (never sudo). Read the script first if you like; the loop cuts v* releases from its own wrap. Or build from source ↓

2 · BUILD

$ git clone https://github.com/tampajohn/chug
$ cd chug
$ cargo build
$ cargo install --path .   # puts the chug binary on PATH (~/.cargo/bin)

Zero setup if you have Claude Code configured: chug reads its auth from ~/.claude/settings.json when the process env doesn't have it.

3 · RUN SOMETHING REAL

$ chug run --spec SPEC.md --goal "Build X and make the check pass" \
    --model anthropic-system.ai.kimi-k3 --max-iters 40 --max-minutes 120

$ chug run --tui ...   # same, with the live dashboard
$ chug run --resume    # continue an aborted run from .chug/transcript.jsonl
$ chug chat            # interactive mode (TUI)
$ chug ledger          # print current LEDGER.md

4 · RUN THE LOOP ON ITSELF

$ nohup ./loopd.sh > /dev/null 2>&1 &   # start (detached)
$ ./loopd.sh status                     # liveness + recent cycle activity
$ ./loopd.sh stop                       # exits after the current cycle

LOOP-SPEC.md is the one-command self-improvement loop; the loopd supervisor runs it continuously — cycle after cycle, no human per-phase prompting. The live count is in the stats band.

5 · DEVELOP

$ cargo build && cargo clippy --all-targets -- -D warnings && cargo test

All three must stay green — the loop's own gates enforce the same bar on every change it lands.