## autonomous coding harness

chug

Given a spec and a goal, it keeps on chugging.

design rule #1 — the loop is code, not conversation. the model never decides whether to continue: the driver does.

live queue · tests · cycles → self-improving in public adversarial validation built in
$ chug run --spec SPEC.md \ --goal "Build X and make the check pass" chug <version> (<commit>) cwd=/srv/app spec=SPEC.md model=… ▪ iter 01 read SPEC.md · seeded LEDGER.md ▪ iter 07 edit_file src/api.rs · bash: cargo test ✓ ▪ iter 12 chug: budget low — 8 iteration(s) remain ✓ goal_complete accepted — check exited 0
one run, sketched — the real driver appends every event to .chug/events.jsonl

## 01 · how it works

The loop is code, not conversation

A run takes a spec file with a check: line and a goal, and drives a fixed loop until the check exits zero. Goal and progress live in files on disk, re-read every iteration — so context trimming can never kill the run.

SPEC + GOAL

SPEC.md with a check: line, plus --goal. Plain files — the driver re-reads them every iteration.

→

DRIVER LOOP

Continue/stop is the driver's call, never the model's. Anti-stall kick, stuck tripwire, budget-low warnings, same-cwd driver lock.

→

THE TOOLS

read_file, write_file, edit_file, bash, grep, tgrep, glob, list_dir, update_ledger, decision_log, delegate, web_fetch, goal_complete — sandboxed to --cwd.

→

GOAL_COMPLETE

Verified, not taken on faith: the spec's check: re-runs and the claim is accepted only on exit 0.

✗ check fails → keep chugging

The failure goes back to the model with the loop still in charge. It fixes, re-runs, re-claims. Rejected claims never end the run — the budget does, and a budget death prints a resume line and leaves the ledger intact.

LEDGER.md → external memory

Updated every iteration, injected into every turn, so transcript trimming never loses progress. A fresh run archives and reseeds it; --resume keeps it.

## 02 · proof it runs itself

Built by its own loop, measured by its own gates

chug improves chug. The loopd supervisor runs the self-improvement cycle back to back, and at every wrap the synced stats band below is regenerated from the repo's own ledgers — done rows in the queue ledger, gate counts in commit messages, the git log. Numbers live there, quoted and never rounded up; this section is the story behind them.

The tgrep arc: validation round after round

When the loop built its own token-budgeted context search, the adversarial validator refused to pass it — again and again: a merge-radius contradiction, symbols mode dropping declarations, vacuous pins, then one flaky perf pin with every mutant red. Each refusal forced a fix-up round; the final round closed it with a fresh mutant batch, all dead. A full artifact set came out of it, and the post-merge gates ran green — the counts are in the stats band.

The bug the pipeline existed to catch

Building its own hooks system, the implementer passed gates — and the validator still failed it. First refusal: PostToolUse was firing on calls that PreToolUse had vetoed, a real semantic bug no test had pinned. The fix-up swept the whole class — the sibling risk-gate-block path had the same hole — with killing tests proven red first. Second round: PASS, zero blocking findings.

Dogfood, verbatim

Its interactive chat mode: written for it, by it. Its delegate tool — a Rust module of its own, launching child chug runs in git worktrees: written by a child run, validated by another model family, then grown through wait_secs long-polling, resume, collect, and terminal-wait mode across later cycles. Its decision_log, hooks, plan mode, TUI, MCP client, Langfuse tracing: same story. And this very page — the one you are reading — was written by chug running its own loop, with this site's check gate re-running on every claim.

Honest failure books: one cycle hit its iteration ceiling mid-wrap and the wrap was lost — the next cycle reconstructed it from git history and ledgers alone. Failed validation rounds are the norm, not the exception; the tgrep arc alone logged refusal after refusal before its pass.

→ live counts: queue, tests, cycles — rewritten at every wrap

## live stats · rewritten by scripts/site-sync.sh at every cycle wrap

106/106queue items landed — done rows in TODO.md, 106 item rows total
870tests green at the newest full-suite gate count in a commit message (cb12234)
41cycles completed by the loopd supervisor — "cycle OK" lines in .chug/loopd/loopd.log
2026-09-28last EVALUATION.md write (git log, chug repo)

Last five items landed — item · ref · date, quoted from the chug repo's git log:

## 03 · the timeline

Seven days, first commit to chug.sh

Every date and hash below is quoted straight from git history — the harness repo and this site's own repo. The loop wrote this page too, so this timeline is the loop reading its own commit log back to you.

2026-09-20

Day one — the core harness f911488

First commit at 17:42: a driver-in-code loop, LEDGER.md as external memory, and goal completion that only counts when the check exits zero. Thirty-one commits land before midnight.

2026-09-20

The TUI, eight minutes later 76ac019

A ratatui dashboard at 17:50 — live activity stream, LEDGER panel, steering input, operator abort.

2026-09-20

The risk gate db6fea6

glob, list_dir and replace_all join the tools, and laya starts classifying every bash command before it runs — destructive calls are blocked with an error the model can route around.

2026-09-20

Chat mode, dogfooded fb842f2

Interactive mode lands with the claim in the commit message itself: "dogfooded by chug itself".

2026-09-20

The meta-loop era begins 7deb7f0

META-SPEC writes down chug-orchestrating-chug; SELF-SPEC makes improvement continuous. That evening "muse implements + kimi validates" becomes the default round policy (0611c33).

2026-09-21

MCP — stdio, then streamable HTTP 3386d2b

SPEC-7 merges: config discovery, handshake, tools/list registry, seven review fixes. The streamable-HTTP/SSE transport (SPEC-9) lands the same day — remote servers appear as mcp__name__tool too.

2026-09-21

Langfuse observability 25a36c1

SPEC-8 merges: a trace per run, a generation per LLM call, spans and scores per tool call. Fire-and-forget — telemetry never changes run behavior.

2026-09-21

The doctrine becomes code 672d04a · aac3629 · 3c795b3

Mutation testing enters the validation template — every landing must now prove its tests can fail. LOOP-SPEC writes the whole evaluate→queue→implement→validate→merge→push loop into one file. glm-5-3-flash takes the implementer seat; kimi-k3 keeps the verdicts.

…and 81 earlier milestones (T1–T88)

2026-09-25

loopd goes continuous 548d494

The supervisor lands after lunch and starts its first cycle the same minute — the first line of the loopd log. Cycles run back-to-back, each wrap the next cycle's input. It is still running today.

2026-09-25

delegate — chug spawns chug 1012dca

T23: launch and status for bounded child runs in git worktrees. Later cycles grow wait_secs long-polling, zombie reaping, resume, collect, and terminal-wait mode.

2026-09-26

Plan mode 87fe53f

chug plan (T73): read-only exploration under a five-tool contract — read_file, grep, glob, list_dir — with submit_plan as the single write and exit path.

2026-09-27

Per-phase model routing c1daaca

T81: glm orchestrates routine cycles, kimi takes the judgment calls — a mechanical freshness predicate in loopd.sh, never model vibes.

2026-09-27

Hooks ccb828a

PreToolUse veto + PostToolUse advisory (T83). The first validation round FAILED the merge — it caught PostToolUse firing on calls that never executed, a real bug — and the class sweep fixed the sibling risk-gate path with it.

2026-09-27

delegate status terminal-wait mode (wake only on goal/abort/liveness/deadline) + LOOP-SPEC adoption (d2b402a)

2026-09-27

Permissions e9afed9

The deny-list lands (T90): .chug/permissions.json, first-match-wins, fail-closed on match, first in the policy chain before hooks and the risk gate. Its FEATURES.md roadmap row flips to LANDED.

2026-09-27

LOOP-SPEC impl-child template --max-iters 50→65 (9db86bc)

2026-09-27

permissions mcp__ matcher-fit canary accepts server-specific globs (aeb12ea)

2026-09-27

F5 phase 1: read_file image input (base64 content blocks + endpoint-reject degrade) (ae7ff5f)

2026-09-27

get_str error names received keys (alias self-correction) (06f3b9e)

2026-09-27

site-sync: chug.sh stats update at wrap (deterministic) (8721c83)

2026-09-27

README Development layout gains permissions + module set-equality guard (e80b3c5)

2026-09-27

META-META-SPEC spec quality bar: check filter must run every test the change adds (962830d)

2026-09-27

README delegate paragraph → per-action sub-bullets (b0c6041)

2026-09-27

site-sync v2: timeline + feature-grid regions (deterministic) (d13a253)

2026-09-27

GitHub releases: tag-triggered prebuilt binaries + notes (9d1182a)

2026-09-27

site-sync timeline ordering + curation bugs (bcd0b66)

2026-09-28

LOOP-SPEC impl-child template --max-iters 65→80 (ffebdaf)

2026-09-28

delegate launch asserts spec + cwd exist (fail-fast on corrupted fields) (80d4a14)

2026-09-28

F6 phase 1: session fork slots (chug fork save/list/restore) (dc29137)

2026-09-28

driver.rs test-module family split (pre-declared trip line crossed again) (ee3943e)

2026-09-28

README Install names the pending first release (quickstart truth) (e5cdebd)

2026-09-28

LOOP-SPEC child goal template: worktree-discipline commitment clause (793a0fc)

2026-09-28

F7 phase 1: streaming responses + console text deltas (2a51cc5)

2026-09-28

delegate.rs test-module family split (T104-shaped, pre-emptive) (75025c9)

2026-09-27

The move to K7 — always-on com.tampajohn.chug-loopd.plist

loopd moves under launchd on the operator's host (K7HC2K125R): the plist wraps loopd.sh in caffeinate -dims with RunAtLoad. The improvement loop no longer lives in a terminal.

2026-09-27

chug.sh — this site e998d45

chug builds its own public site: eight commits in one evening, every section landed against the verify.sh check gate, GitHub Pages serving a CNAME. You are reading the dogfood.

The quiet patch stays on the chart: zero commits for two days between the last early-cycle landing and loopd's first start. Honest books include the blank pages.

## 04 · feature grid

What's in the box

Every capability below is quoted from the repo's README and FEATURES roadmap — including the phases still deferred with written reasons, because a harness that keeps honest books about what it hasn't done yet is the whole point.

builtin tools

read_file write_file edit_file bash grep tgrep glob list_dir update_ledger decision_log delegate web_fetch goal_complete — all paths sandboxed to the run's working directory.

delegate sub-agentslanded

Launch, observe, and collect bounded child chug runs in git worktrees: launch / status / collect, long-poll wait_secs, terminal-wait mode, and resume of an aborted child.

plan modelanded

chug plan explores the repo and drafts an implementation plan with a read-only contract: exactly five tools plus submit_plan, its only write and exit path.

hookslanded

.chug/hooks.json policy-as-config: PreToolUse hooks can veto a tool call before it executes; PostToolUse hooks advise. Fails open, per-checkout, never fires a hook from a hook.

permissionslanded

.chug/permissions.json deny rules evaluated first in the policy chain — before hooks and the risk gate; fail-closed on match, first-match-wins. The first phase is landed; ask-mode and settings.json unification are deferred phases with written reasons.

risk gate (Laya)

Every bash command is classified destructive / risky / safe by a local Laya judge server before it runs; destructive is blocked with an error the model can route around. Verdicts logged to .chug/risk_verdicts.jsonl.

MCP servers

Consumes tools from MCP servers over stdio or streamable HTTP with Claude Code-compatible config. Server tools appear as mcp__name__tool alongside the builtins; per-server fail-soft.

Langfuse observability

Optional tracing to self-hosted Langfuse v3: a trace per run, a generation per LLM call, a span per tool call, outcome and iteration scores. Fire-and-forget — off means zero cost, telemetry never changes run behavior.

TUI

--tui: live activity stream, LEDGER.md panel, and a status bar with model, iteration/budget, elapsed time, and cumulative tokens. i steers, q aborts gracefully.

web_fetch

Read-only HTTP(S) GET with hard bounds: 5 redirects, 10s connect / 30s total, output capped and HTML stripped to text, binary types refused. Bounded and audited where raw curl is neither.

tgrep

Token-budgeted ranked context search: ranked match clusters with tight context windows, quoted phrases, a symbols mode for Rust signatures. Deterministic scoring — no embeddings, no LLM.

chat

Interactive TUI sessions: @file attachments, tab autocomplete for commands and paths, slash commands, esc to interrupt. Steering notes pass into autonomous runs at iteration boundaries.

decision_loglanded

Every loop judgment — validation verdicts, routing calls, recoveries — emits a structured record with class, inputs, options, choice, and a confidence score, building the corpus for confidence-gated routing.

loop hardening

Driver lock against concurrent runs, transcript trimming that keeps the cached prefix byte-stable, jq-mineable events.jsonl, budget-low warnings, a stuck tripwire, and an anti-stall kick.

Image inputlanded

read_file on png/jpg returns image blocks (vision); chat accepts pasted/dragged screenshots. UI work needs eyes.

Session forkin-flight

Clone transcript+ledger at iteration N into a new session id; explore two approaches from one state.

Streaming UXlanded

Text deltas to sinks as they arrive (TUI live typing, headless progress); watchdog gets byte-level liveness for free.

Structured todo toolin-flight

todo_add/update/list driver-visible tools (statuses enforced) as an alternative to freeform LEDGER edits — the orchestration ledger becomes queryable.

Slash-command packsin-flight

.chug/commands/*.md repo-local commands invocable from chat (/review, /triage) and as run goals. Community-extensible without code.

chug as MCP serverqueued

chug mcp-serve: expose run/delegate/status as MCP tools so Claude Code, the bridge fleet, or another chug can drive it. The fleet primitive.

MCP resources+promptsqueued

Consume MCP resource/prompt capabilities (today: tools only).

Web searchqueued

Provider-pluggable search tool complementing web_fetch.

## 05 · the loop doctrine

One cycle, start to push

evaluate→ queue→ implement · glm→ validate · kimi + mutation testing→ merge→ push

The evaluator reads the corpus — events, ledgers, specs, code — and writes the assessment plus the next queue rows, each with its own spec before any child is dispatched. Children implement in git worktrees, one at a time, under iteration and token budgets; glm-5-3-flash implements. Then validation is always kimi-k3 — a different model family, so the verdict is an independent second opinion, and it is adversarial by design: gates re-run from a clean checkout, mutants are injected to prove the tests can actually fail, weak pins get called out. A FAIL verdict is not a crisis; it is the pipeline working. Only then: merge, flip the queue row, push. The doctrine is written down, versioned, and pinned by tests — the loop edits its own rules in the open, never by vibes.

The full protocol — phases, budgets, recovery recipes, hard rules — is one file: LOOP-SPEC.md.

## 06 · get started

Clone, build, chug

Everything below is the repo's own quickstart, verbatim. Put a check: line in your spec — goal_complete is only accepted when the check exits 0.

1 · INSTALL

$ curl -fsSL https://chug.sh/install.sh | sh

Prebuilt binaries for macos-arm64, linux-x86_64 and linux-aarch64 — detects your platform, verifies the sha256, installs to ~/.local/bin (never sudo). Read the script first if you like; the loop cuts v* releases from its own wrap. Or build from source ↓

2 · BUILD

$ git clone https://github.com/tampajohn/chug
$ cd chug
$ cargo build
$ cargo install --path .   # puts the chug binary on PATH (~/.cargo/bin)

Zero setup if you have Claude Code configured: chug reads its auth from ~/.claude/settings.json when the process env doesn't have it.

3 · RUN SOMETHING REAL

$ chug run --spec SPEC.md --goal "Build X and make the check pass" \
    --model anthropic-system.ai.kimi-k3 --max-iters 40 --max-minutes 120

$ chug run --tui ...   # same, with the live dashboard
$ chug run --resume    # continue an aborted run from .chug/transcript.jsonl
$ chug chat            # interactive mode (TUI)
$ chug ledger          # print current LEDGER.md

4 · RUN THE LOOP ON ITSELF

$ nohup ./loopd.sh > /dev/null 2>&1 &   # start (detached)
$ ./loopd.sh status                     # liveness + recent cycle activity
$ ./loopd.sh stop                       # exits after the current cycle

LOOP-SPEC.md is the one-command self-improvement loop; the loopd supervisor runs it continuously — cycle after cycle, no human per-phase prompting. The live count is in the stats band.

5 · DEVELOP

$ cargo build && cargo clippy --all-targets -- -D warnings && cargo test

All three must stay green — the loop's own gates enforce the same bar on every change it lands.