Hands bounded coding work to cheaper worker agents: verified answers, exact diffs, cost receipts.
Cheaper and faster: hand the reading, fixes, tests and reviews to workers that run in parallel. Keep the decisions.
Model-agnostic · Verified answers · Receipts, not vibes · OpenCode / Codex / Claude Code / Gemini CLI workers
Quick start · Dashboard · Benchmarks · Workflow · Skills · Commands · Workers · Responsibility · Write an adapter
Your agent decides · Cheap workers do the typing · A second model reviews
pitroom dash, with sample data
Claude Code, Codex and friends spend premium tokens reading: grepping, opening files, scanning code they'll never quote.
Pitroom goes one step further:
Your agent stays in the driver's seat. A pit crew of cheap workers does the legwork in parallel, while your agent keeps going, and hands back only the answer, the exact diff and a receipt.
Your model, not oursNo model is hardcoded. Pitroom uses your worker CLI's default model: a free tier, a local MLX/Ollama model or a cheap paid one. Switch there and Pitroom follows. |
Any worker CLIOpenCode, Codex CLI, Claude Code and Gemini CLI. One group can mix them, and a fallback chain can cross them (how adapters work). |
Verified answersEvery Your agent trusts what checks out and skips re-reading it. |
Safe by constructionNo commit, push, reset or bulk delete. A git guard stops the tricks. In isolate mode nothing reaches your tree until you apply it. |
pitroom savings --card card.svg makes a shareable card.fallback) and a run that hits "model not found", 429 or quota errors moves to the next model automatically. Still no model hardcoded: the chain is yours.pitroom crew starts several workers at once, pitroom watch --json streams one line per change so your agent follows them live, pitroom wait -g collects every answer, pitroom apply -g lands isolated patches in order. Up to 20 workers run at once by default (maxParallel, 30 at most), the rest queue, and the fallback chain absorbs free-tier rate limits; parallel writers are refused.cheap tier, OpenCode's free model; standard (Codex) and capable (Claude Code) are for tasks that need them, and an optional review tier names who reviews. Gemini CLI with a free Google AI Studio key is one more cheap worker (small daily quota per model). pitroom models shows what each worker offers, the effort levels (low … xhigh) each model accepts, what you say it costs (costs) and what it used; --effort sets the level per run. Details: Models, costs and effort.pitroom revert <id>.sudo, no publishing, no secrets. A worker cannot land a deletion either: pitroom apply refuses a patch that deletes files unless your agent passes --allow-delete.--isolate runs the worker in a private copy of your current state (its own repository, sharing your objects read-only), uncommitted and untracked files included, then hands you a patch: pitroom apply <id> (checked, refuses on conflict) or pitroom discard <id>..pitroom/ folder, no .gitignore edits, no branches. Records live in ~/.local/state/pitroom.using-pitroom, pitroom-research, -crew, -implement), a CLI, and a Claude Code / Codex plugin whose session-start hook loads the workflow. Claude Code, Codex, Gemini CLI, Cursor, or anything that can run a shell command.pitroom dash opens a local page of every run (the task, what the worker did step by step, its result and diff) with history and statistics, for the agent apps that show neither hooks nor a status line; pitroom history searches everything Pitroom ever ran. See Dashboard.Each run gets the worker's own safety mechanism set to the mode (for OpenCode, a permission profile injected through OPENCODE_CONFIG_CONTENT): no commit/push/reset/checkout/stash/clean/rebase, no bulk deletes, no sudo, a file-reading tool that refuses .env and key files (OpenCode and Claude Code workers), no web tools unless you pass --web, no subagents, no recursive delegation. Only allow/deny rules, so a headless run never stalls on a prompt. A git guard on the worker's PATH also stops sh -c "git push", env git reset, aliases and scripts from committing, pushing, resetting, stashing or touching your index. Your opencode.json is never touched. These safeguards reduce risk; they are not a security boundary (see Responsibility and SECURITY.md).
A real run on a free OpenCode Zen model: your agent reads ~554 tokens instead of 163,000.
$ pitroom run "Which function builds the read-only permission profile, and which bash commands does it allow?"
pitroom ✔ done · read · 30s · run 20260930-003902-3270
session ses_… · model opencode/muse-spark-1.3-contributor-free
SUMMARY: The read-only permission profile is built by `readProfile()` in src/profiles.ts:89.
Its bash allow-list is `READ_ONLY_BASH` in src/profiles.ts:29-46.
DETAILS:
- Allowed (src/profiles.ts:31-41): `git status*`:31, `git diff*`:32, … `rg *`:40, `wc *`:41.
- Explicit sub-denies: `rg *--pre*`:43, `git *--output*`:44, `git *--ext-diff*`:45.
FILES CHANGED: none
── receipt: worker processed 163k tokens in 6 steps (7 tool calls) · worker cost $0.000
· returned ~554 tokens, 294× compression · est. saved $0.152 vs Claude Sonnet
Every run prints a receipt, and pitroom savings adds them up (--card writes this shareable card; the numbers here are sample data, and a real measurement follows in Benchmarks):
pitroom dash --detach prints the address of a live page on 127.0.0.1 (this machine only; it reads your runs and changes nothing but the two actions below). It is for the places where hook messages and status lines do not reach, such as the Claude Code and Codex apps: open it in a browser or in the app's own browser pane. The skills tell your agent to start it and give you the address when it runs workers in the background.
The screenshots below show sample data (an imaginary shop-api project), not a real one.
Live. What is running now, and the latest runs. Running cards show a timer, the worker's last words and its animated pixel mascot (a finished card shows the worker's logo); finished ones show the result and, for reviews, the findings.
Click a card and it grows into a panel with the task, what the worker did step by step (files read, searches, commands, edits, with times, failed ones marked), its result, the files and diff it changed, and the details (model, tokens, cost, fallbacks, reference check).
Audits and skipped workers show on the card. A research answer that another worker re-checked (Audits) carries an audited · agrees / partly agrees / disagrees badge and, when it disagrees, the claims it disputes; a worker left out because its quota ran out is listed under the details ("Skipped … cooling down until 14:05", see pitroom cooldown).
History searches everything Pitroom ever ran (the full text of tasks, answers and steps), with filters by state, model and period; each run opens as the same card as on Live. Stats shows runs, success rate, time, tokens and savings, per worker and model and per day.
Stop and Discard. A running card has a Stop button, and a finished --isolate card that is not applied yet has Discard (the same as pitroom stop and pitroom discard; the patch file stays). Each asks once more. Nothing else can be done from the page: it never starts a run and never applies a patch. A POST needs the secret the page was served with (it changes every time the dashboard starts) and the page's own origin, so a web page on another site cannot press them.
The dashboard is a React app built once into dist/ui and served as two static files, so the CLI still has no runtime dependencies. It follows your system's light or dark theme (a button switches it).
Measured on a real, public repository: PI-Desktop at commit c2bfe35, 2,349 tracked files and about 419,000 lines of TypeScript and Rust. Eight questions of the kind an agent asks before it changes code, run as one read-only crew (pitroom crew, at most four workers at a time) on OpenCode's free muse-spark-1.3-contributor-free model.
In short: on that repository eight questions took 156 s of wall clock, and the workers read 5.5 million tokens and handed the agent 7,684 (715× less); a later audit found some of the answers incomplete. On React, Django and Kubernetes, four free models matched gpt-6-sol on nine bounded questions, and in a planted-bug review test almost every model found almost every bug. The tables, and what they do not show, are below.
| Question | Worker time | Tokens processed | Returned to the agent | Compression | Refs verified |
|---|---|---|---|---|---|
| Where is the plugin runtime, and which functions activate and deactivate a plugin? | 37 s | 405,380 | 767 | 529× | 10/11 |
| When does the agent runtime compact a conversation? (function and exact condition) | 90 s | 1,250,719 | 763 | 1,639× | 17/17 |
| Where is session state persisted to disk in the Rust host? | 71 s | 854,969 | 1,426 | 600× | 29/45 ¹ |
| Which IPC channels expose window management to the renderer? | 32 s | 165,180 | 716 | 231× | 3/28 ¹ |
| How does a tool call travel from the agent runtime to the Rust host? | 119 s | 1,818,524 | 866 | 2,100× | 14/15 |
| Which packages and apps depend on the i18n package? | 33 s | 271,297 | 656 | 414× | 15/16 |
| Where is auto-update implemented, and what does the Windows portable target change? | 29 s | 215,693 | 743 | 290× | 20/21 |
| Which tests cover the RPC layer? | 59 s | 511,209 | 1,747 | 293× | 6/7 |
| All eight | 156 s wall clock | 5,492,971 | 7,684 | 715× | 114/160 |
Workers read 5.5 million tokens of code and docs (96 steps, 179 tool calls) and handed the agent 7,684 tokens, about 960 per answer. The eight runs add up to 470 s of worker time but finished in 156 s of wall clock, because four ran at a time (the limit then; the default is now 20): about 3× faster than one after another. That compares the same workers with themselves, not with your agent doing the reading alone, which I did not measure. pitroom savings estimates $3.86 saved for these eight at Claude Sonnet list prices.
Were the answers right? Checked by hand against the repository: 46 key claims across all eight answers (the cited line and what it says) were all correct, and two completeness checks held (exactly 6 invoke and 3 event window channels; no tests in config_sync_rpc.rs). The i18n answer lists the 12 source files that import the package and leaves out 11 test and fixture files that do too, which the question did not ask about. Pitroom's own checker marks fewer references as verified than that: ¹ those two answers cite bare file names (transcripts.rs:651) that it cannot resolve to a path, and it counts them as unverified although the lines are right.
An audit found more than the hand check did. Each of the eight answers was later re-checked by a second free worker (pitroom audit, on space-bunny): 3 agreed and 5 were marked partly right. The disputes are about answers that are incomplete or slightly off, not about the lines they cite: lists that leave things out (session state: more places that write it; i18n: a second, deep-import style of consumer; RPC tests: 14 tests missing), counts that are off, and a few cited lines some lines away from the declaration. One was checked by hand and the auditor was right: the RPC mod tests block holds 59 tests, the answer said about 40. The others were not checked by hand, and an auditor is itself a free model that can be wrong. So treat a worker's list or count as a good first answer that your agent confirms when it matters, which is what --audit is for.
A change in an isolated copy: "add a one-line comment above resolveUpdateMode" finished in 16 s and 4 steps (142,491 tokens). The patch was one file and one line, the comment matched the code, and your working tree is untouched until you apply.
gpt-6-sol on React, Django and KubernetesFree models against gpt-6-sol, on three large repositories. The same kind of work on React (7,252 files), Django (7,091) and Kubernetes (31,353 files), at pinned commits. Nine bounded questions, three per repository: the file and line that define a function, how many files contain a word, and which files contain it. Every answer is checked against git grep, not by opinion: an exact path:line (half a point for the right file on the wrong line), an exact count, and for the lists an F1 score that punishes both missed and invented files. Each worker ran the nine questions once, read-only, with no fallback, so a model that cannot do it fails instead of being swapped for another.
| Worker and model | Score | Median time | Median tokens |
|---|---|---|---|
Codex gpt-6-sol (effort medium) | 100% | 16 s | 50k |
Codex gpt-6-sol (effort low) | 100% | 20 s | 50k |
OpenCode muse-spark-1.3-contributor-free | 100% | 23 s | 44k |
OpenCode space-bunny-free | 100% | 35 s | 63k |
OpenCode mimo-v2.6-flash-free | 100% | 37 s | 42k |
OpenCode big-pickle | 100% | 42 s | 91k |
OpenRouter nemotron-3-ultra-550b-a55b:free | 89% | 27 s | 42k |
OpenCode longcat-2.5-preview-free | 89% | 29 s | 20k |
OpenCode nemotron-3-ultra-free | 89% | 35 s | 21k |
OpenRouter inkling:free | 83% | 14 s | 41k |
OpenCode nemotron-3.5-lightning-free | 72% | 377 s | 69k |
NVIDIA gpt-oss-20b (7 questions) | 57% | 24 s | 42k |
Four free models matched gpt-6-sol on these questions. Every model found every definition; the misses are count questions (a wrong number) and, for two models, list questions, one of them by adding docs/ files that are not under django/. gpt-6-sol was as fast as the quickest free models, and lowering its effort did not cost accuracy here. nemotron-3.5-lightning-free wrote an unrelated text for one list question and ran into the 10-minute limit on another. The scoring harness and every raw answer are in benchmarks/multi-repo.
Only workers that ran are in the table. Models that could not answer at all (a provider error, or the shared daily quota of OpenRouter's free tier) are left out, and so are the two gpt-oss-20b runs that ended in a provider error; a timeout is kept.
Does a review find a planted bug? A bug was changed into real code (an inverted check, a swapped &&/||, a flipped return) in React, Django, Kubernetes and Pitroom itself, committed as a bare "tidy", and pitroom review was asked to review the commit. 29 packages (14 with the bug alone, 15 with the bug and three comment-only edits around it). A review counts as a find when a Critical or Important finding names the changed file and either cites a line within 5 or names the changed identifier or its function; "exact line" is the strict version (within 3).
| Worker and model | Bug found | Exact line | Reviews | Mean time | Slowest |
|---|---|---|---|---|---|
OpenCode space-bunny-free | 100% | 100% | 29 of 29 | 2 min 19 s | 5 min 46 s |
OpenCode nemotron-3-ultra-free | 97% | 79% | 29 of 29 | 1 min 45 s | 3 min 29 s |
OpenCode longcat-2.5-preview-free | 83% | 83% | 29 of 29 | 2 min 53 s | 8 min (1 timed out) |
Codex gpt-6-sol | 100% | 100% | 9 of 29 | 38 s | 52 s |
OpenCode muse-spark-1.3-contributor-free | 100% | 57% | 14 of 29 | 48 s | 2 min 0 s |
OpenCode big-pickle | 100% | 92% | 13 of 29 | 2 min 28 s | 5 min 27 s |
OpenCode mimo-v2.6-flash-free | 100% | 100% | 10 of 29 | 4 min 6 s | 7 min 40 s |
Almost every model found almost every bug, so this test separates them on speed and precision, not on whether they can review: muse-spark described the right bug but often quoted a wrong line number (57% exact), and gpt-6-sol and muse-spark were the fastest by a wide margin. Only three models finished all 29: the other four are scored on the reviews that ran before a provider limit (the free tier's usage limit, and the Codex plan's, which gpt-6-sol reached), and the rest are left out, not counted as misses, so their rows rest on 9 to 14 reviews. Every change is only 2 to 5 lines, which is easy; there is no false-alarm rate (the comment-only "clean" packages turned out to contain comments that were really wrong, so they cannot show one); and React, Django and Kubernetes may be in a model's training data. The harness and every raw review are in benchmarks/review-bugs.
What this does not show
grep can answer, so they do not show how a model handles design questions or large edits, and four free models and gpt-6-sol all scoring 100% says the test is easy at the top, not that they are equal.docs/") did not finish: it was stopped after 19 minutes and 57 tool calls. Give workers bounded tasks.To repeat it on your own repository: pitroom crew -d <repo> -g bench "<question 1>" "<question 2>" …, then pitroom wait -g bench, pitroom show <run> and pitroom savings --models.
macOS and Linux are the main platforms. Windows runs the same test suite in CI (Windows Server, Node 22 and 24, with the fake workers; the worker CLIs themselves are not run there), with these differences:
sh script and is not available on Windows; pitroom doctor says so. Read-only runs are still limited by each worker CLI's own rules and --isolate keeps edits in a copy, so prefer --isolate and check the patch before apply.pitroom.cmd (%USERPROFILE%\.local\bin), started with the Node that ran pitroom install or the one in PITROOM_NODE; it does not search for another Node.core.autocrlf off (or an eol=lf in .gitattributes, as this repository has); a checkout converted to CRLF has not been tested with apply.--verify command runs in cmd.exe, which reports a missing command as is not recognized.You need Node.js 22.13+ and at least one worker CLI: OpenCode v2+ (the default worker, with free models), Codex CLI, Claude Code or Gemini CLI (beta; it needs an API key from Google AI Studio, a Google account sign-in no longer works for it).
Claude Code
claude plugin marketplace add ATASTECH/pitroom
claude plugin install pitroom@pitroom
Inside a session: /plugin marketplace add ATASTECH/pitroom, then /plugin install pitroom@pitroom.
Codex
codex plugin marketplace add ATASTECH/pitroom
codex plugin add pitroom@pitroom
npm i -g pitroom
Gemini CLI
gemini extensions install https://github.com/ATASTECH/pitroom
npm i -g pitroom
Cursor
npm i -g pitroom
pitroom install --mcp --client cursor
This adds pitroom mcp to ~/.cursor/mcp.json and leaves the other servers in that file alone; restart Cursor if it does not show up. Cursor gets the ten MCP tools described below. The skills are linked into ~/.agents/skills and ~/.claude/skills (--no-skills skips that), and Cursor's documentation lists both as places it loads skills from; we have not run that in a Cursor session, and a Cursor forum report says its CLI (cursor-agent) does not load ~/.agents/skills. pitroom install --mcp --dry-run shows what it would change first.
Any agent that speaks MCP (Claude Desktop, and the ones above too)
Pitroom is also an MCP server on stdio: ten tools instead of shell commands and skills (pitroom_run, pitroom_wait, pitroom_show, pitroom_info, pitroom_review, pitroom_audit, pitroom_apply, pitroom_discard, pitroom_revert, pitroom_stop), the runs as resources, and four prompts. The tool definitions sit in the client's context for the whole session, so they are kept small: about 1.9k tokens. It is listed in the MCP Registry as io.github.ATASTECH/pitroom, so clients that browse the registry can find it. After npm i -g pitroom, let Pitroom register itself in the clients it finds (Claude Code, Codex, Gemini CLI, Cursor, Claude Desktop):
pitroom install --mcp # also links the skills; add --no-skills to skip that
pitroom install --mcp --dry-run # only say what it would do
pitroom install --mcp --client cursor,claude-desktop
It registers the launcher by its full path (apps like Claude Desktop start without your PATH), changes Claude Code, Codex and Gemini CLI with their own mcp add command (user scope where it has one), merges the JSON of Cursor and Claude Desktop without touching their other servers (keeping a .bak-pitroom backup, and leaving a file that is not valid JSON alone), skips what is already registered, and pitroom uninstall removes it again. pitroom doctor shows where it is registered. Restart the client afterwards. Or do it by hand:
claude mcp add pitroom -- pitroom mcp
codex mcp add pitroom -- pitroom mcp
gemini mcp add pitroom pitroom mcp
or, for Cursor (~/.cursor/mcp.json), Claude Desktop (claude_desktop_config.json) and others:
{ "mcpServers": { "pitroom": { "command": "pitroom", "args": ["mcp"] } } }
Each tool runs the matching pitroom command, so every rule of the CLI applies unchanged (permission profiles, git guard, isolation, read snapshots). A run is waited for up to waitSeconds (default 50), then comes back as "still running" with its id for pitroom_wait; keep it below your client's tool timeout. The server works in the directory it is started in: the client's for stdio, or the one -d DIR names (Claude Desktop starts it elsewhere, so give it "args": ["mcp", "-d", "/path/to/project"]); the tools' dir argument picks another per call.
"mcpDash": false in the config (or PITROOM_MCP_DASH=0) turns it off.notifications/progress every few seconds ("running · read · 12s · 3 steps, 2 tool calls · last: …"), and a client that resets its timeout when progress comes in (the protocol allows it; not every client does) does not give up on a long run. Cancelling a pitroom_run, pitroom_review or pitroom_audit call stops the run it started (a cancelled pitroom_wait only stops waiting).pitroom://run/<id> (the report) and, for runs that changed files, pitroom://run/<id>/patch (the exact diff), so a client can attach one to a conversation. A client can subscribe to a run (resources/subscribe) and is told when it changes state (notifications/resources/updated), and every client is told when the newest run changes (notifications/resources/list_changed), so it need not poll; over HTTP these arrive on the session's event stream (a GET).research, implement, review and crew say how to use Pitroom for that job (in Claude Code they appear as /mcp__pitroom__research, …), and every skill is a prompt of the same name (using-pitroom, pitroom-research, …), so a client without skills gets the same guidance. The server's instructions tell the agent to use the skills alongside the tools.pitroom_run takes tasks (independent tasks in parallel, one worker each) and continue (a follow-up in the same worker session); pitroom_info reports without changing anything, by topic: runs, history (find an earlier answer before asking again), stats and models (pick a worker), savings, cooldown, config, doctor; pitroom_stop with cooldowns tries rate-limited models again.pitroom mcp --http [--port N] (default 7117) serves the same tools at http://127.0.0.1:7117/mcp for clients that connect to a URL, for example claude mcp add --transport http pitroom http://127.0.0.1:7117/mcp --header "Authorization: Bearer $(cat ~/.local/state/pitroom/mcp-token)" (the command is printed when it starts). It is a local service: it listens on 127.0.0.1 only, needs the bearer token (made on first start, kept at <state dir>/mcp-token with owner-only permissions, or set with PITROOM_MCP_TOKEN), and refuses a Host or Origin that is not local, so a web page cannot use it. Anyone who has the token can run workers as you, so keep it private. It works in the directory it was started in (printed when it starts; -d DIR for another). Each client has its own session; a call that asks for progress is answered as an event stream. Ending a session or stopping the server ends the waiting, not the runs: collect them with pitroom_wait or the CLI. pitroom install --mcp registers the stdio command, which needs no token and no running process.Any other agent (a plain shell) or just the CLI
npm i -g pitroom
pitroom install
| Path | You get |
|---|---|
| Claude Code plugin | The 14 skills, the session-start hook that introduces Pitroom, and a card after each pitroom command. |
| Gemini CLI extension | The 14 skills and a short context file that introduces Pitroom (Gemini asks you to confirm the extension). Gemini CLI also works as a worker, which is separate: see Workers. |
| MCP server | The nine tools above, for any agent that can use MCP servers. No skills: the tools describe themselves, and the server tells the agent how to start. |
| Codex plugin | The 14 skills, a session-start hook that introduces Pitroom (so it is known even when Codex drops skill descriptions because many are installed), and a card after each pitroom command. Codex asks you to trust the plugin's hooks once. The hooks call pitroom, which comes from npm (npm i -g pitroom). |
npm + pitroom install | The CLI, and the skills linked into ~/.agents/skills and ~/.claude/skills. |
Pick one path, not several: doctor warns if the skills load twice. The plugins alone do not put pitroom on your PATH, so install it from npm too if you want to run it yourself (watch, models, savings, the status line).
pitroom doctor --probe
It checks the worker CLIs, models, permissions and skills, and runs one live round trip.
Pitroom is a toolbox, not a procedure: your agent uses it when it helps, or you ask explicitly:
Use pitroom to map how sessions are created and invalidated, then propose a fix. Split this into a pitroom crew: audit auth, billing and uploads for missing input validation.
Or start a worker yourself and read its receipt:
pitroom "Which files read the session cookie? Cite file:line."
pitroom -W gemini:gemini-3.8-flash "…" # the same on another worker (see Workers below)
pitroom savings # what all your runs saved so far
pitroom models # what each worker offers, what it costs you, what it used
# update
claude plugin update pitroom@pitroom
codex plugin marketplace upgrade pitroom
npm i -g pitroom@latest
# uninstall
claude plugin uninstall pitroom@pitroom
codex plugin remove pitroom@pitroom
pitroom uninstall && npm rm -g pitroom
| Flag | For | Your working tree | |
|---|---|---|---|
| read | (default) | research, locating code, reviews | untouched; edit tools are denied |
| isolate | -i | code changes you want to review first | untouched until pitroom apply |
| write | -w | quick in-place edits | edited now; pitroom revert undoes exactly |
Read snapshots. With secret-looking files in the directory, a read run reads a snapshot of your project's current state (uncommitted and untracked files included) without them and without git-ignored files, made from git objects, kept once per state and shared by every read run on it (twenty workers, one snapshot), and cleaned up after three days or by pitroom clean. Nothing is copied back: a read run has nothing to land. Its files are read-only, and the paths a worker cites are turned back into your project's. With no secret-looking files nothing changes: the worker runs in place. --in-place (or "readIn": "project" in the config) reads the directory itself, e.g. for a task that needs build output; "readIn": "snapshot" always uses one.
Answer cache. Ask the same read question on the same code and you get the earlier answer back at once, with no worker run: pitroom ⟲ cached answer · the same question on the same code as run … · no worker ran. "The same code" is exact: the commit and every working file, uncommitted and untracked ones included (git-ignored files are not code), so any edit makes it a new question. Only answers that held up are reused (the run finished, every file:line reference checked out, no audit disputed it) and only for 7 days ("cacheDays" in the config, 0 turns it off, or PITROOM_CACHE_DAYS). Follow-ups, web runs, plan tasks, reviews, audits, --verify runs, parallel work (crews) and changes are never answered from the cache, and --fresh asks a worker anyway (its answer is the one reused from then on). Not part of "the same code": git-ignored files (build output, installed dependencies) and the worker or model that answered; when a question depends on them, use --fresh.
pitroom run "Find every place we build SQL strings by hand; file:line and risk"
pitroom run -i --link node_modules --verify "npm test" "Fix the off-by-one in paginate() and add a test"
pitroom run -w "Rename getUserById to findUserById in src/ and fix imports"
pitroom run --continue last "Now handle the empty-page case too"
Long jobs: --bg returns immediately; pitroom wait <id> blocks for up to 9 minutes (made for agents with 10-minute tool limits; exit code 75 means "call wait again").
Pitroom ships a full development methodology as skills, adapted from superpowers so that its subagents are cheap workers: your agent brainstorms and plans with you, then executes the plan task by task while workers do the typing and a second model does the reviewing.
| Step | Skill | Command |
|---|---|---|
| Design | pitroom-brainstorming | writes the spec to docs/pitroom/specs/ |
| Plan | pitroom-writing-plans | writes the plan to docs/pitroom/plans/, a worker tier on every task |
| Implement | pitroom-driven-development | pitroom run -i --plan PLAN --step N |
| Review | pitroom-review | pitroom review <run> (a read-only reviewer on another model) |
| Fix | pitroom-receiving-review | pitroom run --continue <run> "…" (the re-review sees only the fix) |
| Land | pitroom-driven-development | pitroom apply <run>, your tests, your commit |
| Branch review | pitroom-review | pitroom review --range main..HEAD --plan PLAN |
| Finish | pitroom-finishing | merge, pull request or keep: asked, never automatic |
A review package holds the diff under review with 10 lines of context, and the reviewer's CLI sends it to that worker's model provider. Pitroom does not filter it: secrets committed in a reviewed range go along as they are.
Tiers map plan tasks to workers: "tiers": {"cheap": "opencode", "standard": "codex", "capable": "claude"}. pitroom plan status PLAN rebuilds where a plan stands from the run records (it survives context compaction), and pitroom plan note keeps completions and rulings outside the repo.
Pitroom checks every path:line a worker cites, which proves a line exists, not that the claim about it is true. An audit asks a second worker to verify an answer's key claims against your project and to say AGREE, PARTIAL or DISAGREE, with the claims it disputes.
pitroom audit <run> # now: re-check one read run (-W picks the auditor)
pitroom run --audit "…" # re-check this run's answer when it finishes (--no-audit skips it)
or set "audit": 0.1 in the config (or PITROOM_AUDIT=0.1) to have about one read run in ten audited, in the background, without your agent asking. The same run is always in or out of the sample.
audit tier (else cheap, i.e. a free model if that is your cheap tier). It never delays the run or fails it.pitroom doctor says so). Only read runs are audited: a change has pitroom review.pitroom show carry the verdict and the disputed claims; the Stats tab counts, per worker, how many audited answers were confirmed. Audits are not counted as runs and save nothing.DISAGREE as a reason to look, and AGREE as one more signal.A real crew over this repository, four questions at once on a free OpenCode Zen model:
pitroom crew -g v04-demo \
"Where does Pitroom enforce the maxParallel queue, and how does a waiting run react to pitroom stop? Cite file:line." \
"How does the git guard shim resolve git aliases? Cite file:line." \
"Which fields of OpenCode v2 events does the parser read for token usage, refused tool calls and errors? Cite file:line." \
"How does pitroom apply -g decide which patches to apply and in what order? Cite file:line."
pitroom watch -g v04-demo --json # what the primary agent follows, one line per change:
{"event":"started","run":"20260930-102812-a65d","mode":"read","task":"Where does Pitroom enforce the maxParallel queue, and how d…"}
{"event":"progress","run":"20260930-102812-a65d","steps":3,"tools":4,"last":"grep SIGINT|SIGTERM|stop.*queued|stopped.*queued"}
{"event":"done","run":"20260930-102812-a563","time":"21s","worker":"opencode (default model)","summary":"The shim resolves `alias.<sub>` via `real git config --get` in a loop …","refs":"5/5"}
{"event":"all-done","runs":4,"ok":4,"failed":0}
$ pitroom status -g v04-demo
RUN STATE MODE TIME STEPS WORKER NOW / RESULT
20260930-102812-0c0d done read 30s 5 opencode (default model) Parser `src/backends/opencode/events.ts` reads …
20260930-102812-7642 done read 27s 5 opencode (default model) `pitroom apply -g NAME` applies finished isolat…
20260930-102812-a563 done read 21s 2 opencode (default model) The shim resolves `alias.<sub>` via `real git c…
20260930-102812-a65d done read 39s 7 opencode (default model) maxParallel is enforced by file-based slots (`s…
Humans get the same table live with pitroom watch -g NAME; pitroom wait -g NAME --brief prints one line per worker, pitroom show <run> the full answer.
For changes, pitroom crew -i … gives every worker its own isolated copy of your current state and pitroom apply -g NAME lands the patches in order, stopping at the first conflict. Up to maxParallel workers (default 20, at most 30) run at once, the rest queue; only one --write run per repository is allowed. pitroom stop -g NAME stops running and queued workers.
| Skill | Your agent uses it to |
|---|---|
using-pitroom | see what Pitroom offers and when it pays off (injected at session start by the plugin; optional) |
pitroom-brainstorming | turn an idea into an approved design before any code |
pitroom-writing-plans | write a task-by-task plan with a worker tier per task |
pitroom-driven-development | execute a plan: worker per task, review on another model, fix loop, apply, commit |
pitroom-worktrees | isolate its own branch for the work |
pitroom-tdd | test first, for itself and the workers it briefs |
pitroom-debugging | find root causes, with workers gathering the evidence |
pitroom-verification | run the checks before claiming anything is done |
pitroom-review | get a change, a fix round, a branch or a PR reviewed by another model |
pitroom-receiving-review | weigh review findings with technical rigor |
pitroom-finishing | merge, open a PR or keep the branch, as the user chooses |
pitroom-research | find, map or explain code through a read-only worker |
pitroom-implement | get a one-off change made in an isolated copy |
pitroom-crew | split independent work across parallel workers and merge the results |
pitroom run [-r|-w|-i] [-W WORKER] [-m MODEL] [-d DIR] [-f FILE]… [-t 30m] [--verify CMD] [--link a,b] [--bg] "task"
pitroom run --continue <run|last> "follow-up"
pitroom crew [-i] [-g NAME] "task 1" "task 2" … (or --task-file with --- separators)
pitroom run -i --plan PLAN --step N [--tier T] ["notes"]
pitroom review [run | --range A..B [--plan PLAN]] [--tier T | -W T] [--bg]
pitroom audit RUN [-W worker]
pitroom mcp [-d DIR] [--http [--port N]] # serve Pitroom as MCP tools on stdio, or over HTTP on 127.0.0.1 (bearer token)
pitroom cooldown [--clear]
pitroom plan status PLAN [--json] · pitroom plan note PLAN "Task N: …"
pitroom status|wait|watch [run… | -g NAME] wait: --any --brief --timeout · watch: --json|--brief --interval
pitroom dash [--detach] [--port N] [--open] [--stop] a live page of the runs on 127.0.0.1
pitroom history [TEXT] [--model M] [--state S] [--since 30d] [--limit N] [--json] finished runs, searchable
pitroom history stats [--since 30d] · pitroom history import per worker and model · take in older runs
pitroom show [run] --patch --events --full --json
pitroom apply [run | -g NAME] · pitroom discard|revert [run] · pitroom stop [run | -g NAME]
pitroom ls [--running] [-g NAME] · pitroom clean [--days 14] [--yes]
pitroom savings [--since 7d|30d|all] [--models] [--card file.svg] [--badge]
pitroom models [worker] [--all] [--json] models, effort levels, your costs, your usage
pitroom statusline [--then CMD] · pitroom hook-card
pitroom doctor [--probe] · pitroom config · pitroom init [--model ID] [--fallback A,B] [--yes] [--force] · pitroom install [--copy] [--force] [--mcp [--client A,B] [--no-skills] [--dry-run]] · pitroom uninstall
Exit codes: 0 ok · 1 worker failed · 2 usage · 3 refused/setup · 4 timeout · 5 read-only violation · 6 verify failed · 75 still running.
Every finished run is also written to a SQLite database (history.db in Pitroom's state directory; Node's built-in node:sqlite, nothing to install). It keeps the task, worker and model, time, tokens, savings, the steps the worker took, its answer and the diff, and it is searchable:
pitroom history "login bug" # full-text search over the task, the answer and the steps
pitroom history --state problem --since 7d
pitroom history stats --since 30d # runs, success rate, average time and tokens per worker and model
While a run runs, its raw event stream is a plain file (the simplest thing that survives a crash). When it ends, the stream is compressed, and the stderr log is kept only for runs that did not succeed. pitroom clean removes old run directories but the history keeps them: pitroom show <run> still prints their report, steps, and patch. Runs from before the history existed are taken in by pitroom history import (it also runs on the first pitroom history and pitroom dash). The files remain the source: the database can be deleted and rebuilt with history import for the runs still on disk.
Two settings make every delegation visible, whether or not the agent mentions it. A status line shows running workers and this week's savings (--then keeps your own status line first); a card appears after each pitroom command the agent runs, once per phase (started, finished with its result, applied). The plugin registers the card hook itself; with pitroom install, add both to ~/.claude/settings.json:
{
"statusLine": { "type": "command", "command": "pitroom statusline --then 'your-own-statusline'" },
"hooks": {
"PostToolUse": [{ "matcher": "Bash", "hooks": [{ "type": "command", "command": "pitroom hook-card", "timeout": 10 }] }]
}
}
The Claude Code and Codex apps show neither hook messages nor a status line. For them there are two things that need no setup:
pitroom dash --detach prints the address of a live dashboard with three tabs: Live (every running and recent run as an animated card; click one for the task, what the worker did step by step, its result, the diff and the details), History (search and filter everything Pitroom ever ran; every run opens as the same card) and Stats (success rate, time, tokens and savings per worker and model). It is read-only, listens on 127.0.0.1 only and stops itself after an hour without a request (pitroom dash --stop ends it sooner). Open it in a browser or in the app's own browser pane. The skills tell the agent to start it and give you the address when it runs workers in the background.pitroom watch -g NAME --brief prints one card line when a worker starts and one when it ends. In Claude Code, the agent runs it through the Monitor tool and the lines appear in the app; in Codex the command's output block fills as it goes.Handled for every worker: opencode run blocks forever on an open stdin pipe; ask permissions stall headless runs; default output carries ANSI and banners; and in OpenCode v2 a plain run attaches to the shared background service, where per-run permissions and the git guard would not apply. Pitroom closes stdin, uses allow/deny-only profiles, parses --format json and always runs a private --standalone server.
A worker is named by a target: backend[:model].
pitroom run "…" # preferred worker (config "worker", d
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y pitroomMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-atastech-pitroom": {
"command": "npx",
"args": [
"-y",
"pitroom"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referencePitroom works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.