ddflow

Work-queue kernel for AI coding agents: dependencies, worktree isolation, quality gates, recovery

AI & MLPythonv0.2.0

ddflow

A portable, agent-agnostic work-queue kernel for AI coding agents.

You keep a queue of phases and tasks with declared dependencies. You say "implement phase P2". Independent tasks fan out to parallel agents in isolated git worktrees; dependent ones wait. Every task passes a quality pipeline whose gates cannot be passed by assertion. If an agent crashes, its work is found rather than lost. If everything except the log is destroyed, the project's decision history rebuilds from the log alone.

One dependency beyond python3 and git (Jinja2, for the prompt templates; see Extending it by writing text, not code). Works with Claude Code, Gemini CLI, Codex, Copilot, Cursor, Kimi, opencode, Aider, a CI job, a Makefile, or a human at a terminal — over a CLI and an MCP server that are the same implementation.


Instruction for an agent reader

If you are an AI agent and were asked to read this README, this section is your shorthand. Everything below it is detail.

What ddflow is. A work queue with rules the tools enforce, for projects that agents build. Work is phases and tasks with dependencies and declared file globs. Each task gets a lease and its own git worktree, then passes a gate pipeline: a check that could not run is unavailable, never passed; one reviewer must be from another model family; a bug closes only with a regression test that failed first. Everything is recorded in an append-only log in .ddflow/events/, committed with the code. ddflow writes none of your project's code — you do; it adds only its own config, a rules block and git hooks. Every operation is both a shell command (ddflow …) and an MCP tool (ddflow_…).

If the project already uses it (.ddflow/ exists): call ddflow_brief (or run ddflow brief) first and follow it. It replaces reading the project's rule and lesson files. Then: ddflow_next → ddflow_claim → work only in the worktree it returns → ddflow_gate_status, satisfy each gate → ddflow_merge → ddflow_complete. The full driver is docs/ddflow/drivers/implement-phase.md in that project.

To onboard a project, ask the user first: it installs a tool, registers an MCP server and writes files in their repository.

  1. Install: uv tool install git+https://github.com/delian/ddflow-mcp (not on PyPI yet).
  2. In the project root: ddflow adopt --agents claude (your agent's name; see Wiring it into your agent). It writes .ddflow/, the MCP registration, the commit hook and a rules block in AGENTS.md. Restart the agent session so it loads the MCP server.
  3. ddflow import lists the plans, lessons, decisions and unmerged branches the repository already has, and writes nothing. Run ddflow import --apply only once the user agrees with that list.
  4. ddflow config --set gate.unit_tests.command "<their test command>". Then ddflow companions lists the tools the gates expect that are missing: propose them, never install one unasked.
  5. Put the user's plan in the queue (ddflow phase add …, ddflow task add … --globs …) and finish with ddflow doctor.

Exit codes, everywhere: 0 ok · 1 failure · 2 could not run / nothing to do · 3 refused. Never treat 2 as 0.


Introduction

The problem it solves

An AI coding agent is good at a task and weak at a project. One agent in one session mostly works. Run it for weeks, or run three at once, and the same failures come back:

  • Work disappears. A session crashes or is closed mid-task, and the half-finished change sits in a directory nobody remembers.
  • Agents collide. Two of them edit the same file, and the second merge quietly undoes the first.
  • Checks that never ran look like checks that passed. The linter was missing, the reviewer endpoint was down, the tests were "run" in a summary. The agent reports done, and nothing on record says otherwise.
  • The project forgets. Last week's hard-won lesson, the reason behind a design choice, the bug that was already fixed once — gone at the next session, or at the next context compaction in this one.
  • Nobody can say what happened. Which instruction led to which change, and which review looked at it, lives in a chat transcript that no longer exists.

ddflow is the layer between you and your agents that makes those failures structurally hard rather than a matter of discipline. It does not write code and it is not an agent. It is a queue, a set of rules the tools enforce, and a log of everything that happened.

What changes for you

Without itWith ddflow
You decide what each agent does next, and keep the plan in your head or a chat.The plan is a queue of phases and tasks with dependencies. ddflow next says what can start now and why everything else is blocked.
Parallel agents step on each other.ddflow claim gives each task a lease and its own git worktree; tasks that declare overlapping files are refused, not merged over.
"Done" means the agent said so.Every task passes a gate pipeline you configure. A gate that could not run is recorded unavailable, never passed; at least one reviewer must come from a different model family than the author; a bug cannot be closed without a regression test that failed first.
A crash loses work.ddflow recover finds orphaned worktrees and reports what each holds. It never deletes work.
Every session starts from zero.Lessons, decisions, research verdicts, bugs and your own prompts are recorded as you go. ddflow brief hands the agent the ones relevant to this task in a bounded amount of context, and ddflow recall searches all of it.
History is a transcript.An append-only event log, committed in git. The board, the index and the reports are rebuilt from it; ddflow replay reconstructs the project's decisions from the log alone.

Who it is for

  • One developer with one agent. A plan that survives the session, a memory that survives compaction, and a record of which checks really ran. The queue is useful even with no parallelism at all.
  • Several agents in parallel — subagents, several terminals, several vendors. The dependency graph says which tasks are independent, worktrees keep them apart, and the merge step lands them without anyone switching the main checkout's branch.
  • A team or a CI pipeline. The log is committed with the code, so a fresh clone knows the queue and its history. Read commands have a machine-readable --json form and every command returns the same four exit codes, so a Makefile or a CI job can drive it exactly as an agent does.

It works with the agent you already use, because everything it does is reachable both ways: as a shell command and as an MCP tool. Use whichever your agent, script or CI job has. Your workflow is text, not code — the gate pipeline, the reviewer instructions and the agent-facing prompts are files in your repository that you can edit.

What it is not

  • Not an agent or a model. Your agent does the work; ddflow decides what may start, checks what was claimed, and remembers.
  • Not a hosted service. Everything is files in your repository and a disposable local cache. No account, no server to run beyond the local MCP process.
  • Not a replacement for your tests or CI. It runs the commands you configure and records their real exit codes and output.

A first run

# ddflow-mcp is not on PyPI yet; until the first release, install from the repository:
uv tool install git+https://github.com/delian/ddflow-mcp   # or: pipx install git+https://github.com/delian/ddflow-mcp
cd /path/to/your/project
ddflow adopt                        # registers the MCP server with your agents, writes .ddflow/ and a block in AGENTS.md
ddflow phase add P1 --title "Password reset"
ddflow task add P1.T1 --phase P1 --title "Reset-token endpoint" --globs 'src/auth/**'
ddflow next                          # what can start now, and why the rest is blocked

adopt registers the server you just installed, by its full path: an install that did not come from a package index (from git, a local directory or an archive) carries a direct_url.json in its metadata (PEP 610), and for one of those adopt writes the ddflow-mcp installed beside its interpreter rather than uvx ddflow-mcp, which would fetch from PyPI. Only an install from an index gets uvx. --launch python still forces the interpreter-plus-PYTHONPATH form.

Then tell your agent "implement phase P1". The driver adopt installed tells it to start with ddflow_brief, claim the task, work in its own worktree, satisfy each gate and land the change. ddflow cannot make an agent follow instructions, but it makes skipping them visible: the commit hook adopt installs flags a commit made without a lease (or refuses it, if you set [enforce].commit_without_lease = "block"), and ddflow complete refuses an item whose gates carry no outcome. Watch it with ddflow board, and ask ddflow doctor at any point whether the project is healthy.


How do I…?

Every row is a command you can run in a terminal and a tool an agent can call over MCP — the same implementation, so neither drifts from the other.

I want to…CLIMCP tool
see what the workflow isddflow workflowddflow_workflow
change the workflowddflow workflow pipeline task … · workflow gate <id> … · workflow drop <id>ddflow_workflow_pipeline · _gate · _drop
change any settingddflow config --explain · --set <key> <value>ddflow_configure
add a phase / a taskddflow phase add P1 --title … · ddflow task add P1.T1 --phase P1 --globs 'src/**'ddflow_phase_add · ddflow_task_add
get a plan into the queuesee From plan mode to the queuesame
know what to work onddflow nextddflow_next
start a taskddflow claim <id> → work → ddflow gate … → ddflow merge → ddflow completeddflow_claim, ddflow_gate_*, ddflow_merge, ddflow_complete
see everything about one item or bugddflow show <id> (phase, task or bug id)ddflow_show
see progress / effortddflow progress · ddflow status · ddflow boardddflow_progress · ddflow_status · ddflow_board
find out if we're going in circlesddflow loopsddflow_loops
record a lesson / decision / research / bugddflow lesson add · decision add · research · bug found|fixedddflow_lesson_add · ddflow_decision_add · ddflow_research_add · ddflow_bug_*
search everything the project remembersddflow recall '<regex>'ddflow_recall
check a text against what is already filed (read-only)ddflow similar '<text>' [--kind bug,task,...] [--json] -- exit 0 with candidates, 2 with noneddflow_similar
record what happened this sessionddflow session start|prompt|note|endddflow_session_*
read the engineering logddflow historyddflow_history
check the tooling around the gatesddflow companionsddflow_companions
find work a crashed agent leftddflow recoverddflow_recover
check the project's integrityddflow doctorddflow_doctor
rebuild everything from the logddflow replay --verifyddflow_replay
invoke a workflow / a mode of your ownddflow prompts list · prompts show <name>prompts/list · prompts/get
see what this project left undoneddflow doctor · ddflow statusthe footer on tool results
ask the tool to explain itselfddflow help [topic]ddflow_help

Every read command takes --json. Every exit code means the same thing everywhere: 0 healthy · 1 real failure · 2 could not run / nothing to do · 3 coordination refused. 2 is never collapsed into 0 — "nothing is ready" and "everything is fine" are different facts, and an agent that cannot tell them apart invents work.


Table of contents


Help: what it can do, and the workflow

$ ddflow help                 # what this is, the loop, every capability grouped
$ ddflow help workflow        # workflow · import · gates · parallel · memory · recovery · config

Reachable as ddflow_help over MCP, and that is the point: an agent connecting had 59 tool descriptions and a state-aware handshake, neither of which answers "what is this, and how am I meant to work here". A tool description explains one tool to someone who already picked it; the handshake describes this repository right now.

Two halves, deliberately:

  • The narrative is a template under ddflow/templates/prompts/help/, so ddflow prompts eject-style overriding applies — put your own .ddflow/prompts/help/workflow.md in place and the tool teaches your workflow.
  • The capability inventory is generated from the live tool table. A hand-kept command list in a second place is the documentation-drift class, and this project has paid for it twice.

Three ratchets keep the prose honest, because a page recommending a flag that was renamed is worse than no page — whoever finds nothing reads the code, and whoever finds a wrong answer trusts it. Every command a page names must exist as a CLI leaf or an MCP tool; every topic the index offers must resolve; and every tool must fall into a group, so a new capability has to be classified rather than quietly dropped from an inventory that claims to be complete.


The default workflow at a glance

Three pictures of what ships by default. Every step is configurable (see The workflow, and changing it); ddflow workflow prints what this project actually runs.

The agent's loop. One item at a time: pick it, lease it, clear its gates, land it. Each arrow labelled with an exit code is what the tool says, not a convention — 2 means nothing to do, 3 means refused, and neither is ever read as success.

flowchart TD
    S["Session start<br/>ddflow brief · recover"] --> R{"Crashed agent's<br/>work left over?"}
    R -- yes --> SV["Inspect the worktree, salvage,<br/>ddflow release"] --> N
    R -- no --> N["ddflow next"]
    N -- "exit 2: nothing ready" --> W["ddflow wait<br/>sleeps until a holder lets go"] --> N
    N -- "exit 0: items ready" --> C["ddflow claim<br/>lease + isolated git worktree"]
    C -- "exit 3: refused<br/>(lease or file-glob conflict)" --> N
    C --> P["Task pipeline<br/>satisfy every gate in order"]
    P --> M["ddflow merge<br/>lands the branch from the primary checkout"]
    M --> D["ddflow complete<br/>checks every gate has an outcome<br/>+ a different-family review"]
    D -- "exit 3: refused" --> P
    D -- done --> N

With [flow].integration = "pr", merge opens a pull request and parks the item in review instead; next completes it when the request merges.

The task pipeline. Ten gates, in order. Heavy borders are the gates gates.required names by default (implement, unit_tests, merge); the rest still need some outcome — passed, failed, unavailable, partial, or an explicit skip with a reason — because silence is not a pass.

flowchart LR
    subgraph A["You, the agent"]
        direction TB
        g1["1 · research<br/>falsifiable claim + probe"] --> g2["2 · rules<br/>ddflow brief --item"] --> g3["3 · implement<br/>in the item's worktree"]
    end
    subgraph X["A different model family"]
        direction TB
        g4["4 · rubber_duck<br/>try to refute it"] --> g5["5 · critic<br/>diff vs. stated intent"]
    end
    subgraph T["Tooling"]
        direction TB
        g6["6 · standards<br/>linters, architecture"] --> g7["7 · unit_tests<br/>the suite, executed"]
    end
    subgraph B["You, again"]
        direction TB
        g8["8 · bug_hunt<br/>recurring bug classes"] --> g9["9 · dedupe<br/>already exists?"]
    end
    g10(["10 · merge<br/>ddflow lands it"])
    A --> X --> T --> B --> g10

    classDef req stroke-width:3px
    class g3,g7,g10 req

The phase pipeline. A phase wraps its tasks and checks the whole before it lands:

flowchart LR
    p1["research<br/>the phase's open questions"] --> p2["tasks<br/>each runs its own<br/>ten-gate pipeline,<br/>in parallel where<br/>globs allow"]
    p2 --> p3["unit_tests"] --> p4["bug_hunt"] --> p5["dedupe"]
    p5 --> p6["live_test<br/>run the real thing"] --> p7["corrections"] --> p8["docs<br/>README and docs<br/>match the change"] --> p9(["merge"])

Details: The task pipeline · The phase pipeline · Parallelism and coordination · Crash recovery.


Two ways to drive it

The CLI is the whole product. The MCP server is a second surface over the same commands, and tests/test_mcp_parity.py fails if the two diverge — every subcommand has a tool, every flag is reachable, and each exemption carries a written reason.

Standalone: a terminal, a Makefile, CI

$ ddflow init
$ ddflow config --set gate.unit_tests.command "python -m pytest -q"
$ ddflow phase add P1 --title "Billing" --globs "src/billing/**"
$ ddflow task add P1.T1 --phase P1 --title "Tax rules" --globs "src/billing/tax.py"

$ ddflow next                          # exit 2 = nothing actionable; exit 1 = unknown --phase
$ ddflow claim P1.T1                   # exit 3 = refused, with the reason
leased P1.T1 · worktree .ddflow-worktrees/P1.T1 · branch ddflow/P1.T1

$ cd .ddflow-worktrees/P1.T1 && ...    # do the work
$ ddflow gate status P1.T1             # what the pipeline wants next
$ ddflow gate run P1.T1 unit_tests     # runs it; the exit code IS the evidence
$ ddflow gate record P1.T1 implement --outcome passed --evidence "added tax.py"
                                       # failed/unavailable/partial/skipped also need --reason
$ ddflow complete P1.T1                # exit 3 lists whatever is unsatisfied
$ ddflow merge P1.T1

You get everything except the judgement. Command gates run themselves; agent gates wait for a human to record an outcome, and ddflow gate skip <id> <gate> --reason "..." is the escape hatch — recorded as a skip, never as a pass.

In CI, the exit codes are the interface:

check:
	ddflow doctor        # 1 = integrity problems, each named
	ddflow workflow      # 1 = the pipeline does not hang together
	ddflow cadence       # 2 = no periodic pass is due

2 is never "no problem". A job that treats it as success reports a green build for a suite that never ran.

As an MCP server

ddflow mcp speaks newline-delimited JSON-RPC over stdio. You rarely run it by hand — ddflow adopt writes the launch entry into each agent's own config and leaves existing servers alone:

22 agents are supported. The full table, with what each one gets, is in Wiring it into your agent.

It also copies the driver to docs/ddflow/drivers/, and installs the pre-commit hook that enforces claim-before-you-edit.

What an agent sees the moment it connects, with no call to make:

  • Instructions, returned inside the initialize result itself — and state-aware: what is ready, what is in flight, which setup is missing, whether this project has history worth importing, whether an import was left unfinished.
  • Tools — one per CLI command.
  • Resources — ddflow://board, ddflow://brief, ddflow://lessons, ddflow://research.
  • Prompts — which a client turns into slash commands. Tools are things an agent calls; prompts are things you invoke.

A refused call says so first. Over MCP a tool that did not do what was asked — a claim refused for an overlap (exit 3), or any call whose result would otherwise be its success shape in nulls — returns a JSON body whose first key is "refusal": {"reason": ..., "outcome": ..., "exit": ...}, followed by whatever the operation actually said (a refused claim's alternatives). Exit 2 ("nothing") keeps its declared keys after that lead; a result that fills its declared shape is left as the CLI's --json prints it, with the reason in the second content block.

Two tools exist so an agent can orient itself without being told: ddflow_help (what is this, what is the loop) and ddflow_workflow (what are the rules here).

What goes in AGENTS.md / CLAUDE.md

ddflow adopt writes it as a managed block between and. Your own prose around it is preserved; re-running updates only what is inside. If you write it by hand, four things have to be in it:

  1. Start every session with ddflow_brief (or ddflow brief in a shell).
  2. Claim before you edit — ddflow_next → ddflow_claim → work in the worktree it creates.
  3. The loop — ddflow_gate_status → satisfy each gate → ddflow_complete → ddflow_merge.
  4. The exit codes, and that 2 is not success.

Without that block an agent sees the tools and has no reason to reach for them before editing. The block is what makes the queue authoritative rather than optional — and it is 232 words, because an instruction file nobody finishes reading is one nobody follows.


The workflow, and changing it

$ ddflow workflow
# The workflow this project runs

   1. research      agent
   2. rules         agent
   3. implement     agent       (required)
   4. lint          command     (required, NOT proven able to fail)
      $ ruff check .
   ...

## The rules, and where each came from

  gates.require_outcome                  True              [default]
  gates.enforce_order                    block             [file]
  schedule.max_parallel_tasks            4                 [default]

One answer to "what are the rules here": every gate in order, which are commands and which you perform, which are required, which need evidence, which need a different-family reviewer, which have been proven able to fail — plus the completion rules, the caps, the reviewers, and where each value came from, so a deliberate choice is distinguishable from a default nobody touched.

Changing it

$ ddflow workflow gate lint --command "ruff check ." --into task --after implement --required
$ ddflow workflow pipeline task research,implement,lint,unit_tests,merge
$ ddflow workflow drop dedupe

All four reach MCP — ddflow_workflow, ddflow_workflow_pipeline, ddflow_workflow_gate, ddflow_workflow_drop — so an agent can change the workflow with the operator's agreement. Their descriptions say to ask first and offer dry_run, because a pipeline governs every future item, not the one in hand.

Nothing is written until it is checked, and the order is the point: compose the change, validate the result, then replace the file atomically.

  • A pipeline naming an undefined gate is refused, naming the near miss. That one is otherwise silent and permanent: the outcome folds to empty, completion refuses it forever, and gate record rejects the id as unknown — so the item can never be completed at all, and nothing says why.
  • An unknown section or knob is refused, with a suggestion. [gatez] is valid TOML and used to be written happily, breaking every later command — the write path validated the merged text for syntax and then validated the config already on disk, which is a writer checking the state it is replacing.
  • Dropping a gate takes it out of required too, or it becomes a requirement that quietly requires nothing.

ddflow workflow and ddflow doctor both re-run those checks against what is on disk. Everything is a file you can also edit by hand: gates in [gate.<id>], reviewers in [[reviewer]], companions in .ddflow/companions.toml, and every prompt — including the instructions your agent receives at connect — under .ddflow/prompts/.

One caveat with MCP: the connection instructions are computed once, when the server starts. A workflow changed mid-session is live for every tool call immediately, but the text the agent was handed is stale. Tell it to call ddflow_workflow, or restart.


Why it is built this way

The append-only event log is the source of truth; everything else is a projection that can be deleted and re-derived. The SQLite index, the markdown boards, the search index, the recovery bundle — all disposable, all rebuilt by ddflow rebuild.

That single inversion is what makes the four hard properties fall out for free rather than needing to be engineered:

You getBecause
Two agents on two branches never conflictEach appends to its own file. Measured: a real two-branch merge resolves clean.
A crashed agent loses nothingState is folded, never written. Nothing is half-updated.
The project rebuilds from the logOperator prompts are events.
An edited history is detectableEvent ids are content addresses.

The design decisions, with the probes that settled each, are in docs/RESEARCH.md. The two that most shaped it:

  • SQLite-on-NFS is correct here but 36× slower than local (measured, 12 processes × 40 increments). So the log is authoritative and the database is a disposable cache — which also happens to be the choice that stays correct on filesystems where locking is broken.
  • An expired lease must never be reclaimed automatically. A crashed agent's worktree is sometimes irreplaceable work and sometimes a superseded draft, and nothing in the metadata distinguishes them. Recovery measures and advises; it never deletes.

Install into any project

One line in your agent's MCP config. Nothing else.

Until the first release is on PyPI, uvx ddflow-mcp has nothing to fetch: install from the repository and run ddflow adopt, which registers that installation instead, as in A first run.

{ "mcpServers": { "ddflow": { "command": "uvx", "args": ["ddflow-mcp"] } } }

uvx fetches and runs the published package in an ephemeral environment on first use — no clone, no virtualenv, no PYTHONPATH, no install step for an operator to forget, and no vendored copy to drift from upstream. ddflow needs one runtime dependency beyond python3 and git (Jinja2), which is what lets it install inside sandboxes, CI images and other tools' ephemeral containers.

Then, from the agent, with no shell at all:

CallWhat it does
ddflow_setupcreates .ddflow/, writes the driver and the AGENTS.md section
ddflow_configure with toml: '[gate.unit_tests]\ncommand = "pytest -q"'sets your test command
ddflow_reviewers_detect with write: truefinds a local model server and registers it as a cross-family reviewer
ddflow_phase_add, ddflow_task_addfill the queue
ddflow_briefstart every session here

That is the whole adoption. The per-project instruction text is 232 words — a managed block in AGENTS.md, because the MCP tool descriptions already carry the how, and a second copy of that would drift from the one the model actually reads.

Shell / CI installation, and running from a source checkout
uv tool install ddflow-mcp        # or: pipx install ddflow-mcp
cd /path/to/your/project
ddflow adopt            # every supported agent
ddflow adopt --agents claude,cursor,vscode,kimi   # or name the ones you use

After upgrading ddflow, ddflow doctor notes any driver doc (implement-phase.md, an adopted agent's delta) that differs byte-for-byte from the template the running ddflow ships. ddflow adopt --refresh-docs (MCP: ddflow_setup with refresh_docs) rewrites only those docs, the AGENTS.md/CLAUDE.md blocks and the agents' native rules -- never the MCP launch, hooks or command files -- and refuses a project that was never adopted.

adopt is idempotent and writes managed blocks, so re-running after an upgrade updates them and leaves your own prose alone. It writes the MCP registration into each agent's own config location, merged with whatever servers are already there. From a source checkout it points the config at that checkout instead of the published package, so developing ddflow does not silently configure your project against the released version.

Docker — for operators with no Python toolchain

{ "mcpServers": { "ddflow": { "command": "docker", "args": [
    "run", "-i", "--rm",
    "-v", "${workspaceFolder}:/repo",
    "--add-host=host.docker.internal:host-gateway",
    "ghcr.io/delian/ddflow-mcp:latest" ] } } }

ddflow adopt --launch docker writes exactly that. The image is 107 MB (Alpine; ddflow is pure standard library, so there is no compiled dependency to worry musl about) and behaves identically on Linux, macOS and Windows.

Four things go wrong when a containerised tool touches a bind-mounted git repo. All four are silent, one of them loses work, and all four are handled:

TrapWhat it looks likeHandled by
Worktrees land outside the mountworktree.root defaults to ../.ddflow-worktrees, a sibling of the repo. In a container only the repo is mounted, so worktrees go to the ephemeral layer and are destroyed on exit with the agent's uncommitted work inside them.container.default_worktree_root relocates a sibling root to .ddflow-worktrees inside the repo, and adopt gitignores it
Root-owned filesOn a Linux bind mount the operator needs sudo to edit their own project afterwardsthe entrypoint reads the mount's uid/gid and su-execs down to it
git refuses the mount"detected dubious ownership", surfacing as an unexplained ddflow failuresafe.directory set in the entrypoint
No git identitygit commit fails with "Please tell me who you are"entrypoint prefers GIT_AUTHOR_*, then the repo's own config, then a clearly-marked placeholder

And one that cannot be fully handled, so it is reported: 127.0.0.1 inside a container is the container. A model server on your own machine is not reachable from there. ddflow rewrites loopback reviewer URLs to host.docker.internal, and ddflow doctor tells you that on Linux you must also pass --add-host=host.docker.internal:host-gateway, because unlike Docker Desktop the Linux engine does not provide that name.

The related portability fix: worktree paths are stored in the event log relative to the repo root. The log is committed and shared, so an absolute path is true only on the machine that wrote it — false for a teammate who cloned elsewhere, for CI, and for a container where the repo is /repo. Pinned by test_the_event_log_carries_no_absolute_paths.

Extending it by writing text, not code

Every prompt is an external template, resolved config → project → shipped:

ddflow prompts list              # where each template currently comes from
ddflow prompts eject             # copy the shipped ones into .ddflow/prompts/
$EDITOR .ddflow/prompts/review_system.md

Adding a mode of your own: [[macro]]. Overriding a shipped workflow needs no code, and neither does adding one. A macro is a named, parameterised prompt — "enter debugger mode" — that appears everywhere the shipped workflows do: prompts/list and prompts/get over MCP, which is what a client turns into a slash command, and ddflow prompts list|show in a terminal.

# .ddflow/config.toml   (or .ddflow/macros.toml, if you prefer to split it out)
[[macro]]
name = "debugger"
title = "Enter debugger mode"
description = "Reproduce first, then bisect. No fix without a failing probe."
params = ["symptom"]                                  # required, not optional
tools  = ["ddflow_bug_found", "ddflow_gate_run", "ddflow_bug_fixed"]
prompt = """
You are debugging: {{ symptom }}

Reproduce it before you theorise. Paste the command and its output.
"""

Use prompt_file = "docs/modes/debugger.md" instead for anything long enough that TOML quoting gets in the way.

When to reach for a macro rather than a gate. A gate is a step every item passes through, recorded against that item and blocking its completion. A macro is a MODE an operator enters, belonging to no item and recorded nowhere — "audit this release", "handle this incident". If the thing should hold up a task until it is done, it is a gate; if it is a way of working you want to name and re-enter, it is a macro. Putting a mode in the pipeline makes every task wait for something that was never about that task.

tools is declarative, not a sandbox. It is rendered into the prompt as the ordered set the mode expects, so the agent is told what the mode is for and the next reader can tell what it was supposed to do. It does not restrict what the agent may call — MCP has no mechanism for that, and claiming a security property this cannot honour would be worse than not having it. This is the deliberate departure from dx-zero/mcpn, whose toolMode: situational lets the model pick freely from a bound set with no recorded ordering: a session you cannot replay is a session you cannot review, which is the property the event log exists to give you.

Three things a macro refuses, because each alternative fails quietly: a missing parameter (a prompt with a hole in it reads as a complete instruction), a name that belongs to a shipped command (silent shadowing leaves you editing a block that does nothing), and both prompt and prompt_file (two sources for one body means one is dead and looks live). A refused macro is refused alone and by name — the others still load — and ddflow prompts list, prompts show, doctor and the MCP prompts/list say which one and why; an undecodable prompt_file is a named problem too.

Including the one the agent actually reads first. mcp_instructions.md is the block an MCP client injects into the model's context on connect — the workflow, the reporting duties, and which companion tools to reach for. It is the file to edit when you want this project to work differently:

ddflow prompts eject mcp_instructions
$EDITOR .ddflow/prompts/mcp_instructions.md      # or [prompts] mcp_instructions = "..."

It renders against the live state — adopted, task_pipeline, setup_todo, companions, missing_companions, gate_gaps, recoverable, loops — so the instruction is the next concrete action rather than a fixed blurb the model learns to skip. A broken override says so in the instruction block itself instead of falling back to the default: this is the one surface where nobody would ever notice their edit was not live.

Templates render with Jinja2, which is ddflow's one runtime dependency, and with a strict standard-library renderer when it is absent — a stripped deployment with no reachable package index still starts. The shipped templates use the subset both engines agree on, and tests/test_template_engines.py walks the template REGISTRY, rendering every entry through both engines and asserting the outputs are byte-identical.

That test is iterated rather than hand-listed for a reason. Its predecessor named three templates in a dict, mcp_instructions.md was never added, and in 0.1.1 the largest and most important template rendered correctly under Jinja2 and failed under the fallback — so the entire MCP handshake for an unadopted repository, the first thing a new user ever sees, degraded to ddflow's instruction template could not be loaded. Jinja2 was not a declared dependency at the time, so developers had it and the project venv did not: python -m pytest was green and uv run pytest was red on the same commit.

The fallback now raises on any construct it does not implement rather than copying it through. The old regex engine emitted what it could not parse, so a condition as ordinary as {% if a or b %} — which its single-name pattern never matched — reached the client as literal template source.

Both renderers are strict about undefined variables: a prompt silently missing the diff it was supposed to carry is the vacuous review in template form — the model dutifully reviews nothing and reports no findings.

The rest is TOML: gates and their pipelines ([gate.*], gates.task_pipeline), reviewers ([[reviewer]]), companions ([[companion]]), enforcement ([enforce]), cadences, and the rest of the 141 knobs. ddflow config --set <key> <value> edits one key in place, preserving comments.

What is committed, and what stays on your machine

ddflow recommends services; it never ships one person's configuration. Two layers:

LayerFilesHolds
committed.ddflow/config.toml, .ddflow/gates.tomlgeneric project policy: the test command, the pipelines, the gates (human checkpoints included) — what every clone must agree on
machine-local, git-ignored.ddflow/local/config.toml, .ddflow/local/gates.toml, .ddflow/local/reviewers.tomlyour services: reviewer endpoints, model names of a private deployment, API-key variable names, LAN hosts, a worker count sized to this machine

The local files are read last, so they win; ddflow config --explain reports such a value's source as local. Every writer of an operator-specific value targets the local layer by default:

ddflow reviewers add --preset ollama --model qwen3:8b   # -> .ddflow/local/reviewers.toml
ddflow reviewers detect --write                         # -> .ddflow/local/reviewers.toml
ddflow config --local --set gate.unit_tests.command "pytest -q -n 48"
ddflow config --local --append-toml "$(cat my-reviewer.toml)"

--shared on reviewers add / reviewers detect --write commits the block to .ddflow/config.toml instead — only for a service every clone reaches at the same address. config --set and --append-toml stay committed unless you pass --local (over MCP: ddflow_configure with local=true, ddflow_reviewers_detect with shared=true), because a test command or a pipeline is project policy. The human-gate guards hold on both layers: no writer sets gate.<id>.human, and none can drop a human gate from a pipeline. .ddflow/local/ carries its own * .gitignore, so it stays uncommitted even in a project whose .ddflow/.gitignore predates it. API keys are never written anywhere — only the name of the variable that holds one.

Publishing and registry

Nobody should have to paste JSON into an IDE to use this. server.json is the MCP registry manifest (io.github.delian/ddflow-mcp), and publishing it is what makes ddflow findable in the VS Code and Cursor marketplaces rather than something you configure by hand. It offers three ways to run the same server, so a client picks whichever it supports:

PackageIdentifierFor
pypiddflow-mcp, runtimeHint: uvxAnything with uv — no clone, no install step
ocidocker.io/delian/ddflow-mcp:<version>Operators with no Python toolchain
ocighcr.io/delian/ddflow-mcp:<version>The same image, no Docker Hub account needed

The image is built for amd64 and arm64, because an Apple-silicon operator running it under emulation pays that cost on every tool call, and tool calls are all this server does.

How CI authenticates — four mechanisms, one stored secret:

TargetMechanismStored secret?Setup
PyPIOIDC trusted publishing (id-token: write)NoAdd a trusted publisher on PyPI, once
ghcr.ioGITHUB_TOKEN, injected per run, expires with the jobNonone
Docker HubDOCKERHUB_USERNAME + DOCKERHUB_TOKENYesCreate an access token, add both secrets
MCP registryGitHub OIDC — proves control of the account that owns the io.github.delian/* namespaceNonone
tag + releaseGITHUB_TOKEN (contents: write)Nonone

Docker Hub is the only one that needs a long-lived credential, because it has no OIDC equivalent. Use an access token scoped to this repository, never an account password. If that is one secret too many, delete the Docker Hub login and its two tags — ghcr.io alone satisfies the OCI entries a marketplace needs, and server.json lists both so a client picks whichever resolves.

environment: release on the publishing jobs is a control worth knowing about: point it at a GitHub environment with required reviewers and every release waits for a human, with no change to the workflow.

Order matters and the workflow encodes it. mcp-publisher validates that every package named in the manifest exists, so the registry step runs after both PyPI and Docker — publishing the manifest first would advertise a version nobody can fetch.

The registry verifies ownership, and the workflow checks it first. The registry accepts a package only when the artefact itself names the server: the PyPI package's README carries `` (the first lines of this file), and each image carries the label io.modelcontextprotocol.server.name with the same name. The verify job runs tests/test_registry_ownership.py first, so a release the registry would reject — a missing marker, or a label that does not match — is refused before anything is uploaded; a server.json description over the registry's 100-character limit fails the suite the same way. So do the registry's per-package rules, which the JSON schema does not express and the registry enforces only at publish: an oci package must not carry a version field (or registryBaseUrl/fileSha256) — the tag in its identifier is the version — while the pypi package must carry one. publish #40 failed on exactly that, at the last job, after PyPI and both images had shipped; tests/test_registry_manifest_rules.py now encodes the rules offline and verify runs it first. The publish step also retries a transient registry failure (a 504, a 408/429, a network error) with backoff, checking the registry for the exact version after every attempt, because the publish behind a 504 may have committed; a 4xx that retrying cannot fix fails at once, and only after the budget does it fail with a "Re-run failed jobs" hint.

Four things gate a release, and each exists because the failure it catches is public and irreversible:

  • the tag, ddflow/__init__.py (the one declared version), server.json's version and every OCI identifier's tag must agree — a :0.1.0 left behind while version moved on publishes a manifest pointing at the previous image, installable and wrong;
  • the full suite, plus the slow end-to-end scenarios, which -m 'not slow' otherwise excludes from every ordinary run;
  • the wheel must install into a clean venv and run, and carry its templates — uv build succeeding proves the metadata parses, not that ddflow help works;
  • the image must answer initialize over stdio. A built image that cannot is a broken release every marketplace will happily offer.

Cutting a release

$ git push origin main           # that is the whole release

Every push to main that changes shipped code releases, with the PATCH version bumped. Major and minor move only when you move them. "Shipped" means ddflow/, pyproject.toml, uv.lock, Dockerfile, docker-entrypoint.sh, .dockerignore, server.json, server.template.json or README.md (the PyPI page, and the line the registry verifies): a push of other docs, tests or the ddflow event log releases nothing, because PyPI keeps every version forever and one identical to the last is noise nobody can withdraw. CI runs scripts/bump.sh patch, commits release 0.1.2 to main, publishes PyPI, Docker Hub, ghcr.io and the MCP registry, then creates v0.1.2 and a GitHub release — last, and only once every publish succeeded, because a tag pointing at a half-release is worse than no tag: it looks authoritative.

Pull after a release. The release commit is CI's, so your main is one commit behind it and the next push is refused until you git pull.

How the gate picks the version: it publishes the declared version if PyPI does not have it yet, and bumps patch only if it does. So moving major or minor is yours to do —

$ scripts/bump.sh minor          # 0.1.4 -> 0.2.0 (or: major, or an exact 1.0.0)
$ git commit -am 'release 0.2.0' && git push origin main

— and CI publishes exactly 0.2.0; the next push that bumps nothing releases 0.2.1.

The same rule makes a release that failed before its PyPI upload retry its number with the next push. One that failed after it — Docker Hub, ghcr.io, the MCP registry — does not: PyPI has the number, so the next push moves past it; re-run that run's failed jobs from the Actions page instead. A bump never lands on a number PyPI already holds (a v* tag can publish one out of band): CI skips to the next free patch before it commits anything. The bump is pushed to main before anything publishes: a push that loses a race with another commit fails the run and publishes nothing, where pushed last it would leave PyPI holding a version main does not declare. Runs are serialized, and only main or a v* tag releases.

The version is declared in one place: __version__ in ddflow/__init__.py, the only line a bump edits. pyproject.toml reads it (a hatch dynamic version, so uv.lock records no version for the project and a bump never makes the lock stale), SERVER_INFO (what the server tells every client

Installation

Source-derived launch command. Check the maintainer’s required arguments and credentials before running:

bash
uvx ddflow-mcp

Set up in your AI client

Merge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.

json
{
  "mcpServers": {
    "io-github-delian-ddflow-mcp": {
      "command": "uvx",
      "args": [
        "ddflow-mcp"
      ]
    }
  }
}

Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.

Claude Desktop setup reference

Package

ddflow-mcppypi

Compatible MCP Clients

ddflow works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.

  • Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.
  • Cursor~/.cursor/mcp.jsonRestart Cursor for changes to take effect.
  • VS Code.vscode/mcp.jsonReload VS Code window for changes to take effect.
  • Windsurf~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect.
  • Claude Code.mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.

Learn More