Every tool promises discipline.
Spec kits, skill packs, virtual teams, task graphs — each hands your agent a process to follow. One question separates them: when the model skips a step at hour six, what stops it? For most tools the answer is nothing — the process is a recommendation. team.management makes it a gate. Every number below carries a date, and every source is linked.
That is not a criticism of any of them. A well-written instruction is genuinely useful, and most of the time the model does follow it. The question is what happens the rest of the time.
Anything the model reads is advice. It sits in the context window alongside everything else, and it competes with everything else. Late in a long session, with the window full and the goal three steps away, advice is the first thing to go. Nothing announces that it happened.
A gate is different because it does not live in the context window at all. It runs in the harness. Some gates answer a tool call before it happens — the edit tools stay locked until the work has been agreed. Others refuse to let the job reach its next step until a check has actually run. Either way the model does not get a say.
That distinction is the axis every page below is measured on, and it cuts both ways — it is also the limit. A gate binds the agent, not the person running it. You can always change your own config. What it removes is the silent kind of failure, where a step was skipped and the summary said otherwise.
The pages come in three families. Workflow tools are the ones you would run instead of this, or alongside it — most compose better than they compete. Models and subscriptions are about what you pay and what you get for it, with a date on every number, because those numbers move month to month. Process and verification compares methods rather than products: what actually holds when an agent goes off-script, and what checks the work it claims to have done.
Workflow tools
Every tool here promises your agent discipline. One question for all of them: what happens when the model stops listening?
Spec documents structure what to build. A protocol engine enforces how work proceeds.
Skills that ask the model to be disciplined vs hooks that make discipline non-optional.
A virtual team of role-skills vs one team on enforced rails.
Task decomposition and tracking vs a task lifecycle the agent can’t skip.
A simulated agile team with heavy ceremony vs enforced discipline at plugin weight.
Agent memory and issue graphs vs process enforcement — different layers, use both.
Discipline without ceremony, on two different layers: spec artifacts vs runtime gates.
The DAIC ancestor, dormant since late 2025 — and the maintained engine that grew past it.
Karpathy’s overnight experimentation loop, run as an enforced protocol: free exploration, frozen paths.
Anthropic’s automation recommender vs an engine that enforces the workflow.
Both gate your agent with hooks. The difference is what the hooks protect.
Models & subscriptions
Prices, credits, and limits for the major AI coding subscriptions — dated and sourced. Whichever one you pick, the process layer stays yours.
Subscriptions, credits, and limits — sourced and dated. And why you don’t have to choose.
Google’s agentic stack vs Anthropic’s — quotas, credits, history, sources.
Claude, ChatGPT/Codex, Google AI, Cursor, Copilot, Windsurf — one sourced, dated table.
Process & verification
Not tool against tool — method against method. What actually holds when an agent skips a step or calls unfinished work done.
Dozens of tools structure the work. Almost none check it happened. The map, dated.
Rules files get read, then disregarded — on agent after agent. Four responses, compared.
Files never written, tests never run, “done” anyway. What actually catches it.