# team.management > A more capable model isn't a more trustworthy one — it's a more convincing liar. team.management is an open-source protocol engine for Claude Code that puts every process on rails your AI doesn't leave — not just code — enforced as it happens, provable after. ## Overview - [team.management — The operating system for human + agent teams.](https://team.management/index.md): A more capable model isn't a more trustworthy one — it's a more convincing liar. team.management is an open-source protocol engine for Claude Code that puts every process on rails your AI doesn't leave — not just code — enforced as it happens, provable after. - [Docs](https://team.management/docs.md): Install, the DAIC loop, TDD discipline, the spec-review gate, customization, the MCP reference, and AI providers. - [Protocols](https://team.management/protocols.md): One engine, six shipped protocols. Each drives a multi-step workflow with DAIC modes, agents, and gates managed automatically. - [Compare](https://team.management/compare.md): team.management vs GitHub Spec Kit, Superpowers, gstack, Claude Task Master, BMAD Method, Beads, OpenSpec, cc-sessions, autoresearch — and the AI coding subscriptions side by side, with a date on every number. ## Getting started - [Install in five minutes](https://team.management/docs/install.md): A native Claude Code plugin: add the marketplace, install, run init — your team follows via one committed file. - [DAIC — the loop on rails](https://team.management/docs/daic.md): Discuss → Align → Implement → Document → Commit, enforced by hooks that can’t be bypassed. - [Context preservation & auto-compact](https://team.management/docs/context-preservation.md): A PreCompact hook checkpoints the session before compaction; on restart, task, branch, protocol step, and DAIC mode all reload. ## Discipline - [Test-driven, when it applies](https://team.management/docs/tdd.md): Write the test, watch it fail, then implement. The applicability matrix, honestly drawn. - [The spec-compliance gate](https://team.management/docs/spec-review.md): A SPEC_REVIEW: PASSED sentinel the agent can’t advance past until the diff matches the spec. - [Your protocols, your rules](https://team.management/docs/customization.md): Fork any protocol into custom/ — it overrides the system copy and is never touched on upgrade. - [LLM Wiki — a compounding knowledge base](https://team.management/docs/llm-wiki.md): Claude writes and maintains a structured, interlinked wiki from your raw sources — knowledge compounds instead of being re-derived on every query. ## Reference - [MCP tools reference](https://team.management/docs/mcp.md): 42 tools across 8 modules — the primary tool surface, with guaranteed discoverability. - [Engine functions reference](https://team.management/docs/functions.md): The 45 pre_funcs / post_funcs you wire into protocol steps — branch setup, gates, issue sync, archiving, the optimize loop — grouped by what they do. - [Issue tracking — four providers, one interface](https://team.management/docs/issue-tracking.md): GitLab, Jira, GitHub, Gitea — import issues as tasks, sync status bidirectionally, ship MRs and PRs from the completion step. - [AI providers — Codex & Antigravity](https://team.management/docs/ai-providers.md): Codex and Antigravity join Claude as parallel collaborators across six workflow phases. ## Protocols - [task](https://team.management/protocols/task.md): The standard implementation lifecycle. - [brainstorm](https://team.management/protocols/brainstorm.md): Parallel-specialist ideation → planned tasks. No code. - [research](https://team.management/protocols/research.md): Spikes, PoCs, evaluations. No branch, no code. - [refactoring](https://team.management/protocols/refactoring.md): Test-baseline-gated restructuring. - [optimize](https://team.management/protocols/optimize.md): Metric-driven optimization, interactive batched. - [optimize-unattended](https://team.management/protocols/optimize-unattended.md): Autonomous twin of optimize. For overnight runs. ## Comparisons - [team.management vs GitHub Spec Kit](https://team.management/compare/team-management-vs-github-spec-kit.md): GitHub Spec Kit vs team.management: spec-driven documents vs runtime process enforcement for Claude Code — what each one actually stops, and why the two compose well together. - [team.management vs Superpowers](https://team.management/compare/team-management-vs-superpowers.md): Superpowers vs team.management: both promise engineering discipline for Claude Code. Superpowers works through structured instructions; team.management gates tool calls in the runtime. - [team.management vs gstack](https://team.management/compare/team-management-vs-gstack.md): gstack vs team.management: Garry Tan’s 23 role-skills turn Claude Code into a virtual team; team.management makes the team’s process enforceable. Which failure mode is yours? - [team.management vs Claude Task Master](https://team.management/compare/team-management-vs-claude-task-master.md): Claude Task Master vs team.management for Claude Code task management: PRD-to-task-graph decomposition vs an enforced task lifecycle with DAIC gates, branches, and review. - [team.management vs BMAD Method](https://team.management/compare/team-management-vs-bmad-method.md): BMAD Method vs team.management: 12+ agent personas and PRD ceremony vs an enforced protocol lifecycle inside Claude Code. Which discipline model fits your work? - [team.management vs Beads](https://team.management/compare/team-management-vs-beads.md): Beads vs team.management: Steve Yegge’s git-backed issue graph gives agents memory across sessions; team.management enforces the process within them. Use both. - [team.management vs OpenSpec](https://team.management/compare/team-management-vs-openspec.md): OpenSpec vs team.management: lightweight spec-driven workflow vs enforced protocol lifecycle for Claude Code. Same instinct — discipline without ceremony — different layer. - [team.management vs cc-sessions](https://team.management/compare/team-management-vs-cc-sessions.md): cc-sessions vs team.management: the original DAIC hooks harness is dormant since October 2025. team.management carries the same DNA forward as a maintained protocol engine. - [team.management vs autoresearch](https://team.management/compare/team-management-vs-autoresearch.md): autoresearch by Andrej Karpathy vs team.management’s optimize protocols: autonomous overnight experimentation — and what changes when the loop runs on an enforced, engine-measured lifecycle instead of Markdown instructions. - [Claude Code vs Codex (July 2026)](https://team.management/compare/claude-code-vs-codex.md): Claude Code vs OpenAI Codex in July 2026: subscription tiers, session and weekly limits, credit systems, and API fallbacks compared with primary sources — plus the case for running both on one process. - [Antigravity vs Claude Code (July 2026)](https://team.management/compare/antigravity-vs-claude-code.md): Google Antigravity vs Claude Code in July 2026: Google AI subscription tiers and credit bundles vs Claude plans and session limits — sourced, dated, including the quota history. - [AI coding subscriptions compared (July 2026)](https://team.management/compare/ai-coding-subscriptions.md): Every major AI coding subscription in one table, July 2026: Claude, ChatGPT/Codex, Google AI/Antigravity, Cursor, GitHub Copilot, Windsurf — prices, credit models, limits, and run-out behavior, with primary sources. --- # Your AI says it’s done. Prove it. The operating system for human + agent teams. An open-source protocol engine that puts every process on rails — enforced as it happens, so ‘done’ isn’t something you take on faith. Claude — Codex & Antigravity when you turn them on — **on the subscriptions you already pay for.** ## The model getting more capable does not make it more trustworthy. It makes it a more convincing liar. The answer isn’t a smarter assistant you hope did the work. It’s a harness you own — it shows the work, blocks the shortcuts, and hands you the proof. ## Watch the rails work. A protocol is your process, written down — the steps every task must follow. The engine holds the AI to it, step by step. Pick one below and watch it run. **Your process becomes a protocol — and a protocol can’t be skipped.** Nothing about a protocol is specific to code: it’s just JSON — steps and gates — so you drop your own in `custom/` and it runs on the same rails, dev or not. vendor-onboarding, employee-onboarding, incident-response — a few your business might run [All six protocols, in depth](https://team.management/protocols.md) ## See it. Enforce it. Prove it. Make the AI’s claims verifiable instead of trusting them. The work runs on a track you can watch, behind gates it can’t skip, leaving a record you can prove. See Watch every step as it runs — which stage it’s on, what’s allowed, what’s done. No black box, no “trust me, it’s handled.” Enforce The gate stays shut until the step’s criteria are met — no skipping ahead, no editing the tests to pass, no talking its way through. Prove Every run leaves a record of exactly what happened — the protocol, the steps, the gates it had to pass. Local, version-controlled, yours to keep. The record also says what checked what: review gates are AI judgment; the test gate is deterministic. ## Prompts ask. team.management enforces. A prompt is a request an agent can ignore, forget, or quietly route around. The interesting parts of this system aren’t requests — they’re walls. A prompt-based system can - claim it’s done when the tests never ran - edit the tests to make them pass - alter the database or fixtures to fit - skip a step it was told to follow team.management makes sure the agent - can’t advance past the gate without `SPEC_REVIEW: PASSED` - can’t edit frozen tests or fixtures mid-run — the runtime blocks it - can’t touch edit tools in discussion mode — DAIC blocks them - can’t skip ahead to a step it hasn’t earned **Forward is earned; backward is always open.** The engine won’t let the agent jump ahead — but it can always step _back_, re-plan, and re-earn the path when review finds a real problem. Your part is the alignment up front — the later gates run without you until one needs a decision. [How the four enforcement layers work→](https://team.management/docs/daic.md) ## Tests first. Specs checked. Evidence required. - tests first The task protocol expects a failing test before code — and optimize’s frozen-paths hook physically blocks edits to your tests or metric script mid-run. - specs checked A spec-review gate compares the diff against the task’s success criteria — the step can’t complete until it passes. - evidence required A step can’t complete on “looks good” — the advance must carry literal verification output, or name why none applies. ## Your memory. Your standards. Your skills. Your models. No one should be locked in — the agents follow your policy, and everything they run on stays yours. Your memory The LLM wiki lives in your repo, versioned in your git. What agents learn stays with the project, not with a vendor. Your standards Protocols are JSON you fork or author — your `custom/` directory is never touched on upgrade. Your process is the policy the engine enforces. Your skills Your slash-commands, MCP tools, and subagents keep working. The engine orchestrates them; it doesn’t replace them. Your models Claude by default — or any model, including the one you host yourself. Bring your own model — as the main one. team.management runs inside your harness — self-host the main model and nothing leaves your network — the engine is local files in your repo, and the cloud reviewers don’t join until you invite them. ## Compose the team each process needs. Protocols fan work out to specialist sub-agents — each in its own context window, each returning structured results. A roster ships in the box, but it isn’t fixed: make as many as you want, and point them at whatever a step needs — a code review, a security pass, a research dive. analysts: code-architect, code-explorer, critic, risk-security-analyst, scope-strategist, user-perspective, code-cleanliness, + your own reviewers: code-review, spec-compliance-reviewer, codex-cli, agy-cli, + your own context & docs: context-gathering, context-refinement, logging, service-documentation, + your own ## Plug into the tools you already run. Beyond the agent roster, team.management wires into the systems around your work — issue trackers and version control — so a protocol step can sync, file, or open the merge request for you. [GitHub](https://github.com) [Sync issues and open pull requests straight from a task.](https://github.com) [GitLab](https://gitlab.com) [Sync issues and merge requests as the work moves.](https://gitlab.com) [Jira](https://www.atlassian.com/software/jira) [Mirror tasks to Jira tickets and keep status in step.](https://www.atlassian.com/software/jira) \+ your own Write a connector for any system your team runs. ## Everyone on the same rails. Agents are only half the team. The same protocols and the same wiki make a group of people — and their agents — work like one. Enabling it is one commit — merge it, and every teammate is on the rails. One shared brain The LLM wiki is project-local and version-controlled. What one person’s agent learns, everyone’s agent reads next. Knowledge stops living in one head. Consistent output Same protocols, same gates, same definition of done — whoever is driving. A teammate’s task looks like yours because it ran the same rails. Onboarding for free A new hire doesn’t need the tribal knowledge. The protocols teach the workflow and the wiki carries the context — they ship correctly on day one. No lone-wolf drift No one quietly skips review or invents their own flow. The engine enforces the shared process, so the codebase stays coherent across the whole team. ## Five minutes from now, you’re shipping. `/plugin install team-management` `/plugin marketplace add TeamManagementPlugin/claude-plugin` `/plugin install team-management@team-management` `/team-management:init` \# commit .claude/settings.json — your whole team is enabled You→"create a task for implementing user authentication" [Full install guide→](https://team.management/docs/install.md) open source, by default ## Shared knowledge becomes a common good. The rails, the protocols, the wiki — yours to read, fork, and build on. MIT-licensed. ## Standing on shoulders. team.management draws on the projects and ideas that shaped its design. - [cc-sessions](https://github.com/GWUDCAP/cc-sessions) by GWUDCAP Origin of the DAIC methodology and the sessions / hook-enforcement model team.management is built on. - [superpowers](https://github.com/obra/superpowers) by Jesse Vincent (obra) Composable skills for coding agents — inspiration for the skill/protocol-driven workflow. - [get-shit-done](https://github.com/open-gsd/gsd-core) by open-gsd Meta-prompting, context engineering, and spec-driven development for Claude Code. - [LLM Wiki](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) by Andrej Karpathy The compounding, AI-maintained knowledge-base pattern behind the LLM Wiki feature. - [autoresearch](https://github.com/karpathy/autoresearch) by Andrej Karpathy Autonomous, metric-driven overnight experimentation — inspiration for the optimize protocols. ## How it compares. Every alternative hands your agent advice and hopes. See what each one does the moment the advice gets ignored — tool by tool, number by number. [vs GitHub Spec Kit](https://team.management/compare/team-management-vs-github-spec-kit.md) [vs Superpowers](https://team.management/compare/team-management-vs-superpowers.md) [vs gstack](https://team.management/compare/team-management-vs-gstack.md) [Claude Code vs Codex](https://team.management/compare/claude-code-vs-codex.md) [All comparisons→](https://team.management/compare.md) ## On your terms with AI. # Everything, in one place. Start with install, learn the DAIC loop and the gates that enforce it, then reach for the reference pages when you need the exact tool surface. For the workflows themselves, see the [protocols overview](https://team.management/protocols.md). ## Getting started [Install in five minutes](https://team.management/docs/install.md) [A native Claude Code plugin: add the marketplace, install, run init — your team follows via one committed file.](https://team.management/docs/install.md) ## Concepts [DAIC — the loop on rails](https://team.management/docs/daic.md) [Discuss → Align → Implement → Document → Commit, enforced by hooks that can’t be bypassed.](https://team.management/docs/daic.md) [Context preservation & auto-compact](https://team.management/docs/context-preservation.md) [A PreCompact hook checkpoints the session before compaction; on restart, task, branch, protocol step, and DAIC mode all reload.](https://team.management/docs/context-preservation.md) [Test-driven, when it applies](https://team.management/docs/tdd.md) [Write the test, watch it fail, then implement. The applicability matrix, honestly drawn.](https://team.management/docs/tdd.md) [The spec-compliance gate](https://team.management/docs/spec-review.md) [A SPEC_REVIEW: PASSED sentinel the agent can’t advance past until the diff matches the spec.](https://team.management/docs/spec-review.md) [LLM Wiki — a compounding knowledge base](https://team.management/docs/llm-wiki.md) [Claude writes and maintains a structured, interlinked wiki from your raw sources — knowledge compounds instead of being re-derived on every query.](https://team.management/docs/llm-wiki.md) [Your protocols, your rules](https://team.management/docs/customization.md) [Fork any protocol into custom/ — it overrides the system copy and is never touched on upgrade.](https://team.management/docs/customization.md) ## Reference [MCP tools reference](https://team.management/docs/mcp.md) [42 tools across 8 modules — the primary tool surface, with guaranteed discoverability.](https://team.management/docs/mcp.md) [Engine functions reference](https://team.management/docs/functions.md) [The 45 pre_funcs / post_funcs you wire into protocol steps — branch setup, gates, issue sync, archiving, the optimize loop — grouped by what they do.](https://team.management/docs/functions.md) [Issue tracking — four providers, one interface](https://team.management/docs/issue-tracking.md) [GitLab, Jira, GitHub, Gitea — import issues as tasks, sync status bidirectionally, ship MRs and PRs from the completion step.](https://team.management/docs/issue-tracking.md) [AI providers — Codex & Antigravity](https://team.management/docs/ai-providers.md) [Codex and Antigravity join Claude as parallel collaborators across six workflow phases.](https://team.management/docs/ai-providers.md) # On your terms with AI. One engine. As many protocols as you need. A protocol is a JSON config that drives a multi-step lifecycle. Each step declares a DAIC mode, a block of engine functions that run automatically, and the prompt the agent follows. Protocols are the only sanctioned way to do implementation work — they manage mode transitions, task files, git branches, and completion so the agent never has to flip modes by hand. built-in — six ship in the box, ready to `run` [task](https://team.management/protocols/task.md) [5 steps](https://team.management/protocols/task.md) [The standard implementation lifecycle.](https://team.management/protocols/task.md) [brainstorm](https://team.management/protocols/brainstorm.md) [5 steps](https://team.management/protocols/brainstorm.md) [Parallel-specialist ideation → planned tasks. No code.](https://team.management/protocols/brainstorm.md) [research](https://team.management/protocols/research.md) [4 steps](https://team.management/protocols/research.md) [Spikes, PoCs, evaluations. No branch, no code.](https://team.management/protocols/research.md) [refactoring](https://team.management/protocols/refactoring.md) [6 steps](https://team.management/protocols/refactoring.md) [Test-baseline-gated restructuring.](https://team.management/protocols/refactoring.md) [optimize](https://team.management/protocols/optimize.md) [7 steps](https://team.management/protocols/optimize.md) [Metric-driven optimization, interactive batched.](https://team.management/protocols/optimize.md) [optimize-unattended](https://team.management/protocols/optimize-unattended.md) [7 steps](https://team.management/protocols/optimize-unattended.md) [Autonomous twin of optimize. For overnight runs.](https://team.management/protocols/optimize-unattended.md) custom — ask Claude to run `/team-management:custom-protocol-create`, or author JSON in `custom/`; overrides the system copy on name match, untouched on upgrade hotfix example Triage → patch → verify → ship. Skips planning for urgent fixes. spike example Timeboxed exploration. Throwaway code, captured learnings, no review gate. release example Changelog → version bump → tag → publish. Your shipping ritual, codified. vendor-onboarding example Intake → due diligence → security review → sign-off before go-live. employee-onboarding example Provision access → agent runs the training → first task → manager sign-off. incident example Detect → triage → mitigate → post-mortem — each step in order. # Every protocol, the same skeleton. Steps as JSON, prompts as markdown. Step orchestration — the DAIC mode, which functions run, what arguments advance requires — lives in compact JSON. The verbose human-facing instructions live in separate sub-protocol markdown, so the long prompts can be edited or forked without touching engine logic. Functions, not hardcoded steps. All side effects — branch creation, issue sync, archiving, squashing — are named functions in a step’s `pre_funcs` / `post_funcs`, all catalogued in the [engine functions reference](https://team.management/docs/functions.md). Adding behaviour is a JSON edit, not an engine change. Gates that can’t be talked past. A post-function marked stop-on-failure becomes a hard precondition for advancing — that’s how spec-compliance, completion-evidence, and the optional test gate become structural walls rather than advisory reminders. # Your workflows, in the same engine. You don’t hand-author JSON to start — ask Claude to fork any shipped protocol: ``` /team-management:custom-protocol-create task ``` This runs `protocol_customize("task")`, copies the system config into `team-management/protocol-configs/custom/task.json`, and opens it for editing — then you describe the change in plain language and Claude edits it with you. The custom copy takes precedence on name match and is never overwritten on upgrade. To create a net-new protocol, drop any uniquely-named JSON into `custom/` — it appears in `protocol_list()` immediately. The engine functions you can wire into steps (branch setup, issue sync, archiving, frozen paths) are discoverable via `protocol_available_funcs()`. After an upgrade, run `/team-management:custom-protocol-update-after-reinstall` to diff your copies against the new system ones and decide what to merge. See [customization](https://team.management/docs/customization.md) for the full story. # Two-stage code review. The `task` protocol and the optimize pair share a structurally-gated review step — the framework’s strongest quality gate. ### Stage 1 — spec compliance A read-only reviewer audits the diff against your Success Criteria. On pass, the engine records a `SPEC_REVIEW: PASSED` sentinel. Advance is impossible without it — it’s checked both on entry and as a hard block on exit. ### Stage 2 — code quality One message dispatches the Claude `code-review` agent plus one Task per configured AI provider, all in parallel. Findings are aggregated with equal weight; provider output is advisory and never blocks. Forward is earned; backward is always open. A step advances only when its completion criteria are met — but it can always step back. If review surfaces a real problem, the protocol returns to investigation, re-plans, fixes, and re-earns the path forward. [Ready to run one? Install in five minutes→](https://team.management/docs/install.md) # Every tool promises discipline. Spec kits, skill packs, virtual teams, task graphs — each hands your agent a process to follow. One question separates them: when the model skips a step at hour six, _what stops it?_ For most tools the answer is nothing — the process is a recommendation. [team.management makes it a gate](https://team.management/docs/daic.md). Every number below carries a date, and every source is linked. ## Workflow tools Every tool here promises your agent discipline. One question for all of them: what happens when the model stops listening? [team.management vs GitHub Spec Kit](https://team.management/compare/team-management-vs-github-spec-kit.md) [Spec documents structure what to build. A protocol engine enforces how work proceeds.](https://team.management/compare/team-management-vs-github-spec-kit.md) [team.management vs Superpowers](https://team.management/compare/team-management-vs-superpowers.md) [Skills that ask the model to be disciplined vs hooks that make discipline non-optional.](https://team.management/compare/team-management-vs-superpowers.md) [team.management vs gstack](https://team.management/compare/team-management-vs-gstack.md) [A virtual team of role-skills vs one team on enforced rails.](https://team.management/compare/team-management-vs-gstack.md) [team.management vs Claude Task Master](https://team.management/compare/team-management-vs-claude-task-master.md) [Task decomposition and tracking vs a task lifecycle the agent can’t skip.](https://team.management/compare/team-management-vs-claude-task-master.md) [team.management vs BMAD Method](https://team.management/compare/team-management-vs-bmad-method.md) [A simulated agile team with heavy ceremony vs enforced discipline at plugin weight.](https://team.management/compare/team-management-vs-bmad-method.md) [team.management vs Beads](https://team.management/compare/team-management-vs-beads.md) [Agent memory and issue graphs vs process enforcement — different layers, use both.](https://team.management/compare/team-management-vs-beads.md) [team.management vs OpenSpec](https://team.management/compare/team-management-vs-openspec.md) [Discipline without ceremony, on two different layers: spec artifacts vs runtime gates.](https://team.management/compare/team-management-vs-openspec.md) [team.management vs cc-sessions](https://team.management/compare/team-management-vs-cc-sessions.md) [The DAIC ancestor, dormant since late 2025 — and the maintained engine that grew past it.](https://team.management/compare/team-management-vs-cc-sessions.md) [team.management vs autoresearch](https://team.management/compare/team-management-vs-autoresearch.md) [Karpathy’s overnight experimentation loop, run as an enforced protocol: free exploration, frozen paths.](https://team.management/compare/team-management-vs-autoresearch.md) ## Models & subscriptions Prices, credits, and limits for the major AI coding subscriptions — dated and sourced. Whichever one you pick, the process layer stays yours. [Claude Code vs Codex (July 2026)](https://team.management/compare/claude-code-vs-codex.md) [Subscriptions, credits, and limits — sourced and dated. And why you don’t have to choose.](https://team.management/compare/claude-code-vs-codex.md) [Antigravity vs Claude Code (July 2026)](https://team.management/compare/antigravity-vs-claude-code.md) [Google’s agentic stack vs Anthropic’s — quotas, credits, history, sources.](https://team.management/compare/antigravity-vs-claude-code.md) [AI coding subscriptions compared (July 2026)](https://team.management/compare/ai-coding-subscriptions.md) [Claude, ChatGPT/Codex, Google AI, Cursor, Copilot, Windsurf — one sourced, dated table.](https://team.management/compare/ai-coding-subscriptions.md) # Install in five minutes. team.management is a native Claude Code plugin — you install it from the plugin marketplace, inside Claude Code. No pip or npm package, no separate installer, and nothing is written into your repo beyond the `team-management/` directory you own. ## 1 · Install the plugin ``` /plugin marketplace add TeamManagementPlugin/claude-plugin /plugin install team-management@team-management ``` On first use the plugin's MCP server cold-starts: it builds its own isolated venv under Claude Code's managed plugin-data directory — no system Python packages — and the tools appear once it connects, usually the next turn. Once per version. (A full URL or a local checkout path works as the marketplace source too — or run bare `/plugin` and install from the picker.) Requirements: Claude Code, Python 3.10+ on your `PATH`, and git. ## 2 · Enable it for the project — and your team ``` /team-management:init ``` Run it inside the project you want managed. It merges `enabledPlugins` into the project's `.claude/settings.json` — commit that file, and every teammate auto-enables the plugin when they open the project. No secrets are written there. Windows: this step is required, not optional — it provisions the `python3` runtime the hooks and MCP server launch with (via the `py` launcher). Run it once, then fully quit and reopen Claude Code. ## 3 · Configure Non-secret settings — developer identity, DAIC options, issue tracking (GitLab / Jira / GitHub / Gitea), AI providers — via `/team-management:config`, which writes `team-management/config.json`. Provider tokens live in the per-project `.claude/state/provider-tokens.json` — git-ignored, owner-only (0600), and unreadable by Claude. It's auto-created with blank keys; open it in your editor and fill in only the tokens you use — never in `config.json`, never in the transcript. ## 4 · Create your first task ``` You → create a task for implementing user authentication ``` Claude discusses the approach with you first (it cannot edit code until you align), then starts a protocol that manages the branch, the task file, and the workflow from there. See the [protocols overview](https://team.management/protocols.md) for what runs next. ## Uninstall Run `/plugin` and uninstall team-management— the plugin's venv and caches live under Claude Code's managed plugin-data directory and are cleaned up with it. To disable it for one project instead, remove its entries from `.claude/settings.json`. Never delete `team-management/` — that directory holds your task history, `config.json`, and custom protocols. Uninstalling leaves it in place; a reinstall picks it right back up. Full source and issues on [GitHub](https://github.com/TeamManagementPlugin/claude-plugin). # DAIC — the loop on rails. DAIC is the methodology at the core of team.management: Discuss → Align → Implement → Document → Commit. A `PreToolUse` hook gates every tool call by the current mode — so Claude cannot edit code until you’ve explicitly aligned on an approach. ## Three modes | Mode | What it allows | | -------------- | ---------------------------------------------------------------------------------------------------------------------------- | | Discussion | Blocks the edit tools (Edit, Write, MultiEdit, NotebookEdit) — Claude can read, explore, and propose, but not change source. | | Implementation | Unlocks the edit tools. | | Documentation | Allows docs-only edits — Markdown and task files — while source edits stay blocked. | Mode lives in `.claude/state/daic-mode.json`. A missing or corrupt file reads as discussion — the safe default. The protocol engine sets the right mode on each step entry, so you never flip it by hand. ## Why the DAIC gate can’t be bypassed The gate is a hook that runs before the tool, in the runtime — not an instruction in the prompt. It signals the harness through its exit code: allow or block. An agent can’t talk its way past a process that intercepts the tool call itself. Read-only Bash commands still run in discussion mode (a separate allowlist gate), so Claude can keep investigating. Specialized subagents and the wiki get explicit whitelists so the framework never blocks its own machinery. ## The four enforcement layers The DAIC gate is one layer of team.management’s four-layer stack: | Layer | What it does | | --------- | ----------------------------------------------------------------------------------------------------------- | | Hooks | Intercept tool calls before they run — the gate is in the runtime, not the prompt. | | MCP | The deterministic layer — state transitions happen through tools the agent invokes but cannot fake or skip. | | Protocols | The workflow rails — which step you’re on, what’s allowed, what comes next. | | Prompts | The knowledge layer — your conventions, your wiki, loaded when they matter. | An agent can talk its way around a prompt; it can’t talk its way past the runtime. ## Branch enforcement rides along For write tools, the same hook also checks the git branch against the task’s required prefix — blocking accidental commits to the wrong branch. The task name’s action prefix picks the branch prefix: | Task prefix | Branch prefix | | ----------------------------------------------------------- | ------------- | | `implement-` · `refactor-` · `migrate-` · `test-` · `docs-` | `feature/` | | `fix-` | `fix/` | | `o-` (optimize) | `optimize/` | | `b-` (brainstorm) | `brainstorm/` | There are four failure modes the hook guards against: wrong branch, no branch, task missing, and branch missing. It fails open on a flaky git call but fails safe on missing task state. # Context preservation & auto-compact. Long sessions hit token limits. team.management intercepts both events — approaching the limit, and the compaction itself — to save a checkpoint and restore it, so the workflow continues exactly where it left off instead of forgetting the task. ## Token monitoring team.management’s `PostToolUse` hook monitors token usage for auto-compact after every tool call; as a fallback when auto-compact is off, the user-message hook warns once at 80% and again at 90%. With auto-compact on, compaction fires first — you rarely see the 90% warning. ## Auto-compact `auto_compact.enabled` (default `true`) and `auto_compact.threshold` (default `85`%) live in `team-management/config.json`. At the threshold the context-compaction protocol triggers; the `PreCompact` hook fires first and writes a checkpoint — active task, branch, protocol step, and DAIC mode. ## Restoration After a session restart, `session-start.py` reads `.claude/state/current_task.json` and injects a session-state summary — the task file, the protocol position, and the DAIC mode are all restored. After an in-session compaction, the user-message hook injects the same restoration on your next message. The agent picks up mid-protocol, not from scratch. Both `.claude/state/current_task.json` and `.claude/state/daic-mode.json` are protected — the protocol engine and hooks own them; never edit by hand. # Test-driven, when it applies. Write the test first. Watch it fail. Write the minimal code to pass. If you didn’t watch it fail, you don’t know it tests the right thing — tests written after code pass immediately, which proves nothing. ## When it applies Bug fixes (a failing test reproduces the bug, the fix makes it pass), behaviour changes, and feature work with clear inputs and outputs. Refactoring of already-tested code, where the existing suite is the safety net. Where it doesn’t — and isn’t forced: exploratory spikes, framework or config scaffolding, one-off scripts, and documentation-only changes. When uncertain, the default is to write the test — the 30 seconds to find out usually costs less than debugging later. ## Red · Green · Refactor Red — one test, one behaviour, then watch it fail for the expected reason. Green — the simplest code that passes; no extra parameters “for the future.” Refactor — clean up only after green, with tests staying green throughout. The `task` protocol’s implementation step expects this discipline; the [optimize](https://team.management/protocols/optimize.md) protocol’s frozen-paths hook physically prevents editing tests mid-experiment. # The spec-compliance gate. Stage 1 of the two-stage code review answers one question: does the diff match the promise? A read-only reviewer compares your working tree against the task’s Success Criteria — and the protocol can’t advance until it passes. ## How the sentinel works The `spec-compliance-reviewer` agent audits the git diff against the `## Success Criteria` section and returns a pass/fail verdict. On pass, the orchestrator records the literal sentinel `SPEC_REVIEW: PASSED` in the audit log. `require_spec_review_passed` runs as both an entry reminder and a hard exit block. Because it’s a stop-on-failure post-function, advancing without the sentinel is impossible by construction — not discouraged, blocked. ## Why two stages Each catches a different failure. Spec compliance catches well-written code that drifted from the spec. Code-quality review (Claude plus any AI providers, in parallel) catches correctness and security bugs. Either alone misses the other class. The advance summary is itself gated: it must carry verification evidence — a fenced output block, `N/N passed`, `exit 0` — or the escape hatch `no-verification-applicable: `. Prose like “looks good” is rejected. See the [task protocol](https://team.management/protocols/task.md). # Your protocols, your rules. Six protocols ship in the box, but they’re yours to change. Drop a JSON config in `custom/` and it overrides the system copy of the same name — with no engine change, and never overwritten on upgrade. ## Load order: custom → system → source When the engine resolves a protocol or a sub-protocol prompt, it walks an ordered search path — `custom/` first, then the installed `system/` copy, then the package source. The first match wins, so a forked `custom/task.json` shadows the shipped one cleanly. ## What you can change Override any step’s DAIC mode, swap in a custom pre- or post-function, reorder steps, or author a whole new protocol. Behaviour lives in named functions discoverable through the engine, so adding to a step is a JSON edit — not a code change. The same rule protects your customizations of shipped agents and the `CLAUDE.tm.custom.md` rules file: team-management only overwrites the names it ships on update, so a renamed copy is always safe. ## Forking a protocol — just ask Claude You don’t hand-author JSON to start. Ask Claude to run `/team-management:custom-protocol-create task` — it calls `protocol_customize`, copies the system config into `custom/task.json`, and opens it for editing. From there you describe the change in plain language and Claude edits the config with you. ``` You → fork the task protocol and add a security-review step ``` ## Staying current after an upgrade When the plugin updates, the shipped protocols may gain steps or change behaviour — but your `custom/` copies are never touched, so they can quietly fall behind. Run `/team-management:custom-protocol-update-after-reinstall` (which drives `protocol_check_drift`) to diff your custom copies against the upgraded system copies and decide what to merge. Nothing is overwritten without your say-so. # LLM Wiki — a compounding knowledge base. RAG re-derives knowledge on every question. The LLM Wiki is different: Claude incrementally builds and maintains a persistent wiki from the sources you curate — cross-references already wired, contradictions already flagged. You source and ask; Claude does the summarizing, filing, and bookkeeping. ## Three layers Raw sources (`wiki/raw/`) — immutable input you curate. Wiki pages (`wiki/pages/`) — LLM-owned markdown with YAML frontmatter; Claude writes every page. Schema (`wiki/schema.md`) — tells Claude the focus areas, page types, and tag vocabulary; auto-loaded at session start. ## Setup Opt in via `/team-management:config` — when `wiki/` is absent it offers to scaffold `wiki/index.md`, `wiki/log.md`, `wiki/schema.md`, and `wiki/raw/README.md`. Toggle with `wiki.enabled` in `config.json` (default off). The `wiki/` directory is unconditionally whitelisted in the DAIC hook, so wiki edits are never blocked regardless of mode. ## Three commands | Command | What it does | | -------------------------------------- | ------------------------------------------------------------------------------------------------------- | | `/team-management:wiki-ingest ` | Read a raw file, discuss takeaways, write a summary page, update related pages, then the index and log. | | `/team-management:wiki-tune [section]` | Interactively evolve `wiki/schema.md` via Q\&A. | | `/team-management:wiki-lint` | Health-check: orphan pages, broken links, missing frontmatter, stale or contradictory content. | ## Security: what not to put in `wiki/raw/` `wiki/raw/` is tracked in git. Never put secrets there — API keys, tokens, PEM/SSH keys, `.env` contents. `git rm` does not remove a file from history; if a secret lands there, rotate the credential and run `git filter-repo` to purge it. There is no automated secret scanning — this warning is the only guard. # MCP tools reference. The MCP server is the primary tool surface of team.management: 42 tools across 8 modules, exposed as native `mcp__plugin_team-management_tm__*` tools. They exist so workflow operations are first-class typed tools with structured errors and guaranteed discoverability — not shell commands the agent must remember to format. ## The 8 modules | Module | Tools | What it covers | | ------------------- | ------ | ---------------------------------------------------------------------------------- | | `issue_tracking.py` | 14 | Import, create, sync, and comment on issues across GitLab / Jira / GitHub / Gitea. | | `protocol.py` | 11 | Drive the JSON protocol engine — start, advance, branch, and customize workflows. | | `code_review.py` | 4 | Automated review runs, plus reading and commenting on MR / PR reviews. | | `git_operations.py` | 4 | Commit, push, and open merge / pull requests linked to the task. | | `daic.py` | 3 | Switch DAIC mode manually (the protocol engine does this per step automatically). | | `notifications.py` | 3 | Reach the user out-of-band when a step blocks on their input. | | `config.py` | 2 | Read and write non-secret settings through the guided config flow. | | `release.py` | 1 | Cut a release. | | **total** | **42** | across 8 modules | ## Complete tool reference Every tool, by module. In a Claude session each is prefixed `mcp__plugin_team-management_tm__` — e.g. `mcp__plugin_team-management_tm__protocol_start`. | issue_tracking.py | Purpose | | -------------------------------- | -------------------------------------------------------------------- | | `issue_status` | Show the active provider, linked task count, and per-task issue IDs. | | `config_issue_tracking_status` | Read the issue-tracking section of config.json. | | `config_code_review_enforcement` | Read the code-review warning-enforcement mode (strict / relaxed). | | `issue_read` | Import an issue by ID or URL and write it as a task file. | | `issue_create` | Create a provider issue from an existing task. | | `issue_update` | Update title, description, status, or labels on a linked issue. | | `issue_sync` | Push the current task status to the linked issue. | | `issue_push` | Push task file content to the linked issue description. | | `issue_link` | Attach an existing issue ID to a task. | | `issue_unlink` | Remove the issue link from a task. | | `issue_comment` | Post a comment on the linked issue. | | `issue_set_status` | Transition the issue to a new status, with an optional comment. | | `issue_api` | Make a raw authenticated API call to the active provider. | | `issue_dependency` | Declare a blocking / blocked-by relationship between issues. | | protocol.py | Purpose | | -------------------------- | ---------------------------------------------------------------------------------- | | `protocol_list` | List all available protocols with their step structures. | | `protocol_start` | Start a named protocol — the first step then drives task-file and branch creation. | | `protocol_current` | Read the active protocol name, current step, and DAIC mode. | | `protocol_advance` | Complete the current step with an evidence summary and move forward. | | `protocol_goto` | Jump back to a named step — the re-planning anchor. | | `protocol_log` | Read the protocol audit log for the current or a named task. | | `protocol_abort` | Abort the active protocol with a reason. | | `protocol_save_note` | Append a note to the protocol audit log. | | `protocol_available_funcs` | List the engine functions usable in pre_funcs / post_funcs. | | `protocol_customize` | Fork a system protocol into custom/ for editing. | | `protocol_check_drift` | Compare custom protocols against the post-upgrade system copies. | | code_review\.py | Purpose | | ----------------------- | ---------------------------------------------------------------------------------------- | | `code_review` | Run an automated review of the diff and post the results to the MR / PR when one exists. | | `fetch_mr_review` | Fetch an existing GitLab MR review by URL or IID. | | `merge_request_comment` | Post a comment on a GitLab merge request. | | `pull_request_comment` | Post a comment on a GitHub / Gitea pull request. | | git_operations.py | Purpose | | ---------------------- | ------------------------------------------------------------------------------------------------ | | `git_commit` | Stage and commit the working tree with a message. | | `git_push` | Push the current branch to origin. | | `merge_request_create` | Create a GitLab MR linked to the task's issue (GitHub / Gitea PRs open via the completion flow). | | `merge_request_update` | Update a GitLab MR's title or description, or close / reopen it. | | daic.py | Purpose | | --------------------------------- | ------------------------------------------------------- | | `daic_mode_switch_discussion` | Switch to discussion mode — blocks edit tools. | | `daic_mode_switch_implementation` | Switch to implementation mode — unlocks edit tools. | | `daic_mode_switch_documentation` | Switch to documentation mode — docs-only edits allowed. | | notifications.py | Purpose | | -------------------------------------- | ----------------------------------------------------------- | | `notify_user` | Send an out-of-band notification (e.g. Telegram). | | `notification_status` | Check whether notifications are configured and enabled. | | `notification_discover_telegram_chats` | List the chats the Telegram bot can see, to pick a chat ID. | | config.py | Purpose | | --------------- | ------------------------------------------------------------------------------------------ | | `config_get` | Read a masked config snapshot plus the schema of every settable key. | | `config_update` | Schema-validated write of non-secret settings — gated to the /team-management:config flow. | | release.py | Purpose | | ---------------- | -------------------------------------------------------------------------------------------------------------- | | `release_create` | Create a versioned release on GitLab and/or GitHub — whichever are enabled, independent of the issue provider. | ## Why MCP, and why stateless Claude agents reliably discover MCP tools in the tool list but often miss slash commands — so exposing the workflow this way makes it dependable. The server is stateless: every tool delegates to the provider utilities and protocol engine in team.management’s `hooks/` directory, so the MCP server and the hooks always run byte-identical code — no version skew. One tool was deliberately not exposed: there’s no way for Claude to disable its own DAIC enforcement. Turning off the framework is a user-only operation. # Engine functions reference. A protocol step’s behaviour isn’t hardcoded — every side effect is a named function listed in the step’s `pre_funcs` or `post_funcs`. The engine ships 45 of them across eight families; this page is the catalogue. The list is authoritative and re-callable in any session via `protocol_available_funcs()`. ## pre_funcs vs post_funcs `pre_funcs` run when a step is entered — detecting the task, loading a baseline, injecting a reminder, or resolving which AI providers to launch. `post_funcs` run when a step tries to advance — creating files, committing, syncing issues, or enforcing a gate. A `post_func` only becomes a hard wall when its step sets `post_funcs_stop_on_failure: true` — that’s what turns `require_spec_review_passed`, `check_completion_evidence`, and `verify_tests_pass` from advisory reminders into structural gates. (Placement in `pre_funcs` is cosmetic for gating — the advance only checks `post_funcs`.) ## Eight families | Family | Count | What it covers | | ------------------------- | ------ | -------------------------------------------------------------------------------- | | Task lifecycle & state | 5 | Detect, create, and track the task and its state file. | | Git & branches | 5 | Branch, commit, merge, push, and open the merge request. | | Gates & verification | 5 | The structural walls — wire as stop-on-failure post_funcs to actually gate. | | Issue tracking | 2 | Mirror task state to the configured provider. | | Completion & cleanup | 7 | Wrap the task up — dispatch the finish flow, archive, and reset state. | | Refactoring test baseline | 3 | Capture a green test baseline as the refactoring safety net. | | Optimize loop | 11 | The metric-driven autoresearch family — measure, experiment, and gate on gaming. | | Knowledge & AI providers | 7 | Wiki upkeep and the per-phase parallel-provider launch instructions. | | **total** | **45** | across eight families | ## Complete function reference Every function, by family. _When_ is where it typically wires — `pre` (on step entry) or `post` (on advance). | Task lifecycle & state | When | What it does | | -------------------------------- | ---- | -------------------------------------------------------------------------- | | `auto_detect_task` | pre | Detect the active task from the current git branch. | | `verify_branch_and_task` | pre | Check the branch matches task state and the task file exists (warns only). | | `create_task_file` | post | Write the task markdown from AI-provided content, validating frontmatter. | | `set_task_state` | post | Set current_task.json and rename the pending protocol log to the task. | | `update_task_status_in_progress` | post | Flip the task file's status frontmatter to in-progress. | | Git & branches | When | What it does | | ---------------------- | ---- | --------------------------------------------------------------------------- | | `git_setup_branch` | post | Create and check out the task branch; can carry uncommitted changes across. | | `git_commit` | post | Stage and commit the working tree with a task-derived message. | | `git_merge_main` | post | Fetch and merge the default branch, reporting any conflicts. | | `git_push` | post | Push the current branch to the remote with upstream tracking. | | `create_merge_request` | post | Open an MR / PR linked to the task's provider issue. | | Gates & verification | When | What it does | | ---------------------------------------- | ---------- | ----------------------------------------------------------------------------------------- | | `require_spec_review_passed` | pre / post | Require a SPEC_REVIEW: PASSED note in the log before leaving code-review. | | `check_completion_evidence` | post | Block the advance unless the summary carries verification evidence (or the escape hatch). | | `verify_tests_pass` | post | Run the configured test command; a non-zero exit blocks the advance. | | `validate_code_review_in_worklog` | post | Confirm the code-review results were appended to the work log. | | `validate_no_critical_issues_in_worklog` | post | Fail if the latest review block reports any critical issues. | | Issue tracking | When | What it does | | ------------------------- | ---- | ------------------------------------------------------- | | `create_issue_if_enabled` | post | Create a provider issue when issue tracking is enabled. | | `update_issue_status` | post | Move the linked issue to completed / closed. | | Completion & cleanup | When | What it does | | ------------------------------ | ---- | ---------------------------------------------------------------------------------- | | `present_completion_options` | pre | With no issue provider set, offer the merge-local / push-PR / keep / discard menu. | | `completion_dispatch` | post | Run the completion chain: archive, commit, merge, push, status, cleanup, checkout. | | `require_discard_confirmation` | post | Two-step typed confirmation before a discard force-deletes the branch. | | `archive_task` | post | Move the finished task file into tasks/done/. | | `cleanup_task_scoped_state` | post | Remove the task-scoped state directory. | | `clear_task_state` | post | Reset current_task.json to the empty state. | | `checkout_default_branch` | post | Switch back to main / master after completion. | | Refactoring test baseline | When | What it does | | ------------------------- | ---- | ------------------------------------------------------- | | `capture_test_baseline` | post | Snapshot the test command and result before any change. | | `load_test_baseline` | pre | Load the saved baseline for regression comparison. | | `cleanup_test_baseline` | post | Remove the baseline file once refactoring completes. | | Optimize loop | When | What it does | | ------------------------- | ---------- | ------------------------------------------------------------------------------------ | | `validate_optimize_setup` | post | Validate optimize settings before any durable side effects are created. | | `write_optimize_setup` | post | Persist the validated optimize settings to state. | | `capture_metric_baseline` | post | Run the metric once on HEAD and record the baseline. | | `validate_metric_script` | post | Pre-flight the metric script — run twice, check the parser and stability. | | `run_metric` | pre / post | Run the metric command N times under a filtered env and aggregate. | | `log_experiment_result` | post | Measure HEAD and append an experiment row — the engine, not the LLM, owns the value. | | `check_cost_estimate` | post | Project the wall-clock cost of the run before experimentation starts. | | `check_termination` | post | After each iteration, test the four stop conditions and end when one matches. | | `update_best_commit` | post | Record the best commit + metric in the task frontmatter when a run improves. | | `batch_checkpoint` | post | Pause at each batch boundary for user approval (interactive optimize). | | `policy_compliance_audit` | post | Heuristic scan for metric-gaming; flags feed the review prompt (never blocks). | | Knowledge & AI providers | When | What it does | | ----------------------------------------------- | ---- | ----------------------------------------------------------------------------- | | `wiki_update_reminder` | pre | During documentation, remind Claude to update the LLM wiki (if enabled). | | `resolve_ai_providers` | pre | Return the parallel AI-provider launch instructions for the code-review gate. | | `resolve_ai_providers_for_investigation` | pre | …for the task investigation phase. | | `resolve_ai_providers_for_implementation` | pre | …for the implementation planning phase. | | `resolve_ai_providers_for_brainstorm` | pre | …for the brainstorm analysis phase. | | `resolve_ai_providers_for_exploration` | pre | …for the research exploration phase. | | `resolve_ai_providers_for_refactoring_planning` | pre | …for the refactoring planning phase. | ## Discoverable, not hardcoded Because behaviour lives in these named functions, adding to a step is a JSON edit, not an engine change — wire a function into a step’s `pre_funcs` / `post_funcs` and it runs. To build your own workflow, fork a protocol (see [customization](https://team.management/docs/customization.md)) and compose these engine functions; for the tools that drive the engine itself, see the [MCP reference](https://team.management/docs/mcp.md). # Issue tracking — four providers, one interface. Connect Claude tasks to your tracker. Import an issue and the task file is created; complete the task and the issue closes, the MR ships, the label flips — all driven by MCP tools, the same way across every provider. ## The four providers | Provider | `issue_tracking.provider` | Notes | | -------- | ------------------------- | --------------------------------------------------------------------------------------------------------- | | GitLab | `gitlab` | Issues + merge-request workflow. | | Jira | `jira` | Jira API with markdown-to-wiki conversion. | | GitHub | `github` | Issues + pull-request workflow. | | Gitea | `github` | Same key — auto-detected from the base URL (`/api/v1` path or a `gitea` host); uses label IDs, not names. | | None | `disabled` | Completion shows a 4-option menu instead of a provider flow (see Completion dispatch). | ## Import, create, sync `issue_read` takes an ID or full URL and writes a task file under `team-management/tasks/`; pass update mode to re-fetch and overwrite an existing one. `issue_create` goes the other way — pushes a task's title and description to a new provider issue. `issue_sync` pushes the current task status to the linked issue, and `auto_sync` does it automatically on task-lifecycle events. ## Completion dispatch With a provider configured, the completion step merges the branch, opens the MR/PR, and transitions the issue. With `provider: "disabled"`, it instead offers a four-option menu — `merge_local` / `push_pr` / `keep` / `discard`. `push_pr` uses `gh pr create` with an idempotency precheck; `discard` is gated behind a typed confirmation (friction, not security). Either way, `HEAD` must match the task's feature branch before any local flow runs. All of this is exposed through the `issue_*` tools on the [MCP reference](https://team.management/docs/mcp.md). # AI providers — Codex & Antigravity. Turn on AI providers and team.management brings Codex and Antigravity in alongside Claude as parallel collaborators at six decision points in the workflow. They run through the `codex` and `agy` CLIs — your existing subscriptions, no extra API keys — each reading the repo read-only in its own sandbox. ## The six phases | Phase | Config flag | What the provider does | | ---------------------- | --------------------------------- | ---------------------------------------------------------------------- | | `code_review` | `include_in_code_review` | Parallel security/quality review of the diff, at the code-review gate. | | `brainstorm` | `include_in_brainstorm` | Analysis alongside the six specialist agents. | | `investigation` | `include_in_investigation` | Independent reading of task scope and risks. | | `implementation` | `include_in_implementation` | Plan review before code is written. | | `research_exploration` | `include_in_research_exploration` | Independent exploration of the research question. | | `refactoring_planning` | `include_in_refactoring_planning` | Review of the refactoring plan. | ## How the providers run team.management runs each enabled provider as a parallel Task agent — not an MCP server — dispatched in the same message as Claude’s own agent for that phase. Output is advisory: a provider failure degrades gracefully and never blocks the workflow. Claude’s pass always runs; the others are additive. Enable providers globally with `ai_providers.enabled_providers` (e.g. `["codex", "agy"]`), then turn participation on per phase with the `include_in_*`flags — so you can run Codex on code-review only, or Antigravity everywhere. The wrappers enforce a fixed call deadline — codex 300 s, agy 330 s watchdog (the `ai_providers.timeout` config key is currently inert). See the [protocols](https://team.management/protocols.md) for where each phase sits. ## Credentials never leave the box Before team.management injects a task description into any provider prompt, it passes through a credential filter — 17 named regex patterns (API keys, PEM blocks, bearer tokens, connection strings, `.env` contents) matched line-by-line. The first match wins; the whole line is replaced with `[REDACTED:]`. It’s defense-in-depth for the task-description channel — the codebase itself is read by the provider’s own read-only CLI sandbox. ## Custom prompt templates Five of the six phases (everything except `code_review`) load their provider prompt from a markdown file, so you can tune it per project without touching code. Drop a `-.md` in `custom/providers/` (phase names hyphenated — e.g. `codex-investigation.md`, `agy-refactoring-planning.md`). A missing template is non-fatal: the engine falls back to its inline default and warns. These variables are available to a template: | Variable | Resolves to | | ------------------ | ------------------------------------------------------------- | | `{task_name}` | The active task's name. | | `{branch}` | The task's git branch. | | `{task_file_path}` | Path to the task markdown file. | | `{phase}` | Human-readable phase name (e.g. “task investigation”). | | `{plan_summary}` | The full task markdown, credential-filtered before injection. | Legacy keys. `include_in_architecture`, `include_in_exploration`, and the old `gemini.*` provider keys (e.g. `gemini.default_model`) are deprecated — Antigravity is configured under `agy.*` now. Their values are never auto-forwarded — session-start emits a one-time warning when they appear so you can migrate deliberately. # task > The standard implementation lifecycle. task carries one unit of work from “understand the request” through to “merged, issue closed, archived.” The engine applies the right DAIC mode on each step and runs the git/issue plumbing automatically — your job is to discuss, write code, and provide verification evidence. Everything else is enforced structurally. ``` protocol_start(protocol_name="task") ``` ## Steps (5 steps) 1. **investigation** — discussion — Read-only. Understand scope, align on an approach, compose the task file. No code is written here. This is also the re-planning anchor — any later step returns here when an assumption breaks. 2. **implementation** — implementation — Full edit access. Write code to satisfy every success criterion — no new scope, no drive-by refactors. Stop-on-blocker rather than guess. Changes are not committed yet. 3. **code-review** — implementation — The two-stage gate. Stage 1: spec-compliance against the success criteria. Stage 2: parallel code-quality review (Claude + any AI providers). Cannot advance without passing both. 4. **documentation** — documentation — Docs-only edits — source edits are hook-blocked. Update affected CLAUDE.md files, finalize the work log, refresh the wiki if enabled. 5. **completion** — discussion — You verify, then the engine does all git/issue/archive work automatically: branch merge, push, MR/PR, issue close, archive. ## What makes it distinctive ### The two-stage code-review gate Stage 1 (“does the diff match the promise?”) catches well-written code that drifted from spec — the spec-compliance-reviewer audits the diff against your Success Criteria and writes a SPEC_REVIEW: PASSED sentinel. Stage 2 (“is the diff correct/secure?”) dispatches the Claude code-review agent plus one Task per configured AI provider, in parallel. Both stages are post_funcs with stop-on-failure — un-bypassable by construction. ### Commits happen only at completion The LLM never commits prematurely. The engine creates the branch up front (an investigation post-func), but every commit, merge, push and MR/PR lives in the completion step’s engine-driven funcs, which branch on your issue-tracking provider (or present a 4-option menu when none is configured). ### Evidence-gated advance The advance summary out of code-review must carry literal verification evidence — a fenced output block, “N/N passed”, “exit 0” — or the escape hatch no-verification-applicable: \. Prose like “looks good” is rejected. branch: `feature/ · fix/ — set by task prefix` # brainstorm > Parallel-specialist ideation → planned tasks. No code. brainstorm turns a fuzzy “should we build X?” into a justified results document plus a set of concrete, coverage-audited implementation tasks — without writing a line of production code. Its centerpiece is a 6-specialist parallel analysis with a conflict-resolution loop. Every step runs in documentation mode by design. ``` protocol_start(protocol_name="brainstorm") ``` ## Steps (5 steps) 1. **topic** — documentation — Define the topic in one sentence, why it’s being considered, which modules are affected. Compose the brainstorm task file with its scaffold (Decisions, Expert Analysis, Conflicts). 2. **discussion** — documentation — Deep discussion recording every accepted decision. Doubles as the conflict-resolution step — re-entered from analysis when specialists disagree. 3. **analysis** — documentation — The defining step: 6 specialist subagents fan out in a single parallel-dispatch message (plus AI providers when configured). Conflicts loop back to discussion. 4. **results** — documentation — Synthesize the analysis into a justified results document — the recommendation and its rationale. 5. **planning** — documentation — Turn the results into concrete, coverage-audited implementation tasks, handed off to separate task protocol runs. ## What makes it distinctive ### Six parallel specialists Analysis launches six specialist subagents in one message — architecture, code-impact, critique, user-perspective, risks & security, scope & phasing — each in its own context window. When AI providers are configured they join in the same parallel dispatch. ### Conflict-resolution loop When specialists disagree, the protocol loops back to discussion: it presents the competing viewpoints, you decide, resolutions are recorded as new decisions, and every open question must be cleared before analysis re-runs. ### No production code The product is knowledge and a plan. Source edits are hook-blocked throughout; the actual build happens in a separate task run. branch: `brainstorm/ — b-brainstorm-` # research > Spikes, PoCs, evaluations. No branch, no code. research is the protocol you reach for when you don’t yet know enough to write a task. It produces a knowledge artifact — findings plus a recommendation — rather than shippable code. There’s no commit/push/MR step and no git branch; the only writes are to the research task file. ``` protocol_start(protocol_name="research") ``` ## Steps (4 steps) 1. **scoping** — discussion — Formulate a specific, answerable research question. Classify the type (spike / architecture / evaluation / exploration), bound the scope, set a time box, compose the task file. 2. **exploration** — discussion — Deep, read-only investigation — gather evidence, analyze code, run read-only experiments. code-explorer agents and AI providers do the heavy lifting. 3. **synthesis** — documentation — Distill the evidence into findings and a recommendation in the task file. 4. **conclusion** — discussion — You review; the engine archives the task. No branch to merge, no dispatcher menu. ## What makes it distinctive ### Four research types Scoping classifies the work as a spike (feasibility → PoC), architecture (design analysis → ADR), evaluation (tool comparison → weighted matrix), or exploration (codebase understanding → documented patterns). Each gets a tailored investigation strategy. ### No branch by design Research tasks declare branch: none. There’s no git_setup_branch in scoping and no completion dispatcher in conclusion — research output is knowledge, not a diff. branch: `none — r-, always branchless` # refactoring > Test-baseline-gated restructuring. refactoring changes the shape of the code without changing its behaviour, with a test suite as the safety net at both ends. Its defining feature: it captures a test baseline on the default branch before any change, then formally compares against it after — so a regression cannot slip through unnoticed. ``` protocol_start(protocol_name="refactoring") ``` ## Steps (6 steps) 1. **test-baseline** — discussion — Run the suite on the default branch and snapshot the result before any change. Pre-existing failures are stored so they aren’t later mistaken for regressions. 2. **planning** — discussion — Define the refactor scope, affected files, success criteria, and a risk pass. The branch is created NOW, cleanly separating baseline from work. 3. **refactoring** — implementation — Restructure incrementally, testing as you go. 4. **test-verify** — discussion — Run the suite on the refactored branch and compare against the baseline — a regression gate. Read-only by design: a failing comparison sends you back, not deeper in. 5. **code-review** — implementation — The same two-stage gate as task’s — spec-compliance sentinel included. The one deliberate difference: the clean-tests gate (verify_tests_pass) is omitted, because the baseline comparison — not a clean exit 0 — is the behavioural guarantee. 6. **completion** — discussion — Engine-driven git/issue/archive, same dispatcher as task. ## What makes it distinctive ### Baseline-first ordering Step 1 runs before the branch exists, on main/master, so the baseline reflects the true starting point — not a state already perturbed by the refactor. The branch is created in step 2. ### Formal regression comparison test-verify compares the post-refactor suite against the captured baseline. Behaviour is the contract; the diff is allowed to be large as long as the baseline holds. branch: `feature/ — -refactor-` # optimize > Metric-driven optimization, interactive batched. optimize turns “make this number better” into a reproducible hypothesis loop. You define a single scalar metric and a measurement script; the engine — not the LLM — measures every commit and logs a TSV leaderboard, so the result cannot be gamed by a hallucinated number. It pauses for your approval at batch boundaries. At completion it squashes the branch down to baseline..best_commit and ships an MR/PR carrying the leaderboard. ``` protocol_start(protocol_name="optimize") ``` ## Steps (7 steps) 1. **setup** — discussion — Interactive metric elicitation: the metric + direction, termination conditions, noise control, batch size, frozen paths. Writes the validated settings to optimize-state.json (the baseline step later adds the measured baseline to it). 2. **metric-script** — implementation — Author the measurement script — prints exactly one float, exit 0, deterministic. Validated by a double-run stability gate before you can advance. 3. **baseline** — implementation — Capture the baseline metric on HEAD; project total cost from real timing. 4. **experimentation** — implementation — The hypothesis loop: one hypothesis = one focused commit. The engine measures each commit. Checkpoints between batches return to discussion for your approval. 5. **synthesis** — documentation — Findings document, plus a non-blocking metric-gaming audit. 6. **code-review** — implementation — Full two-stage gate on the cumulative baseline..best_commit diff. 7. **completion** — discussion — Squash from best_commit; ship the leaderboard as the MR/PR description. ## What makes it distinctive ### Engine-owned measurement The engine runs the metric script on every commit and writes the TSV leaderboard. The LLM’s reported numbers are untrusted — the leaderboard is the single source of truth, so it cannot be gamed. ### Frozen paths During experimentation, the metric script and test files are frozen — a hook physically blocks edits to them even in implementation mode. You can’t move the goalposts mid-experiment. ### Batch checkpoints After each batch of hypotheses the protocol returns DAIC to discussion for your approval before continuing — keeping a human in the loop on an otherwise mechanical loop. branch: `optimize/ — o-` # optimize-unattended > Autonomous twin of optimize. For overnight runs. optimize-unattended is the autonomous twin of optimize: the same metric-driven hypothesis loop, but with no batch checkpoints. The experimentation step runs unattended to a termination condition (or a manual interrupt), making it suitable for overnight and weekend runs. Six of its seven steps are byte-identical to optimize — only experimentation differs. ``` protocol_start(protocol_name="optimize-unattended", max_iterations=200, max_duration="8h") ``` ## Steps (7 steps) 1. **setup** — discussion — Same interactive metric elicitation as optimize. batch_size is still elicited but inert — there’s no checkpoint to gate. 2. **metric-script** — implementation — Same metric-script contract + double-run stability gate. 3. **baseline** — implementation — Same baseline capture + cost projection, plus an unbounded-mode typed acknowledgment. 4. **experimentation** — implementation — The autonomous loop — no DAIC switches mid-run, no per-iteration user gate. Runs until a termination condition fires or the operator interrupts. 5. **synthesis** — documentation — Same findings document + non-blocking metric-gaming audit. 6. **code-review** — implementation — Same full two-stage gate. 7. **completion** — discussion — Same squash-from-best_commit + leaderboard MR/PR. ## What makes it distinctive ### No checkpoints — runs to termination The experimentation step drops the batch_checkpoint post_func. It keeps going until check_termination fires (max_iterations, max_duration, regression-halt, or target metric reached) or you interrupt manually. ### Termination conditions Set bounds up front: a max iteration count, a wall-clock duration (e.g. “8h”), a consecutive-regression halt, or a target metric value. The first to fire ends the run. ### Same engine, same guarantees Engine-owned measurement, frozen paths, and the squash+leaderboard completion are shared verbatim with optimize. Only the human-in-the-loop pause is removed. branch: `optimize/ — o-` # team.management vs GitHub Spec Kit. [GitHub Spec Kit](https://github.com/github/spec-kit) structures what to build — specs, plans, and task documents. team.management enforces how work proceeds — gated steps the agent can’t skip. They solve different problems, and they compose. ## What Spec Kit is Spec Kit is GitHub’s toolkit for spec-driven development: `/speckit.specify` turns an idea into a spec, then plans and task lists follow, all as markdown artifacts that work across two dozen coding agents. It ships from GitHub’s own org, releases near-daily, and sits at ≈122k stars as of July 2026. The idea: give the agent a precise, agreed description of the work, and better output follows. It often does — and the documents are useful to humans too. ## What team.management is [team.management](https://team.management) is an open-source protocol engine that runs _inside_ Claude Code. A spec describes the destination. A protocol is the track that gets you there: investigation, alignment, implementation, review. The engine tracks each step, and a `PreToolUse` hook blocks the wrong tools at the wrong time — in the runtime, not in the prompt. See [the DAIC loop](https://team.management/docs/daic.md) for the mechanics. ## The core difference — documents vs gates A spec is advice the agent reads. A protocol is a state machine the agent is _inside_. Spec Kit’s artifacts live in the context window, where the model weighs them against everything else it has read — and on a long session it drifts. team.management’s steps live outside the context window. The engine knows which step you’re on, and the hook blocks implementation tools until discussion is aligned, whatever the model currently “thinks.” Can you rip the rails out? Of course — you own the repo. Edit the config and the gates are gone. But the _agent_ can’t do that mid-task. A document can’t promise even that. | | GitHub Spec Kit | team.management | | --------------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | What it structures | The work product: spec → plan → tasks documents | The working process: gated protocol steps with DAIC modes | | Runtime enforcement | None — artifacts are context, the agent can drift from them | PreToolUse hook blocks edit tools until the step allows them | | Workflow lifecycle | Linear document pipeline per feature | Named protocols: task, brainstorm, research, refactoring, optimize | | Git & task automation | None (documents only) | Branch per task, task files, status transitions — automated | | Code review | Not part of the kit | Enforced review step; can panel Codex and Antigravity as reviewers | | Agent support | 24+ agents, agent-agnostic | Claude Code native (reviewers via Codex/Antigravity subscriptions) | | Repository | [github.com/github/spec-kit](https://github.com/github/spec-kit) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free, MIT | Free, MIT | ## Use the kit — on rails You don’t have to choose. Spec Kit is the best tool around for pinning down what “done” means; team.management makes the path to it non-negotiable. Draft the spec with `/speckit.specify`, then run the implementation as a team.management task protocol: the spec defines the work, the gates hold the process. If unclear requirements are your only pain, the kit alone will carry you far. The day the model ships around the spec, add the rails under it. ## FAQ **Is GitHub Spec Kit enough to keep an AI agent on process?** Spec Kit produces excellent spec, plan, and task documents, but they are inputs to the agent’s context — nothing at runtime stops the agent from drifting away from them mid-session. If you need the process itself enforced (no code before alignment, no completion without review), that is the layer team.management adds. **Can I use Spec Kit and team.management together?** Yes, and it’s a genuinely good combination: use Spec Kit to write the spec and plan, then run the implementation through a team.management protocol so the agreed process is enforced while the agent builds against that spec. **Does team.management replace Claude Code?** No. team.management is a plugin that runs inside Claude Code — it adds a protocol engine, DAIC tool-gating, task files, git branch automation, and code-review gates on top of the harness you already use. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # team.management vs Superpowers. [Superpowers](https://github.com/obra/superpowers) and team.management make the same promise — disciplined AI development — with opposite mechanisms. Superpowers, by its own README, works “not through tool gating but via structured instructions.” team.management is the tool gating. ## What Superpowers is Superpowers is Jesse Vincent’s composable skills framework and full development methodology — brainstorm → plan → TDD → subagent execution → mandatory review — installed as a plugin across Claude Code, Codex, Cursor, Copilot CLI and others. It is the most-starred workflow plugin in the space (≈256k stars, v6.x, as of July 2026) and its TDD discipline in particular is genuinely deep. Its README states the mechanism plainly: the methodology holds _“not through tool gating but via structured instructions.”_ The skills trigger automatically and tell the model what a disciplined engineer would do next. ## What team.management is [team.management](https://team.management) is the other mechanism: an open-source protocol engine inside Claude Code where the process lives in a state machine, not in the instructions. A `PreToolUse` hook checks every tool call against the current protocol step — in discussion mode the edit tools are blocked before they execute. Nothing needs to be remembered, because nothing is being asked. The mechanics are documented in [the DAIC loop](https://team.management/docs/daic.md). ## The core difference — asking well vs not asking Structured instructions work most of the time — the model usually complies, and Superpowers’ results show it. “Usually” is exactly the problem enforcement exists to solve. Instructions compete with everything else in context: a long debugging session, a user in a hurry, an eager model that’s sure it’s right. A `PreToolUse` hook doesn’t compete for attention. In discussion mode the edit tool comes back blocked — every time, at hour one and at hour eleven. Could you disable it? Sure — it’s your machine and your config, and the diff would show it. The model doesn’t get that option. That’s the point: you stay free; the agent stays on the rails. | | Superpowers | team.management | | --------------------- | ------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------ | | Mechanism | Skills — structured instructions the model follows | Hooks — runtime gates the model can’t override | | Own words | “not through tool gating but via structured instructions” | “A hook that runs before the tool, in the runtime — not an instruction in the prompt” | | Methodology | Brainstorm → plan → TDD → subagent execution → review | Protocols: task, brainstorm, research, refactoring, optimize | | Process state | In the conversation and skill state | In the engine: step, mode, task file, audit log — outside the context window | | Git & task automation | Via skills and workflows | Branch per task, task files, status transitions — engine-managed | | Harness reach | Claude Code, Codex, Cursor, Copilot CLI and more | Claude Code native (Codex/Antigravity join as reviewers) | | Repository | [github.com/obra/superpowers](https://github.com/obra/superpowers) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free, open source | Free, MIT | ## Run the skills — on rails Superpowers’ methodology and skill library are excellent, and skills run happily _inside_ protocol steps: let the brainstorm and TDD skills do what they do best, and let the engine hold the line those skills assume — implementation locked until alignment, review actually happening before “done.” Working solo, with your eyes on the session, Superpowers alone may be all you need. But if you answer for agents you weren’t watching, “the model was instructed to” is not an answer you can give a client. Put the same skills on rails. ## FAQ **What is the actual difference between skills and hooks?** A skill is an instruction set the model loads and follows — it improves behavior but the model can still deviate, especially deep into a long session. A hook runs in the harness before each tool call and allows or blocks it by exit code. One shapes the model’s intent; the other constrains its actions regardless of intent. **Is Superpowers’ methodology compatible with team.management?** Largely yes. Superpowers’ brainstorm-plan-TDD-review methodology maps naturally onto team.management’s protocol steps, and skills can run inside protocol steps. If you rely on Superpowers’ multi-agent patterns, check them against DAIC mode gating — subagents get explicit whitelists in team.management. **Which one should a team standardize on?** If you want a rich skill library and multi-harness reach (Claude Code, Codex, Cursor and more), Superpowers is the broader ecosystem. If you need a guarantee that every member’s agent followed the process — not a well-worded request that it do so — only the hook layer can give you that, and that layer is team.management. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # team.management vs gstack. [gstack](https://github.com/garrytan/gstack) staffs Claude Code with a virtual team — CEO, eng manager, designer, QA — as 23 role-skills. team.management asks a different question: not who the agent plays, but whether the process those roles describe actually gets followed. ## What gstack is gstack is Garry Tan’s open-source Claude Code setup — “23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA,” in the repo’s own words. Released March 2026, it hit ≈122k stars by July 2026 and is updated near-daily. Everything is slash commands and Markdown skills: `/plan-ceo-review` rethinks the product, `/review` hunts production bugs, `/qa` drives a real browser, `/ship` lands the PR. It also ships opt-in safety guardrails — `/careful` warns before destructive commands (and can be overridden), `/freeze` locks edits to one directory, `/guard` does both. That’s more safety machinery than most skill packs carry. ## What team.management is [team.management](https://team.management) is an open-source protocol engine inside Claude Code that treats your _process_ — not your staffing — as the thing to get right. Named protocols sequence the steps, a `PreToolUse` hook gates the tools at each one, and every task gets a branch, a work log, and an audit trail ([the DAIC loop](https://team.management/docs/daic.md) has the mechanics). The roles you hire can change by the day; the rails under them don’t. ## The core difference — personas vs process state gstack’s roles are personas the model performs when you invoke them: excellent prompts, voluntarily followed, one command at a time. Between commands, nothing holds — the model can ship without `/review`, skip `/qa`, or “fix one more thing” after the review passed. In team.management the engine holds the lifecycle: it knows you’re in code-review, keeps you there until the gate passes, and writes the audit log either way. gstack’s guardrails are scoped to a session and overridable by a sentence; a DAIC gate is scoped to the task and isn’t. Can the rails be torn up too? By you — any time; it’s MIT-licensed config in your own repo. By the model mid-session — no. That difference is what a gate means. | | gstack | team.management | | --------------------- | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | Mechanism | 23 role-skills + slash commands (all Markdown) | Protocol engine + PreToolUse hooks (runtime gating) | | Safety features | /careful warnings (overridable), /freeze directory lock — opt-in, per session | DAIC modes per protocol step — always on, per task | | Workflow lifecycle | You invoke the right role at the right time | The engine sequences steps; skipping isn’t offered | | Review & QA | Strong skills (/review, /qa) run on request | Review is a gated step; AI providers can join as reviewers | | Git & task automation | /ship, /land-and-deploy skills | Branch per task, task files, status transitions — engine-managed | | Tuned for | Solo founders and small teams shipping fast | Teams needing the process provable per agent | | Repository | [github.com/garrytan/gstack](https://github.com/garrytan/gstack) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free, MIT | Free, MIT | ## gstack hires the team; team.management puts it on rails Adopt, don’t choose: gstack’s roles are the best off-the-shelf staff your Claude Code can hire, and they slot into protocol steps naturally — `/plan-ceo-review` during investigation, `/review` and `/qa` inside the code-review gate. gstack brings the roles; team.management makes the sequence non-optional. Shipping solo, for yourself, gstack alone is a joy. The moment someone else needs to trust what shipped — a client, a cofounder, a teammate — keep the same roles and add rails under them. ## FAQ **Doesn’t gstack’s /guard and /freeze count as enforcement?** They’re real and useful: /careful warns before destructive commands (and is overridable by design), /freeze restricts edits to one directory while debugging. They’re session-scoped safety conveniences you toggle on demand. team.management enforces a workflow lifecycle — which step you’re in, what tools that step allows, what gates stand between implementation and done — with state that persists across the whole task. **Can gstack and team.management run together?** Yes in principle — gstack’s review and QA skills are the kind of thing you’d run inside a team.management code-review step. Both ship opinionated CLAUDE.md guidance, so expect to reconcile instructions; the protocol engine itself doesn’t conflict with gstack’s slash commands. **Which is better for a team rather than a solo founder?** gstack is tuned for the solo builder shipping like a team — that’s its origin story. team.management is built around the accountability problem teams have: every member runs the same protocols, and the process is enforced per agent rather than suggested per prompt. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # team.management vs Claude Task Master. [Claude Task Master](https://github.com/eyaltoledano/claude-task-master) turns a PRD into a dependency-aware task graph and serves it to your editor over MCP. team.management wraps each task in an enforced lifecycle — investigation, implementation, review, completion — with gates between the steps. ## What Claude Task Master is Task Master (npm `task-master-ai`) is the original AI task manager for coding agents: feed it a PRD, get a dependency-aware task graph, and work through it from Cursor, Claude Code, or Windsurf over MCP. It earned its ≈28k stars (July 2026) by solving decomposition well — big fuzzy goals become ordered, sized, trackable units. Worth knowing before you commit to it: as of July 2026 its release cadence has slowed — last release March 31, 2026; last push April 28, 2026. Maintained, but decelerating. ## What team.management is [team.management](https://team.management) is an open-source protocol engine inside Claude Code where a task is more than a row in a list: starting one cuts a git branch, loads a context manifest, and walks the agent through investigation → implementation → review, with a `PreToolUse` hook gating the tools at every step ([the DAIC loop](https://team.management/docs/daic.md)). “Done” becomes a state the engine grants — not a status the model self-reports. ## The core difference — what a task is In Task Master a task is a _tracking record_: title, dependencies, status the agent updates. In team.management a task is a _running process_: it owns a git branch, a context manifest, a work log, and a position in a protocol — and the DAIC hook decides which tools are available at that position. The difference shows when the agent gets eager: a tracker ends up with stale statuses; an enforced lifecycle ends up with a blocked tool call and a paper trail. You can switch all of this off — enforcement binds the agent, never the owner. What it adds while it’s on is the one thing a tracker can’t give you: statuses that were earned, not self-reported. | | Claude Task Master | team.management | | ----------------------- | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------ | | Core job | PRD → dependency-aware task graph, tracked over MCP | Task → enforced lifecycle (investigate → implement → review → complete) | | Runtime enforcement | None — the agent works freely against the list | PreToolUse hook gates tools by protocol step | | Decomposition | Strong: AI-generated graph with dependencies and sizing | Manual/protocol-guided task creation with context manifests | | Git integration | None built in | Branch per task, enforced branch discipline, automated transitions | | Code review | Not part of the model | Gated review step; Codex/Antigravity can join as reviewers | | Editor reach | Cursor, Claude Code, Windsurf and more (MCP) | Claude Code native | | Maintenance (July 2026) | Slowing — last release 2026-03-31 | Actively developed | | Repository | [github.com/eyaltoledano/claude-task-master](https://github.com/eyaltoledano/claude-task-master) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free — MIT with Commons Clause (no-resale condition) | Free, MIT | ## Task Master builds the graph; team.management runs it on rails Use them in sequence, not in competition. Task Master turns a PRD into an ordered graph, and each node of that graph makes a natural team.management task: a branch, a context manifest, an enforced lifecycle. Decomposition upstream, enforcement downstream. If your only pain is shaping the backlog, Task Master alone covers it. When “done” has to mean something you can stand behind, run each task through the rails. ## FAQ **Is Claude Task Master still maintained?** As of July 2026 it reads as maintained but slowing: the last release (0.43.1) shipped at the end of March 2026 and the last repository push was late April 2026. It remains widely used; factor the cadence into a long-term bet. **Doesn’t a task list already keep the agent on track?** A task list tells the agent what remains; it doesn’t constrain how any task gets done. Task Master’s agent can implement without discussing, skip review, or mark items done unverified — nothing checks. In team.management each task runs through protocol steps with tool gating and a review gate, so “done” has a defined, enforced meaning. **Claude Code now has native task tracking — where does that leave both tools?** Claude Code’s 2026 native tasks absorb plain cross-session task tracking. What the harness doesn’t absorb is Task Master’s PRD decomposition (a real feature) and team.management’s process enforcement (a different layer entirely). Judge each tool on the part the harness can’t absorb, not on the overlapping checklists. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # team.management vs BMAD Method. [BMAD Method](https://github.com/bmad-code-org/BMAD-METHOD) simulates an entire agile team — analyst, PM, architect, scrum master, dev, QA — that plans before anyone codes. team.management gets you the discipline without hiring twelve personas: the process is enforced by the engine, not acted out by a cast of agents. ## What BMAD Method is BMAD (≈51k stars as of July 2026) staffs your project with 12+ agent personas that run a full agile ceremony: an Analyst researches, a PM writes the PRD, an Architect designs, a Scrum Master slices stories, then Dev and QA implement and check. It’s the “enterprise ceremony” end of AI-assisted development, and for large greenfield efforts the ritual produces genuinely thorough plans. ## What team.management is [team.management](https://team.management) is an open-source protocol engine inside Claude Code built from a handful of concepts — protocols, steps, modes, gates — each enforced by the runtime rather than performed by a persona. A `PreToolUse` hook checks every tool call against the current step ([the DAIC loop](https://team.management/docs/daic.md)). That one hook does the whole job. ## The core difference — performed process vs enforced process BMAD’s discipline is a performance: roles act out a process, and your guarantee is only as good as the performance. Every artifact is advice to the next persona, and the model playing “Dev” is as free to drift as any other agent. team.management drops the cast and keeps the process: steps, gates, modes, and an audit log, enforced by hooks no matter how many agents are involved. You lose the ceremony. You gain a guarantee that survives a bored model at hour six. Can you remove it? Any time — it’s config in your repo. Can the “Dev” persona remove it at 2am, three hours into a session? No. No document can make that promise. | | BMAD Method | team.management | | --------------------- | ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------ | | Model | 12+ personas performing agile ceremony | One protocol engine gating real steps | | Discipline via | Generated documents reviewed by other personas | Runtime tool gating + gated review step | | Weight per feature | Heavy — hours of multi-agent planning on complex work | Light — protocol steps over your normal Claude Code session | | Best at | Greenfield planning depth (PRD, architecture docs) | Everyday execution integrity (no skipped steps, no unreviewed merges) | | Runtime enforcement | None | PreToolUse hook blocks tools by DAIC mode | | Git & task automation | Story files; no branch enforcement | Branch per task, status transitions, work logs — engine-managed | | Repository | [github.com/bmad-code-org/BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free, open source | Free, MIT | ## Let BMAD plan — let the rails hold it The two combine well. BMAD’s personas produce PRDs and architecture documents that make superb context manifests for team.management tasks: planning depth upstream, enforced execution downstream — and the “Dev” persona can no longer drift from the architecture it was handed. If a greenfield push only needs planning depth, BMAD alone delivers it. When execution matters as much as the plan, run the implementation on rails. ## FAQ **Is BMAD’s ceremony worth it?** For complex greenfield work where requirements genuinely need PRD-grade thinking, BMAD’s up-front planning earns its cost. Community comparisons consistently place it at the slow, thorough end of the spectrum — hours of agent time per feature. For day-to-day tasks that weight is the main complaint. **Do BMAD’s personas enforce anything?** No — the personas produce documents and review each other’s outputs, but it’s all generated advice. The dev agent can diverge from the architecture doc mid-implementation and nothing intervenes. team.management enforces at the tool-call layer instead: fewer documents, harder guarantees. **Can I combine BMAD-style planning with team.management?** Yes — BMAD’s PRD/architecture outputs make strong inputs for a team.management task’s context manifest, and the implementation then runs under enforced DAIC gates. Teams that love BMAD’s planning but not its unguarded execution do exactly this. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # team.management vs Beads. [Beads](https://github.com/gastownhall/beads) solves agent amnesia — a git-backed issue graph that survives across sessions. team.management solves agent drift — enforcement within them. They are different problems, and you can fix both at once. ## What Beads is Beads is Steve Yegge’s answer to the “50 First Dates” problem: coding agents that wake up every session with no memory. It keeps a graph-based issue tracker in git — issues, dependencies, status — that agents read and write, so context compounds instead of evaporating. It moved to the gastownhall org and sits at ≈25k stars (July 2026), with Gas Town (≈17k stars) building a multi-agent workspace manager on top. ## What team.management is [team.management](https://team.management) is an open-source protocol engine inside Claude Code working the other axis: not what the agent _knows_, but what it may _do_ right now. Steps, modes, and gates are held by the engine — outside the context window — and a `PreToolUse` hook applies them to every tool call ([the DAIC loop](https://team.management/docs/daic.md)). Perfect recall doesn’t grant permission. ## The core difference — remembering vs constraining Beads makes the agent _know more_: what was decided, what’s blocked on what, what happened last week. team.management makes the agent _able to do less_ at any given moment: no edits during discussion, no completion without review. A perfect memory doesn’t stop an agent from cutting corners, and a gate doesn’t tell it what happened last week. Different problems, different layers. And yes — you can uninstall the gates whenever you like. They bind the agent per tool call, never you. Advice in a prompt binds no one at all. | | Beads | team.management | | ------------------- | -------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | Core job | Persistent memory: git-backed issue graph across sessions | Enforced process: gated protocol steps within work | | Data model | Graph of issues + dependencies, owned locally in git | Task files + protocol state; issues linked to GitHub/GitLab/Jira | | Runtime enforcement | None — it informs, it doesn’t gate | PreToolUse hook gates tools by DAIC mode | | Cross-session story | Excellent — that’s the product | Process-scoped: manifests, work logs, audit logs per task | | Multi-agent | Gas Town builds workspaces on Beads | AI providers (Codex/Antigravity) join reviews; one lead agent per task | | Repository | [github.com/gastownhall/beads](https://github.com/gastownhall/beads) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free, open source | Free, MIT | ## Beads is the memory; team.management is the rails Run them as one stack: Beads carries the backlog and the cross-session context; each item of that backlog runs as a team.management task with enforced steps. The graph remembers, the gates hold — neither replaces the other. If amnesia is your only pain, Beads alone cures it. The day you stop watching every session, add the rails. ## FAQ **Do Beads and team.management overlap?** Less than the category suggests. Beads is a memory substrate: issues, dependencies, and context that persist between sessions. team.management is a process layer: what the agent may do right now, given the step it’s in. The overlap is issue tracking, where team.management links tasks to GitHub/GitLab/Jira rather than owning a graph database. **Can they run together?** Yes, naturally: Beads carries the backlog and cross-session context; each item of work then runs as a team.management task with enforced steps. Neither tool fights the other’s layer. **Doesn’t team.management also preserve context?** It preserves process context — task files, context manifests, work logs, protocol audit logs — which covers the lifecycle of a task. Beads goes further as a general memory graph across everything the agent touches. If cross-session memory is your primary pain, Beads is the stronger dedicated answer. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # team.management vs OpenSpec. [OpenSpec](https://github.com/Fission-AI/OpenSpec) is the light option among the spec-driven tools — discipline without BMAD’s ceremony. team.management takes the same instinct one layer down: not lighter documents, but gates that don’t depend on the documents being obeyed. ## What OpenSpec is OpenSpec (≈61k stars as of July 2026) is a spec-driven workflow for Claude Code, Codex, Cursor and Copilot that deliberately trims the ceremony: a change starts as a proposal, becomes GIVEN/WHEN/THEN specs and a design, executes as tasks, and archives. “Discipline without ceremony” is its explicit pitch, and in the widely repeated “BMAD vs Spec Kit vs OpenSpec” comparisons it wins on speed. ## What team.management is [team.management](https://team.management) is an open-source protocol engine inside Claude Code, built from a handful of runtime concepts — steps, modes, gates — instead of a document pipeline. One `PreToolUse` hook does the enforcing ([the DAIC loop](https://team.management/docs/daic.md)); the protocols stay as light as the specs. ## The core difference — thin documents still aren’t gates Lighter documents make the workflow nicer. They don’t change the mechanism: an OpenSpec spec is still advice in the context window, and the agent honoring it is still voluntary. team.management is minimal in a different way — few concepts, each one backed by the runtime. OpenSpec trimmed the ceremony and kept the trust model. team.management keeps the lightness and changes the trust model. And yes — you can strip the gates out as easily as you added them; they bind the agent, not you. But a light document holds no better than a heavy one when the model stops cooperating. A gate holds either way. | | OpenSpec | team.management | | ------------------- | ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------ | | Philosophy | Discipline without ceremony — light spec artifacts | Discipline that doesn’t rely on trust — enforced steps | | Pipeline | Proposal → spec → design → tasks → archive | Investigation → implementation → review → documentation → completion | | Runtime enforcement | None — documents guide the agent | PreToolUse hook gates tools by DAIC mode | | Speed on a change | Fastest of the spec-driven cluster (community benchmarks) | Adds gate friction by design — that friction is the product | | Agent support | Claude Code, Codex, Cursor, Copilot | Claude Code native (Codex/Antigravity join reviews) | | Repository | [github.com/Fission-AI/OpenSpec](https://github.com/Fission-AI/OpenSpec) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free, open source | Free, MIT | ## OpenSpec writes the specs; team.management holds the gates Use both and the stack stays light: OpenSpec’s proposal → spec flow defines the change, and a team.management protocol carries the implementation through gates the model can’t skip. If unclear intent is your only pain, OpenSpec alone is the fastest cure. If the process itself needs to hold, add the gates under the specs. ## FAQ **How does OpenSpec differ from GitHub Spec Kit and BMAD?** OpenSpec is the lightweight corner of that triangle: proposal → GIVEN/WHEN/THEN specs → design → tasks → archive, with far less generated ceremony than BMAD and a slimmer pipeline than Spec Kit. Community benchmarks in 2026 repeatedly place it as the fastest of the three on the same task. **If OpenSpec is already lightweight, what does team.management add?** The guarantee. OpenSpec’s discipline still travels as documents the agent reads and may drift from; nothing at runtime holds the line. team.management adds the enforcement layer — tool gating per step, a review gate before done — under whatever spec workflow you keep. **Together or either/or?** Together works: OpenSpec structures the change proposal, team.management enforces the implementation lifecycle. If you only take one: pick by your dominant failure mode — unclear intent (OpenSpec) vs unenforced process (team.management). Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # team.management vs cc-sessions. [cc-sessions](https://github.com/GWUDCAP/cc-sessions) pioneered the idea team.management is built on: DAIC modes enforced by hooks, not prompts. It has been dormant since October 2025; team.management carries the same DNA forward — credited on our landing page — as a maintained, full protocol engine. ## What cc-sessions is cc-sessions (GWUDCAP) is the opinionated Claude Code harness that proved the core mechanism: block Edit/Write in discussion mode with a hook, flip modes on explicit trigger phrases, keep tasks in files, enforce git discipline. Modest in stars (≈1.5k) but big in influence: it proved that runtime gating works, and a generation of “make Claude behave” tools followed. As of July 2026 the repository is dormant: last commit October 17, 2025. For a tool whose value lives in hooks tracking a fast-moving harness, dormancy is the one status that matters. ## What team.management is [team.management](https://team.management) carries cc-sessions’ core insight — gate the tools, not the prompts — into a maintained, open-source engine: named protocols sequence whole lifecycles, the `PreToolUse` hook sets the mode per step ([the DAIC loop](https://team.management/docs/daic.md)), and every task gets a branch, a work log, and an audit trail. ## The difference — a mechanism vs an engine cc-sessions is the mechanism: DAIC gating plus conventions. team.management builds it out: a protocol engine (over MCP) that sequences whole lifecycles — task, brainstorm, research, refactoring, optimize — sets the DAIC mode per step, manages branches and task state, runs gated code review with optional Codex/Antigravity reviewer panels, writes audit logs, and syncs issues with GitHub/GitLab/Jira. If you liked cc-sessions, nothing needs unlearning — the concepts are the same, held by an engine instead of by convention. The limit is inherited too: hooks bind the agent, never the operator. That was the original design’s point, and it still is. | | cc-sessions | team.management | | ------------------ | ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------ | | Status (July 2026) | Dormant — last commit 2025-10-17 | Actively developed | | DAIC gating | Yes — the original implementation | Yes — inherited, set per protocol step by the engine | | Workflow model | Conventions + trigger phrases | Named protocols with step state, goto, and audit logs | | Beyond the loop | — | AI-provider review panels, issue tracking, notifications, wiki | | Task management | Hand-maintained task files | Engine-created tasks, branches, status transitions | | Repository | [github.com/GWUDCAP/cc-sessions](https://github.com/GWUDCAP/cc-sessions) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free, open source | Free, MIT | ## When cc-sessions is still the right call If you want the smallest possible thing that gates tools by mode — a few hooks you can read in an afternoon and own outright — cc-sessions still delivers exactly that, dormancy accepted. If you want the idea maintained and grown into a team-grade engine, that is why team.management exists. cc-sessions is credited on [the landing page](https://team.management): team.management started from its idea. ## FAQ **Is cc-sessions abandoned?** As of July 2026 it appears dormant: the last commit landed October 17, 2025. It still works as published, but there’s no active maintenance to track Claude Code’s changes — which, for a hooks-based tool, matters. **Is team.management a fork of cc-sessions?** It’s a descendant in design rather than a fork: the DAIC discussion/implementation gating, task files, and git discipline carry over as concepts. Around them team.management adds what cc-sessions never had — an MCP protocol engine with multiple named protocols, step state and audit logs, AI-provider review panels, issue tracking, notifications, and a wiki system. **I use cc-sessions today — what does migrating look like?** Install the team.management plugin, run its init, and your working habits mostly transfer: discussion-first remains the default, trigger-phrase-style alignment becomes protocol steps. Task files are created by the engine per task rather than by hand. The enforced feel of DAIC is the same — that’s the inherited part. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # team.management vs autoresearch. [autoresearch](https://github.com/karpathy/autoresearch) is Andrej Karpathy’s overnight-experimentation pattern — the inspiration for the optimize protocols, credited on [our landing page](https://team.management). The difference: the loop still explores freely, but the path around it is frozen. ## What autoresearch is autoresearch (≈91k stars as of July 2026, released March 2026) is Karpathy’s demonstration of a research org run by agents: give an AI a small but real LLM training setup and let it experiment autonomously overnight — modify `train.py`, train for a fixed five minutes, keep the change if `val_bpb` improved, repeat. You wake up to a log of experiments and, hopefully, a better model. The repo is deliberately small — three files — because the point is the pattern. The README is clear about how you steer it: you don’t touch the Python — you edit `program.md`, the Markdown instruction file that sets up your “autonomous research org.” And the training harness, `prepare.py`, is protected by one line of instructions: _“Do not modify prepare.py.”_ ## What team.management is [team.management](https://team.management)’s [optimize](https://team.management/protocols/optimize.md) and [optimize-unattended](https://team.management/protocols/optimize-unattended.md) protocols take that idea and put an enforced lifecycle around it — for any repo and any scalar metric. Setup asks for the metric, the stopping rules, and the frozen paths. The measurement script is checked for determinism before the loop may start. A baseline is captured. Then the loop runs — one hypothesis, one commit, each measured _by the engine, not the model_. Completion squashes the baseline-to-best diff and ships the leaderboard with the MR/PR. Two protocols share that lifecycle. optimize is the interactive twin: checkpoints between batches return control to you for approval. optimize-unattended removes the checkpoints and runs to a termination condition — the autoresearch shape, built for exactly the overnight run. ## The core difference — what holds the loop In autoresearch, everything that holds the loop is prose. `program.md` tells the agent how to run it. “Don’t touch `prepare.py`” is a convention the agent honors because it was asked. Even the scoreboard runs on trust: the agent creates and updates `results.tsv` itself. In a reference repo, that’s a feature — you iterate on the instructions and watch what happens. In a repo you answer for, running unwatched at 4am, it’s a gap: instructions hold until they don’t. team.management moves those load-bearing parts into the engine — the step sequence, the untouchable paths, the measurement, the stopping rules — and leaves the model free where freedom is the point: inside the experiment. Overnight, that gap has a precise name: nothing stops a fabricated result. autoresearch’s agent runs the training, greps its own `run.log` for the metric, writes its own row to `results.tsv`, and marks it keep or discard. A run that hallucinated a better number — or quietly edited the evaluation — produces the same morning log as one that earned it. optimize closes that path. The engine runs the measurement itself on every commit — “LLM-passed metric values are not trusted,” in its own words. The metric script and eval data sit behind frozen paths. And at synthesis, an audit sweeps the whole run for gaming patterns: frozen-path edits, hardcoded best-metric constants, hand-edited leaderboard rows. The audit is a heuristic; the weight sits on the engine-owned measurement. Could you unfreeze the paths mid-run? You — yes; it’s your config and your repo. The agent — no. Overnight, that difference decides what you wake up to: a leaderboard, or a surprise. | | autoresearch | team.management | | ------------------- | ---------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | | What it is | A deliberately small reference repo (nanochat training) | Protocols built into the team.management engine | | The loop is held by | program.md — Markdown instructions the human iterates | Seven enforced steps: setup → metric script → baseline → experimentation → synthesis → review → completion | | Metric integrity | The agent checks its own result — and maintains results.tsv itself | The engine measures every commit and writes the leaderboard; the model never touches the scoreboard | | A faked result | Indistinguishable from a real one — the morning log is self-reported | Can’t reach the leaderboard — the engine runs the measurement; a synthesis audit flags gaming patterns | | Untouchable files | “prepare.py — not modified”: a stated convention | Frozen paths elicited at setup, enforced during the loop | | Overnight stopping | Let it run; read the log in the morning | max_iterations / max_duration / regression halts / target metric — or interrupt | | What you wake up to | A log of experiments and (hopefully) a better model | A squashed baseline→best diff + leaderboard, gated by code review | | Domain | Single-GPU LLM training, val_bpb | Any scalar metric with a measurement script | | Repository | [github.com/karpathy/autoresearch](https://github.com/karpathy/autoresearch) | [github.com/TeamManagementPlugin/claude-plugin](https://github.com/TeamManagementPlugin/claude-plugin) | | Price & license | Free to read; no license file as of July 2026 | Free, MIT | ## autoresearch is the reference; team.management runs it on rails The idea is Karpathy’s, and the credit is on [our landing page](https://team.management). To study the pattern in its purest form, read autoresearch — it’s three files, and a joy to read. When the same pattern has to run against a repo you answer for — with a metric nobody can hallucinate and a diff somebody will review — run it as [optimize-unattended](https://team.management/protocols/optimize-unattended.md). Free experimentation, frozen paths. ## FAQ **Is autoresearch a product I can point at my own codebase?** By design, no — it’s a deliberately small reference implementation for one domain: an agent iterating on single-GPU nanochat training against a fixed 5-minute budget and a val_bpb metric. Its value is the pattern, not the packaging. Generalizing that pattern to any repo and any scalar metric — with enforcement around it — is exactly what team.management’s optimize and optimize-unattended protocols are. **What did team.management change from autoresearch’s pattern?** Three things moved from prose into the engine. The loop’s structure: program.md instructions became seven enforced protocol steps (setup → metric script → baseline → experimentation → synthesis → code review → completion). The metric: instead of the agent checking whether its own result improved, the engine measures every commit and keeps a leaderboard the model can’t game with a hallucinated number. The ending: instead of a morning log, you get a squashed baseline-to-best-commit diff with the leaderboard attached, gated by code review. **Is it actually safe to leave optimize-unattended running overnight?** That’s what the guardrails are for: frozen paths declared at setup and enforced during the loop, engine-side measurement of every commit (self-reported numbers never reach the leaderboard), termination conditions (max iterations, max duration, halt after N consecutive regressions, or target metric reached), manual interrupt at any time, a metric-gaming audit at synthesis — and a full code-review gate on the cumulative diff before anything ships. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # Claude Code vs Codex (July 2026). The teams getting the most out of the Claude Code vs Codex choice increasingly don’t make it — they run both, on one enforced process. Here it is with sourced, dated numbers, the most-asked comparison in agentic coding. ## Subscriptions, side by side Figures below are as of July 2026, from [claude.com/pricing](https://claude.com/pricing) and [OpenAI’s pricing docs](https://learn.chatgpt.com/docs/pricing). Prices move fast in this space — before you budget, check the linked sources rather than anyone’s summary, including ours. | | Claude (Claude Code) | ChatGPT (Codex) | | ------------ | --------------------------------------------- | ------------------------------------------------- | | Budget entry | — | Go — $8/mo, Codex included | | Standard | Pro — $17/mo billed annually, $20 monthly | Plus — $20/mo | | Power | Max 5x — $100/mo · Max 20x — $200/mo | Pro — from $100/mo, 5x or 20x rate-limit variants | | Team | Team — $20/seat/mo annual (premium seat $100) | Business — $20/user/mo annual ($25 monthly) | ## How the limits actually work | | Claude (Claude Code) | Codex (ChatGPT) | | ------------------------ | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | | Metering | Rolling 5-hour session windows, weekly caps on top | Per-model message ranges per 5-hour window + a weekly cap, token-based credits (since April 2026) | | Weekly limits | Max plans carry two — one across all models, one for Sonnet | One weekly cap, shared with ChatGPT’s other agentic features | | Published message counts | None — Anthropic doesn’t publish fixed counts | ≈15–90 messages/window on the top model (Plus), more on lighter ones | Claude shares usage between the Claude app and Claude Code; the [session-and-weekly model is documented here](https://support.claude.com/en/articles/11145838). Because Anthropic deliberately doesn’t publish fixed message counts, any page quoting exact numbers is recycling stale 2025 data. ## Where team.management fits team.management doesn’t compete with either — it’s the process layer that runs _inside_ Claude Code. Your protocols enforce the same investigation → implementation → review lifecycle whichever model does the work, and during review steps Codex joins as an independent second opinion, on the ChatGPT subscription you already pay for (see [AI providers](https://team.management/docs/ai-providers.md)). Pick your primary on model taste and budget. The process stays provider-neutral — so when the pricing changes under you, the choice stays reversible. ## FAQ **Which subscription is cheaper for heavy agentic coding?** As of July 2026 the entry points are comparable — Claude Pro from $17–20/month, Codex included in ChatGPT Plus at $20 (and even Go at $8) — and both escalate to $100–200 power tiers (Claude Max 5x/20x; ChatGPT Pro with 5x/20x variants from $100). Real cost depends on your usage shape: Anthropic caps by rolling 5-hour sessions plus weekly limits; OpenAI meters per-model message windows backed by token-based credits. **What happens when I hit the limits?** Claude: wait for the session or weekly reset, upgrade, or opt in to usage credits / pay-as-you-go at API rates — always with explicit consent. Codex: buy additional credits or drop to a lighter model; API-key usage bills at standard API rates. Neither silently charges you. **Do I have to choose between Claude Code and Codex?** No — and this is team.management’s whole angle: it runs your process on Claude Code and can bring Codex in as an independent reviewer during code-review steps, using the ChatGPT subscription you already pay for, no extra API keys. The models check each other; the process stays one thing. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # Antigravity vs Claude Code (July 2026). Google’s agentic coding stack matured fast in 2026 — Antigravity 2.0, a CLI, an SDK, and a rebuilt $100 Ultra tier at I/O. It’s a real alternative to Claude Code now, with a history worth knowing before you commit a team to it. ## The two stacks Antigravity is Google’s agentic development environment — desktop IDE, CLI, and SDK since I/O 2026 — gated by consumer Google AI subscriptions rather than a standalone developer SKU. Access to Vertex Model Garden models (including non-Google ones) comes with the [Google AI Pro tier](https://support.google.com/googleone/answer/14534406); the [$100/month Ultra tier (I/O 2026)](https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2026-developer-highlights/) carries 5x Antigravity usage. Claude Code has the simpler pricing: [Pro at $17–20/month, Max at $100 (5x) or $200 (20x)](https://claude.com/pricing), usage in rolling 5-hour sessions with weekly caps, and an explicit, opt-in path to API pay-as-you-go when you need more. | | Antigravity (Google AI) | Claude Code (Claude) | | ---------------- | --------------------------------------------------------------------------- | -------------------------------------------------------------- | | Free tier | Yes — public preview, rate-limited baseline | No meaningful Claude Code use on free | | Paid tiers | Pro (reported \~$20/mo) · Ultra $100/mo (5x); 20x variant unpriced publicly | Pro $17–20/mo · Max $100 (5x) · Max $200 (20x) | | Limits model | Tier baseline quota + purchasable credit bundles (2,500/5,000/20,000) | Rolling 5-hour sessions + weekly caps, shared app + Code | | Transparency | Credit→token mapping unpublished — the top developer complaint | Message counts unpublished; windows and reset rules documented | | When you run out | Buy credit bundles or wait for refresh | Wait, upgrade, or opt in to API-rate usage (explicit consent) | | Track record | Dec 2025–Mar 2026 quota cuts and protests; reset at I/O 2026 | Weekly caps added in 2025; the limits scheme has held since | Quota history documented by [DevClass (March 2026)](https://devclass.com/ai-ml/2026/03/13/users-protest-as-google-antigravity-price-floats-upward/5209219). All figures as of July 2026 — verify against the linked sources before budgeting. ## Where team.management fits team.management runs on Claude Code and treats Antigravity as a _reviewer_: during protocol code-review steps, the engine can dispatch your changes to Antigravity through the Google subscription you already hold — no API keys, no second process to govern (see [AI providers](https://team.management/docs/ai-providers.md)). If you’re Antigravity-primary, one thing to know before you plan around it: the protocol engine is Claude Code-native today, so enforcement applies to the Claude Code side of your work. ## FAQ **What does Antigravity actually cost?** As of July 2026: individual use remains free in public preview at a rate-limited baseline. Higher limits come from Google AI subscriptions — the Pro tier (reported around $20/month) and Google AI Ultra at $100/month with 5x Antigravity usage, announced at I/O 2026. A 20x variant is confirmed to exist; Google hasn’t published its price. When quota runs out, you can buy AI-credit bundles (2,500 / 5,000 / 20,000 credits) — Google doesn’t publish the credit-to-token mapping. **What was the Antigravity quota controversy?** Between December 2025 and March 2026, free and Pro quotas were repeatedly cut — documented cases showed a Pro user’s weekly throughput dropping by more than 30x — and refresh windows moved from 5-hourly to weekly, prompting public protest and a Google statement in March 2026. I/O 2026 reset the story with the $100 Ultra tier. The episode is worth weighing: quota policy is part of the product. **Can I use Antigravity and Claude Code together?** Yes — that’s team.management’s normal mode: Claude Code runs the enforced process, and Antigravity joins code-review steps as an independent reviewer through your existing Google subscription. You get Google’s models checking Anthropic’s work (and vice versa) without the process fragmenting. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources. # AI coding subscriptions compared (July 2026). The subscriptions are capacity; the process is the part you keep. That rule is what keeps sane the two or three AI coding subscriptions most developers now juggle at once. Here is the whole landscape in one dated, sourced table. ## The landscape, one table All figures as of July 2026, from each vendor’s pricing page (linked). Where a vendor doesn’t publish a number, the cell says so. | Provider | Standard tier | Power tier | Limits model | When you run out | | --------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | ---------------------------------------------------- | --------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- | | [Claude](https://claude.com/pricing) (Claude Code) | Pro — $17/mo annual, $20 monthly | Max — $100 (5x) / $200 (20x) | Rolling 5-hour sessions + weekly caps; counts unpublished by design | Wait, upgrade, or opt-in usage credits / API rates | | [ChatGPT](https://learn.chatgpt.com/docs/pricing) (Codex) | Plus — $20/mo (Go — $8) | Pro — from $100, 5x/20x variants | Per-model messages per 5-hour window + weekly cap; token-based credits since Apr 2026 | Buy credits or drop to a lighter model; API at standard rates | | [Google AI](https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2026-developer-highlights/) (Antigravity) | Pro — reported \~$20/mo; free preview baseline exists | Ultra — $100/mo (5x); 20x variant unpriced publicly | Tier baseline quota; credit bundles 2,500/5,000/20,000; token mapping unpublished | Buy credit bundles or wait for refresh | | [Cursor](https://cursor.com/pricing) | Pro — $16/mo billed annually, $20 monthly ($20 usage included) | Pro+ $60 ($70 included) · Ultra $200 ($400 included) | First-party pool (Auto is metered, not unlimited) + frontier usage in dollars at API rates | Opt-in on-demand usage at the same API rates, billed in arrears | | [GitHub Copilot](https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/) | Pro — $10/mo ($10 in AI Credits) | Pro+ — $39/mo ($39 in credits) | AI Credits consumed by token usage since Jun 2026; no rollover; completions stay free on paid plans | Purchase additional usage at published API rates | | [Windsurf](https://devin.ai/pricing) (Cognition) | Pro — $20/mo | Max — $200/mo | Daily + weekly auto-refreshing quotas since the Mar 2026 overhaul | Overage at API pricing | ## The pattern behind the table Three things converged in 2026. Standard tiers clustered at $17–20 and power tiers at $100–200. Metering converged on tokens, under increasingly similar wrappers. And every vendor kept one lever to itself: the quota policy — the same lever that moved against users in the Antigravity quota cuts of early 2026. The conclusion for a team: treat subscriptions as interchangeable capacity, and keep what you can’t afford to re-buy — your process, its enforcement, its audit trail — in a layer you own. ## That layer is the point of team.management [team.management](https://team.management) is the open-source, MIT-licensed process layer under whichever subscriptions you run: protocols enforce your lifecycle in Claude Code, and [AI providers](https://team.management/docs/ai-providers.md) lets Codex and Antigravity join code review through those same subscriptions — no extra API keys, no per-seat fee, nothing to re-buy when the subscription pricing changes. Which it will. ## FAQ **Why do all these products meter differently?** 2026 converged on token-based credits under the hood — OpenAI moved Codex to token metering in April, GitHub Copilot replaced premium requests with AI Credits in June, Google sells credit bundles, Cursor meters frontier models in dollars at API rates. What differs is the wrapper: session windows (Anthropic), per-model message ranges (OpenAI), included-dollar pools (Cursor), monthly credit grants (Copilot), refresh quotas (Windsurf). **What’s the cheapest way to run agents seriously?** As of July 2026, standard tiers sit at $17–20/month almost everywhere (Claude Pro, ChatGPT Plus, Google AI Pro, Cursor Pro, Windsurf Pro), with Copilot Pro at $10 and ChatGPT Go at $8 below them; power tiers land at $100–200 (Claude Max, ChatGPT Pro, Google Ultra, Cursor Ultra, Windsurf Max). The real cost driver isn’t the sticker price — it’s whether the way you work fits the metering. Long daily sessions and burst weekends stress session windows and credit pools differently; check that fit before upgrading. **Where does team.management fit in this landscape?** It’s not a subscription and not a model — it’s the open-source process layer that stays constant while you shuffle providers. Protocols run in Claude Code; Codex and Antigravity join reviews through subscriptions you already pay for. Whichever column you pick from this table next year, the process — and its audit trail — doesn’t change. Facts and figures on this page are as of July 2026, verified against the sources linked inline. If you’re reading this much later — check the sources.