SPEC-DRIVEN DELIVERY FOR AI CODING AGENTS

AI writes the code. The spec decides what ships.

Test Maze holds your user stories and acceptance tests, keeps your coding agent inside them, verifies the work independently and gates the release. Works with Claude Code, Cursor, Cline, Gemini CLI, Codex CLI and any MCP client.

Free Basic plan Two commands to connect Never reads your repo
~/my-app · claude code + test maze
> Build a password reset flow. Tests first.
SPEAKS MCP · WORKS WITH✦Claude Code✦Cursor✦Cline✦Gemini CLI✦Codex CLI✦Any MCP client
HOW THE LOOP WORKS

You ask. Your agent builds. The spec decides.

The story and its acceptance tests are written first and reviewed by a person. Your agent builds inside them, a separate verifier grades the work, and the release ships on the results.

You
Coding agent
Test Maze MCP
1 · Spec first
“Build the password reset flow. Plan it and write the tests first.”
Saves a user story with acceptance cases across 8 test facets. The criteria are the test cases.
feature.implement · userstory.create · case.create_batch
You review the story, ask for changes or mark it Reviewed, and tick Definition of Ready. Agents cannot.
2 · Every change traced
Declares the story it is implementing on this branch. Until then, the hook blocks edits.
spec.start_work
Writes the code. Every commit carries a Testmaze-Ref trailer; CI re-checks the range.
3 · Verified independently
A separate verifier agent runs the acceptance cases against the running app and records the results, pinned to the commit.
testrun.create · testrun.record_results
Fixed rules decide. Pass: ship. Fail: which case broke, and whether to fix the code or the test.
pdlc.verify
4 · Shipped on evidence
The gate checks the linked runs and open defects. Green and clear: the release ships. Otherwise it is refused, and only a person can override, with a reason.
release.readiness · release.ship
1 · Spec first
  1. You → Coding agent
    “Build the password reset flow. Plan it and write the tests first.”
  2. Coding agent → Test Maze MCP
    Saves a user story with acceptance cases across 8 test facets. The criteria are the test cases.
    feature.implement · userstory.create · case.create_batch
  3. Test Maze MCP → You
    You review the story, ask for changes or mark it Reviewed, and tick Definition of Ready. Agents cannot.
2 · Every change traced
  1. Coding agent → Test Maze MCP
    Declares the story it is implementing on this branch. Until then, the hook blocks edits.
    spec.start_work
  2. Coding agent (on its own)
    Writes the code. Every commit carries a Testmaze-Ref trailer; CI re-checks the range.
3 · Verified independently
  1. Coding agent → Test Maze MCP
    A separate verifier agent runs the acceptance cases against the running app and records the results, pinned to the commit.
    testrun.create · testrun.record_results
  2. Test Maze MCP → Coding agent
    Fixed rules decide. Pass: ship. Fail: which case broke, and whether to fix the code or the test.
    pdlc.verify
4 · Shipped on evidence
  1. Coding agent → Test Maze MCP
    The gate checks the linked runs and open defects. Green and clear: the release ships. Otherwise it is refused, and only a person can override, with a reason.
    release.readiness · release.ship
FOUR PILLARS

Spec it. Build it. Prove it. Ship it.

Each one is a thing that exists in the product today, not a promise.

01 · SPEC FIRST

No story, no code.

Your agent drafts user stories and acceptance criteria from a feature title; people review them and tick Definition of Ready. The acceptance criteria are the test cases.

feature-spec · review queue · Ready / Done · 8 test facets
02 · EVERY CHANGE TRACED

No reference, no commit.

Turn on spec-based development and the agent cannot edit until it declares the story, every commit carries a Testmaze-Ref, and CI re-checks the range. The traceability matrix shows requirement → cases → commits → defects for every story.

Claude Code hooks · git hooks · spec check-range · Traceability
03 · VERIFIED INDEPENDENTLY

Stop letting AI grade its own homework.

A separate verifier runs every acceptance test against your running app and records the results pinned to the commit. The verdict is fixed rules over those results; no model decides pass or fail.

testmaze-verifier · pdlc.verify · recordedBy · git sha + branch + clean tree
04 · SHIPPED ON EVIDENCE

The gate opens on results, not confidence.

Failed tests become defects with a regression test and a GitLab issue. A release ships only when its runs are green and no blocker is open; a person can override, with a reason, on the record.

ship gate · release.readiness · defects · release notes from results

See it in your workflow

Pick a workflow and watch every tool call travel between you, your agent, Test Maze and your app.

Best for: Vibe coders & founders

“Set up Test Maze for this project”

You ask in plain English.No test framework to learn.

Step 1 / 11
0
lines of your repository read
Your agent sends results, not source.
2
commands to connect
One token per project, from npm.
8
test facets per feature
From happy path to accessibility.
66
MCP tools for your agent
Ask in plain English; it picks.
THE GAP

AI coding agents are great writers and poor judges. And nobody wrote the spec down.

WITHOUT

The agent writes the spec, the code and the grade

No spec to build against
Requirements live in a chat window. Each session the agent rebuilds its own idea of “done”.
Nothing stops an unasked-for change
Rules files and plan modes are suggestions. Code lands that no story asked for.
Self-grading drift
The agent edits the assertion until the test passes. The bug ships.
Shipping on confidence
A tidy summary says it is ready. Nobody can point at the run, the commit or the open bugs.
WITH TEST MAZE

Test Maze holds the spec and the verdict

Spec first
User stories and acceptance tests are on record and reviewed by a person before any code, so the target cannot move.
Every change traced
No edit without a declared story, no commit without a Testmaze-Ref. CI re-checks the range.
Verified independently
A separate verifier runs the tests. Fixed rules decide; every run carries the git sha, branch and a clean-tree flag.
Shipped on evidence
Defects hold the release. The gate opens on green runs, and an override needs a person and a reason.
BUILT FOR

Everyone who ships with AI.

Same product, different first pillar. Whether you vibe-code a weekend app, own the acceptance criteria for a team or have to show an auditor the chain, start where it hurts.

YOUR REPO STAYS YOURS

We hold the spec and the verdict. We never read your repo.

What reaches Test Maze: the product summary your agent writes, stories, test cases and results, any code snippet an agent chooses to paste into a code-quality check, and — only when spec-based development is on — commit subject lines and Testmaze-Ref trailers. Never diffs, never files.

Stays on your machine

Your repository.env & secretsYour IDE’s LLMRendered promptsGit diffs & history

Reaches Test Maze

Product summary your agent writesStories & test casesRun resultsGit sha · branchCommit subjects + Testmaze-Ref trailers (only with spec-based development on)Screenshots (opt-in)Snippets you ask us to review

Exactly what we see, and what we never do →

WIRE IT UP

Two commands. Your agent now has a spec, a verifier and a gate.

Create an MCP token in your workspace, then run these in your project folder. The token lands in a git-ignored .env.testmaze; init also installs the Claude Code and git hooks and the testmaze-verifier subagent.

Full setup guide
terminal · zsh
npx -y @testmaze/mcp init tmt_xxx
claude mcp add tm --scope project \
  -- npx -y @testmaze/mcp

Spec it. Build it.
Prove it. Ship it.

Five minutes to install. Your agent drafts the story, a person approves it, a separate verifier grades the work and the release ships on the results. Test Maze never reads your repository. Start on the free Basic plan; upgrade when your team scales.