Define the destination
- Describe the goal in plain language
- List acceptance criteria that decide when it's done
- Approve the plan, if you want a checkpoint
- Review the pull request and merge
NH Wheel turns a goal and its acceptance criteria into a tested, reviewed pull request. Write tasks in NH Wheel or sync them from GitHub and GitLab. A team of AI agents plans the work, breaks it down, writes the code, tests it and opens the PR. You review and merge.
Most AI coding tools still need someone at the keyboard. NH Wheel works on a different contract: you describe the outcome and how it will be judged, and agents handle everything between the ticket and the pull request.
wheel/ branchesFollow a real task, WHL-231 · Add Apple Pay to checkout, through all six stages.
Write the goal in plain language and list the acceptance criteria. Each criterion is typed (unit, end-to-end, performance or manual), so the agents know how to prove it. Attach designs, docs or links as context, or let a GitHub or GitLab issue become the task.
The main agent (Opus, high effort by default) explores the repository, follows your project's conventions and proposes an implementation plan with risks and an estimate. Turn on plan approval to review it before any code is written.
<ApplePayButton> with device feature detectionPOST /api/checkout to accept payment tokensEach subtask gets an agent, its dependencies and the acceptance criteria it has to satisfy. Independent subtasks are scheduled in parallel, so two developers can work at the same time without stepping on each other.
Developer agents (Sonnet, medium effort by default) work on wheel/WHL-231-apple-pay, follow your lint rules and existing patterns, and commit in small, readable steps. Your protected branches are never touched.
12 export function ApplePayButton({ total }: Props) {
13+ const supported = useApplePaySupport();
14+ if (!supported) return null; // C1
15+
16+ return (
17+ <button className="apple-pay"
18+ onClick={() => startPayment(total)} />
19+ );
20 }
api/checkout.ts…The Tester writes and runs unit, end-to-end and load tests, and links evidence to each criterion. When something fails, the failure goes back to the right developer with full context and the loop repeats, up to the retry limit you set.
$ pnpm test --filter checkout
✓ ApplePayButton › hidden on unsupported devices 12 ms C1
✓ checkout.api › token creates a PAID order 48 ms C2
✓ e2e › declined card keeps cart, shows error 2.1 s C3
✗ load › p95 latency 342 ms > 300 ms C4
↻ sent back to Developer #2 · cache gateway session
✓ load › p95 latency 212 ms @ 500 rps C4
Acceptance criteria verified: 4 / 4
The Reviewer agent (Opus) checks the diff, then opens a PR or MR with a written summary, the test results and a checklist that maps each criterion to its evidence. You approve and merge, or comment and send it back.
wheel/WHL-231-apple-pay → main · +318 −41 · 11 filesFollow every task from goal to merged pull request: what the agents planned, what they changed and how they proved it.
See every task in flight, which agents are working on what, and what shipped this week.
Every role is a sub-agent you configure: choose the model, set the effort, decide how many run in parallel. Put Opus-level reasoning where it matters and Sonnet speed everywhere else. Try it below.
Opus where judgement matters: orchestrating, planning and reviewing.
Sonnet for high-volume coding and testing. Haiku for docs and chores.
Owns the plan, delegates subtasks and enforces the acceptance criteria.
Connect a GitHub organization or a GitLab group in a few clicks. Issues flow in as tasks, and finished work flows back as pull or merge requests, with CI checks, labels and links kept in sync.
wheel/ branchesnh-wheel[bot] wants to merge 7 commits into main from wheel/WHL-231-apple-pay
Adds Apple Pay as a payment method at checkout on supported devices. Other browsers keep the card form. Gateway sessions are cached to keep checkout latency under budget.
apple-pay-button.spec.tsxcheckout.e2e.tscheckout.e2e.tsk6: 212 msA task isn't done until every acceptance criterion is verified, and each one links to the test or check that proves it.
Require approval for the plan, the merge, or both. Or let trusted projects run fully on autopilot.
Each task runs on its own branch and workspace, so parallel tasks never collide and nothing lands without review.
Failing tests go back to the right developer agent with full context, up to a retry limit, before anyone is paged.
Watch every plan, tool call, commit and test run as it happens. Pause or redirect a task at any time.
Spend per task, per agent and per model, with budgets, alerts and hard stops that end runaway work.
Project instructions, lint rules and existing patterns guide every agent, so the code reads like your team wrote it.
Who approved what, which agent changed which file, and why. Every decision is kept for later review.
One orchestrator, a team of specialist agents and the right model for each job. Here is how the AI turns your goal into a pull request, and where you stay in control.
You stay in the loop: approve the plan before coding starts, and review the PR before it merges.
Clear the long tail of well-defined tickets without pulling engineers off roadmap work.
Delegate with precise acceptance criteria, then review outcomes instead of keystrokes.
Deliver client tasks across many repositories with one consistent, auditable process.
Short answers to the questions teams ask first.
No. Agents only write to branches prefixed with wheel/. Every change reaches your default branch through a pull request (GitHub) or merge request (GitLab) that you, your CI and your branch protection rules control.
No. You can create tasks in NH Wheel, or connect a repository and turn GitHub or GitLab issues into tasks. Either way you add the goal and the acceptance criteria, and progress syncs back to the issue.
Each sub-agent can use Claude Opus, Sonnet or Haiku, with its own effort level from low to max. A common setup is Opus for the orchestrator and reviewer and Sonnet for developers and testers, but any mix works.
Yes. Turn on plan approval per project or per task. The orchestrator stops after planning and waits for you to approve or request changes.
The tester sends the failure and its context back to the developer who owns that subtask, and the loop repeats. If it still fails after the retry limit you set, the task pauses and you're notified.
Pick the model and effort per role, set a budget per task, and get alerts as spend approaches it. The usage dashboard shows tokens and cost by task, agent and model.