Requirements in. Reviewed pull requests out. By morning.
Describe what the business needs. Nightshift writes the cards: acceptance criteria, scope, and the paths an agent may touch, at a depth nobody has the afternoon to write thirty times. Load them before you leave. Coding agents build them, separate agents review them, and a versioned ruleset decides what merges without you.
One board free, with no time limit. $10 per board after that. Your model spend stays yours.
That is the whole failure mode, and it multiplies rather than adds. An agent handed a card that never says what done looks like returns something that compiles, satisfies nothing anybody asked for, and costs more to read than it would have cost to write. Thirty of those is not thirty times the output. It is a morning gone.
So the depth goes in at the top, where it is cheapest. Hand Nightshift a requirement and it decomposes it into cards carrying acceptance criteria, a category, an estimate and a path lock: what a good engineer would write given the afternoon, on all thirty rather than the first one.
A paragraph of intent, a ticket, a customer complaint. Whatever you actually have.
Into cards small enough for one agent to finish and specific enough to be judged when it says it has.
Depth at grooming time is what makes a night run worth starting. It is also what the gate checks the diff against afterwards.
Cards, a house style every agent is handed, and a roster you shape. Every figure below is an illustration rather than anybody's night.
Order total, with GST
Retry-After on a 429
Session cookie on callback
Rate limit the invites
Backlog empty state
Grades come from a human-editable, diffable ruleset that lives in your repository. Highest matching rule wins.
Cosmetic, docs, dependency bumps, styling.
Code, review and merge, unattended.
Feature work, components, scoped refactors.
Code and review; a human merges.
Core services, data models, auth, billing.
AI and human review both required.
Infrastructure, migrations, security, CI/CD.
Never autonomous. Named approver only.
Thirty cards running at once is only worth having if you can leave. The interesting part of an autonomous loop is therefore not what it does: it is what it declines to do, and on whose authority.
Title, description, acceptance criteria, and the paths an agent is allowed to touch.
Computed from the paths the card declares, by the same policy that will judge the diff afterwards. Both numbers coming from one ruleset is what makes the comparison mean anything.
Clones with a short-lived token scoped to that one repository, works in a throwaway checkout, opens a draft pull request. Nothing an agent writes is presented as ready.
A different machine, on a different lease. No agent ever reviews its own work.
The actual diff is re-graded and checked against the card’s path lock. If it grades riskier than it was dispatched as, the autonomy decision was made on false information and auto-merge is blocked.
Anthropic, OpenAI, Google, Bedrock, Azure or OpenRouter. The coding agents and the review agents can run on different models, because reading a diff and writing one are not the same job and do not cost the same.
Encrypted on the way in and never readable back. The only path that decrypts one is building a work order for a runner that has asked for work: nothing shows it to a browser, an operator, or a log.
Model cost lands on your account at your rate, not marked up through ours. What you pay us is for the board and the governance around it.
Underneath the headline: what merged, what is waiting on you, what the gate stopped, what it cost, and whether you were paying to write the code or to have it read.
23 merged overnight.
5 need you.
23 merged with nobody watching.
4 passed review and are waiting for you to merge.
1 was stopped by the gate and needs a look.
2 are still running and will finish before nine.
30 cards loaded. Spend: $47.80.
A styling card that quietly edited something outside its scope. Here is what happened to it.
The review agent approved it: it only ever saw a tidy diff. The gate refused it, because the change landed outside the paths the card declared. Review and authorisation are different jobs, and only one of them is allowed to be fooled.
Three things to set up, once. After that the loop runs on whatever you groom during the day.
A GitHub App with three permissions, named on the screen. Short-lived tokens, scoped to the one installation.
Acceptance criteria, a scope, an estimate. The grade decides how much autonomy the card is allowed, before anything runs.
The night run works the ready cards within your caps. Every branch arrives with the gate's verdict on it.
$0
One board, one agent, one person.
$10 per board / month, or $80 a year
Up to 5 people. Unlimited agents.
Model spend is yours and is not included: bring your own Anthropic or OpenAI key. See pricing