Our engineers and AI build by day. The moment they log off, AI takes the night shift — reviewing, hardening, regression-testing, and teaching itself across every commit of the day. By the time anyone returns, every change from the day has already been reviewed, hardened and regression-tested — with bounded fixes or decisions prepared for the morning gate. The repository is never idle.
The fast work and the slow work don't belong in the same hours. Writing tests and building features is interactive — it wants a person in the loop. Heavy review, full regression, security analysis and lesson-extraction are slow, expensive, and need nobody watching.
So we split the day. Cheap, scoped checks run live while we build. The expensive judgment runs overnight, batched across the whole day's work — run in hours that would otherwise be idle, on a bounded, governable overnight budget.
It's not "work longer." It's a second shift that costs us no daytime and no extra people.
We study the project once, up front, and write the contract the loops run against. This is what lets the night shift work unsupervised without going off the rails.
The work, broken into cards with dependencies marked. The day loop pulls straight from this.
Every key decision answered up front, with the reasoning. It's the single source of truth the overnight reviewers check against — so they never have to guess "was this intentional?"
The final sign-off checklist — written now, before the build can bias it. We verify what we needed, not just what we shipped.
One continuous run of the /loop automation, self-paced. It pulls the next card, runs the full per-card pipeline, commits, and moves on until the day's queue is drained.
Pulls answers from key_decisions.md
Failing tests from the acceptance criteria
Code until the tests pass
Types · lint · scoped tests · secrets — stays live
Guards against integration breaks stacking
Log any assumption, flag it, move on
/loop drives this: take next card → spec → test → build → scoped CI → scoped E2E → commit → repeat, self-paced, until empty.
Kicked off in the evening on a schedule. Inside it, /loop walks the day's commits one by one and runs the heavy judgment that needs nobody present.
Architecture & correctness on each commit
When the change touches a security surface
A second, independent model family — built to disagree
Check key_decisions.md; prepare the safe next step, else record & escalate
Once, over the whole day's diff
Extract the day's lessons, once
By dawn, a single report is waiting: what the night confirmed, the bounded patches or decisions it prepared for sign-off, and the high-stakes calls it escalated untouched for you to make.
Findings are table stakes. The next layer of leverage is letting the night shift prepare the safe decisions and escalate the dangerous ones by a published authority matrix. For the pilot, this means ready-to-review patches and recommendations by morning, not unsupervised merge authority. Two axes set every call: how reversible a change is (its blast radius), and how confident the automated review is. Blast radius sets a hard ceiling; confidence decides how far up to it the agent may go.
Low blast radius and high confidence — reversible, local, test-covered, consistent with an answered decision. The agent prepares the patch, records the reasoning in the decision log, and rolls it into the morning batch accepted by default: you skim the batch and pull the rare outlier, you don't sign off each one. Obvious fixes, local refactors, naming, added coverage.
Moderate blast radius or moderate confidence. The agent drafts the most reversible option and records it, but it's held — nothing lands until you approve it, item by item. This is the short list that earns real attention; the safe bulk is already handled in Tier 1. Nothing that can't be backed out before merge.
High blast radius or low confidence — architecture, data model, auth, tenant isolation, migrations, money, anything irreversible. The agent does not change code. It records the question, its recommended option and the trade-offs, and waits for the morning gate.
The night shift will hit decisions nobody wrote down — that's the normal case, not the exception. When it does, it files the question as a ticket instead of guessing.
Each morning we triage those tickets, answer them, and append the answer back into the document with the date and the reasoning. Skip that one step and we'd re-litigate the same question three cards later.
The morning gate is built to stay light, not to become a queue: the Tier 1 batch is accepted with a skim, the short Tier 2 list gets a real look, and only the rare Tier 3 escalations need a genuine decision. Effort scales with the risk, not with how many commits the night touched — so the day opens from a verified baseline, with no sign-off backlog.
It's the same discipline that turns a process into an auditable, self-documenting system — every decision, every review, every security check leaves a trail.
This isn't a tooling tweak. It changes the economics of how a team ships software — and the gap compounds every single day.
Competitors' code sits dark from 6 PM to 9 AM — ~15 idle hours a day. Ours is being reviewed, hardened and tested in exactly those hours. We get a second shift for a bounded, governable overnight compute-and-model spend.
Most teams trade speed for review depth. We refuse the trade: full review on every commit, absorbed by a bounded overnight budget. Daytime throughput never slows.
Two independent AI model families review each commit. Single-reviewer blind spots — human or AI — get caught before they ever reach a customer. A defect a rival ships, we intercept overnight.
Lessons are extracted every night and decisions are written down, not held in people's heads. The team improves while it sleeps — and the knowledge can't walk out the door.
| A typical team | Our model | |
|---|---|---|
| Hours 6 PM–9 AM | Dark. Nothing happens. | Reviewing, securing, regression-testing |
| Overnight output | Nothing | Verified findings, prepared patches, and safe recommendations — within published bounds |
| Review coverage | Sampled — what time allows | Every single commit |
| Independent second opinion | Rare | A second AI model family, every change |
| Institutional memory | In people's heads | Extracted & written down nightly |
| Decision rationale | Tribal knowledge | key_decisions.md — audit-ready |
| Speed vs. quality | A trade-off | Both — verification runs on a bounded overnight budget |
The 24-hour cycle is only possible because AI does the parts that don't need a person. We don't bolt one chatbot onto the side — we run a coordinated fleet.
Specialist agents — spec, test, build, review, security, teacher — each doing one job well, handing off down the line.
Frontier models on the hard reasoning (spec, review); fast, cheap models on the mechanical work (tests, E2E). We pay for intelligence only where it pays back.
The same automation drives the day loop self-paced and the night loop on a schedule. It works the second shift so nobody has to.
Two vendors' models checking each other. Diversity of reasoning is our hedge against any single model's failure mode.
For a financial-operations and security company, an unsupervised night shift only earns trust if it's tightly bounded. Four guardrails make it safe.
Every resolved question is appended back, dated, with rationale. The source of truth never goes stale.
Findings stay attributable to a card; the expensive suite runs once. Overnight budget spent where it pays.
When a builder must assume, it logs and flags it. The day loop builds one card at a time; overnight, independent commits are reviewed in parallel while dependent chains are gated in order — so a flawed foundation never hides under the work built on top of it.
The overnight pass prepares work only inside its authority tier — high-stakes calls are recorded for you, never executed, and it never merges. The standard reviewed-PR gate into the protected branches is unchanged: agents don't approve or merge. We triage the night's escalations before the next build day, and run the full human acceptance plan before anything ships. A security block always waits for a person.
Give us one build to prove it: same scope, this operating model. We'll show the throughput, the review coverage, and the defects caught overnight that a normal team would have shipped.
Approve a pilot →