Proposal for leadership · 2026-06-16

The team that
never clocks out.

Our engineers and AI build by day. The moment they log off, AI takes the night shift — reviewing, hardening, regression-testing, and teaching itself across every commit of the day. By the time anyone returns, every change from the day has already been reviewed, hardened and regression-tested — with bounded fixes or decisions prepared for the morning gate. The repository is never idle.

24/7
The codebase is worked around the clock. No 15-hour overnight dead zone.
2nd
shift
Review, hardening, regression and bounded follow-up work — added each day with no hit to daytime throughput.
Every
commit
Reviewed by two independent AI model families — not a sampled few.
The insight

Most teams go dark at 6 PM. Ours goes to work.

The fast work and the slow work don't belong in the same hours. Writing tests and building features is interactive — it wants a person in the loop. Heavy review, full regression, security analysis and lesson-extraction are slow, expensive, and need nobody watching.

So we split the day. Cheap, scoped checks run live while we build. The expensive judgment runs overnight, batched across the whole day's work — run in hours that would otherwise be idle, on a bounded, governable overnight budget.

It's not "work longer." It's a second shift that costs us no daytime and no extra people.

24/7never idle
Dayhumans + AI build
6 PMhand to the night shift
NightAI reviews · secures · tests · teaches
Dawnreport waiting
How it works · step one

Before a line of code: three documents.

We study the project once, up front, and write the contract the loops run against. This is what lets the night shift work unsupervised without going off the rails.

01 · Plan

Development plan

The work, broken into cards with dependencies marked. The day loop pulls straight from this.

02 · The keystone

key_decisions.md

Every key decision answered up front, with the reasoning. It's the single source of truth the overnight reviewers check against — so they never have to guess "was this intentional?"

03 · Acceptance

Human testing plan

The final sign-off checklist — written now, before the build can bias it. We verify what we needed, not just what we shipped.

The day shift

Build, task by task — with a person in the loop.

One continuous run of the /loop automation, self-paced. It pulls the next card, runs the full per-card pipeline, commits, and moves on until the day's queue is drained.

1

Spec

Pulls answers from key_decisions.md

2

Tests first

Failing tests from the acceptance criteria

3

Build

Code until the tests pass

4

Scoped CI gate

Types · lint · scoped tests · secrets — stays live

5

Scoped E2E

Guards against integration breaks stacking

6

Commit

Log any assumption, flag it, move on

!
What deliberately stays in the day loop: the scoped CI gate (the builder's own steering signal) and scoped E2E. We batch the full suite overnight — never the per-build checks.

/loop  drives this: take next card → spec → test → build → scoped CI → scoped E2E → commit → repeat, self-paced, until empty.

The night shift · unattended

While you sleep, AI puts every commit on trial.

Kicked off in the evening on a schedule. Inside it, /loop walks the day's commits one by one and runs the heavy judgment that needs nobody present.

1

Reviewer

Architecture & correctness on each commit

2

Security review

When the change touches a security surface

NEW
3

Cross-AI review

A second, independent model family — built to disagree

4

Route by matrix

Check key_decisions.md; prepare the safe next step, else record & escalate

5

Full regression

Once, over the whole day's diff

6

Teach

Extract the day's lessons, once

+
The cross-AI review is the new edge. Claude builds the code and reviews it — so it shares its own blind spots. A second model family from a different vendor, prompted to refute, catches the class of mistake a single reviewer never will. Same bar applies: a finding is blocking only with a concrete reproduction — otherwise it's advisory.
i
Review runs per commit; the full suite and teacher run once per batch. Per-commit review attributes findings to a card. Running the full suite N times overnight would burn the budget for no extra signal — so we don't.

By dawn, a single report is waiting: what the night confirmed, the bounded patches or decisions it prepared for sign-off, and the high-stakes calls it escalated untouched for you to make.

The upgrade · pilot extension

Beyond review: bounded autonomy, if we choose it.

Findings are table stakes. The next layer of leverage is letting the night shift prepare the safe decisions and escalate the dangerous ones by a published authority matrix. For the pilot, this means ready-to-review patches and recommendations by morning, not unsupervised merge authority. Two axes set every call: how reversible a change is (its blast radius), and how confident the automated review is. Blast radius sets a hard ceiling; confidence decides how far up to it the agent may go.

Tier 1 · prepare · default-accept

Prepare — accepted by default

Low blast radius and high confidence — reversible, local, test-covered, consistent with an answered decision. The agent prepares the patch, records the reasoning in the decision log, and rolls it into the morning batch accepted by default: you skim the batch and pull the rare outlier, you don't sign off each one. Obvious fixes, local refactors, naming, added coverage.

Tier 2 · draft · hold for approval

Draft — held for your yes

Moderate blast radius or moderate confidence. The agent drafts the most reversible option and records it, but it's held — nothing lands until you approve it, item by item. This is the short list that earns real attention; the safe bulk is already handled in Tier 1. Nothing that can't be backed out before merge.

Tier 3 · record only

Escalate — a human decides

High blast radius or low confidence — architecture, data model, auth, tenant isolation, migrations, money, anything irreversible. The agent does not change code. It records the question, its recommended option and the trade-offs, and waits for the morning gate.

!
Blast radius is a hard ceiling, not a vote. No amount of confidence promotes an architectural or security decision out of Tier 3. Confidence only moves work up to the ceiling its reversibility allows — which is exactly where the value is, and exactly where the risk isn't.
i
The Tier 3 list is the bar we already hold. Auth, tenant isolation, migrations, money movement, infra and security have always required a human at Airiam. The matrix doesn't loosen that — it just lets everything below it be pre-processed overnight without waiting on one.
What keeps it honest

key_decisions.md is alive, not frozen.

The night shift will hit decisions nobody wrote down — that's the normal case, not the exception. When it does, it files the question as a ticket instead of guessing.

Each morning we triage those tickets, answer them, and append the answer back into the document with the date and the reasoning. Skip that one step and we'd re-litigate the same question three cards later.

The morning gate is built to stay light, not to become a queue: the Tier 1 batch is accepted with a skim, the short Tier 2 list gets a real look, and only the rare Tier 3 escalations need a genuine decision. Effort scales with the risk, not with how many commits the night touched — so the day opens from a verified baseline, with no sign-off backlog.

It's the same discipline that turns a process into an auditable, self-documenting system — every decision, every review, every security check leaves a trail.

Why this puts us ahead

A moat the market can't cheaply copy.

This isn't a tooling tweak. It changes the economics of how a team ships software — and the gap compounds every single day.

Velocity

Your clock runs 24 hours

Competitors' code sits dark from 6 PM to 9 AM — ~15 idle hours a day. Ours is being reviewed, hardened and tested in exactly those hours. We get a second shift for a bounded, governable overnight compute-and-model spend.

Quality

Quality with no velocity tax

Most teams trade speed for review depth. We refuse the trade: full review on every commit, absorbed by a bounded overnight budget. Daytime throughput never slows.

Defect escape rate

Two minds on every change

Two independent AI model families review each commit. Single-reviewer blind spots — human or AI — get caught before they ever reach a customer. A defect a rival ships, we intercept overnight.

Compounding

A team that gets smarter nightly

Lessons are extracted every night and decisions are written down, not held in people's heads. The team improves while it sleeps — and the knowledge can't walk out the door.

 A typical teamOur model
Hours 6 PM–9 AMDark. Nothing happens.Reviewing, securing, regression-testing
Overnight outputNothingVerified findings, prepared patches, and safe recommendations — within published bounds
Review coverageSampled — what time allowsEvery single commit
Independent second opinionRareA second AI model family, every change
Institutional memoryIn people's headsExtracted & written down nightly
Decision rationaleTribal knowledgekey_decisions.md — audit-ready
Speed vs. qualityA trade-offBoth — verification runs on a bounded overnight budget
The engine

Powered by frontier AI — used the way it should be.

The 24-hour cycle is only possible because AI does the parts that don't need a person. We don't bolt one chatbot onto the side — we run a coordinated fleet.

Orchestration

A multi-agent pipeline

Specialist agents — spec, test, build, review, security, teacher — each doing one job well, handing off down the line.

Efficiency

The right model for each job

Frontier models on the hard reasoning (spec, review); fast, cheap models on the mechanical work (tests, E2E). We pay for intelligence only where it pays back.

Autonomy

/loop runs it unattended

The same automation drives the day loop self-paced and the night loop on a schedule. It works the second shift so nobody has to.

Resilience

Adversarial cross-model review

Two vendors' models checking each other. Diversity of reasoning is our hedge against any single model's failure mode.

"What's the catch?"

Rigorous by design — not reckless.

For a financial-operations and security company, an unsupervised night shift only earns trust if it's tightly bounded. Four guardrails make it safe.

1

The decision log stays alive

Every resolved question is appended back, dated, with rationale. The source of truth never goes stale.

2

Review per commit, full suite per batch

Findings stay attributable to a card; the expensive suite runs once. Overnight budget spent where it pays.

3

Assumptions are visible, dependencies don't stack

When a builder must assume, it logs and flags it. The day loop builds one card at a time; overnight, independent commits are reviewed in parallel while dependent chains are gated in order — so a flawed foundation never hides under the work built on top of it.

4

A human gate every morning & at the end

The overnight pass prepares work only inside its authority tier — high-stakes calls are recorded for you, never executed, and it never merges. The standard reviewed-PR gate into the protected branches is unchanged: agents don't approve or merge. We triage the night's escalations before the next build day, and run the full human acceptance plan before anything ships. A security block always waits for a person.

The ask

Let's run one project on this model and measure the gap.

Give us one build to prove it: same scope, this operating model. We'll show the throughput, the review coverage, and the defects caught overnight that a normal team would have shipped.

Approve a pilot →