A step-by-step guide to stand up the round-the-clock operating model on a developer machine. You create a handful of files, drop in one API key when you have it, and optionally register a nightly trigger. It pulls work from your team's Linear project and runs on your machine — no CI changes, nothing to install beyond what you already run.
Confirm these are on the machine, then make three config calls up front. Everything after assumes them.
claude --version.fetch + node --test are built in). Verify with node --version.spec-writer, ci-gate, e2e-gate, code-review, security-review, and TDD — confirm each exists in the target repo (oscar-backend) before you start; the loops assume them.pending and the night loop still runs. Sends diffs to an external vendor — clear data-handling first (Step 5).| Decision | Default for v1 |
|---|---|
| Linear project | The project mapped to this repo. The day loop pulls the next open issue with no open blockers. Set the team/project id in day-loop.md if matching by name is ambiguous. |
| Security-surface trigger | A path-glob + keyword list checked into the harness (auth, RLS, secrets, SQL, infra, external calls). Deterministic, not a judgment call. |
| Night-loop authority | Bounded by the decision authority matrix (see Beyond review). It prepares/drafts patches on a night/* branch but never merges; high-stakes changes are record-only. Start with everything at Tier 3 and open tiers up as trust grows. |
Two roots: Claude Code artifacts under .claude/, and everything else under a new harness/. Create the tree, then fill it in the steps that follow.
.claude/skills/harness-bootstrap/ SKILL.md # FIRST run on a new project — builds the 2 docs (STEP 2) .claude/commands/ day-loop.md # one card's pipeline; run under /loop (STEP 3) night-loop.md # headless entry → runs the night Workflow (STEP 4) morning-triage.md # the human morning gate (STEP 6) .claude/agents/ harness-reviewer.md # architecture & correctness — per commit harness-security.md # security review — conditional harness-teacher.md # lesson extraction — per batch harness/ README.md # quick reference + smoke test templates/ # key_decisions · human-testing-plan (plan lives in Linear) night-loop.workflow.js # the night pipeline (STEP 4) authority-matrix.json # tiers, ceilings & confidence floors (STEP 4B) cross-ai/ refute.mjs refute.test.mjs .env # (STEP 5, .env gitignored) scripts/ run-night-loop.ps1 run-night-loop.sh # launchers · Windows + macOS (STEP 5) install-scheduled-task.ps1 com.airiam.nightshift.plist # schedulers (STEP 6 / macOS) reports/ # night reports + logs (gitignored) decisions-inbox/ # tickets the night files (gitignored)
.gitignore# night-shift harness — local runtime artifacts harness/reports/ harness/decisions-inbox/ harness/cross-ai/.env harness/.last-night-sha # per-machine cursor — run the night shift from one designated host
The very first action on any new project is a dedicated skill — /harness-bootstrap. Rather than dropping empty templates, it studies the project and interviews you, then writes populated key_decisions.md and human-testing-plan.md. The plan itself lives in Linear.
CLAUDE.md, docs, existing code) and the mapped Linear project to learn the domain and scope.decisions-inbox/, reports/) and won't overwrite existing docs unless you confirm.dev, and Done only after release/acceptance — so "Done" never means "committed but unverified."# Key decisions — {{PROJECT}} (single source of truth · living doc) ## D1 — <question> - decision: <the answer> - rationale: <why> - date: {{DATE}} # The morning gate appends new D-entries here — never edits history.
human-testing-plan.template.md follows the same shape: a dated checklist of acceptance steps, written before the build so it can't be biased by what shipped.A single slash command that processes one card. The built-in /loop drives it self-paced until the queue drains.
In Review; take the highest-priority one and move it to In Progress. None eligible → print "queue drained" and stop the loop.key_decisions.md (invoke spec-writer). Missing a decision → post it as a comment on the issue and continue; do not block.ci-gate). This stays live — it's the builder's steering signal.e2e-gate).> /loop /day-loop # self-paced: pick Linear card → spec → test → build → scoped CI → scoped E2E → commit → repeat # stops on its own when the Linear project has no open, unblocked issues
A Workflow script does the deterministic fan-out; a thin slash command is the headless entry point that computes the batch and runs it.
harness/.last-night-sha to HEAD (all of them if the marker is absent).harness/authority-matrix.json here and pass it in; the Workflow script has no filesystem access, so its pure route() needs the matrix as args. (The review agents do have file access — they read key_decisions.md themselves.)scriptPath: harness/night-loop.workflow.js, passing { commits, matrix } as args, and wait for it to finish — the tool returns a task id and runs in the background.harness/reports/YYYY-MM-DD-night.md from its return value, then advance .last-night-sha to HEAD.Architecture & correctness per commit (harness-reviewer).
Only if the diff hits a security surface (harness-security).
Shell out to refute.mjs; pending if no key.
Classify by blast radius + confidence → prepare · draft · escalate (see Beyond review).
Whole suite over the combined batch diff.
Extract the day's lessons (harness-teacher).
One dated report, grouped by tier: prepared · drafted · escalated.
export const meta = { name: 'night-shift', description: "Per-commit review + batch regression/teach over the day's commits", phases: [{ title: 'Review' }, { title: 'Verify' }, { title: 'Batch' }], } // Inputs come from /night-loop. The script has NO filesystem access, so the // matrix is read by the command and passed in — never read from disk here. const commits = args?.commits ?? [] // shas, oldest → newest const matrix = args?.matrix ?? { tier3_surfaces: [], tier1_paths: [], confidence_floor: { prepare: 0.85, draft: 0.5 } } // Structured output the reviewers must return (validated by the tool). const FINDINGS = { type: 'object', required: ['confidence', 'files', 'findings'], properties: { confidence: { type: 'number' }, // 0..1 files: { type: 'array', items: { type: 'string' } }, // paths the commit touched findings: { type: 'array', items: { type: 'object' } }, // { blocking, reproduction, note } } } const LESSONS = { type: 'object', required: ['lessons'], properties: { lessons: { type: 'array', items: { type: 'string' } } } } // Pure routing — no I/O, just the matrix passed in. Blast radius is a hard ceiling. const rx = p => new RegExp('^' + p.replace(/[.]/g, '\\.').replace(/\*\*/g, '§').replace(/\*/g, '[^/]*').replace(/§/g, '.*') + '$') // illustrative glob const hits = (files, pats) => files.some(f => pats.some(p => rx(p).test(f))) function route({ sha, review, sec, cross }) { const files = review?.files ?? [] const conf = Math.min(review?.confidence ?? 0, sec?.confidence ?? 1) if (hits(files, matrix.tier3_surfaces)) return { sha, tier: 3, review, sec, cross } // ceiling → escalate if (conf >= matrix.confidence_floor.prepare && hits(files, matrix.tier1_paths)) return { sha, tier: 1, review, sec, cross } // prepare return { sha, tier: 2, review, sec, cross } // draft, held } // Per commit — pipelined; each commit verifies as soon as its review lands. const perCommit = await pipeline(commits, sha => agent(`Review commit ${sha} for architecture & correctness.`, { agentType: 'harness-reviewer', phase: 'Review', schema: FINDINGS }), (review, sha) => parallel([ () => agent(`Security-review ${sha} IF it touches a security surface.`, { agentType: 'harness-security', phase: 'Verify', schema: FINDINGS }), () => agent(`Run: git show ${sha} | node harness/cross-ai/refute.mjs — return its JSON.`, { phase: 'Verify' }), ]).then(([sec, cross]) => route({ sha, review, sec, cross })), // classify → prepare / draft / escalate ) // Once per batch — skip the heavy suite if the overnight budget is already spent. phase('Batch') const regression = budget.remaining() > 50_000 ? await agent('Run the full suite over the batch diff; summarize failures.', { phase: 'Batch' }) : { skipped: 'budget' } const lessons = await agent("Extract the day's lessons, once.", { agentType: 'harness-teacher', phase: 'Batch', schema: LESSONS }) return { perCommit, regression, lessons }
Findings are table stakes. The upgrade is letting the night shift prepare the safe fixes and escalate the dangerous ones by a published authority matrix — ready-to-review patches by morning, never unsupervised merge. Two axes set every call: a change's blast radius (how reversible it is) sets a hard ceiling; the review's confidence decides how far up to that ceiling the agent may act.
Low blast radius and high confidence — reversible, local, test-covered, consistent with an answered decision. The agent commits the patch to a night branch, logs the reasoning, and rolls it into the morning batch accepted by default. Obvious fixes, local refactors, naming, added coverage.
Moderate blast radius or moderate confidence. The agent drafts the most reversible option on a branch and records it, but it's held — nothing lands until you approve it item by item. Nothing that can't be backed out before merge.
High blast radius or low confidence — architecture, data model, auth, tenant isolation, migrations, money, anything irreversible. The agent does not touch code. It files a ticket with its recommended option and the trade-offs, and waits for the morning gate.
{
"tier3_surfaces": [ // hard ceiling → always escalate, never edit
"**/migrations/**", "**/auth/**", "**/rls/**", "infrastructure/**",
"**/*payment*", "**/*billing*", "sql/**", "**/tenant*"
],
"tier1_paths": [ // eligible for 'prepare' IF confidence is high
"**/*.test.*", "docs/**", "**/README*"
],
"confidence_floor": { "prepare": 0.85, "draft": 0.5 }
}
route() step/night-loop loads authority-matrix.json and passes it to the Workflow as args.matrix; route() is a pure function of that matrix (the script itself can't touch the filesystem).tier3_surfaces glob, it's Tier 3, full stop, regardless of confidence.tier1_paths match → Tier 1; otherwise → Tier 2.key_decisions.md — a change consistent with an answered decision raises confidence; a silent decision drops it and files a ticket.night/<date> branch only. The protected-branch PR gate is unchanged; agents never approve or merge.tier1_paths empty so nothing is auto-prepared beyond drafts — everything is Tier 2 or 3. Open the tiers up as trust grows; the matrix is the dial. The Tier 3 list is the bar Airiam already holds — auth, tenant isolation, migrations, money, infra always wait for a human.One small Node file, vendor-agnostic. It reads three env vars, POSTs the diff to any OpenAI-compatible endpoint, and prints a JSON verdict. No key → it prints pending and exits cleanly.
// Reads a unified diff on stdin, prints a JSON verdict on stdout. const { CROSS_AI_BASE_URL: BASE, CROSS_AI_MODEL: MODEL, CROSS_AI_API_KEY: KEY } = process.env if (!BASE || !MODEL || !KEY) { process.stdout.write(JSON.stringify({ status: 'pending', reason: 'CROSS_AI_* not configured' })) process.exit(0) // never fails the night loop } const diff = await readStdin() const system = 'You are an adversarial reviewer from a different model family. ' + 'REFUTE this change. A finding is blocking ONLY with a concrete reproduction. ' + 'Default refuted=false unless you have a specific, real failure.' let res try { res = await fetch(`${BASE.replace(/\/$/,'')}/chat/completions`, { method: 'POST', headers: { 'content-type': 'application/json', authorization: `Bearer ${KEY}` }, signal: AbortSignal.timeout(60_000), // never let a hung endpoint stall the night body: JSON.stringify({ model: MODEL, temperature: 0, messages: [ { role: 'system', content: system }, { role: 'user', content: 'Respond ONLY as JSON {"refuted":bool,"reproduction":string|null,"notes":string}.\n\n' + diff }, ]}), }) } catch (e) { // timeout / network error → advisory, not fatal process.stdout.write(JSON.stringify({ status: 'error', reason: String(e?.name ?? e) })) process.exit(0) } if (!res.ok) { // non-2xx → advisory, not fatal process.stdout.write(JSON.stringify({ status: 'error', code: res.status })) process.exit(0) } const data = await res.json() const text = data.choices?.[0]?.message?.content ?? '{}' process.stdout.write(JSON.stringify({ status: 'ok', verdict: safeJson(text) })) function safeJson(t){ try{ return JSON.parse(t) }catch{ return { refuted:false, reproduction:null, notes:t } } } function readStdin(){ return new Promise(r=>{let d='';process.stdin.on('data',c=>d+=c);process.stdin.on('end',()=>r(d))}) }
CROSS_AI_BASE_URL=https://api.openai.com/v1 # or Azure OpenAI / Gemini-compat / local server CROSS_AI_MODEL=<a-non-anthropic-model-id> CROSS_AI_API_KEY=<paste-key-here>
| Env var | Purpose | |
|---|---|---|
| CROSS_AI_BASE_URL | req | Any OpenAI-compatible base URL. |
| CROSS_AI_MODEL | req | The model id — the "different family" built to disagree. |
| CROSS_AI_API_KEY | req | Bearer key. Absent → step reports pending. |
git diffs to a second, non-Anthropic vendor; a diff can carry secrets, credentials, or customer data in fixtures. For a fin-ops/security shop that needs a real answer, not a default: a signed DPA, a no-training-on-inputs guarantee, known data residency, and ideally a secret-scan on the diff before it leaves the machine. Prefer a vendor Airiam already has terms with.$ node --test harness/cross-ai/refute.test.mjs # asserts: key-absent → {"status":"pending"}; key-present → parses a mock verdict
# Headless launcher for the unattended night shift. $ErrorActionPreference = 'Stop' $repo = Split-Path -Parent (Split-Path -Parent $PSScriptRoot) Set-Location $repo # Load the cross-AI key from the local, gitignored .env if present $envFile = "$repo\harness\cross-ai\.env" if (Test-Path $envFile) { Get-Content $envFile | Where-Object { $_ -match '^\s*[^#].*=' } | ForEach-Object { $k,$v = $_ -split '=',2; [Environment]::SetEnvironmentVariable($k.Trim(), $v.Trim()) } } $stamp = Get-Date -Format 'yyyy-MM-dd_HHmmss' $log = "$repo\harness\reports\night-$stamp.log" claude -p '/night-loop' --dangerously-skip-permissions *>> $log # hands-off (see Step 5); scope with --allowedTools where you can
acceptEdits auto-approves edits but still prompts for shell and git, so a scheduled run hangs on the first test/commit (see Troubleshooting) — that's why the launcher uses --dangerously-skip-permissions. Be clear-eyed about what that grants: Tier 1/2 write patches and run git/tests. What bounds the risk is where it can act — every write lands on a throwaway night/* branch, it never touches protected branches, and it never merges (the PR gate is unchanged). Prefer the narrowest scope that still finishes unattended — a --allowedTools/settings allowlist over blanket skip — and decide before you schedule.Kick the night loop off by hand first. Once you trust it, register a nightly trigger. Then the morning gate keeps the source of truth alive.
PS> powershell -NoProfile -ExecutionPolicy Bypass -File harness\scripts\run-night-loop.ps1 # watch harness\reports\ for the .log and the dated night report
# OPTIONAL — run once to register a nightly Task Scheduler job. Never auto-run. $script = "$PSScriptRoot\run-night-loop.ps1" $action = New-ScheduledTaskAction -Execute 'powershell.exe' ` -Argument "-NoProfile -ExecutionPolicy Bypass -File `"$script`"" $trigger = New-ScheduledTaskTrigger -Daily -At 6:00PM Register-ScheduledTask -TaskName 'OSCAR Night Shift' -Action $action -Trigger $trigger ` -Description 'Unattended nightly review loop'
Unregister-ScheduledTask -TaskName 'OSCAR Night Shift'. The task only calls the launcher — all behavior stays in the repo.launchd and cron run with a minimal environment: claude and node may not be on PATH, and Claude's login must be reachable for the scheduled user (on macOS, keychain access differs under launchd). Use absolute paths or source the profile in the launcher, make sure the machine is unlocked/awake at the trigger time, and run the scheduled job once by hand to prove the environment before trusting the timer.night/* branch; pull the rare outlier. Effort scales with risk, not commit count.key_decisions.md (a dated D- entry with rationale, never editing history) so it's never re-litigated.human-testing-plan.md before anything ships.Only the launcher and the scheduler are platform-specific — the Node adapter, the Claude commands, and the Linear queue are identical. On macOS, swap the PowerShell pair (Steps 5–6) for a bash launcher and a launchd agent.
#!/usr/bin/env bash # Headless launcher for the unattended night shift (macOS / Linux). set -euo pipefail repo="$(cd "$(dirname "$0")/../.." && pwd)" cd "$repo" # Load the cross-AI key from the local, gitignored .env if present env_file="$repo/harness/cross-ai/.env" [ -f "$env_file" ] && set -a && . "$env_file" && set +a stamp="$(date +%Y-%m-%d_%H%M%S)" log="$repo/harness/reports/night-$stamp.log" claude -p "/night-loop" --dangerously-skip-permissions >>"$log" 2>&1 # hands-off (see Step 5); scope with --allowedTools where you can
$ chmod +x harness/scripts/run-night-loop.sh $ ./harness/scripts/run-night-loop.sh # manual run; watch harness/reports/
<?xml version="1.0" encoding="UTF-8"?> <!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd"> <plist version="1.0"> <dict> <key>Label</key> <string>com.airiam.nightshift</string> <key>ProgramArguments</key> <array> <string>/bin/bash</string> <string>/Users/you/oscar-backend/harness/scripts/run-night-loop.sh</string> </array> <key>StartCalendarInterval</key> <dict><key>Hour</key><integer>18</integer><key>Minute</key><integer>0</integer></dict> <key>StandardOutPath</key> <string>/tmp/nightshift.out</string> <key>StandardErrorPath</key> <string>/tmp/nightshift.err</string> </dict> </plist>
$ launchctl load ~/Library/LaunchAgents/com.airiam.nightshift.plist $ launchctl unload ~/Library/LaunchAgents/com.airiam.nightshift.plist # to remove
launchd over cron on macOS. If the Mac is asleep at 6 PM, launchd runs the missed job on the next wake; cron silently skips it. The plist calls the launcher only — all behavior stays in the repo.# m h dom mon dow run the night shift at 18:00
0 18 * * * /bin/bash ~/oscar-backend/harness/scripts/run-night-loop.sh
--dangerously-skip-permissions (or a scoped allowlist). The loop writes to a night/* branch, so it isn't read-only — the risk is bounded by that branch scope and the unchanged PR gate, not by the loop being unable to act.A one-card dry run proves the whole chain before you trust it with real work.
1. /harness-bootstrap # studies the project, writes the 2 populated docs 2. create one trivial Linear issue # e.g. "add a /health route", no blockers 3. /loop /day-loop # builds it, commits, moves the issue to In Review 4. node --test harness/cross-ai/refute.test.mjs # adapter green 5. run the launcher # Windows: powershell -File harness\scripts\run-night-loop.ps1 · macOS: ./harness/scripts/run-night-loop.sh 6. open harness/reports/<today>-night.md # report exists, marker advanced 7. /morning-triage # triage any filed tickets
| Symptom | Cause & fix |
|---|---|
| Cross-AI always "pending" | One of the three CROSS_AI_* vars is unset, or .env isn't being loaded. Confirm .env exists and the launcher's parse ran; test with the vars exported in the shell. |
| Night loop stops on a prompt | Permission mode too strict for headless — the loop needs edits, git and shell (on a night/* branch). acceptEdits still prompts for shell/git; use --dangerously-skip-permissions or a scoped --allowedTools allowlist (Step 5). |
| Day loop never stops | A Linear issue's blocker never reaches In Review, or issues aren't being moved to In Review on commit. Check the project's "blocked by" relations and workflow states. |
| Day loop picks nothing | Wrong project scope, or no issue is both open and unblocked. Confirm the Linear project mapping in day-loop.md and that at least one issue has all blockers at In Review or later. |
| Report re-reviews old commits | harness/.last-night-sha didn't advance. Confirm /night-loop writes it to HEAD at the end. |
| Workflow won't run headless | The Workflow tool needs the command's instructions to invoke it. Make sure night-loop.md explicitly tells Claude to run it with the script path. |
| Report/marker never written | The Workflow runs in the background (returns a task id immediately). If the claude -p session exits before it finishes, the "Finish" step never runs. night-loop.md must wait for completion, then write the report and advance the marker. |
route() / matrix ignored | The Workflow script can't read files. Confirm /night-loop reads authority-matrix.json and passes it in args.matrix; a missing arg falls back to the empty-tier default (everything → Tier 2/3). |
| Re-running re-reviews the batch | Safe by design: the marker only advances on success, so a failed/half run re-reviews the same commits next time. Delete the stray night/* branch first, or resume the Workflow from its run id. |
| Command | When | What it does |
|---|---|---|
| /harness-bootstrap | first, new project | Studies the project and writes key_decisions.md + human-testing-plan.md. |
| /loop /day-loop | the work day | Builds Linear cards one at a time, person in the loop. |
| run-night-loop.ps1 | evening | Launches the unattended night review over the day's commits. |
| /morning-triage | each morning | Answers the night's tickets; appends to key_decisions.md. |
Backlog → In Progress → In Review → Merged In Dev → Done lifecycle mapping; the exact tier3_surfaces globs and confidence_floor values; how far to open Tier 1 for the pilot; and the permission scope for the unattended run. All have working defaults above, but the data-handling and permission calls should get an explicit yes before the first scheduled night.Everything here is buildable locally today with the dependencies noted above. Once the team signs off on the open items — data-handling for cross-AI, the Linear dependency, and the permission scope — this becomes the checklist we execute.