Wave 10 · Design Desk RAG

It designs the plan with you — it no longer refuses.

The Design Desk assistant now produces a phased project plan and level-of-effort estimate even when there is no comparable internal document — grounding Airiam-specific facts in your documents, filling the rest from engineering best practice, and clearly labeling everything it assumes so an engineer can adjust it.

Full plan + labeled estimates Airiam facts stay grounded & cited Directions editable live Now defaulting to GPT-5.6 Terra
What changed

From “not found” to a decision-ready deliverable.

The same request that used to stall now returns a phased plan you can review — because the fix was the assistant's governance prompt, not the model.

Before
  • “No comparable internal estimate found” → no answer.
  • Re-asked scope you had already decided.
  • Turning on web search changed almost nothing.
  • Treated templates and processes as hard rules.
Now
  • Builds a phased plan with labeled labor estimates.
  • Adopts the scope you state; assumes the rest and says so.
  • Reaches the web as first-class context when internal is thin.
  • Uses templates as guides; never fabricates Airiam's real rates.

New engine, same discipline. This release also upgrades the reasoning model: the Design Desk now runs on GPT-5.6 Terra by default — Airiam's newest model, at high reasoning effort — with GPT-5.6 Sol selectable per chat. The grounding reframe is prompt-level, so everything above holds across models.

The one idea to remember

It separates what it knows from what it assumes.

Every number is one of two kinds — and the assistant treats them very differently.

Grounded facts · cited

What a document actually says

Anything attributed to a specific Airiam record, plus Airiam's actual rates, margins, quantities, and SOW/Quote wording. These come only from retrieved documents or your uploads, and are cited. The assistant will not invent them.

Assumption-based · labeled

Best-practice engineering estimates

Task breakdowns, phases, and hour ranges. Produced from engineering best practice even with no internal match — each one flagged as an assumption pending Design Desk engineering review, so you know exactly what to tighten.

1

You state the scope

“Consolidate the 3 local domains into the single Entra ID tenant; ~70 endpoints; include local profile migration.” Decisive scope is adopted as decided — no re-asking.

2

It grounds facts, assumes the rest

Airiam-specific facts are pulled from your documents and cited; gaps are filled from best practice (and the web, when enabled) — each assumption clearly labeled.

3

It delivers the plan

A phased task list with per-phase labor ranges — the decision-ready deliverable, not a deflection.

4

It surfaces assumptions & risks

Every estimate ends with the assumptions it made and the open questions a reviewer should tighten before quoting.

Getting the best out of it

Point it at the outcome; it does the rest.

State scope decisively

Say what's decided. Decisive scope is adopted instead of turned into a clarifying question.

Ask for the deliverable

“Produce a phased project plan with projected labor.” You'll get the plan; assumed numbers are labeled.

Turn on web search

For best-practice enrichment when you don't have a close internal example. Results are first-class, cited distinctly.

Web search seems to do nothing? It's gated by both the per-turn toggle and your account's web_search group. If toggling changes nothing, ask an admin to confirm your group membership — it may be the gate, not the model.

Admins · tune it live

Edit the assistant's standing directions — no redeploy.

Under Admin → Prompt Sections, a write-admin can change how the assistant behaves at runtime. Each section has a placement:

overrideoperating_brief overrideorg_context overridetool_addendum additivecapabilities_note additivetail

Empty = the default

Overrides replace their built-in block; additive placements append. Delete or disable a row to reset that placement to its built-in default.

Preview before you trust it

The assembled-prompt preview shows the full system prompt exactly as it will be sent — toggle web/tools to see the conditional parts.

Favor additive notes

Standing rules like “assume Windows 11 Pro unless told otherwise” fit a tail note. Reach for overrides only to replace a whole block.

Handle with care. These are the model's literal instructions — a careless override changes every answer. Edits propagate across servers within ~60 seconds, are write-admin only, and every change lands in the Audit Log. Never paste a specific customer's numbers or names here.

How we know it works

Measured, not asserted.

0.92
live, LLM-judged · was 0.22

An offline “did it produce a real, grounded estimate?” score

Before the reframe, the scenarios that matter scored 0.22 — mostly refusals. Scored again on live gpt-5.4 output after the reframe, the same rubric gives 0.92 (LLM-judged): complete, assumption-labeled plans instead of refusals, with no fabricated Airiam-specific rates. One flagship scenario still partly re-opens settled scope — measured, not hidden. “It looks the same” is a failing result established by evidence, not a subjective call.

In one sentence

It grounds Airiam facts in your documents, produces the plan and estimate you asked for, labels everything it assumes, and lets you tune its directions live.

Airiam facts cited Assumptions labeled Honors your scope Web first-class when internal is thin