The Design Desk assistant now produces a phased project plan and level-of-effort estimate even when there is no comparable internal document — grounding Airiam-specific facts in your documents, filling the rest from engineering best practice, and clearly labeling everything it assumes so an engineer can adjust it.
The same request that used to stall now returns a phased plan you can review — because the fix was the assistant's governance prompt, not the model.
New engine, same discipline. This release also upgrades the reasoning model: the Design Desk now runs on GPT-5.6 Terra by default — Airiam's newest model, at high reasoning effort — with GPT-5.6 Sol selectable per chat. The grounding reframe is prompt-level, so everything above holds across models.
Every number is one of two kinds — and the assistant treats them very differently.
Anything attributed to a specific Airiam record, plus Airiam's actual rates, margins, quantities, and SOW/Quote wording. These come only from retrieved documents or your uploads, and are cited. The assistant will not invent them.
Task breakdowns, phases, and hour ranges. Produced from engineering best practice even with no internal match — each one flagged as an assumption pending Design Desk engineering review, so you know exactly what to tighten.
“Consolidate the 3 local domains into the single Entra ID tenant; ~70 endpoints; include local profile migration.” Decisive scope is adopted as decided — no re-asking.
Airiam-specific facts are pulled from your documents and cited; gaps are filled from best practice (and the web, when enabled) — each assumption clearly labeled.
A phased task list with per-phase labor ranges — the decision-ready deliverable, not a deflection.
Every estimate ends with the assumptions it made and the open questions a reviewer should tighten before quoting.
Say what's decided. Decisive scope is adopted instead of turned into a clarifying question.
“Produce a phased project plan with projected labor.” You'll get the plan; assumed numbers are labeled.
For best-practice enrichment when you don't have a close internal example. Results are first-class, cited distinctly.
Web search seems to do nothing? It's gated by both the per-turn toggle and your account's web_search group. If toggling changes nothing, ask an admin to confirm your group membership — it may be the gate, not the model.
Under Admin → Prompt Sections, a write-admin can change how the assistant behaves at runtime. Each section has a placement:
operating_brief
overrideorg_context
overridetool_addendum
additivecapabilities_note
additivetail
Overrides replace their built-in block; additive placements append. Delete or disable a row to reset that placement to its built-in default.
The assembled-prompt preview shows the full system prompt exactly as it will be sent — toggle web/tools to see the conditional parts.
Standing rules like “assume Windows 11 Pro unless told otherwise” fit a tail note. Reach for overrides only to replace a whole block.
Handle with care. These are the model's literal instructions — a careless override changes every answer. Edits propagate across servers within ~60 seconds, are write-admin only, and every change lands in the Audit Log. Never paste a specific customer's numbers or names here.
Before the reframe, the scenarios that matter scored 0.22 — mostly refusals. Scored again on live gpt-5.4 output after the reframe, the same rubric gives 0.92 (LLM-judged): complete, assumption-labeled plans instead of refusals, with no fabricated Airiam-specific rates. One flagship scenario still partly re-opens settled scope — measured, not hidden. “It looks the same” is a failing result established by evidence, not a subjective call.