A walkthrough of the platform's four capabilities through one architecture model. The example domain is claims submission. All numbers, traces, and code shown are realistic — produced under the same rules the live system follows.
No syntax. No schema. The user describes the system the way they think about it.
{ "ir_version": "1.0", "system": "Claims Submission Pipeline", "components": [ { "id": "patient_portal", "kind": "client", "name": "Patient Portal" }, { "id": "api_gateway", "kind": "gateway", "name": "API Gateway", "capacity": { "service_time": { "value": {"kind":"lognormal","p50_us":8000,"p95_us":20000}, "provenance": "inferred" }, "concurrency": { "value": 4, "provenance": "default" } } }, { "id": "claims_svc", "kind": "service", "name": "Claims Service", "capacity": { "service_time": { "value": {"kind":"lognormal","p50_us":25000,"p95_us":80000}, "provenance": "inferred" }, "concurrency": { "value": 8, "provenance": "default" } } }, { "id": "claims_db", "kind": "datastore", "name": "Claims Postgres", /*...*/ }, { "id": "rcm_bus", "kind": "topic", "name": "RCM Service Bus" }, { "id": "eligibility_svc", "kind": "external", "name": "Eligibility Check", "failure": { "error_rate": { "value": 0.05, "provenance": "specified" }, "timeout_us": { "value": 3000000, "provenance": "inferred" }, "retry": { "count": {"value":3,"provenance":"specified"}, /*...*/ }, "circuit_breaker": { /* per user request */ } } } ], "workloads": [{ "id": "peak_traffic", "kind": "open", "arrival": { "kind": "poisson", "rate_per_sec": { "value": 200, "provenance": "specified" } } }] }
The user sees, in priority order, which guesses most affect downstream simulation. They can correct any before review.
The token color encodes outcome: ● in flight, ● completed, ● error path. Queue depth labels update from real event data.
Every metric is rendered as a range, not a point. Provenance is displayed alongside every value.
eligibility_svc external — when it fails, the 3s timeout fires before retry kicks in. That single component contributes 78% of p99 latency.claims_svc is 42% under peak load. There is significant headroom; no scaling action is needed.
| Metric | Baseline | Variant | Delta |
|---|---|---|---|
| p50 latency | 142 ms | 144 ms | +2 ms (+1%) |
| p95 latency | 412 ms | 438 ms | +26 ms (+6%) |
| p99 latency | 3,200 ms | 1,180 ms | −2,020 ms (−63%) |
| Success rate | 94.7% | 94.2% | −0.5% |
| Eligibility retries | 8.4% | 11.2% | +2.8% |
| Eligibility breaker trips | 0.7/min | 0.4/min | −0.3/min |
The tradeoff is visible: aggressive timeout means more retries (slightly raising p95) but dramatic improvement in p99 (where the platform's SLO lives). The team can now make this decision with evidence, not guess.
# @ir-generated:eligibility_svc_client # Generated from IR canonical hash: sha256:a91f4c... # DO NOT EDIT this region. Make changes via the IR or in @ir-extension-point regions below. from tenacity import retry, stop_after_attempt, wait_exponential_jitter from purgatory import AsyncCircuitBreaker import httpx @ir_failure( timeout_us=800_000, # after what-if; provenance=specified retry_count=3, # provenance=specified retry_backoff="jittered_exponential", # provenance=default circuit_breaker_threshold=5, # provenance=specified runtime_effect="active" ) class EligibilityClient: def __init__(self, base_url: str): self._client = httpx.AsyncClient(timeout=0.8) self._breaker = AsyncCircuitBreaker(name="eligibility_svc", threshold=5) @retry(stop=stop_after_attempt(3), wait=wait_exponential_jitter()) async def check_eligibility(self, claim_id: str) -> EligibilityResponse: async with self._breaker: response = await self._client.post( "/eligibility/check", # @ir-extension-point:eligibility_request_transform json={"claim_id": claim_id} # @ir-extension-point-end ) response.raise_for_status() return EligibilityResponse.parse_obj(response.json()) # @ir-generated-end:eligibility_svc_client
Active fields (timeout, retry, circuit breaker) are real middleware. Service-time distributions and other simulation-only fields appear as documentation annotations but do not affect runtime. The user sees the full IR projected into their code, with full traceability.
eligibility_svc is 580 ms; the IR models it as p95=320 ms. Observed distribution suggests lognormal with p50=210 ms, p95=580 ms — significantly heavier-tailed than the inferred IR value.eligibility_svc.capacity.service_time → lognormal(p50=210ms, p95=580ms) with provenance changed from inferred to reconciled.
eligibility_svc.failure.error_rate → 0.082, provenance reconciled. The user may want to investigate the payer system before accepting; this could be a real degradation.
audit-log-service, called by the claims service after every write. No corresponding component exists in the IR.audit_log_svc component (kind=service) and a sync connection from claims_svc; binding observed at p50=12ms.
All proposed patches are reviewable structured diffs, not opaque suggestions. Accepting a patch updates the IR with provenance reconciled and a link to this reconciliation run — a full audit trail.