Transaction flow simulation platform

A walkthrough of the platform's four capabilities through one architecture model. The example domain is claims submission. All numbers, traces, and code shown are realistic — produced under the same rules the live system follows.

Stage 1 of 9
Describe the system in plain language
A user describes what they want to model. The platform does not require a structured input — prose works. Importers exist for AsyncAPI / OpenAPI / OpenTelemetry, but the front door is language.
Natural language input
When a patient submits a claim through the patient portal, the request goes through an API gateway that does auth and rate limiting, then to the claims service. The claims service writes the claim to Postgres and publishes a ClaimSubmitted event to the RCM service bus. Two subscribers consume that event: an eligibility checker that calls an external payer system, and a history service that reads from the claims database. The eligibility check can be slow and is known to fail occasionally — we want to retry with backoff and have a circuit breaker. We expect about 200 requests per second at peak.

No syntax. No schema. The user describes the system the way they think about it.

Stage 2 of 9
Generated IR with honest provenance
The platform produces an Architecture IR — the single canonical truth artifact every other capability reads from. Every value is tagged with where it came from: specified (the user said it), inferred (the LLM guessed), or default (a documented engine default). Honesty is structural, not optional.
Architecture IR (excerpt)
{
  "ir_version": "1.0",
  "system": "Claims Submission Pipeline",
  "components": [
    {
      "id": "patient_portal", "kind": "client", "name": "Patient Portal"
    },
    {
      "id": "api_gateway", "kind": "gateway", "name": "API Gateway",
      "capacity": {
        "service_time": { "value": {"kind":"lognormal","p50_us":8000,"p95_us":20000},
                          "provenance": "inferred" },
        "concurrency": { "value": 4, "provenance": "default" }
      }
    },
    {
      "id": "claims_svc", "kind": "service", "name": "Claims Service",
      "capacity": {
        "service_time": { "value": {"kind":"lognormal","p50_us":25000,"p95_us":80000},
                          "provenance": "inferred" },
        "concurrency": { "value": 8, "provenance": "default" }
      }
    },
    { "id": "claims_db", "kind": "datastore", "name": "Claims Postgres", /*...*/ },
    { "id": "rcm_bus", "kind": "topic", "name": "RCM Service Bus" },
    {
      "id": "eligibility_svc", "kind": "external", "name": "Eligibility Check",
      "failure": {
        "error_rate": { "value": 0.05, "provenance": "specified" },
        "timeout_us": { "value": 3000000, "provenance": "inferred" },
        "retry": { "count": {"value":3,"provenance":"specified"}, /*...*/ },
        "circuit_breaker": { /* per user request */ }
      }
    }
  ],
  "workloads": [{
    "id": "peak_traffic", "kind": "open",
    "arrival": { "kind": "poisson",
      "rate_per_sec": { "value": 200, "provenance": "specified" } }
  }]
}
Inferred-values manifest (sorted by impact on headline metrics)
inferred
eligibility_svc.failure.timeout_us = 3,000,000 μs (3 seconds)
The user said "the eligibility check can be slow" but didn't name a timeout.
Source span: "the eligibility check can be slow and is known to fail occasionally"
inferred
claims_svc.capacity.service_time (lognormal, p50=25ms, p95=80ms)
The user didn't specify timing for the claims service.
Source span: "the claims service writes the claim to Postgres"
specified
workload.arrival.rate_per_sec = 200
Directly from the user's "about 200 requests per second at peak".
default
api_gateway.capacity.concurrency = 4
Engine default for gateway components when not specified.

The user sees, in priority order, which guesses most affect downstream simulation. They can correct any before review.

Stage 3 of 9
Visual canvas — the same IR, drawn
The IR renders as an architecture diagram with full traceability. Editing the canvas edits the IR; the visualization is a view, not a separate truth. Layout is auto-arranged but adjustable.
Patient Portal client API Gateway gateway · 4 workers Claims Service service · 8 workers Claims Postgres datastore RCM Bus topic · pub/sub Eligibility Check external · retries · breaker History Service service · 4 workers Claims Postgres datastore (read) REST REST REST async async
service
gateway
datastore
topic/queue
external
─── sync
- - - async
Stage 4 of 9
Storyboard — tokens move through the system
A single transaction is animated through the topology. This is a replay of an actual discrete-event simulation trace — not a decorative timer. Each token is a real claim submission; the timing reflects modeled service times and queueing.
Patient Portal client API Gateway q: 0 Claims Service q: 0 Postgres RCM Bus q: 0 Eligibility History Svc
Speed: t = 0.0 ms

The token color encodes outcome: ● in flight, ● completed, ● error path. Queue depth labels update from real event data.

Stage 5 of 9
Stochastic results — predicted behavior at 200 rps
30 replications, warm-up discarded, 95% confidence intervals reported as ranges of outcomes. Numbers are queueing-theoretically valid — the engine is validated against M/M/1, M/M/c, and M/G/1 closed-form solutions.
End-to-end latency, claim submission flow
p50
142 ms
95% CI: 138–146 ms
From specified workload + inferred service times
p95
412 ms
95% CI: 396–428 ms
Driven by Eligibility Check retries
p99
3.2 s
95% CI: 2.8–3.6 s
Driven by 3s timeout on Eligibility
success rate
94.7%
95% CI: 94.1–95.2%
Retries recover most failures

Every metric is rendered as a range, not a point. Provenance is displayed alongside every value.

Findings — observation, cause, recommended next experiment
p99 latency is dominated by Eligibility Check timeouts
Under the specified workload, p99 falls between 2.8s and 3.6s. The dominant contributor is the eligibility_svc external — when it fails, the 3s timeout fires before retry kicks in. That single component contributes 78% of p99 latency.

Smallest mitigation: lower the timeout to 800ms (still longer than p95 of normal responses at 320ms) and rely on retry to recover. Predicted new p99: 1.2s (a 62% improvement). Try this in what-if comparison next.
The Claims Service is operating well within capacity
Utilization at claims_svc is 42% under peak load. There is significant headroom; no scaling action is needed.
RCM Bus subscriber buffer approaching saturation
The history-service subscription buffer reaches 78% capacity during burst periods. Operating near saturation — small changes in load will produce large changes in latency. The load level at which behavior stabilizes is approximately 165 rps.
Stage 6 of 9
What-if — compare two variants of the design
The platform reuses the same arrival realizations across both runs (common random numbers), so the comparison is deterministic. The primary unit of analysis is the delta, not the absolute.
Variant: lower eligibility timeout from 3000 ms to 800 ms
Metric Baseline Variant Delta
p50 latency 142 ms 144 ms +2 ms (+1%)
p95 latency 412 ms 438 ms +26 ms (+6%)
p99 latency 3,200 ms 1,180 ms −2,020 ms (−63%)
Success rate 94.7% 94.2% −0.5%
Eligibility retries 8.4% 11.2% +2.8%
Eligibility breaker trips 0.7/min 0.4/min −0.3/min

The tradeoff is visible: aggressive timeout means more retries (slightly raising p95) but dramatic improvement in p99 (where the platform's SLO lives). The team can now make this decision with evidence, not guess.

Stage 7 of 9
Generate scaffolded code from the same IR
The IR is consumed by the codegen target to produce a runnable FastAPI project. Every field appears as a structured annotation; the failure model (timeout, retry, circuit breaker) becomes real middleware. The user's business logic goes in marked extension points.
Generated: app/externals/eligibility_svc/client.py
# @ir-generated:eligibility_svc_client
# Generated from IR canonical hash: sha256:a91f4c...
# DO NOT EDIT this region. Make changes via the IR or in @ir-extension-point regions below.

from tenacity import retry, stop_after_attempt, wait_exponential_jitter
from purgatory import AsyncCircuitBreaker
import httpx

@ir_failure(
    timeout_us=800_000,                     # after what-if; provenance=specified
    retry_count=3,                          # provenance=specified
    retry_backoff="jittered_exponential",    # provenance=default
    circuit_breaker_threshold=5,            # provenance=specified
    runtime_effect="active"
)
class EligibilityClient:
    def __init__(self, base_url: str):
        self._client = httpx.AsyncClient(timeout=0.8)
        self._breaker = AsyncCircuitBreaker(name="eligibility_svc", threshold=5)

    @retry(stop=stop_after_attempt(3), wait=wait_exponential_jitter())
    async def check_eligibility(self, claim_id: str) -> EligibilityResponse:
        async with self._breaker:
            response = await self._client.post(
                "/eligibility/check",
                # @ir-extension-point:eligibility_request_transform
                json={"claim_id": claim_id}
                # @ir-extension-point-end
            )
            response.raise_for_status()
            return EligibilityResponse.parse_obj(response.json())

# @ir-generated-end:eligibility_svc_client

Active fields (timeout, retry, circuit breaker) are real middleware. Service-time distributions and other simulation-only fields appear as documentation annotations but do not affect runtime. The user sees the full IR projected into their code, with full traceability.

Stage 8 of 9
Reconcile prediction against production reality
After the system ships, OpenTelemetry traces from production are compared to the predicted IR-derived traces. Drift is decomposed causally — the platform identifies which component's behavior diverged most from its IR model, with proposed structured patches to update the IR.
Reconciliation run: 14 days, ~8.4M observed transactions
observed p99
1,580 ms
predicted: 1,180 ms
+33% drift
observed success
93.4%
predicted: 94.2%
−0.8 pp drift
retry rate
17.6%
predicted: 11.2%
+57% drift
unmapped spans
2.3%
in OTel, not IR
audit-log calls
Drift findings, ranked by contribution × confidence
magnitudeeligibility_svc service-time distribution is wider than modeled
Observed p95 at eligibility_svc is 580 ms; the IR models it as p95=320 ms. Observed distribution suggests lognormal with p50=210 ms, p95=580 ms — significantly heavier-tailed than the inferred IR value.

Contribution to total drift: 64%. Confidence: high (8.4M samples).
Proposed IR patch: eligibility_svc.capacity.service_time → lognormal(p50=210ms, p95=580ms) with provenance changed from inferred to reconciled.
frequencyHigher real-world failure rate than specified
Observed eligibility error rate is 8.2%; the IR specifies 5%. This explains the higher retry rate (17.6% vs predicted 11.2%).

Contribution to total drift: 22%. Confidence: high.
Proposed IR patch: eligibility_svc.failure.error_rate → 0.082, provenance reconciled. The user may want to investigate the payer system before accepting; this could be a real degradation.
topologyAn audit-log service is in production but not modeled
2.3% of observed spans come from audit-log-service, called by the claims service after every write. No corresponding component exists in the IR.

Contribution to total drift: 8% (small but real). Confidence: medium.
Proposed IR patch: add a new audit_log_svc component (kind=service) and a sync connection from claims_svc; binding observed at p50=12ms.

All proposed patches are reviewable structured diffs, not opaque suggestions. Accepting a patch updates the IR with provenance reconciled and a link to this reconciliation run — a full audit trail.

Stage 9 of 9
Four capabilities, one architecture model
The gap nobody else has filled: connecting requirements, picture, simulation, and code through a single IR — then continuously reconciling that IR against running production. Existing tools cover one or two of these capabilities in isolation. This platform unifies all four with the closed loop.
Requirements → IR
NL prose, AsyncAPI/OpenAPI imports, architecture images, OTel service graphs — all produce the same IR. Honest provenance tracks every value's origin.
IR → visual diagram
The canvas is a view of the IR. Editing the canvas edits the IR. Hierarchical zoom and overlays make 1-component or 1000-component systems navigable.
Architecture IR
single canonical truth
IR → animated simulation
Discrete-event engine produces validated, queueing-theoretically sound results. Animation is a replay of the trace — never a decorative timer. Same engine produces storyboard demos and stochastic SLO predictions.
IR → scaffolded code
FastAPI reference target produces runnable, structured code with the IR's failure model as real middleware. Soft round-trip preserves user edits across regenerations.
The closed-loop differentiator
Production OpenTelemetry traces are reconciled against predicted IR-derived traces. Drift is decomposed causally; the platform proposes structured IR updates to close the gap between model and reality. This is the world-class capability that distinguishes this platform from every visualization or simulation tool currently in the market.