Implemented flagship

Autonomous Work Order Intelligence & Operations Platform

An evidence-aware, multi-agent architecture for construction payment-protection operations — processing work orders through specialist agents for intake, research, evidence, and discrepancy detection, with controlled correction and human-in-the-loop escalation for complex cases.

Work order flow · intake → research → evidence → discrepancy → QC → human review

120Synthetic work orders
24Scenario families
154/156Recorded intake-agent evaluation (~99%)
75Automated tests passing
DeployedAzure Container Apps

Recorded on the synthetic evaluation dataset (v2.1, seed 42). The ~99% figure is the intake-agent evaluation baseline, not a whole-system or real-usage accuracy claim.

Project film · 2.5 min

AI proposes. Software validates. Rules control. Humans review.

Evidence provenance, specialist agents, the deterministic state machine, correction overlays, bounded retries, the three golden paths, evaluation, and the cross-cloud deployment. Music only, with no narration; all work orders are synthetic.

Project overview

A production-oriented, evidence-aware multi-agent architecture for construction payment-protection workflows. The system is designed to process work orders through specialist AI agents for intake, research, evidence analysis, discrepancy detection, and quality control, with controlled correction and human-in-the-loop escalation for complex cases.

The key architectural question isn't connecting multiple agents — it's deciding where AI autonomy should stop, and where deterministic software should take control.

Key capabilities

Multi-agent AI

Specialist agents for intake, research, evidence, and discrepancy detection.

Evidence provenance

Separate customer claims from independently researched information.

Controlled autonomy

A deterministic Python control plane holds workflow authority, not the model.

Immutable corrections

Proposes corrections without modifying original source records.

Human-in-the-loop

Safe escalation for cases with unresolved material conflicts.

Deployed on Azure

Runs on Azure Container Apps, deployed by GitHub Actions with OIDC.

System architecture

Cross-cloud architecture, from work order to operational recommendation.

View architecture in repository
Synthetic data layer (AWS)
S3Document storage
LambdaData processing
API GatewaySecure access
MCP integration
MCP adapterTool integration
Secure communicationAuthenticated access
Tool governanceControlled access
Multi-agent system (Microsoft Foundry)
Intake agentUnderstand work-order input
Research agentInvestigate information
Evidence agentConstruct evidence view
Discrepancy agentDetect matches, conflicts, gaps
QC agentFinal quality review
Deterministic control plane (Python/FastAPI)
Structured parsingValidate agent outputs
State machineControl workflow routing
Correction logicManage correction limits
API layerRESTful asynchronous API
Frontend & deployment
Next.jsOperations console
DockerContainerized deployment
Azure Container AppsHosting
GitHub ActionsOIDC-based CI/CD

Golden path scenarios

Three illustrative scenarios demonstrate the system's intended safety boundaries.

Illustrative synthetic scenarios — for demonstration only, not measured results from a deployed system.

01Straight-through autonomySYN-WO-000001

IntakeResearchEvidenceDiscrepancyQCComplete

Human review: noCorrections: 0
02Controlled autonomous correctionSYN-WO-000116

IntakeResearchEvidenceDiscrepancyCorrectionResearchQCComplete

Human review: noCorrections: 1
03Safety escalationSYN-WO-000111

IntakeResearchEvidenceDiscrepancyQCHuman review

Human review: yesCorrections: 0

Performance & evaluation

Recorded intake-agent evaluation

154/156 (~99%)Correct evaluations
100%Tool selection accuracy
15.1sP50 latency
18.8sP95 latency
83%Tool output utilization
120 / 24Work orders / scenario families

Recorded on 120 synthetic work orders across 24 scenario families. Synthetic evaluation only; these are not results from real customer workloads.

Tech stack

Built with

PythonFastAPIMicrosoft FoundryGPT-5-miniMCPAWS S3AWS LambdaAPI GatewayNext.jsReactTypeScriptTailwind CSSDockerAzure Container AppsGitHub ActionsOIDCpytest

Project documentation

Read more

Assumptions & limitations

  1. The deployed application runs on synthetic work-order data. It demonstrates the architecture end to end; it is not serving real customer workloads under production monitoring.
  2. Evaluation figures come from the synthetic dataset and cover the intake agent. They are not measured results from real usage.
  3. The repository link points to this project's own directory and documentation inside the shared portfolio repository.
  4. Execution state is held in memory, so the backend runs as a single replica; durable shared state is future work.