Autonomous Work Order Intelligence & Operations Platform
An evidence-aware, multi-agent architecture for construction payment-protection operations — processing work orders through specialist agents for intake, research, evidence, and discrepancy detection, with controlled correction and human-in-the-loop escalation for complex cases.
Work order flow · intake → research → evidence → discrepancy → QC → human review
Recorded on the synthetic evaluation dataset (v2.1, seed 42). The ~99% figure is the intake-agent evaluation baseline, not a whole-system or real-usage accuracy claim.
Project film · 2.5 min
AI proposes. Software validates. Rules control. Humans review.
Evidence provenance, specialist agents, the deterministic state machine, correction overlays, bounded retries, the three golden paths, evaluation, and the cross-cloud deployment. Music only, with no narration; all work orders are synthetic.
Project overview
A production-oriented, evidence-aware multi-agent architecture for construction payment-protection workflows. The system is designed to process work orders through specialist AI agents for intake, research, evidence analysis, discrepancy detection, and quality control, with controlled correction and human-in-the-loop escalation for complex cases.
Key capabilities
Multi-agent AI
Specialist agents for intake, research, evidence, and discrepancy detection.
Evidence provenance
Separate customer claims from independently researched information.
Controlled autonomy
A deterministic Python control plane holds workflow authority, not the model.
Immutable corrections
Proposes corrections without modifying original source records.
Human-in-the-loop
Safe escalation for cases with unresolved material conflicts.
Deployed on Azure
Runs on Azure Container Apps, deployed by GitHub Actions with OIDC.
System architecture
Cross-cloud architecture, from work order to operational recommendation.
Golden path scenarios
Three illustrative scenarios demonstrate the system's intended safety boundaries.
Illustrative synthetic scenarios — for demonstration only, not measured results from a deployed system.
Intake → Research → Evidence → Discrepancy → QC → Complete
Intake → Research → Evidence → Discrepancy → Correction → Research → QC → Complete
Intake → Research → Evidence → Discrepancy → QC → Human review
Performance & evaluation
Recorded intake-agent evaluation
Recorded on 120 synthetic work orders across 24 scenario families. Synthetic evaluation only; these are not results from real customer workloads.
Tech stack
Built with
Assumptions & limitations
- The deployed application runs on synthetic work-order data. It demonstrates the architecture end to end; it is not serving real customer workloads under production monitoring.
- Evaluation figures come from the synthetic dataset and cover the intake agent. They are not measured results from real usage.
- The repository link points to this project's own directory and documentation inside the shared portfolio repository.
- Execution state is held in memory, so the backend runs as a single replica; durable shared state is future work.