Research
How intelligent physical systems fail.
We study failure as a first-class research object — not an edge case to patch over, but the primary signal for building AI that can be trusted in the physical world.
Physical AI
We study the gap between simulated training environments and physical deployment — sensor drift, actuator wear, latency budgets, and the compounding errors that only appear after thousands of real-world cycles.
Manufacturing Intelligence
Production floors are adversarial to AI in ways clean datasets never capture — lighting changes, material variance, and machine wear. We build the evaluation methodology that catches these before they cause downtime.
Retail Intelligence
Retail environments generate long-tail visual scenarios at a scale few other domains do. We benchmark how perception systems handle occlusion, restocking chaos, and adversarial packaging.
Robotics
From gripper slip to localization drift in warehouse fleets, we catalog the failure modes that separate a demo from a dependable system running three shifts a day.
Vision Systems
We stress-test vision pipelines against the conditions that actually occur on-site: dust, glare, motion blur, and camera degradation over a hardware lifecycle.
Agentic Automation
One inventory-reconciliation agent we studied retried the same malformed record for eleven hours, because its own status check couldn't tell 'still working' from 'stuck on the same step forever.' We treat that distinction — progress observability — as a first-class requirement, not a logging nicety.
Failure Analysis
Most incident reports describe symptoms, not causes. We develop reproducible methods for tracing a failure back to its origin — data, model, or system integration.
Reliability Engineering
One manufacturing partner turned a failure catalog into three specific guardrails on their production line and cut their false-accept defect rate by 64%. That's the bar we hold this work to — not a taxonomy someone reads, but a number that moves.