Reliability for a distributed device network

Managed connectivity provider

99.99% uptime · 1000% fulfillment throughput · 85% smaller firmware, 12x faster startup

The problem

Hundreds of thousands of remote devices under contractual uptime obligations. Failures meant expensive field dispatches, and the firmware delivery process couldn't keep up with sales growth, with concurrent updates causing instability.

The approach

Established a Site Reliability Engineering practice with a clear charter and reliability metrics. Re-architected the firmware flashing system for parallel, reliable updates over wired and wireless connections, and guided Kubernetes adoption with zero-downtime deployments and shared observability.

The outcome

Core service uptime reached 99.99%, sharply reducing field visits. Firmware update success improved dramatically and fulfillment throughput scaled with revenue.