Fleet problems
4 posts
Uptime is not reliability: the gray failures green dashboards miss (Problem Scope 1/5)
A server can pass every health check and still run below spec. Why uptime misses gray failure, and the three properties that make infrastructure reliable.
What happens when the loop stops: why remediation is the mission (3/3)
Finding faults faster does not end the firefighting loop. One BIOS incident, told two ways, shows why the fix itself has to be planned, approved and verified.
Where the loop always starts: the hardware signals your monitoring misses (2/3)
PCIe replays, BIOS drift and disks that SMART calls healthy: the hardware signals behind silent degradation, and what 35,000 servers at Flipkart taught us.
Your team is trapped in a loop: why infra teams keep firefighting (1/3)
Hardware telemetry split across tools, runbooks that age with every firmware change, servers stuck in maintenance: the loop that keeps infra teams firefighting.