Back
AI handles incidents, engineers lose touch with their systems
SiTech AI Team3 წთ. საკითხავი

AI handles incidents, engineers lose touch with their systems

Incident-response automation resolves routine alerts reliably enough that engineers get less practice with their own systems, a paradox researchers warned about decades ago.

AI incident-response tools, commonly called "AI SREs", can inspect alerts, form hypotheses, query telemetry, correlate recent deployments and sometimes implement a fix themselves. In a blog post, Rootly's AI Labs lead Sylvain Kalache argues that this progress carries a cost: engineers are losing touch with the systems they are responsible for.

Routine incidents are how intuition is built

Kalache, a former LinkedIn SRE and co-founder of Holberton School, writes that automated resolution of routine alerts removes the practice human responders need. Routine incidents are exactly how engineers develop an intuition for how their systems behave and fail. When a hard, never-seen-before incident arrives that automation cannot solve, responders will have to take over with less practice than they would have had before.

He cites human-factors researcher Lisanne Bainbridge, whose 1983 paper "The Ironies of Automation" described the paradox: automation reduces operators' opportunities to practise routine work while leaving them responsible for abnormal situations, so operators need more skill and more training than before automation. Kalache predicts that the average MTTR for most incidents will fall, but that resolution time for complex incidents will rise because responders have lost touch with their systems.

What aviation does about rare failures

Aviation offers a model. Automation handles much of the flying, yet pilots remain responsible for engine failures, unreliable instruments, rejected takeoffs and stalls. Modern turbine engines experience fewer than one in-flight shutdown per 100,000 engine flight hours — rare enough that a commercial pilot may complete an entire career without experiencing one outside a simulator. Under US FAA rules, captains must complete recurrent training or a proficiency check every six months, including engine failure during takeoff. The price of losing that skill shows up in accidents: on TransAsia Airways Flight 235, the crew misidentified which engine had failed, and the aircraft crashed 117 seconds after the first warning.

Simulations, trainers and comprehension debt

Kalache's employer, incident-management company Rootly, partnered with Uptime Labs to apply the idea through realistic incident simulations: engineers take the incident commander's seat during a simulated e-commerce outage, using observability tools while coordinating with LLM-powered stakeholders in Slack. AI can also serve as a trainer — a responder can ask an agent to explain the steps it took, the signals it examined and the evidence behind its diagnosis. But explanation and observation are not substitutes for practice, Kalache insists, comparing it to learning tennis by watching rather than playing.

His conclusion is that engineering teams risk accumulating "comprehension debt": a growing gap between how their systems work and how well responders understand them. Regular hands-on interaction with the system, unfamiliar failures, tabletop exercises and chaos engineering should be treated as part of on-call readiness rather than optional extras.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.