Back
The rise of agentic AI on Kubernetes: a new infrastructure layer
SiTech AI Team3 წთ. საკითხავი

The rise of agentic AI on Kubernetes: a new infrastructure layer

In a post published on The New Stack, Rhys Oxenham of SUSE argues that AI agents can now observe Kubernetes clusters, reason about them and act within set limits, but their value depends on the context they can see and the boundaries teams draw around them.

AI agents can now observe Kubernetes clusters, reason about what they find and act within predefined limits. In a post published on The New Stack, Rhys Oxenham, vice president and general manager of AI at SUSE, argues that their value depends less on the models than on the context they can see and the boundaries teams draw around them. Without cluster state, policy and access rules, an agent can only guess.

What changes when AI reaches the infrastructure layer

Teams once treated AI as an application concern: models sat on top of existing systems and the stack underneath stayed mostly unchanged. As AI workloads move into production, they place new demands on the infrastructure layer covering compute, storage and networking, with resource needs that swing with training and inference bursts. A recent Forrester report describes the modern AI computing stack as stretching from the models into the infrastructure beneath them, where Kubernetes has become a control point for scheduling workloads, applying policy and connecting data centers, clouds and edge sites.

Why context decides whether agents help

Traditional automation runs the same script whether the environment has changed or not; an agentic system observes, reasons about what it finds and then acts, usually after human sign-off. The signals it receives, the policy and access context it holds and the limits on what it may change separate it from a generic assistant.

Manual management holds up on a handful of clusters but becomes unreliable as the estate grows; each new cluster adds lifecycle work such as upgrades, patching and renewals. Configuration drift is a high risk, and policies can apply unevenly from team to team. Clusters spread across data centers, clouds and edge sites leave teams without a single view, while knowledge is fragmented across logs, metrics and runbooks.

Oxenham describes the repetitive side of this work as “toil”: repeated triage, manual signal correlation, alert follow-up and routine checks. These tasks are not difficult, but they consume time and stall modernization. A recent survey reached the same conclusion: reducing toil is a clear opportunity. Kept under human review, agents can gather signals, correlate them and propose a likely cause.

Four principles for agentic Kubernetes operations

The post sets out four principles: start with observable context, so agents see cluster state, policy and history before they act; separate suggestions from actions, so any change waits for human approval and a defined scope; connect agents to existing access rules, identity and audit paths; and keep the ecosystem open. SUSE points to Rancher Prime, which it calls the industry’s first context-aware agentic AI ecosystem, where specialized agents work behind an intelligent router and act through existing access controls. Where change control must stay fully manual, agents may be limited to observation and suggestion.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.