- Joel Zamboni

Agentic AI DevOps: What AI Agents Do, and Don't Do, in Operations

AI DevOps in practice: the operations work AI agents can own, such as watching, routing and first analysis, and where engineers keep judgment and approval.

AI DevOps has become one of those phrases that means whatever the vendor needs it to mean. For some it is a chat assistant that writes Terraform. For others it is an autonomous system that fixes production while everyone sleeps. If you run engineering at a SaaS company, neither picture helps you decide what to trust an AI agent with on Monday morning. This post is a practical account of what AI agents do well in operations, what they should not do on their own, and where a human has to stay in the loop.

We run nine AI agents alongside senior engineers at Webera, so we have a stake in this. We will try to be specific about the limits as well as the uses.

What AI DevOps and agentic DevOps actually mean

Two different things get grouped under the same label.

AI-assisted DevOps is a person using a model as a tool. An engineer asks for a GitHub Actions workflow, a CloudWatch query or an explanation of a stack trace, then reads the answer and decides what to do with it. The human drives every step.

Agentic DevOps is a system that takes actions in a loop. An agent receives a goal or an event, calls tools (read the logs, query metrics, look at the last deploy, open a pull request), looks at the results and decides on the next step. The agent drives some of the steps, and the design question becomes which steps.

Most of the interesting decisions in AI agents for DevOps live in that second category. An assistant that gives a wrong answer wastes a few minutes. An agent with write access to production that takes a wrong action can cause an incident.

Where an AI DevOps agent earns its place

Operations work has a large share of tasks that are repetitive, time-sensitive and well defined. Those are the tasks where agents help most.

Watching

Production generates more signal than people can read. An agent can watch metrics, logs and deploy events continuously, correlate a latency increase with the release that went out twenty minutes earlier, and open an incident with that context attached. A person doing the same thing at 3am first has to wake up, log in and find the right dashboard.

Routing

Once something looks wrong, someone has to decide how bad it is and who should know. Agents are good at applying a severity policy consistently, paging the right on-call engineer, attaching the matching runbook and suppressing duplicate alerts from the same root cause. This is also where alert fatigue gets fought: an agent that filters noise well means the page that does arrive is worth reading.

First analysis

Before anyone fixes anything, someone gathers evidence: recent changes, error rates by endpoint, slow queries, resource saturation, similar past incidents. An agent can do this in parallel and hand the engineer a short summary with a hypothesis, instead of a blank terminal. It can also draft the fix as a pull request, so the engineer’s first action is reviewing a proposal rather than writing one from scratch.

Repetitive upkeep

A lot of DevOps work is checking things nobody has time to check. Is every alert linked to a runbook? Which instances have been idle for two weeks? Which IAM policies grant more than they use? Which dependencies have known vulnerabilities? Agents can run these checks on a schedule and report what changed, which keeps small problems from piling up.

What AI agents should not do on their own

The same qualities that make agents useful (speed, persistence, willingness to act) become risks when the task needs context the agent does not have.

Judgment calls that depend on business context. An agent can see that a database is oversized for its average load. It cannot know that the size is there because your largest customer runs a batch import on the first of every month, or that finance has already committed to a Reserved Instance. A recommendation is fine. An automatic downsize is not.

High-impact or irreversible changes. Deleting resources, changing IAM permissions, running data migrations, rotating credentials used by other systems and changing DNS for production domains all have a large blast radius. The OWASP project lists Excessive Agency among its top risks for LLM applications and names three causes: too much functionality, too many permissions and too much autonomy. Its main mitigation is to require a human to approve high-impact actions before they happen.

Novel failures. Agents are strongest on problems that resemble ones they have seen or that a runbook describes. When a failure is genuinely new, such as an upstream provider behaving in an undocumented way, an experienced engineer reasoning from first principles is still the better tool. The agent’s job there is to gather evidence fast and get out of the way.

Accountability. When an incident affects customers, someone has to own the decision, explain it to the customer and lead the review afterwards. That is a human role. “The agent did it” is not an answer a customer, an auditor or your board will accept.

Where humans stay in the loop

A useful way to design AI DevOps is to sort actions by risk and assign each class a default owner. This is the model we think works for most SaaS teams:

ActionAgentHuman
Read metrics, logs, traces and configDoes it continuouslyReviews summaries
Open an incident and page on-callDoes it, using the severity policyOwns the policy
Investigate and propose a causeDoes the first passConfirms or rejects
Draft a fix as a pull requestWrites itReviews and approves
Apply a change to productionRuns it through the pipeline after approvalApproves
Delete resources, change IAM, migrate dataRecommends onlyDecides and approves
Talk to customers about an incidentDrafts notesCommunicates and owns

Two principles sit behind the table. First, agents get read access broadly and write access narrowly, with every change going through the same pipeline, code review and audit trail as a human change. Second, the bar for autonomy rises with the cost of being wrong. Paging someone unnecessarily costs a few minutes. Dropping a table costs a company.

The risks worth planning for

More change, less stability. Google’s DORA research program studied this directly. Its 2025 report on AI-assisted software development found that AI adoption still has a negative relationship with software delivery stability, and that without strong automated testing, version control practices and fast feedback loops, more change volume leads to more instability. The report’s summary is blunt: AI “doesn’t fix a team; it amplifies what’s already there.” Agents in operations are subject to the same rule. If your deploys are fragile, agents that speed up change will surface that fragility faster.

Trust has to be earned. The same DORA report found that 30% of respondents had little or no trust in AI-generated code. That skepticism is healthy. Treat agent output the way you would treat a capable new hire’s work: review it, measure it and widen its scope only as it proves reliable.

Permissions sprawl. Give each agent its own role with the minimum permissions for its job, log every action (AWS CloudTrail covers API calls in your accounts), and review those roles like any other privileged access.

Noise instead of signal. An agent with skills but no judgment produces dashboards nobody reads and alerts nobody trusts. We wrote about this failure, and how we addressed it, in Intent Over Instructions: our agents start from a shared set of beliefs about what matters to the client before they get any technical instructions.

Questions to ask about AI DevOps tools and providers

  1. What can the agents change without a human approving it? Get a specific list, not a reassurance.
  2. How are agent permissions scoped and audited? Ask to see the roles and the logs.
  3. Do agent changes go through code review and the same pipeline as human changes? If agents can click in the console, that is a red flag.
  4. Who is accountable when something goes wrong? There should be a named person, not a product.
  5. What happens when the agent is unsure? Good systems escalate. Bad ones guess.
  6. Who owns the code, runbooks and configuration the agents produce? It should be you.

If you are also tightening the human side of incident response, our post on incident management for SaaS teams covers runbooks, on-call and post-mortems.

Frequently asked questions

What is AI DevOps?

AI DevOps is the use of AI models in software delivery and operations. It ranges from AI-assisted work, where an engineer asks a model for a workflow or a query and decides what to do with the answer, to agentic DevOps, where an agent calls tools in a loop to watch, investigate and propose changes. The useful question is which steps the agent is allowed to take on its own.

Will DevOps get replaced by AI?

The evidence so far points to AI changing DevOps work rather than replacing it. Agents take over much of the watching, routing, first analysis and routine checks, while judgment calls, high-impact changes, novel failures and accountability stay with people. DORA’s 2025 research also found that AI amplifies a team’s existing practices, so strong DevOps foundations matter more, not less.

Which AI tool is best for DevOps?

There is no single best tool, because the right choice depends on what you want the agent to do and what it is allowed to touch. Judge AI DevOps tools by what they can change without approval, how their permissions are scoped and audited, whether their changes go through code review, and what they do when unsure. A tool that answers those questions well is worth more than one with the longest feature list.

How Webera uses AI agents for DevOps

Webera is a DevOps subscription for SaaS teams on AWS. Senior engineers work alongside nine AI agents, each with one job: Sentinel for observability, Guardian for security, Conductor for CI/CD, Dispatcher for incidents, Optimizer for FinOps, Navigator for networking, Keeper for databases, Forge for application readiness and Warden for standards.

The split follows the model above. The agents do the watching, routing and first analysis around the clock. The engineers make the calls and ship the fixes. In the example night on our home page, Sentinel ties a latency spike to the evening deploy and opens an incident, Dispatcher pages the on-call engineer with the runbook attached, Keeper traces the slow queries to a missing index and proposes the change, and the engineer reviews it and builds the index online. The cost recommendation that comes out of that night waits for the client to decide.

Every plan includes all nine agents:

  • Pro, $4,999 a month. Unlimited requests with two active at a time, 24/7 monitoring and incident response, and an average 36-hour turnaround.
  • Elite, $9,999 a month. A named engineer and a customer success manager, a compliance concierge for SOC 2, HIPAA and PCI, and an average 24-hour turnaround.
  • Platform, $14,999 a month. A team working inside your sprints, multi-account and Kubernetes operations, and knowledge transfer to your engineers.

Each plan has a three-month minimum, billed monthly, then continues month to month, and you own all the code and configuration. You can read more about the agents on the agents page, or compare the plans on our pricing page.

Need DevOps expertise?

Our team of senior engineers can help you implement these practices.

Book a Discovery Call