- Agents pay off on high-volume, rules-heavy work with clear success criteria — intake, triage, reconciliation, and evidence gathering.
- Small error rates compound: ten steps at 95% reliability each succeed only about 60% of the time end to end.
- Let the LLM decide and a deterministic workflow execute. Put a human in front of every irreversible action.
- Treat every tool an agent can call — including MCP servers — like a third-party integration: scoped, reviewed, and logged.
In 2024 the enterprise AI conversation was about chatbots. In 2026 it is about agents — software that reads a request, plans the steps, calls other systems, and gets work done with little or no human input. The promise is real, but so is the backlash. Gartner has predicted that more than 40% of agentic AI projects will be cancelled by the end of 2027 because of rising costs, unclear business value, or weak risk controls.
So the useful question is not “should we use agents?” but “which of our workflows are agent-shaped, and what has to be true before we let software act on our behalf?” This guide answers both, based on what we see building AI and automation systems for clients in insurance, healthcare, and technology.
What actually makes something an agent
The word is used loosely, so it helps to separate three patterns that are often sold under the same name:
| Pattern | What it does | Who acts | Typical risk |
|---|---|---|---|
| Chatbot / assistant | Answers questions from documents or data | A person reads the answer | Wrong answer |
| Copilot | Drafts an email, summary, code, or form for review | A person approves and sends | Missed review |
| Agent | Plans steps and calls tools or APIs to complete a task | The software, with checkpoints | Wrong action in a real system |
The jump from copilot to agent is the jump from “suggests” to “does.” That is where most of the value is — and where most of the engineering effort has to go.
Where agents work today
The workflows that succeed share four traits: high volume, written rules, systems with APIs, and an outcome you can measure. In practice, that points to a handful of categories:
- Document-heavy intake. Claims, invoices, onboarding packets, and referrals arrive in many formats. An agent can classify each document, extract the fields, check them against your systems, and route only the exceptions to people. Our guide to intelligent document processing covers this pattern in depth.
- Support and IT triage. Tickets already carry categories, history, and resolutions. An agent can gather diagnostics, suggest the fix, and resolve routine requests, escalating anything unusual. Support portals like the one we built for Goodfellows are a natural foundation because the data is already structured.
- Knowledge work with a follow-up action. “What does our policy say about this, and open the request if it qualifies.” This builds on a private retrieval (RAG) system and adds one or two safe actions.
- Reconciliation and reporting. Matching records across systems, flagging mismatches, and preparing the first draft of a recurring report.
- Compliance evidence collection. Pulling access reviews, configuration snapshots, and logs into an audit-ready package on a schedule.
Where agents fail
Most failed agent projects we review fall into one of these traps:
- Open-ended goals. “Grow our pipeline” is not a task an agent can verify it has completed. “Enrich these 200 leads with firmographic data” is.
- No clean system access. If the only way into a system is clicking through screens, the agent inherits all the fragility of screen scraping.
- Expensive mistakes without review. Payments, denials, deletions, and customer communications need a human checkpoint until the error rate is proven.
- No ground truth. Without a set of past cases with known correct outcomes, nobody can say whether the agent is getting better or worse.
- Long chains of steps. Reliability multiplies. If each step is right 95% of the time, a ten-step chain completes correctly only about 60% of the time (0.9510 ≈ 0.60). Shorter chains with validation between steps win.
An architecture for agents you can trust
The pattern we use is bounded autonomy: the model makes judgement calls, but everything around it is ordinary, testable software.
- Deterministic orchestration. The workflow (steps, retries, timeouts) is code. The LLM is called at specific decision points — classify, extract, choose the next action — not left to roam.
- Tool allow-lists and least privilege. Each agent gets its own service identity with only the permissions that workflow needs. Read access is the default; write access is the exception.
- Human approval for irreversible actions. Payments, external emails, record deletion, and adverse decisions go to an approval queue with the agent’s reasoning attached.
- Validation between steps. Extracted values are checked against business rules and source systems before the next step runs.
- Full tracing. Every prompt, tool call, input, and output is logged with a run ID so any decision can be replayed and explained to an auditor.
- Evaluation and budgets. A test set of real cases runs before every change, and each run has a cost and step limit — plus a kill switch.
If you cannot draw the agent’s workflow as a flowchart with clear checkpoints, it is not ready for production. The flowchart is what your security team, your auditors, and your operations staff will review.
MCP and how agents reach your systems
The Model Context Protocol (MCP), introduced in late 2024 and now supported across the major AI platforms, has become the standard way to expose tools and data to agents. Instead of custom code for every model and every system, you publish an MCP server once — for your CRM, ticketing system, or document store — and any compatible agent can use it.
That convenience creates new risks worth planning for:
- Prompt injection through content. A document, email, or web page the agent reads can contain instructions. Treat all retrieved content as untrusted input.
- Over-broad servers. An MCP server that exposes “run any query” gives the agent far more power than the workflow needs. Expose narrow, purpose-built tools instead.
- Unvetted third-party servers. Review, pin versions, and scope credentials exactly as you would for any other integration.
Choosing your first agent
Score your candidate workflows from 1 to 5 on each of these, and start with the highest total:
| Criterion | Score high when… |
|---|---|
| Volume | The task happens hundreds or thousands of times a month |
| Rule clarity | The rules are written down, even if exceptions exist |
| System access | The systems involved have APIs or exports |
| Cost of error | Mistakes are cheap to catch and reverse |
| Measurability | You already track time, cost, or error rate for the task |
Then run a short, fixed-scope pilot on real data with a baseline and a success threshold agreed up front. If the numbers are there, extend to production; if not, you have spent weeks, not a year, finding out. We explain that structure in why most GenAI pilots never reach production.
Frequently asked questions
What is the difference between an AI agent and RPA?
RPA follows a fixed script and breaks when a screen or document changes. An AI agent decides which step to take next based on the content in front of it, so it handles variation, such as differently formatted documents or free-text requests. The best designs combine both: deterministic workflows for the steps that never change, and an LLM only at the decision points.
Are AI agents safe to use with regulated or confidential data?
They can be, if the agent runs in your own cloud account or private environment, uses least-privilege credentials, requires human approval for irreversible actions, and logs every tool call. Safety comes from the architecture around the model, not from the model itself.
How long does it take to deploy a first AI agent?
A focused pilot on one workflow typically shows measurable results in four to six weeks. Hardening it for production — single sign-on, monitoring, cost controls, and runbooks — usually adds a few more sprints, depending on the systems it must connect to.
Which LLM should we use for enterprise agents?
Stay model-agnostic. Amazon Bedrock, Azure OpenAI, Google Vertex AI, and self-hosted open-weight models can all power agents. Pick based on data residency, cost per task, and how well the model performs on an evaluation set built from your real cases, and design so the model can be swapped later.
