How does AI move from an isolated tool into a functioning business workflow?
Having an AI tool is not the same as having AI work inside your processes. Integration is the step most organisations underestimate: connecting AI to business systems, defining when a human decides, and measuring the result against the workflow it was meant to improve. This page explains what integration actually requires โ from assistance to automation to agentic workflows โ and what the evidence says about when it works and when it does not.
From assistance to agentic โ four levels of AI integration into business workflows.
AI does not enter a workflow at full autonomy. It enters at the assistance level and moves up as the organisation builds the surrounding infrastructure: connections to business systems, rules for when the AI acts, human decision points, and measurement of what the integration produces. Each level changes who does what โ and what the organisation needs to have in place for the level to work safely.
Assistance โ AI suggests, human decides and acts.
The AI provides suggestions, drafts, summaries or classifications. The human reviews, decides and executes. The AI has no connection to business systems โ it cannot read from them or write to them. The human copies information between the AI and the systems manually. This is where most organisations are today.
Partial Automation โ AI drafts or processes, human approves, system executes.
The AI produces output that flows into a business system after human approval. The AI may read from one system (retrieving context, looking up data) but writes only through a human gate. Example: AI drafts a claims assessment; human reviews and approves; system processes the claim. The AI does not act autonomously โ it prepares. The human remains the decision point.
Conditional Automation โ AI acts within defined boundaries; human monitors and handles exceptions.
The AI can read from and write to business systems, but only within pre-defined rules, thresholds and categories. For standard cases that fall within the defined boundaries, the AI processes end-to-end without human intervention. Cases that fall outside the boundaries โ ambiguous, unusual, high-value, high-risk โ are automatically routed to a human specialist. This is the level at which documented integration cases (Allianz, Milton Keynes) operate.
Agentic โ AI plans, executes and coordinates across multiple systems; human governs.
The AI determines the sequence of actions needed to achieve an outcome, calls multiple systems in coordination, and handles variation without pre-defined paths for every case. The human role shifts from reviewing individual actions to governing the agent's boundaries: what it can access, what it can decide, what it can spend, and what must be escalated regardless of the AI's confidence. This level exists in production in narrow domains (code generation, customer service triage) but is not yet documented at scale in multi-system business workflows with independent evidence.
The gap between L1 and L3 is not primarily about AI capability โ it is about integration infrastructure. The AI model can handle L3-level tasks. The organisation's systems, rules, exception-handling processes and measurement infrastructure often cannot. Moving from L1 to L3 is an integration project, not an AI-purchase project.
The layers that turn an AI call into an integrated workflow step.
An AI model call on its own is not integration. Integration is the set of layers that surround the call: what triggers it, what context it receives, what it is asked to do, how its output becomes action, and what happens when the output is wrong or the situation does not fit. Each layer is necessary. Organisations that skip a layer discover the gap later โ usually through a failure.
The most common integration failure pattern: the AI task works (Layer 3) but the context is incomplete (Layer 2), so the action is wrong (Layer 4), and there is no exception path (Layer 5), so the error reaches a customer or a business decision undetected. The AI was fine. The integration failed.
What documented cases tell us about AI integrated into real workflows.
The cases below are not product announcements. They are documented examples of AI integrated into operational business processes, with measurable before-and-after results, from organisations that have described their integration approach publicly. Each case illustrates a different integration pattern โ and a different set of conditions required for it to work.
Allianz โ Project Nemo: AI-assisted claims triage and processing
Allianz deployed seven AI agents within its claims processing workflow under the internal name "Project Nemo." The agents handle claims intake, document classification, data extraction, coverage verification, and routing decisions. Claims are automatically processed for standard cases; unusual or high-value cases are escalated to human specialists.
Result: approximately 80% reduction in processing time per claim. Claims reaching human review do so in under 5 minutes from submission, with relevant context pre-assembled. The AI does not make final decisions on non-standard claims โ it prepares the file for the human specialist.
Allianz Germany โ Automated pet insurance claims processing
Allianz Germany introduced AI-driven automation for simple pet-insurance claims. The system reads submitted invoices, classifies the treatment type, checks coverage against the policy, calculates the payable amount, and either processes the payment or routes to human review. The automation rate โ the proportion of claims handled without human intervention โ reached 49.7%.
Critically, this is not an AI decision rate โ it is an automation rate for a specific, well-defined subset of claims. The remaining ~50% of claims require human involvement because they are ambiguous, unusual or exceed defined thresholds. The system was designed with the assumption that roughly half of claims would need human review, and the routing logic was built accordingly.
Milton Keynes City Council โ AI-assisted planning application processing
Milton Keynes City Council integrated AI into its planning application workflow to reduce the time from application receipt to validation and from validation to decision. The AI assists with document classification, completeness checking, and information extraction from submitted planning documents.
Before-and-after measurements: average receipt-to-validation time reduced from 15.8 days to 7.6 days. Average validation-to-decision time reduced from 53.1 days to 43.2 days. Estimated officer time saved: approximately 1,360 hours per year. These are end-to-end cycle time measurements โ not task-level time savings on a single step.
Bank of America โ EricaAssist: AI-powered employee contact centre support
Bank of America deployed EricaAssist to support over 18,000 contact centre employees. When a customer calls, the AI provides the employee with real-time guidance โ relevant policy information, suggested responses, next-best-action recommendations โ in under 3 seconds. The employee remains in control of the customer conversation; the AI operates as an invisible assistant.
Result: approximately 1 minute reduction in average call handling time. Across 18,000+ employees handling high volumes of calls, this compounds to significant operational capacity. The AI does not interact with the customer. It supports the employee.
IBM โ AskHR: AI-powered employee service platform across multiple back-end systems
IBM's AskHR platform handles over 80 task types โ from benefits questions to payroll inquiries to IT support requests โ processing more than 2.1 million conversations per year. The platform integrates with Workday (HR), SAP (finance), and Concur (travel and expenses) to retrieve context and execute transactions. Employees interact with a single interface; the AI routes queries, retrieves information from the relevant back-end system, and either answers directly or escalates to a human specialist.
The integration architecture is the core of this case: the AI is not one model doing everything. It is a routing and retrieval layer connected to multiple business systems, each with its own data, rules and access controls. The value is not in the AI's conversational ability โ it is in the AI's ability to determine which system has the answer and retrieve it without the employee needing to know the internal system landscape.
Every case has: defined task boundaries, system integration for context retrieval, a human decision point or exception path, and before-and-after measurement. None of the cases claim: 100% automation, removal of human judgment, "autonomous" operation, or AI replacing the workflow rather than being integrated into it. The pattern is consistent: AI handles the structured, repeatable, high-volume part. Humans handle the exceptions, the judgment calls, and the cases that fall outside the defined boundaries. This is not a limitation of the technology โ it is how reliable business automation has always worked.
Every AI-integrated workflow needs defined exception paths โ or every AI error becomes a business error.
The most important design decision in AI integration is not which model to use or how to prompt it. It is deciding what happens when the AI is wrong, uncertain, or facing a case it was not designed for. The documented integration cases all have this in common: they defined the exception path before they deployed the AI, not after the first failure.
Examples from the evidence
Simple pet insurance claim with complete invoice and clear treatment code (Allianz Germany). Standard planning application with all required documents (Milton Keynes). Routine HR query matching a known task type (IBM AskHR).
Examples from the evidence
Complex claim requiring judgment about policy interpretation (Allianz Project Nemo). Unusual planning application with novel legal questions (Milton Keynes). Employee query that does not match any of the 80+ defined task types (IBM AskHR).
The organisations with the most reliable AI-integrated workflows designed the exception path first: what cannot be automated, what must be escalated, who receives the escalation, what context they need, and how quickly they must respond. They did not deploy AI and then discover exceptions through failures. Building the exception path before deployment is the single design decision that most distinguishes documented integration successes from documented integration failures.
What AI agents are, what they are not, and what documented guidance says about making them safe in a business workflow.
"AI agent" is the most used and least defined term in current AI integration discussions. The technical guidance โ from NIST and OWASP โ provides a more precise vocabulary. An agent is not "an AI that does things." It is a system that uses tools, makes plans, and executes actions across multiple steps without human approval at each step. The distinction between what an agent can do, what it is permitted to do, and what it should decide autonomously is the core of safe agent integration.
An agent may have the functionality to query a database, but its permissions should limit it to read-only access on specific tables. Its autonomy should limit it to running pre-approved query patterns โ not constructing arbitrary queries. An agent may have the functionality to send emails, but its autonomy should require human approval for any external recipient. The integration question is not "what can the agent do?" It is "what is the agent permitted to do, and what must a human approve?" These are integration decisions โ not model-capability decisions.
"An AI agent is an AI that acts autonomously to achieve goals."
This definition โ common in marketing โ treats "autonomous" as a binary property. An AI either is or is not an agent. It ignores the dimensions that matter: what the agent can access, what it can decide, and what happens when it is wrong. Under this definition, "deploying an agent" sounds like giving an AI the keys and hoping it drives well.
"An AI agent is a system that uses tools to execute multi-step tasks within defined functionality, permission, and autonomy boundaries."
This definition โ consistent with NIST and OWASP guidance โ treats agent design as a set of integration decisions. What tools? What permissions? What autonomy? What logging? What rollback? "Deploying an agent" means answering these questions for each workflow the agent is part of. The answers will be different for different workflows โ and that is the point.
NIST explicitly distinguishes between tool-using agent systems (which operate within defined constraints) and the unsupported idea of fully autonomous AI operating without human governance. The documented integration cases โ Allianz's seven agents, IBM's 80+ task types โ are all tool-using agent systems with defined boundaries, not autonomous actors. The phrase "autonomous AI agent" is marketing, not a technical category.
How to measure whether AI integration is improving the workflow โ or just adding complexity.
The documented integration cases share a measurement approach: measure the workflow before AI, integrate AI into specific steps, measure the workflow after, and compare across multiple dimensions โ not just one. Measuring only the AI step tells you whether the model works. Measuring the full workflow before and after tells you whether the integration works. These are different questions.
Workflow without AI integration
Workflow with AI integrated at defined steps
The organisations with the most credible integration results โ Allianz, Milton Keynes, Bank of America โ all measured before-and-after across multiple dimensions. They did not report "AI accuracy" as their primary metric. They reported what changed in the workflow: cycle time, throughput, officer hours, call handling time. The AI is a component. The workflow is the unit of measurement.
Is your AI integrated โ or just present?
The distinction between having AI tools and having AI-integrated workflows is measurable. These questions are designed to help you assess which side of that divide your organisation is on. They are based on the patterns observed across the documented integration cases in Section 4.
Do you know exactly which workflow steps AI is being used for โ and which it is not?
- Have you mapped the specific steps in your business processes where AI currently contributes โ or is AI use happening wherever individual workers decide to use it?
- The documented integration cases all start with a defined task taxonomy: 80+ task types (IBM), specific claims categories (Allianz), a single well-defined process (Milton Keynes). AI use that is not mapped to specific workflow steps is AI activity, not AI integration.
- Do you know which steps AI is explicitly not used for โ and why?
- Defining what is outside scope is as important as defining what is inside. The organisations with reliable AI integration can name the tasks AI does not touch โ and the reason (high error cost, regulatory requirement, judgment required).
Is your AI connected to the business systems it needs โ or are humans copying data between them?
- Does the AI retrieve context from business systems automatically โ or does a human copy the relevant information into a prompt?
- Manual context transfer is the hallmark of L1 assistance. It works for individual productivity. It does not scale to workflow integration. The IBM AskHR case works because the AI retrieves context from Workday, SAP and Concur โ not because employees paste their HR records into a chat window.
- Does AI output flow into business systems automatically after defined checks โ or does a human re-enter the AI's output into the system?
- If a human reads the AI output and retypes it into another system, the integration is manual at both ends. The AI is an assistant. The workflow is not integrated.
Do you have a defined, tested process for what happens when the AI is wrong or uncertain?
- Is there a defined set of conditions under which AI output is automatically escalated to a human โ or does escalation depend on the AI deciding it is uncertain?
- Self-assessment of uncertainty by the AI is useful but insufficient. The integration needs rule-based escalation triggers: claim value above threshold, case type not in defined taxonomy, data completeness below minimum. The AI cannot be the sole judge of when it needs human help.
- Do the people who receive escalated cases have the context they need to decide โ or do they have to start from scratch?
- In the Allianz Nemo case, escalated claims arrive with context pre-assembled โ the AI's role in the exception path is to make the handover fast, not to avoid the handover. If escalation means "human redoes the work from the beginning," the AI step added time, not reduced it.
Are the AI's functionality, permissions, and autonomy three separate, documented things?
- Can you state โ for each AI-integrated workflow โ what the AI can technically do (functionality), what it is permitted to access (permissions), and what it can decide without human approval (autonomy)?
- If these three dimensions are not separated and documented per workflow, the integration has not addressed OWASP LLM06 (Excessive Agency). The most common agent failure pattern: the AI was given the permissions of the human it was meant to assist, without the human's judgment about when to use them.
Are you measuring the workflow โ or the AI?
- Do you have before-and-after measurements of the full workflow โ cycle time, throughput, error rate, rework, escalation rate โ not just the AI's accuracy on its specific task?
- Measuring only the AI step tells you whether the model works. Measuring the full workflow before and after tells you whether the integration works. The Milton Keynes case reports receipt-to-decision time โ not "AI document classification accuracy." The latter matters. The former matters more.
- If a workflow step became faster but the overall cycle time did not improve โ would you know?
- The difference between task time and cycle time is one of the three critical distinctions (Section 5). If you are not measuring the full cycle, you do not know whether the AI integration improved the workflow or just moved the bottleneck to a different step.
That is not a criticism. It is the normal state of AI adoption in 2025โ2026. The documented integration cases โ Allianz, Milton Keynes, IBM, Bank of America โ are notable precisely because they have answered these questions. Most organisations have not. The gap between "we use AI" and "AI is integrated into our workflows" is where most AI value claims fail. The purpose of this self-check is to make that gap visible โ so it can be closed, not ignored.
How to read the cases and claims on this page.
The evidence base for AI integration into business workflows is younger and thinner than the evidence base for AI task-level productivity. The five cases in Section 4 are among the best-documented public examples. They are not a comprehensive sample. They are what is available to cite with a described methodology and a named organisation. The notes below explain what each type of evidence can and cannot support.
Documented integration cases are not randomised controlled trials.
The five cases in Section 4 describe real organisations integrating AI into real workflows with measured before-and-after results. They do not have control groups. They cannot rule out the possibility that other changes โ process improvements, staff changes, seasonal variation โ contributed to the measured improvement. They are evidence that integration can produce these results under these conditions. They are not proof that the AI alone caused the results.
Organisations that achieve measurable integration results self-select for publication.
The organisations that publish their integration results are the ones that succeeded. Organisations that attempted AI integration and saw no measurable improvement โ or saw negative effects โ are underrepresented in the public evidence base. The available cases tell you what is possible. They do not tell you how often attempts produce these results, or what the average result looks like.
"Automation rate" is a design metric, not an AI performance metric.
When Allianz Germany reports a 49.7% automation rate for pet insurance claims, that number reflects a design decision about which claims are simple enough to automate โ not the AI's accuracy on a representative sample of all claims. The AI may be 98% accurate on the subset it processes. The 49.7% is the proportion of all claims that fall into that subset. These are different numbers with different meanings. Conflating them produces misleading conclusions about AI capability.
Cycle time measurements are more credible than task-time measurements โ but harder to attribute.
Milton Keynes reports end-to-end cycle time reduction (15.8โ7.6 days receipt-to-validation). This is a stronger measurement than "AI reduced document classification time by X%," because it captures the effect on the workflow the citizen experiences. But it is also harder to attribute: the improvement may reflect process changes made alongside the AI integration. The most credible integration cases report both: what changed at the AI step and what changed across the full cycle โ with the understanding that the former is easier to attribute and the latter is what matters to the business.
Vendor case studies without independent documentation carry limited evidentiary weight.
The cases in Section 4 are drawn from publicly documented integration examples where the organisation has described its approach in some detail. Vendor-produced case studies that report large-sounding numbers without describing the integration architecture, measurement method, or exception handling are not used as primary evidence. A vendor case study that says "Company X achieved 90% automation with our platform" without describing what was automated, how it was measured, what the boundaries were and what happened to exceptions โ tells you nothing you can verify or learn from.
Agentic AI at scale in multi-system business workflows is not yet documented with independent evidence.
The evidence base for L4 agentic integration โ AI that plans, coordinates across systems and handles variation without pre-defined paths โ consists primarily of vendor announcements, prototype descriptions and narrow-domain examples (code generation, customer service triage). There is no publicly available, independently verifiable documentation of a multi-system agentic business workflow operating at scale with before-and-after measurement as of mid-2025. This does not mean such systems do not exist. It means the evidence to evaluate them does. Claims about agentic business integration should be treated as directional, not as established fact.
NIST and OWASP guidance represents technical consensus โ not regulatory requirement.
The NIST and OWASP guidance cited in Section 6 is technical best practice, not law. Following it does not guarantee regulatory compliance. Not following it does not automatically mean non-compliance. The guidance is useful because it names specific risks and specific mitigations โ functionality, permissions, autonomy โ that any organisation integrating AI agents into business workflows should address, regardless of which regulation applies.
Source policy
For important factual claims, this page prioritises:
- Publicly documented integration cases with described methodology and named organisations
- Technical guidance from established standards bodies (NIST, OWASP)
- Peer-reviewed research and publicly available working papers
- Organisation-published case studies with measurable before-and-after data and described integration architecture
- Vendor-neutral analysis from research organisations with transparent methodology
Vendor-produced case studies without described methodology are not used as primary evidence. "Automation rate" figures are contextualised with what the rate measures. Claims about agentic AI are identified as directional when independent verification is not available.
Documented integration cases
Allianz โ Project Nemo: AI agents in claims processing
Publicly documented deployment of seven AI agents within Allianz's claims processing workflow. Reports ~80% reduction in processing time, with claims reaching human review in under 5 minutes. Described integration architecture: AI handles standard cases within defined boundaries; human specialists handle exceptions with pre-assembled context. Used for: L3 conditional automation pattern in Sections 2, 4 and 5.
Allianz Germany โ Automated processing of simple pet insurance claims
Documented deployment achieving 49.7% straight-through automation for a defined subset of pet insurance claims. The remaining ~50% are routed to human review by design. Illustrates the distinction between automation rate and AI capability. Used for: boundary-design pattern and critical distinctions in Sections 4 and 5.
Milton Keynes City Council โ AI-assisted planning application processing
Publicly documented integration of AI into planning application workflow with measured before-and-after end-to-end cycle times. Receipt-to-validation: 15.8โ7.6 days. Validation-to-decision: 53.1โ43.2 days. Estimated officer time saved: ~1,360 hours/year. Used for: end-to-end measurement pattern in Sections 4, 5 and 7.
Bank of America โ EricaAssist: AI-powered contact centre employee support
Documented deployment supporting 18,000+ contact centre employees with real-time guidance in under 3 seconds. Reports approximately 1 minute reduction in average call handling time. AI supports the employee; the employee remains the customer-facing actor. Used for: L1/L2 assistance pattern in Section 4.
IBM โ AskHR: AI-powered employee service platform
Documented deployment handling 80+ task types and 2.1M+ conversations per year, integrated with Workday, SAP and Concur back-end systems. Illustrates multi-system integration with defined task taxonomy. Used for: multi-system integration and task-routing pattern in Sections 2, 4 and 6.
Technical guidance โ agentic AI security & safety
NIST โ Guidance on tool-using and agentic AI systems
Technical guidance on the risks and mitigations for AI systems that use tools, make plans and execute multi-step actions. Defines the core risk dimensions: unintended goal pursuit, excessive access, insufficient logging, autonomy exceeding detection capability. Used for: agent definition and safety framework in Section 6.
OWASP โ LLM06: Excessive Agency (OWASP Top 10 for LLM Applications)
Defines Excessive Agency as a top-10 risk for LLM-integrated applications. Establishes the functionality-permissions-autonomy distinction as the framework for safe agent integration. Used for: agent vulnerability framework and the three-dimension model in Section 6.
Further reading
McKinsey โ "The state of AI in early 2024: Gen AI adoption spikes and starts to generate value"
Survey data on AI adoption patterns and self-reported value generation. Provides context on the gap between AI adoption and measured business impact. Self-reported executive data โ see Research Notes for caveats on survey-based evidence.