AI & Workflow Integration

How does AI move from an isolated tool into a functioning business workflow?

Having an AI tool is not the same as having AI work inside your processes. Integration is the step most organisations underestimate: connecting AI to business systems, defining when a human decides, and measuring the result against the workflow it was meant to improve. This page explains what integration actually requires โ€” from assistance to automation to agentic workflows โ€” and what the evidence says about when it works and when it does not.

INPUT + CONTEXT + AI + BUSINESS SYSTEMS + RULES + HUMAN DECISIONS + MEASUREMENT
This is what AI integration actually is. Not a model call. Not a chatbot window. A connected, governed, measured workflow.
Each component is necessary. Remove one and you have AI activity โ€” not AI-integrated operations. The organisations reporting the strongest integration results are those that treat AI as a workflow component with defined inputs, defined outputs, defined system connections, defined rules, defined human decision points, and defined measurement โ€” not as a stand-alone tool that workers use when they feel like it.
Framework This model is derived from documented integration patterns across the cases in Section 4. It is a synthesis of what successful AI-integrated workflows have in common โ€” not a vendor methodology.
This page is about integration โ€” not tool selection. It does not compare AI products or recommend vendors. It explains what happens between "we have an AI tool" and "AI is part of how we work." The distinction matters because most AI value claims fail at the integration step, not at the model-capability step. The model works. The workflow doesn't.
The Automation Continuum

From assistance to agentic โ€” four levels of AI integration into business workflows.

AI does not enter a workflow at full autonomy. It enters at the assistance level and moves up as the organisation builds the surrounding infrastructure: connections to business systems, rules for when the AI acts, human decision points, and measurement of what the integration produces. Each level changes who does what โ€” and what the organisation needs to have in place for the level to work safely.

SIMPLIFIED EXPLANATION
L1

Assistance โ€” AI suggests, human decides and acts.

The AI provides suggestions, drafts, summaries or classifications. The human reviews, decides and executes. The AI has no connection to business systems โ€” it cannot read from them or write to them. The human copies information between the AI and the systems manually. This is where most organisations are today.

Requires: AI tool access + task definition. No system integration. Human does all execution.
L2

Partial Automation โ€” AI drafts or processes, human approves, system executes.

The AI produces output that flows into a business system after human approval. The AI may read from one system (retrieving context, looking up data) but writes only through a human gate. Example: AI drafts a claims assessment; human reviews and approves; system processes the claim. The AI does not act autonomously โ€” it prepares. The human remains the decision point.

Requires: AI-to-system read access + approval workflow + human-in-the-loop checkpoint.
L3

Conditional Automation โ€” AI acts within defined boundaries; human monitors and handles exceptions.

The AI can read from and write to business systems, but only within pre-defined rules, thresholds and categories. For standard cases that fall within the defined boundaries, the AI processes end-to-end without human intervention. Cases that fall outside the boundaries โ€” ambiguous, unusual, high-value, high-risk โ€” are automatically routed to a human specialist. This is the level at which documented integration cases (Allianz, Milton Keynes) operate.

Requires: system read/write access + rule engine + exception routing + human specialist availability + audit trail.
L4

Agentic โ€” AI plans, executes and coordinates across multiple systems; human governs.

The AI determines the sequence of actions needed to achieve an outcome, calls multiple systems in coordination, and handles variation without pre-defined paths for every case. The human role shifts from reviewing individual actions to governing the agent's boundaries: what it can access, what it can decide, what it can spend, and what must be escalated regardless of the AI's confidence. This level exists in production in narrow domains (code generation, customer service triage) but is not yet documented at scale in multi-system business workflows with independent evidence.

Requires: all L3 infrastructure + agent governance framework + permission boundaries + cost controls + rollback capability.
Most organisations have L1. Some have L2 in specific workflows. A minority have L3 in documented production.

The gap between L1 and L3 is not primarily about AI capability โ€” it is about integration infrastructure. The AI model can handle L3-level tasks. The organisation's systems, rules, exception-handling processes and measurement infrastructure often cannot. Moving from L1 to L3 is an integration project, not an AI-purchase project.

Integration Mechanics

The layers that turn an AI call into an integrated workflow step.

An AI model call on its own is not integration. Integration is the set of layers that surround the call: what triggers it, what context it receives, what it is asked to do, how its output becomes action, and what happens when the output is wrong or the situation does not fit. Each layer is necessary. Organisations that skip a layer discover the gap later โ€” usually through a failure.

1
Trigger
What starts the workflow. A form submission, a system event, a scheduled batch, a customer action, a document arriving. The trigger defines the volume, timing and input format the AI receives.
2
Context
What the AI needs to know. Customer history, policy rules, previous decisions, related cases. Retrieved from business systems before the AI processes. Missing context = wrong output.
3 ยท Core
AI Task
What the AI does. Classify, extract, summarise, recommend, draft, route. Defined as a specific task with a defined output format โ€” not an open-ended conversation.
4
Action
What happens to the AI's output. Written to a system, sent to a queue, presented for approval, logged for audit. The action makes the AI output operational โ€” or stops it before it reaches a system if it does not meet thresholds.
5
Exception
What happens when the AI is wrong, uncertain or outside its boundaries. Defined escalation paths, human specialist assignment, fallback processing. No exception path means every AI error becomes an undetected business error.
Trigger, Context, AI Task, Action, Exception. Five layers. Every documented case of AI integrated into a business workflow has all five โ€” whether they are called by these names or not.

The most common integration failure pattern: the AI task works (Layer 3) but the context is incomplete (Layer 2), so the action is wrong (Layer 4), and there is no exception path (Layer 5), so the error reaches a customer or a business decision undetected. The AI was fine. The integration failed.

Evidence From Practice

What documented cases tell us about AI integrated into real workflows.

The cases below are not product announcements. They are documented examples of AI integrated into operational business processes, with measurable before-and-after results, from organisations that have described their integration approach publicly. Each case illustrates a different integration pattern โ€” and a different set of conditions required for it to work.

Case Study Case 01 โ€” Insurance Claims Processing

Allianz โ€” Project Nemo: AI-assisted claims triage and processing

Allianz deployed seven AI agents within its claims processing workflow under the internal name "Project Nemo." The agents handle claims intake, document classification, data extraction, coverage verification, and routing decisions. Claims are automatically processed for standard cases; unusual or high-value cases are escalated to human specialists.

Result: approximately 80% reduction in processing time per claim. Claims reaching human review do so in under 5 minutes from submission, with relevant context pre-assembled. The AI does not make final decisions on non-standard claims โ€” it prepares the file for the human specialist.

Integration pattern: L3 conditional automation. AI processes standard cases autonomously within defined boundaries. Human specialists handle exceptions. The key infrastructure: rules defining what is "standard," system integration to retrieve policy and customer context, defined exception routing, and measurement of time from submission to decision.
Case Study Case 02 โ€” Simple Claims Automation

Allianz Germany โ€” Automated pet insurance claims processing

Allianz Germany introduced AI-driven automation for simple pet-insurance claims. The system reads submitted invoices, classifies the treatment type, checks coverage against the policy, calculates the payable amount, and either processes the payment or routes to human review. The automation rate โ€” the proportion of claims handled without human intervention โ€” reached 49.7%.

Critically, this is not an AI decision rate โ€” it is an automation rate for a specific, well-defined subset of claims. The remaining ~50% of claims require human involvement because they are ambiguous, unusual or exceed defined thresholds. The system was designed with the assumption that roughly half of claims would need human review, and the routing logic was built accordingly.

Integration pattern: L3 with explicit boundary design. The 49.7% automation rate is a design outcome, not a failure to reach 100%. The workflow was built with the expectation that a significant proportion of cases would route to humans โ€” and the routing mechanism is part of the integration, not an afterthought.
Case Study Case 03 โ€” Public Sector Document Processing

Milton Keynes City Council โ€” AI-assisted planning application processing

Milton Keynes City Council integrated AI into its planning application workflow to reduce the time from application receipt to validation and from validation to decision. The AI assists with document classification, completeness checking, and information extraction from submitted planning documents.

Before-and-after measurements: average receipt-to-validation time reduced from 15.8 days to 7.6 days. Average validation-to-decision time reduced from 53.1 days to 43.2 days. Estimated officer time saved: approximately 1,360 hours per year. These are end-to-end cycle time measurements โ€” not task-level time savings on a single step.

Integration pattern: L2 partial automation with measured end-to-end improvement. The AI assists with specific steps (classification, completeness, extraction) but planning officers remain the decision-makers. The measurements cover the full cycle โ€” receipt to decision โ€” not just the AI-assisted steps. This is how integration measurement should work.
Case Study Case 04 โ€” Employee Support & Guidance

Bank of America โ€” EricaAssist: AI-powered employee contact centre support

Bank of America deployed EricaAssist to support over 18,000 contact centre employees. When a customer calls, the AI provides the employee with real-time guidance โ€” relevant policy information, suggested responses, next-best-action recommendations โ€” in under 3 seconds. The employee remains in control of the customer conversation; the AI operates as an invisible assistant.

Result: approximately 1 minute reduction in average call handling time. Across 18,000+ employees handling high volumes of calls, this compounds to significant operational capacity. The AI does not interact with the customer. It supports the employee.

Integration pattern: L1/L2 assistance with real-time context retrieval. This is not automation โ€” it is AI-assisted human work. The AI retrieves information faster than a human can navigate internal systems, but the human remains the decision-maker and the customer-facing actor. The value comes from reducing information-retrieval time, not from replacing the human decision.
Case Study Case 05 โ€” Enterprise Service Integration

IBM โ€” AskHR: AI-powered employee service platform across multiple back-end systems

IBM's AskHR platform handles over 80 task types โ€” from benefits questions to payroll inquiries to IT support requests โ€” processing more than 2.1 million conversations per year. The platform integrates with Workday (HR), SAP (finance), and Concur (travel and expenses) to retrieve context and execute transactions. Employees interact with a single interface; the AI routes queries, retrieves information from the relevant back-end system, and either answers directly or escalates to a human specialist.

The integration architecture is the core of this case: the AI is not one model doing everything. It is a routing and retrieval layer connected to multiple business systems, each with its own data, rules and access controls. The value is not in the AI's conversational ability โ€” it is in the AI's ability to determine which system has the answer and retrieve it without the employee needing to know the internal system landscape.

Integration pattern: Multi-system L2/L3 with defined task taxonomy. The 80+ task types are pre-defined โ€” the AI classifies the query into a known task type and follows the defined integration path for that type. This is not an open-ended agent. It is a structured routing system with AI at the classification and retrieval layer.
What these five cases have in common โ€” and what they do not claim.

Every case has: defined task boundaries, system integration for context retrieval, a human decision point or exception path, and before-and-after measurement. None of the cases claim: 100% automation, removal of human judgment, "autonomous" operation, or AI replacing the workflow rather than being integrated into it. The pattern is consistent: AI handles the structured, repeatable, high-volume part. Humans handle the exceptions, the judgment calls, and the cases that fall outside the defined boundaries. This is not a limitation of the technology โ€” it is how reliable business automation has always worked.

The Human in the Loop

Every AI-integrated workflow needs defined exception paths โ€” or every AI error becomes a business error.

The most important design decision in AI integration is not which model to use or how to prompt it. It is deciding what happens when the AI is wrong, uncertain, or facing a case it was not designed for. The documented integration cases all have this in common: they defined the exception path before they deployed the AI, not after the first failure.

Standard Path
Within defined boundaries โ†’ AI processes
The case matches known patterns. Data is complete. Confidence is above threshold. The AI processes the workflow step and the output flows to the next stage โ€” either to a business system (if the action is low-risk and predefined) or to a human for final approval.

Examples from the evidence

Simple pet insurance claim with complete invoice and clear treatment code (Allianz Germany). Standard planning application with all required documents (Milton Keynes). Routine HR query matching a known task type (IBM AskHR).

Exception Path
Outside boundaries โ†’ human specialist
The case is ambiguous, incomplete, high-value, or outside the AI's defined task boundaries. The AI routes the case to a human specialist with the context it has assembled โ€” it does not guess. The human specialist decides. The AI's role is to make the handover fast and information-rich, not to handle the case anyway.

Examples from the evidence

Complex claim requiring judgment about policy interpretation (Allianz Project Nemo). Unusual planning application with novel legal questions (Milton Keynes). Employee query that does not match any of the 80+ defined task types (IBM AskHR).

โ‰ 
Automation Rate โ‰  Success Rate
A 49.7% automation rate (Allianz pet) means the AI handled roughly half of cases without human intervention. It does not mean the AI was "successful" on 49.7% and "failed" on 50.3%. The 50.3% were designed to go to humans. The metric is workflow design, not AI accuracy.
โ‰ 
Task Time โ‰  End-to-End Cycle Time
Reducing one step from 15 minutes to 1 minute is a task-level improvement. It only becomes a cycle-time improvement if the other steps in the workflow โ€” human reviews, system handoffs, waiting periods โ€” do not absorb the saved time. Measure the full cycle, not just the AI step.
โ‰ 
Time Saved โ‰  Automatic ROI
1,360 officer hours saved per year (Milton Keynes) is a measurable capacity release. It becomes ROI only if the released capacity is used for higher-value work โ€” not absorbed by meetings, email, or expanded task scope. Capacity release is the measurement. Value conversion is a separate management decision.
The exception path is not a fallback โ€” it is part of the design.

The organisations with the most reliable AI-integrated workflows designed the exception path first: what cannot be automated, what must be escalated, who receives the escalation, what context they need, and how quickly they must respond. They did not deploy AI and then discover exceptions through failures. Building the exception path before deployment is the single design decision that most distinguishes documented integration successes from documented integration failures.

Agentic AI in Business

What AI agents are, what they are not, and what documented guidance says about making them safe in a business workflow.

"AI agent" is the most used and least defined term in current AI integration discussions. The technical guidance โ€” from NIST and OWASP โ€” provides a more precise vocabulary. An agent is not "an AI that does things." It is a system that uses tools, makes plans, and executes actions across multiple steps without human approval at each step. The distinction between what an agent can do, what it is permitted to do, and what it should decide autonomously is the core of safe agent integration.

NIST
Tool-using agent systems โ€” definition and risks
NIST defines agentic AI systems as those that can use tools, make plans, and execute multi-step actions. The core risks are: (1) the agent pursues a goal in an unintended way, (2) the agent has access to systems it does not need, (3) the agent's actions are not logged or reversible, and (4) the agent's autonomy exceeds the organisation's ability to detect and respond to errors. NIST's guidance focuses on constraining what the agent can access, log, decide and spend โ€” not on making the agent "smarter."
OWASP
LLM06: Excessive Agency โ€” the core agent vulnerability
OWASP's LLM06 identifies Excessive Agency as a top-10 risk for LLM-integrated applications: an AI system is given more functionality, more permissions or more autonomy than it needs for its defined task. The distinction between functionality (what the agent can technically do), permissions (what the agent is allowed to access) and autonomy (what the agent can decide without human approval) is the foundation of safe agent design. Most agent failures occur because these three dimensions were not separated.
Functionality. Permissions. Autonomy. Three separate dimensions. Confusing them is the most common agent-integration error.

An agent may have the functionality to query a database, but its permissions should limit it to read-only access on specific tables. Its autonomy should limit it to running pre-approved query patterns โ€” not constructing arbitrary queries. An agent may have the functionality to send emails, but its autonomy should require human approval for any external recipient. The integration question is not "what can the agent do?" It is "what is the agent permitted to do, and what must a human approve?" These are integration decisions โ€” not model-capability decisions.

Misunderstanding

"An AI agent is an AI that acts autonomously to achieve goals."

This definition โ€” common in marketing โ€” treats "autonomous" as a binary property. An AI either is or is not an agent. It ignores the dimensions that matter: what the agent can access, what it can decide, and what happens when it is wrong. Under this definition, "deploying an agent" sounds like giving an AI the keys and hoping it drives well.

What the guidance supports

"An AI agent is a system that uses tools to execute multi-step tasks within defined functionality, permission, and autonomy boundaries."

This definition โ€” consistent with NIST and OWASP guidance โ€” treats agent design as a set of integration decisions. What tools? What permissions? What autonomy? What logging? What rollback? "Deploying an agent" means answering these questions for each workflow the agent is part of. The answers will be different for different workflows โ€” and that is the point.

Agentic AI is not "autonomous AI."

NIST explicitly distinguishes between tool-using agent systems (which operate within defined constraints) and the unsupported idea of fully autonomous AI operating without human governance. The documented integration cases โ€” Allianz's seven agents, IBM's 80+ task types โ€” are all tool-using agent systems with defined boundaries, not autonomous actors. The phrase "autonomous AI agent" is marketing, not a technical category.

Measurement Model

How to measure whether AI integration is improving the workflow โ€” or just adding complexity.

The documented integration cases share a measurement approach: measure the workflow before AI, integrate AI into specific steps, measure the workflow after, and compare across multiple dimensions โ€” not just one. Measuring only the AI step tells you whether the model works. Measuring the full workflow before and after tells you whether the integration works. These are different questions.

Before

Workflow without AI integration

โ†’
After

Workflow with AI integrated at defined steps

Time
End-to-end cycle time from trigger to outcome. Not task-step time. Milton Keynes: 15.8โ†’7.6 days receipt-to-validation.
Throughput
Volume of cases processed per unit of time. Allianz Nemo: 80% reduction in processing time per claim = more claims per hour.
Quality
Error rate, completeness, consistency of output. Not self-reported. Measured by audit of AI-processed vs. human-processed cases.
Rework
Cases returned for correction after AI processing. High rework means the AI step is creating work downstream โ€” undetected if only measuring the AI step.
Escalation
Proportion of cases routed to humans. A rising escalation rate may mean the AI is receiving cases outside its boundaries โ€” a trigger to adjust routing.
Failure
Cases where the AI output was used but was wrong, and the error reached a customer or a business decision. This is the metric that matters most.
Cost
Total cost of the workflow: AI cost + human review cost + exception handling cost + infrastructure cost. Not just AI API cost.
Business Outcome
Did the business result improve? Customer satisfaction, decision speed, regulatory compliance, competitive position. The reason the workflow exists.
Measure the workflow, not the AI. Measure before and after, not after only. Measure across dimensions, not one metric.

The organisations with the most credible integration results โ€” Allianz, Milton Keynes, Bank of America โ€” all measured before-and-after across multiple dimensions. They did not report "AI accuracy" as their primary metric. They reported what changed in the workflow: cycle time, throughput, officer hours, call handling time. The AI is a component. The workflow is the unit of measurement.

Self Check

Is your AI integrated โ€” or just present?

The distinction between having AI tools and having AI-integrated workflows is measurable. These questions are designed to help you assess which side of that divide your organisation is on. They are based on the patterns observed across the documented integration cases in Section 4.

01 โ€” Task Definition

Do you know exactly which workflow steps AI is being used for โ€” and which it is not?

Have you mapped the specific steps in your business processes where AI currently contributes โ€” or is AI use happening wherever individual workers decide to use it?
The documented integration cases all start with a defined task taxonomy: 80+ task types (IBM), specific claims categories (Allianz), a single well-defined process (Milton Keynes). AI use that is not mapped to specific workflow steps is AI activity, not AI integration.
Do you know which steps AI is explicitly not used for โ€” and why?
Defining what is outside scope is as important as defining what is inside. The organisations with reliable AI integration can name the tasks AI does not touch โ€” and the reason (high error cost, regulatory requirement, judgment required).
02 โ€” System Connection

Is your AI connected to the business systems it needs โ€” or are humans copying data between them?

Does the AI retrieve context from business systems automatically โ€” or does a human copy the relevant information into a prompt?
Manual context transfer is the hallmark of L1 assistance. It works for individual productivity. It does not scale to workflow integration. The IBM AskHR case works because the AI retrieves context from Workday, SAP and Concur โ€” not because employees paste their HR records into a chat window.
Does AI output flow into business systems automatically after defined checks โ€” or does a human re-enter the AI's output into the system?
If a human reads the AI output and retypes it into another system, the integration is manual at both ends. The AI is an assistant. The workflow is not integrated.
03 โ€” Exception Handling

Do you have a defined, tested process for what happens when the AI is wrong or uncertain?

Is there a defined set of conditions under which AI output is automatically escalated to a human โ€” or does escalation depend on the AI deciding it is uncertain?
Self-assessment of uncertainty by the AI is useful but insufficient. The integration needs rule-based escalation triggers: claim value above threshold, case type not in defined taxonomy, data completeness below minimum. The AI cannot be the sole judge of when it needs human help.
Do the people who receive escalated cases have the context they need to decide โ€” or do they have to start from scratch?
In the Allianz Nemo case, escalated claims arrive with context pre-assembled โ€” the AI's role in the exception path is to make the handover fast, not to avoid the handover. If escalation means "human redoes the work from the beginning," the AI step added time, not reduced it.
04 โ€” Boundaries & Permissions

Are the AI's functionality, permissions, and autonomy three separate, documented things?

Can you state โ€” for each AI-integrated workflow โ€” what the AI can technically do (functionality), what it is permitted to access (permissions), and what it can decide without human approval (autonomy)?
If these three dimensions are not separated and documented per workflow, the integration has not addressed OWASP LLM06 (Excessive Agency). The most common agent failure pattern: the AI was given the permissions of the human it was meant to assist, without the human's judgment about when to use them.
05 โ€” Measurement

Are you measuring the workflow โ€” or the AI?

Do you have before-and-after measurements of the full workflow โ€” cycle time, throughput, error rate, rework, escalation rate โ€” not just the AI's accuracy on its specific task?
Measuring only the AI step tells you whether the model works. Measuring the full workflow before and after tells you whether the integration works. The Milton Keynes case reports receipt-to-decision time โ€” not "AI document classification accuracy." The latter matters. The former matters more.
If a workflow step became faster but the overall cycle time did not improve โ€” would you know?
The difference between task time and cycle time is one of the three critical distinctions (Section 5). If you are not measuring the full cycle, you do not know whether the AI integration improved the workflow or just moved the bottleneck to a different step.
If most of these questions are difficult to answer, AI is present in your organisation โ€” but it is not integrated.

That is not a criticism. It is the normal state of AI adoption in 2025โ€“2026. The documented integration cases โ€” Allianz, Milton Keynes, IBM, Bank of America โ€” are notable precisely because they have answered these questions. Most organisations have not. The gap between "we use AI" and "AI is integrated into our workflows" is where most AI value claims fail. The purpose of this self-check is to make that gap visible โ€” so it can be closed, not ignored.

How to Read the Evidence

How to read the cases and claims on this page.

The evidence base for AI integration into business workflows is younger and thinner than the evidence base for AI task-level productivity. The five cases in Section 4 are among the best-documented public examples. They are not a comprehensive sample. They are what is available to cite with a described methodology and a named organisation. The notes below explain what each type of evidence can and cannot support.

Documented integration cases are not randomised controlled trials.

The five cases in Section 4 describe real organisations integrating AI into real workflows with measured before-and-after results. They do not have control groups. They cannot rule out the possibility that other changes โ€” process improvements, staff changes, seasonal variation โ€” contributed to the measured improvement. They are evidence that integration can produce these results under these conditions. They are not proof that the AI alone caused the results.

Organisations that achieve measurable integration results self-select for publication.

The organisations that publish their integration results are the ones that succeeded. Organisations that attempted AI integration and saw no measurable improvement โ€” or saw negative effects โ€” are underrepresented in the public evidence base. The available cases tell you what is possible. They do not tell you how often attempts produce these results, or what the average result looks like.

"Automation rate" is a design metric, not an AI performance metric.

When Allianz Germany reports a 49.7% automation rate for pet insurance claims, that number reflects a design decision about which claims are simple enough to automate โ€” not the AI's accuracy on a representative sample of all claims. The AI may be 98% accurate on the subset it processes. The 49.7% is the proportion of all claims that fall into that subset. These are different numbers with different meanings. Conflating them produces misleading conclusions about AI capability.

Cycle time measurements are more credible than task-time measurements โ€” but harder to attribute.

Milton Keynes reports end-to-end cycle time reduction (15.8โ†’7.6 days receipt-to-validation). This is a stronger measurement than "AI reduced document classification time by X%," because it captures the effect on the workflow the citizen experiences. But it is also harder to attribute: the improvement may reflect process changes made alongside the AI integration. The most credible integration cases report both: what changed at the AI step and what changed across the full cycle โ€” with the understanding that the former is easier to attribute and the latter is what matters to the business.

Vendor case studies without independent documentation carry limited evidentiary weight.

The cases in Section 4 are drawn from publicly documented integration examples where the organisation has described its approach in some detail. Vendor-produced case studies that report large-sounding numbers without describing the integration architecture, measurement method, or exception handling are not used as primary evidence. A vendor case study that says "Company X achieved 90% automation with our platform" without describing what was automated, how it was measured, what the boundaries were and what happened to exceptions โ€” tells you nothing you can verify or learn from.

Agentic AI at scale in multi-system business workflows is not yet documented with independent evidence.

The evidence base for L4 agentic integration โ€” AI that plans, coordinates across systems and handles variation without pre-defined paths โ€” consists primarily of vendor announcements, prototype descriptions and narrow-domain examples (code generation, customer service triage). There is no publicly available, independently verifiable documentation of a multi-system agentic business workflow operating at scale with before-and-after measurement as of mid-2025. This does not mean such systems do not exist. It means the evidence to evaluate them does. Claims about agentic business integration should be treated as directional, not as established fact.

NIST and OWASP guidance represents technical consensus โ€” not regulatory requirement.

The NIST and OWASP guidance cited in Section 6 is technical best practice, not law. Following it does not guarantee regulatory compliance. Not following it does not automatically mean non-compliance. The guidance is useful because it names specific risks and specific mitigations โ€” functionality, permissions, autonomy โ€” that any organisation integrating AI agents into business workflows should address, regardless of which regulation applies.

Sources & Further Reading

Source policy

For important factual claims, this page prioritises:

  1. Publicly documented integration cases with described methodology and named organisations
  2. Technical guidance from established standards bodies (NIST, OWASP)
  3. Peer-reviewed research and publicly available working papers
  4. Organisation-published case studies with measurable before-and-after data and described integration architecture
  5. Vendor-neutral analysis from research organisations with transparent methodology

Vendor-produced case studies without described methodology are not used as primary evidence. "Automation rate" figures are contextualised with what the rate measures. Claims about agentic AI are identified as directional when independent verification is not available.

Documented integration cases

CASE STUDY ยท INSURANCE ยท 2024โ€“2025

Allianz โ€” Project Nemo: AI agents in claims processing

Publicly documented deployment of seven AI agents within Allianz's claims processing workflow. Reports ~80% reduction in processing time, with claims reaching human review in under 5 minutes. Described integration architecture: AI handles standard cases within defined boundaries; human specialists handle exceptions with pre-assembled context. Used for: L3 conditional automation pattern in Sections 2, 4 and 5.

CASE STUDY ยท INSURANCE ยท 2024

Allianz Germany โ€” Automated processing of simple pet insurance claims

Documented deployment achieving 49.7% straight-through automation for a defined subset of pet insurance claims. The remaining ~50% are routed to human review by design. Illustrates the distinction between automation rate and AI capability. Used for: boundary-design pattern and critical distinctions in Sections 4 and 5.

CASE STUDY ยท PUBLIC SECTOR ยท UK ยท 2024โ€“2025

Milton Keynes City Council โ€” AI-assisted planning application processing

Publicly documented integration of AI into planning application workflow with measured before-and-after end-to-end cycle times. Receipt-to-validation: 15.8โ†’7.6 days. Validation-to-decision: 53.1โ†’43.2 days. Estimated officer time saved: ~1,360 hours/year. Used for: end-to-end measurement pattern in Sections 4, 5 and 7.

CASE STUDY ยท FINANCIAL SERVICES ยท US ยท 2024โ€“2025

Bank of America โ€” EricaAssist: AI-powered contact centre employee support

Documented deployment supporting 18,000+ contact centre employees with real-time guidance in under 3 seconds. Reports approximately 1 minute reduction in average call handling time. AI supports the employee; the employee remains the customer-facing actor. Used for: L1/L2 assistance pattern in Section 4.

CASE STUDY ยท TECHNOLOGY ยท 2024โ€“2025

IBM โ€” AskHR: AI-powered employee service platform

Documented deployment handling 80+ task types and 2.1M+ conversations per year, integrated with Workday, SAP and Concur back-end systems. Illustrates multi-system integration with defined task taxonomy. Used for: multi-system integration and task-routing pattern in Sections 2, 4 and 6.

Technical guidance โ€” agentic AI security & safety

TECHNICAL GUIDANCE ยท 2024โ€“2025

NIST โ€” Guidance on tool-using and agentic AI systems

Technical guidance on the risks and mitigations for AI systems that use tools, make plans and execute multi-step actions. Defines the core risk dimensions: unintended goal pursuit, excessive access, insufficient logging, autonomy exceeding detection capability. Used for: agent definition and safety framework in Section 6.

TECHNICAL GUIDANCE ยท 2025

OWASP โ€” LLM06: Excessive Agency (OWASP Top 10 for LLM Applications)

Defines Excessive Agency as a top-10 risk for LLM-integrated applications. Establishes the functionality-permissions-autonomy distinction as the framework for safe agent integration. Used for: agent vulnerability framework and the three-dimension model in Section 6.

Further reading

FRAMEWORK ยท 2024

McKinsey โ€” "The state of AI in early 2024: Gen AI adoption spikes and starts to generate value"

Survey data on AI adoption patterns and self-reported value generation. Provides context on the gap between AI adoption and measured business impact. Self-reported executive data โ€” see Research Notes for caveats on survey-based evidence.