Agentic AI in ERP: What an SAP Architect Actually Measured
Most claims about AI agents inside ERP come from a marketing page. This one comes with a failure rate, a latency curve and a cost per execution.
By Zuzanna — Content & Marketing
Przeczytaj po polskuON THIS PAGE
- 01The interesting number is the one that looks bad
- 02The paper argues for a hybrid system, not an autonomous one
- 03What the dispute management agent did well, and where it failed
- 0481.7% is a triage number, not an execution number
- 05The model is the cheapest part of the system
- 06One argument in the paper does not hold up
- 07The question is which process qualifies, not which product to buy
- 08Ask what the process is currently blocked on
- 09Frequently asked questions
The interesting number is the one that looks bad
On 15 October 2025, IEEE Access published a paper by Siar Sarferaz, a chief software architect in SAP's ERP research and development organisation. It sets out a framework for building agentic AI inside an ERP system, derived from an analysis of 26 agentic AI use cases across ERP modules, and demonstrated on a working implementation in SAP's own ERP.
The architecture diagram is not the useful part. Every vendor has one of those now.
The useful part is a comparison buried in the evaluation. The scenario is dispute management: an agent reads billing complaints, checks them against contract terms and proposes credit notes. It classified disputes correctly 81.7% of the time. The conventional approach scored 97.3% in the same controlled tests.
Dispute classification accuracy in SAP’s own controlled tests.
Open the paper on IEEE XploreThat is a vendor architect publishing a number that undercuts the easiest version of the sales pitch. The paper says so plainly: accuracy was “acceptable but could not outperform” the non-agentic approach. It is worth reading the rest of the paper carefully because of it.
The paper argues for a hybrid system, not an autonomous one
The framing is more conservative than the industry conversation around it. ERP processes span a wide range: some are structured and rule-based, others are ambiguous and data-driven. Traditional programming is reliable and explainable for the first kind. Agents handle unstructured input for the second. The conclusion is a hybrid system, not a replacement of deterministic logic with probabilistic logic.
The paper also notes something that rarely survives a product demo: in ERP contexts, a human is typically in the loop for legal compliance reasons.
So the question when scoping a project is not whether agents beat code. It is which decisions genuinely require judgement over unstructured input, and which are just rules nobody has written down yet.
What the dispute management agent did well, and where it failed
The scenario is small and concrete: a subscriber charged €200 a month sees €211.40 on an invoice and emails to complain, while the contract allows a 2% annual increase to €204. The agent has to read the email, decide whether it is a dispute, open a case, pull billing history and contract terms, calculate the €7.40 discrepancy and propose a credit note.
The reported results split cleanly in two.
Where the agent performed well:
- Throughput. A batch of 1,000 historical dispute records was analysed in 0.9 hours, against an estimated 128 human-hours for the equivalent non-agentic approach. The comparison does not include the time needed to review the agent’s output, which, as section 03 argues, is necessary at this accuracy level.
- Cost. Operational expense per resolved dispute fell by 67%.
- Latency. The paper reports sub-1,000ms processing for standard dispute analysis, but in the same sentence says 81% of simple billing cases were resolved within the target threshold, so read the first claim as typical rather than universal. Complex multi-contract disputes averaged about 3,100ms. The paper describes this as “well within” its 3,000ms target for complex interactions; it is in fact slightly above it.
- Scaling. Load testing with 50 to 500 simulated simultaneous cases showed linear scaling, with average response times rising from 847ms to 2,340ms.
Time to analyse 1,000 historical dispute records against the estimated human equivalent. Excludes time to review the agent’s output.
Where it failed, and why:
- 12.3% of dispute management failures came from incomplete email parsing.
- 5.8% came from invalid contract data references.
- 3.2% came from generative model timeouts under peak load.
The paper labels these as shares of failures, but they add up to 21.3%, not 100%, and the text does not explain the gap. Read them as relative weights, not a complete breakdown.
Read those failure modes again. Only the timeouts are a pure infrastructure problem. Incomplete email parsing is partly a model problem, since reading unstructured email is exactly the agent’s job, but it is also a symptom of messy, incomplete input. Invalid contract references are a data quality problem in the ERP itself. Two of the three failure modes sit where the agent meets the data, which is where integration projects have always failed. The paper’s own mitigations point the same way: stricter data validation, fallback execution paths and continuous retraining on error cases.
81.7% is a triage number, not an execution number
An accuracy figure only means something once you attach it to a consequence.
At 81.7%, an agent that drafts a classification for a caseworker to confirm is genuinely useful. Roughly four in five cases arrive pre-analysed, and the human reviews rather than starts from nothing. This is the mode in which the throughput gain is worth having, although the review time has to be added to the 0.9 hours.
At 81.7%, an agent that issues credit notes without review is a liability. Roughly one in five wrong, at scale, in finance.
The paper’s own design accounts for this. The agent can generate a credit note automatically, but only within predefined approval thresholds. For smaller amounts it may be authorised to release the credit note without human intervention. Complex cases, and those requiring judgement beyond the agent’s defined parameters, are escalated to staff.
That is the actual design pattern, and it is not glamorous: the stakes set the autonomy. The more value is at risk, the less the agent decides alone. Cheap, reversible, high-volume decisions go to the agent. Expensive or irreversible ones get a human. The engineering work is in drawing that line precisely and enforcing it, not in the model.
The model is the cheapest part of the system
The paper puts hard numbers on runtime overhead, measured in controlled ERP test environments, and they are small.
- Inference cost between $0.002 and $0.008 per agent execution, depending on prompt complexity and model tier.
- Around 1.2GB of additional memory per 100 concurrent agents.
- A 1–2% CPU increase for orchestration, state management and API handling.
- 50–80KB of network traffic per execution.
If that were the whole cost, every ERP would already be full of agents.
Inference is the cheapest line item in the project. The twenty requirements below are not.
See the twenty requirementsIt is not, and the paper shows why. Before it reaches the architecture, it lists twenty requirements derived from those 26 use cases. A selection:
- Legal compliance. Data protection, consent, read-access logging, and a way for auditors to inspect how an agent decided something.
- Content validation. Checking what goes into the agent and what comes out, including protection against adversarial input.
- Error handling. Logging, anomaly monitoring, and fallback paths for when the agent produces unreliable output.
- Explainability. Traceable logic and data sources behind every recommendation, retained well enough to survive a regulatory review.
- Extensibility. Customer-specific extensions that must survive product upgrades.
- Lifecycle, scalability, metering, localisation, configuration, mass processing.
That list is the project. The agent is a fraction of it. The paper's own limitations section agrees from the other direction, noting that the added runtimes and orchestration layers increase system complexity, introduce latency risk under heavy load, and create administrative burden that turns into technical debt if nobody manages it.
One argument in the paper does not hold up
The paper argues that because all programming languages are Turing-complete, a framework proven in SAP’s ABAP environment can be reproduced on any other ERP platform. For general-purpose languages the formal claim holds; the practical conclusion does not follow from it. Turing-completeness says the target platform can express the logic. It says nothing about whether the equivalent runtime, permission model, metadata layer and audit infrastructure exist there. A few lines later, the paper itself concedes that practical adaptation requires systematic architectural mapping and careful consideration of platform-specific constraints.
Its stated limitations are worth carrying into any project plan. The evaluation primarily covers short-term feasibility, technical correctness and initial deployment. Long-term maintenance and model drift are flagged as open questions, and change management, user acceptance, AI oversight and governance are explicitly not addressed.
The question is which process qualifies, not which product to buy
Nothing here suggests pointing an agent at an ERP and waiting. The scenario above is not unique to one vendor, and whether it works is determined by a property of the process, not the product.
A process is a reasonable candidate when most of these are true:
- The input is genuinely unstructured: email, free-text notes, documents, scanned forms.
- The volume is high enough that human hours are a real cost.
- The correct answer can be verified against structured data the system already holds.
- A wrong answer is caught before it does damage, by a threshold, an approval or a reconciliation.
- There is already an audit trail the agent's decisions can be written into.
It is a poor candidate when the rules are already known and stable, when a wrong answer moves money irreversibly, when data quality is the actual problem, or when nobody can say what "correct" means well enough to measure it. That last one is the common case: most operations that want an agent first need someone to write down the decision the agent would be making.
Ask what the process is currently blocked on
The paper is useful because it is specific enough to argue with: a working architecture, a measured scenario, a cost profile, an honest accuracy gap and a list of what was not tested.
The conclusion is not "add agents to the ERP." It is that the deterministic parts of your operation are still better served by deterministic software, and an agent earns its cost where a process stalls on a human reading something unstructured before any rule can be applied. That is a much smaller target than the category name suggests. It is also a buildable one.
At TailoredByte, that is how we scope this work: identify the specific decision that is blocked, verify that its outcome can be checked against data the business already has, and put deterministic logic around the parts that never needed a model in the first place. It is also what we are building towards: an integration layer that sits across a company’s systems rather than inside any one of them, so an agent can check its own answer against everything the business already knows before it acts. More on that later this year.
Frequently asked questions
Does an ERP agent need access to data outside the ERP?
Often, yes, and the published error analysis is part of the argument. The leading causes of failure were incomplete email parsing and invalid contract data references. An agent that can cross-check its answer against more of the business’s records, inside and outside the ERP, has more ways to catch its own mistakes before they reach a customer. The paper does not test this directly; it is our reading of its error analysis.
Does agentic AI outperform conventional ERP automation?
Not on accuracy, based on the published evidence. In SAP’s dispute management evaluation, the agentic approach classified cases correctly 81.7% of the time against 97.3% for the non-agentic approach (the paper does not describe that baseline in detail). The agent’s advantages were throughput, with 1,000 records analysed in 0.9 hours against an estimated 128 human-hours, and a 67% reduction in operational expense per resolved dispute.
Where do AI agents in ERP actually fail?
In the reported error analysis, the largest share of failures came from incomplete parsing of incoming email, followed by invalid contract data references, with model timeouts a distant third. Two of the three sit where the agent meets its input and data, rather than in the model’s reasoning.
How much does it cost to run AI agents inside an ERP?
Runtime cost is low: roughly $0.002–0.008 per agent execution, with about 1.2GB additional memory per 100 concurrent agents and a 1–2% CPU increase. The real cost sits in compliance, audit logging, explainability, error handling, fallback paths, lifecycle management and upgrade-safe extensibility.
Can an ERP agent make decisions without human approval?
Within limits, and the limits should be tied to consequence. The published framework releases low-value credit notes automatically under predefined thresholds and escalates complex cases to staff. In regulated processes, a human in the loop is generally a compliance requirement rather than a design preference.
Is this framework specific to SAP?
The implementation is. The paper argues the framework is portable because all programming languages are equivalent in expressive power, but portability in practice depends on whether the target platform offers comparable runtime, permission, metadata and audit infrastructure. The paper itself acknowledges that adaptation requires systematic architectural mapping and attention to platform-specific constraints.
Where is your process waiting on a human to read something?
Bring us the queue: the inbox, the scanned forms, the free-text notes that stop a workflow until someone opens them.
In one 30-minute conversation, we will tell you whether an agent belongs there, whether a rule engine solves it more cheaply, or whether the real problem is the data underneath.
Book a scoping call with TailoredByte