The Most Common Challenges in AI and Automation Implementation

Martijn Zuiderbaan

Author

Last Modified: September 15, 2026
AI Implementation Challenges -feature img

The decision took about ninety seconds. A delivery team was scoping an automation over an order-entry workflow, and someone asked whether they needed to profile the source data before building. The answer was no, because the field list was already in the requirements document and the requirements document had been signed off. Everyone moved on.

The requirements document had been written from the customer relationship management screen. The screen showed a clean, single-value country field. The table behind it held three years of free-text entries, because a form validation rule had been added in the second year and nobody had gone back to fix what came before. The automation was built, tested against recent records, piloted against recent records, and went live. It failed on a fifth of the historic backlog, which is where the volume was.

Nothing about that failure was unknowable in week one. It was findable with about a day of profiling. It got found in week nine instead, by which time the fix was not a day of profiling but a rebuild of the parsing logic, a re-run of everything already processed, and a conversation with a sponsor who had stopped asking about progress and started asking about the plan.

That gap is the actual subject of this article. The challenges below are the ones that come up on nearly every AI and automation programme, and most lists cover them. What most lists leave out is that each challenge has a stage where it is cheap to find and a stage where it is expensive, and the difference is not a rounding error. So rather than walking the challenges in order, this piece sorts them by where they are findable, prices the ladder, and then says what you can buy that moves a problem down it.

First, a working definition

Intelligent automation is not a product. It is a workflow, a set of connections between systems, and artificial intelligence applied at the points where the inputs are not predictable. Robotic process automation, usually shortened to RPA, is the part that drives existing screens and applications when there is no application programming interface, or API, available to call instead. AI is the part that copes with email text, documents, and free-form requests. If you want the longer version of that distinction, we have written it up separately.

Definitions matter here only because they set the boundary of what can go wrong. A workflow can be wrong. A connection can be wrong. A model can be wrong. And all three can be wrong in ways that a demonstration will not show you.

Three Labelled Blocks

Five places a problem can be found, and what each one costs

Take one material defect. Not a typo, not a cosmetic issue, but something that means the automation produces a wrong or unusable result on a real class of input. Here is what clearing that single defect costs at each of five detection points, at a blended delivery rate of EUR 85 per hour.

Found in discovery or design, before anything is built:

about 2 hours, so EUR 170. You change a diagram and a decision.

Found during build:

about 6 hours, EUR 510. You change code that exists but has no dependents yet.

Found in user acceptance testing:

about 14 hours, EUR 1,190. You change code, re-run the test cycle, and update the test evidence.

Found in a pilot:

about 30 hours, EUR 2,550. Add a regression pass, a change record, and a re-brief for the pilot group.

Found live, at full volume:

about 70 hours, EUR 5,950. Add impact analysis, an emergency change window, and a fix that has to work first time.

The live figure carries a second cost that the earlier ones do not. By the time a full-volume defect is visible, work has already been processed wrongly. Assume 40 items at 12 minutes each to identify and correct, at a business-side rate of EUR 40 per hour. That is 8 hours and EUR 320, which takes the true cost of one late defect to EUR 6,270, against EUR 170 for the same defect caught in discovery.

There is a harder version of the live case. On one engagement for a Nordic healthcare provider, the workflow could not be taken offline to be corrected, because the service it fed had no acceptable downtime window. The correction had to be designed, built, and deployed while the automation kept running, with a rollback path proven before anything was applied. Nothing about the underlying defect was unusual. The constraint on when it could be fixed is what made it expensive.

A Horizontal Bar Chart

A worked model: the same twelve defects, found in two different places

This model is illustrative. The rates and hours are plausible mid-market figures, not measurements from a specific client, and the point of publishing the arithmetic is that you can substitute your own numbers and see whether the conclusion survives. It is built to make one comparison only, and it deliberately holds the thing most business cases quietly change.

The programme has twelve material defects. Both versions have all twelve. No readiness activity removes a defect from existence. It only moves the rung at which the defect is found. That constraint is what makes this different from a savings case, because there is no productivity assumption anywhere in it.

Where the twelve surfaced

Version one, thin discovery and no pilot. Two workshops, the requirements document, straight into build, one round of user acceptance testing on recent data, go live.

  • Discovery: 1 defect
  • Build: 2 defects
  • User acceptance testing: 3 defects
  • Pilot: 2 defects
  • Live at full volume: 4 defects

Version two, readiness first. Process discovery run against the actual work rather than the described work, a data profiling pass over the full history rather than the recent slice, and a pilot restricted to one variant of the workflow so that failures are legible.

  • Discovery: 5 defects
  • Build: 3 defects
  • User acceptance testing: 2 defects
  • Pilot: 1 defect
  • Live at full volume: 1 defect

What each version cost

1

Version one.

Rework totals 396 engineering hours, so EUR 33,660, plus four live defects at EUR 320 of operational correction each, so EUR 1,280. Total EUR 34,940, or EUR 2,911.67 per defect.

2

Version two.

The readiness work is 124 hours, being 60 for process discovery, 24 for data profiling, and 40 of pilot overhead, so EUR 10,540. Rework falls to 156 hours, so EUR 13,260. One live defect adds EUR 320. Total EUR 24,120, or EUR 2,010 per defect.

The difference is EUR 10,820 avoided, 30.97 per cent lower, on an identical defect count. Expressed in engineering time, 240 hours of rework did not happen, bought with 124 hours of readiness.

Three things that have to be true

A model that only produces a favourable number is not telling you anything. Here is what has to hold for this one, stated as thresholds rather than as reassurance.

The readiness work has to come in under EUR 21,360. That is the point at which version two stops being cheaper. It is 251 hours, which is 2.03 times the 124 hours assumed. So the case has real headroom, and it also tells you the shape of the failure: a discovery phase that runs to six weeks and 250 hours has spent its own benefit.

At least 49.34 per cent of the intended shift has to actually happen. If readiness moves the defect distribution only part of the way from version one to version two, the benefit scales with how far it moves. Below roughly half, the readiness spend is not recovered. This is the single most useful threshold in the model, because it is the one that fails silently.

Cost has to escalate steeply between build and live. Model the escalation as a slope, where slope 1.0 is the ladder above and a lower slope means a flatter, more forgiving programme. The two versions break even at a slope of 0.432, at which a live defect costs 33.63 hours instead of 70, only 5.61 times a build-stage defect instead of 11.67 times. Running the comparison across the slope:

  • Slope 1.00, live defect 70.00 hours: version one EUR 34,940, version two EUR 24,120, EUR 10,820 avoided, 30.97 per cent lower.
  • Slope 0.80, live defect 57.20 hours: EUR 29,364 against EUR 22,352, EUR 7,012 avoided, 23.88 per cent.
  • Slope 0.60, live defect 44.40 hours: EUR 23,788 against EUR 20,584, EUR 3,204 avoided, 13.47 per cent.
  • Slope 0.432, live defect 33.63 hours: both EUR 19,096.43. No difference.
  • Slope 0.30, live defect 25.20 hours: EUR 15,424 against EUR 17,932. Readiness costs EUR 2,508 more than it saves.

That last line is the honest one. If your automations are genuinely cheap to fix in production, if there is no compliance exposure, no work to re-process, no change window, and no sponsor watching, then front-loading discovery is waste and you should not do it. Most of the programmes we are called into are not that. The test is whether you can name your own slope, and the way to find out is to price the last three defects you fixed after go-live and compare them to three you fixed during build.

What the model does not claim

It does not say readiness prevents defects, because it holds the count fixed at twelve on both sides. It does not price the revenue or reputational consequence of the four live defects in version one, which on a customer-facing workflow would dominate everything above. It does not include licence, infrastructure, or run cost, because those are the same in both versions. And it assumes the twelve defects are equally severe, which no real programme’s are.

Two Stacked Bars

The eight challenges, sorted by where they are findable

The challenges themselves are not controversial. What follows is each one placed at the rung where it can be caught, with the specific check that catches it there.

Findable in discovery, at roughly EUR 170 each

Automating before the workflow is understood. The usual version of this is that the team automates the process as described in a document rather than as performed at a desk. The described version has no exceptions in it, because nobody describes their exceptions. The check is to sit with the work for a day and count how many of the items handled in that day would have completed cleanly under the documented path. If the answer is under 80 per cent, the document is not the process.

People and change management. This gets treated as a communications task scheduled for the week before go-live, which is why it stalls things. It is findable in discovery because the question is answerable then: who currently owns this work, what happens to their day when the automation takes part of it, and who decides the exception cases afterwards. If those three answers do not exist as names, you have a defect, and it is currently costing EUR 170.

Measurement design. Deciding what the programme will be judged on belongs at the start, not at the first review. We have argued the case for outcome metrics over activity counts at length elsewhere outcome metrics over activity counts, and the reason it belongs in discovery is that some outcome metrics require a baseline you can only capture before you change anything.

Findable in data profiling, at roughly EUR 170 to EUR 510 each

Data quality and access. This breaks more automations than any other single cause, and it is the cheapest thing on this list to test. Profile the full history, not the recent slice: distinct values per field, null rate, format variance, and the date on which each pattern changed. The order-entry example at the top of this article is exactly this defect, found at rung five for around EUR 6,270 when a day of profiling at rung one would have cost EUR 170. This is the check that pays for itself most reliably, and it is the one most often cut for time.

The related access question is separate and equally cheap to answer early: can the automation read what it needs to read, through a supported route, with credentials that will not expire mid-quarter. When the only route is a screen rather than an API, that is a decision about brittleness rather than a blocker, and it is one reason RPA remains the right tool in some estates RPA remains the right tool in some estates.

Findable in a one-variant pilot, at roughly EUR 2,550 each

Solutions built outside the real workflow. A tool the team has to leave their working environment to use will be used until the first busy week. Whether that is true is not answerable in a demonstration, because a demonstration removes time pressure. It is answerable in a narrow pilot, which is why the pilot in version two of the model is restricted to one variant. A pilot spread across four variants tells you that adoption is mixed and nothing about why.

Adoption of AI-assisted steps specifically. Where AI drafts or classifies, the failure mode is not rejection but silent over-acceptance, and the pilot is where you can still see it. Instrument the pilot to record how often a suggestion was edited before use, not just how often it was used. AI workflows that take guided steps inside systems need those boundaries defined before the pilot starts, not discovered during it boundaries defined before the pilot starts.

Not findable before go-live, so they have to be designed in

These four cannot be caught earlier, which is the reason they belong in a different category rather than further down the same list. For these, the only available move is to build the mechanism before you need it.

Governance and compliance. Treated as paperwork, this arrives as a delay. Treated as product design, it is a set of build decisions: what gets logged, what is retained and for how long, which decisions require a recorded human approval, and who can change the rules. The regulatory position on this has moved recently and is covered in the next section.

Security and privacy. The specific thing to build is monitoring, and the UK National Cyber Security Centre is direct about what that means for AI systems. Its tell providers to measure outputs and performance “such that you can observe sudden and gradual changes in behaviour affecting security”, covering both intrusion and “natural data drift”, and separately to monitor and log inputs such as inference requests, queries or prompts, “to enable compliance obligations, audit, investigation and remediation in the case of compromise or misuse”. The same section says update processes must reflect that changes to data, models or prompts can change system behaviour, so major updates get treated like new versions. None of that is retrofittable cheaply, which is why it is here rather than in the discovery tier.

The run plan. Scaling breaks when there is no named owner, no monitoring, no exception queue with a service level, and no change process after go-live. This is the challenge that produces the most expensive version of a rung-five defect, because without monitoring the defect is not found by you, it is found by a customer or an auditor. If your existing estate has reached this state, the symptoms are recognisable from the outside the symptoms are recognisable from the outside.

ROI measurement. Leadership loses confidence when the number reported does not match what anyone can observe. The fix is structural: state the denominator, publish the breakeven alongside the benefit, and name the metric that would move in the wrong direction if the programme were going badly. The model above is written that way on purpose.

The Eight Challenges

The counter-case: readiness theatre

The argument above has an obvious failure mode, and it is common enough to have a shape. Call it readiness theatre. The programme buys the discovery phase, runs the workshops, produces the deliverable, and the defect distribution does not move.

The mechanism is specific. A discovery workshop asks people to describe their exceptions, and people describe exceptions as categories rather than as volumes. So the output is a list that reads “missing purchase order number, wrong entity, duplicate submission, legacy format” with no count against any entry and no note of which system each one comes from. That list cannot be used to scope a pilot, because there is no basis for choosing which variant to pilot. So the pilot gets scoped to the happy path, the happy path passes, and the four defects that were going to surface at full volume still surface at full volume. You have paid the 124 hours and bought a document.

In the model, readiness theatre is any outcome where less than 49.34 per cent of the intended distribution shift occurs. Below that line, the discovery spend is not recovered, and it is worth being blunt that a programme in this state is worse off than one that never bought discovery at all.

How to detect it, before the pilot rather than after. Take the exception list your discovery phase produced. Check whether every entry carries a volume and a named source system. Then take the ten most recent items that the current process could not complete without a human decision, and check how many of them appear on that list. If fewer than five appear, discovery produced a description, not a distribution shift, and the pilot scope should not be signed off yet. That check takes under an hour and it is the highest-value hour in the phase.

Two Exception Lists

What the rules now require, and when

Any piece written about implementation challenges before mid-2026 is likely to state a compliance timeline that has since changed, so this is worth getting right rather than gesturing at.

The European Union’s Artificial Intelligence Act, Regulation (EU) 2024/1689, has been formally amended for the first time since adoption. The Digital Omnibus on AI, Regulation (EU) 2026/1744 of 8 July 2026, was published in the Official Journal on 24 July 2026 and , according to the European Commission‘s own announcement. The two dates that matter most for planning both moved: the rules for high-risk AI systems listed in Annex III now apply from 2 December 2027, and the rules for high-risk AI embedded in physical products under Annex I, covering categories such as machinery, toys and lifts, apply from 2 August 2028. The same amendment extends several measures previously reserved for small and medium-sized enterprises to small mid-cap companies, expands regulatory sandboxes and adds an EU-level one, simplifies the AI literacy requirement on companies while giving the Commission and member states a stronger role in it, and permits processing of special categories of personal data where that is needed to detect and correct bias.

Two obligations are worth reading in the original rather than in summary, because both are build decisions rather than paperwork.

requires providers of high-risk systems to establish and document a post-market monitoring system, proportionate to the technology and the risk, which “actively and systematically collect[s], document[s] and analyse[s] relevant data” on performance “throughout their lifetime” so that continuing compliance can be evaluated. That monitoring system has to rest on a documented post-market monitoring plan, and the plan forms part of the technical documentation. Read against the ladder above, Article 72 is a legal requirement to have the mechanism that turns a rung-five defect into something you find rather than something a customer finds.

puts an obligation on both providers and deployers to take measures ensuring, to their best extent, a sufficient level of AI literacy among staff and others operating the systems on their behalf, taking account of their technical knowledge, experience, education and training, the context of use, and the people the system will be used on. That is the change management challenge above, restated as a duty. Note that the Digital Omnibus simplifies how this requirement operates, and that the European Commission’s AI Act Service Desk carries a notice on both articles stating that the displayed text has not yet been updated to reflect the amendments. Where a specific obligation is load-bearing for your programme, have counsel confirm the current consolidated wording rather than relying on a reading of the 2024 text, including ours.

The pattern is worth naming: the regulation is asking for the same artefacts that make a programme cheap to run. A monitoring plan is not a compliance overhead if you were going to need drift detection anyway.

A Timeline

The public-sector version of the same problem

It is fair to ask whether any of this is measured anywhere, or whether it is only ever anecdote. One place it has been examined on the record is the UK Parliament’s Committee of Public Accounts, in its report , published 26 March 2025.

The committee’s first conclusion is that out-of-date legacy technology and poor data quality and data-sharing “is putting AI adoption in the public sector at risk”, because “AI relies on high quality data to learn, but too often government data is of poor quality and locked away in out-of-date legacy IT systems”. The specifics behind that are useful because they carry denominators. Access to good-quality data was identified as a barrier to implementing AI by 62 per cent of the 87 government bodies responding to the National Audit Office’s survey. Difficulties recruiting and retaining staff with AI skills were identified by 70 per cent of those bodies. The department responsible defines a legacy system as one based on “an end-of-life product, out of support from the supplier, impossible to update, no longer cost-effective, or considered to be otherwise above the acceptable risk threshold”, and estimated that 28 per cent of central government systems met that definition in 2024. Of the 72 highest-risk legacy systems prioritised under the 2022 to 2025 digital and data roadmap, 21 still lacked remediation funding.

Two other findings are directly about the ladder. The committee found “no systematic mechanism for bringing together and disseminating the learning from all the pilot activity across government”, which is a description of paying rung-four costs repeatedly without ever moving a defect to rung one. And it noted that only 33 records had been published on the Algorithmic Transparency Recording Standard as at January 2025, which is the run-plan and traceability challenge showing up as a measured gap rather than as an opinion.

The scale is different from a mid-market programme. The mechanism is identical, which is the point of citing it.

Four Stat Cards

Buy stage one separately

Every procurement conversation about this reduces to a single structural question: are you buying a programme, or are you buying the right to find out whether the programme is a good idea. Those are different purchases and they should be priced differently.

The route that works is to buy the discovery and profiling phase as its own engagement, with its own price, its own deliverables, and an explicit walk-away point at the end of it. Not a free pre-sales workshop, which is a sales activity and produces sales artefacts, and not a discovery phase folded into a fixed-price build, which gives the supplier a reason for discovery to conclude that the build should proceed.

What that engagement should be required to hand over, and what to check on receipt:

  • A process record taken from observed work, with a stated count of how many items observed in the sample completed under the documented path. A percentage with a denominator, not an assurance.

  • A data profile over the full history, not a recent extract, listing distinct values, null rates and format variance per field the automation will read, with the dates on which patterns changed.

  • An exception register with a volume and a source system against every entry. This is the artefact that the theatre check above tests, and it is the one most likely to arrive incomplete.

  • A named pilot variant and a stated reason for choosing it, derived from the exception register rather than from convenience.

  • A stated price to stop. What it costs to end the relationship at the end of stage one, and what you keep if you do. If the answer is that you keep nothing usable, you have bought a sales document.

The walk-away option is the part that changes behaviour, and it is worth insisting on even when you have no intention of using it. A supplier who is willing to be paid for stage one and then told no is a supplier whose stage-one findings you can believe. This also gives you a clean way to compare two suppliers on something other than their estimate for work neither of them has scoped yet. Where an implementation has already gone the other way, the recovery pattern is different again, and we have written about how those situations usually unfold how those situations usually unfold.

The number pair that ranks the two programmes in opposite orders

If you take one measurement away from this article, take this one, because it is the single number that distinguishes the two versions of the programme without requiring you to know their total costs.

Share of total programme cost incurred after go-live. In version one, EUR 25,080 of the EUR 34,940 lands after go-live, which is 71.78 per cent. In version two, EUR 6,270 of EUR 24,120 lands after go-live, which is 25.99 per cent. The companion figure is the share of defects found before build begins: 8.33 per cent against 41.67 per cent.

Now notice what happens when you rank the two programmes. On weeks to go-live, version one wins, by the 3.10 person-weeks that version two spent on readiness. On total cost to a working state, version two wins, by EUR 10,820. The two metrics rank the same pair of programmes in opposite orders, and whichever one your reporting uses is the one your delivery teams will optimise for. If the only number on the status report is the go-live date, you have asked for version one and you will get it.

The uncomfortable implication is that a programme showing 70 per cent of its cost after go-live is not necessarily failing. It might be running a workflow so cheap to fix in production that front-loading would be waste, which is the slope 0.30 row of the sensitivity table. But you cannot tell the two apart without pricing your own defects, and almost nobody has.

Two Donut Charts

A readiness check you can run in a week

Not a maturity model. Five questions, each answerable in a day or less, each mapping to a rung on the ladder.

Take the last day's work in the target process and count how many items completed under the documented path.

If under 80 per cent, the document is not the process and build should not start.

Profile the full history of every field the automation will read.

Distinct values, null rate, format variance, and the date each pattern changed. This is the cheapest defect-finder on the list.

Take the ten most recent items a human had to decide, and check how many appear on your exception register with a volume against them.

Under five means the register is a description.

Name the owner of the exception queue after go-live, the monitoring they will see, and the service level they are held to.

Three names or numbers, not a team.

Price your own escalation slope.

Take three defects you fixed after go-live and three you fixed during build, and compare the hours. If the ratio is under about 5.6 to 1, the model in this article does not apply to you and you should front-load less.

If question five is unanswerable because nobody recorded the hours, that is the most useful finding in the set, and fixing it costs nothing except a field on a ticket.

Five Numbered Check Items

Closing

The eight challenges in this article are the ones everyone has. The variable is not which challenges you face, it is which rung you find them on, and the difference between rung one and rung five on a single defect in this model is EUR 170 against EUR 6,270. Whether that ratio holds in your organisation is an empirical question you can answer this week.

So here is the question to take into your next delivery review, and it looks backwards rather than forwards. Name the last material defect your team found after go-live. Then say which earlier stage could have caught it, and what specifically would have had to happen there. If the answer is that nothing earlier could have caught it, that is a genuine design-in problem and it belongs in the fourth tier above. If the answer is a day of data profiling, you have just priced your own discovery phase, and you did not need a business case to do it.

If it would help to have someone else run that assessment on one workflow before you commit to a build, that is exactly what a stage-one engagement is for a stage-one engagement.

FAQs

Why do AI and automation projects fail more often than the technology suggests they should?

Because most failures are not technology failures. They are defects in the workflow understanding, the data, or the operating model that were present from the start and were found late. A defect found in discovery costs a couple of hours to clear. The same defect found live at full volume costs an order of magnitude more, because by then there is code with dependents, a change window to negotiate, and work that has already been processed wrongly and has to be corrected.

It depends on how expensive your production fixes actually are, and that is measurable. In the illustrative model in this article, 124 hours of readiness work avoids 240 hours of rework, but the case breaks even if the readiness spend exceeds about 251 hours, if less than half of the intended shift in when defects are found actually happens, or if a live defect costs less than about 5.6 times a build-stage defect. If your automations are genuinely cheap to fix in production, front-loading is waste. Price three of your own late fixes against three early ones before deciding.

Profiling the full history of every data field the automation will read, rather than a recent extract. Data problems break more automations than any other single cause, and they are among the cheapest to find early. The specific failure to look for is a validation rule added at some point in the past, which makes recent records look clean while years of earlier records remain inconsistent. Recent-data testing passes and full-volume processing does not.

The Artificial Intelligence Act, Regulation (EU) 2024/1689, was amended by the Digital Omnibus on AI, Regulation (EU) 2026/1744, which entered into force on 27 July 2026. The high-risk rules for systems listed in Annex III now apply from 2 December 2027, and those for high-risk AI embedded in physical products under Annex I from 2 August 2028. Two provisions bear directly on implementation: Article 72 requires providers of high-risk systems to run a documented post-market monitoring system for the system’s lifetime, and Article 4 requires providers and deployers to ensure a sufficient level of AI literacy among the people operating them. Because the consolidated text is still being updated in places, confirm the current wording with counsel where a specific obligation is load-bearing.

Buy the discovery and data profiling phase separately from the build, with its own price and an explicit walk-away point. Require five things on delivery: a process record taken from observed work with a stated completion percentage, a data profile across the full history, an exception register with volumes and source systems against every entry, a named pilot variant with a stated reason for the choice, and a stated price to stop along with what you keep if you do. A supplier willing to be paid for stage one and then told no is a supplier whose stage-one findings you can act on.

References

European Commission, “AI omnibus enters into force”, Directorate-General for Communications Networks, Content and Technology, at digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force. Cited for the amendment of the AI Act by Regulation (EU) 2026/1744 of 8 July 2026, its publication in the Official Journal on 24 July 2026 and entry into force on 27 July 2026, and for the revised application dates of 2 December 2027 for Annex III high-risk systems and 2 August 2028 for Annex I high-risk AI embedded in physical products, together with the extension of SME measures to small mid-cap companies, the expansion of regulatory sandboxes, the simplification of the AI literacy requirement, and the permission to process special categories of personal data for bias detection and correction. Read at source; the page carries a last-update date of 31 July 2026. The consolidated legislative text of the amending regulation itself was not retrievable at the time of writing, so only the Commission’s own statements of the timeline are used here.

Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 72, post-market monitoring by providers and post-market monitoring plan for high-risk AI systems, via the European Commission’s AI Act Service Desk at ai-act-service-desk.ec.europa.eu/en/ai-act/article-72. Cited for the requirements in paragraphs 1 to 3: a documented post-market monitoring system proportionate to the technology and risk, active and systematic collection and analysis of performance data throughout the system’s lifetime, and a post-market monitoring plan forming part of the Annex IV technical documentation. The article text displayed is the official version of 13 June 2024, and the page itself carries a notice that it has not yet been updated for the Digital Omnibus amendments, which is why the article is cited for the nature of the obligation rather than for its current wording in detail.

Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 4, AI literacy, at ai-act-service-desk.ec.europa.eu/en/ai-act/article-4. Cited for the obligation on providers and deployers to take measures ensuring, to their best extent, a sufficient level of AI literacy among staff and others operating AI systems on their behalf, having regard to technical knowledge, experience, education and training, the context of use, and the persons on whom the systems are used. Same caveat on the pending Digital Omnibus update applies.

UK National Cyber Security Centre, Guidelines for secure AI system development, “Secure operation and maintenance”, at ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-operation-maintenance. Cited for the guidance to monitor system behaviour so that sudden and gradual changes affecting security can be observed, including natural data drift; to monitor and log inputs such as inference requests, queries and prompts to enable compliance, audit, investigation and remediation; and to run update processes that reflect the fact that changes to data, models or prompts can change system behaviour, treating major updates as new versions. Published 27 November 2023, version 1.0, and the page carries the same review date, which is stated here because the guidance predates current generative AI deployment practice even though the monitoring principles it sets out are unchanged.

Committee of Public Accounts, House of Commons, “Use of AI in Government”, Eighteenth Report of Session 2024-25, published 26 March 2025, at publications.parliament.uk/pa/cm5901/cmselect/cmpubacc/356/report.html. Cited for the committee’s conclusion that out-of-date legacy technology and poor data quality and data-sharing is putting public sector AI adoption at risk; for access to good-quality data being identified as a barrier by 62 per cent of the 87 government bodies responding to the National Audit Office’s survey and AI skills recruitment and retention by 70 per cent; for the departmental definition of a legacy system and the estimate that 28 per cent of central government systems met it in 2024; for 21 of the 72 highest-risk legacy systems still lacking remediation funding; for the absence of any systematic mechanism to disseminate learning from pilot activity across government; and for 33 records having been published on the Algorithmic Transparency Recording Standard as at January 2025. Read in full at source.

Proof & Testimonials

Trusted by teams building scalable automation

FedEx Express Europe

PAteam's deep architectural expertise helps us execute current opportunities while strategically planning for the future. Their flexibility has been key to our shared success.

— Andrzej Srebro

IT Manager

The Wasserstrom Company

PAteam significantly improved our productivity. By handling day-to-day development, they've enabled our employees to focus on high-value exceptions.

Michal T. Slominski

EVP, Information Technology

Healthcare Sweden

When an incident threatened our environment, PAteam restored operations with zero downtime. We rely on partners who deliver the highest level of service.

Director

Healthcare, Sweden

MI Homes

PAteam improved our productivity tremendously. Their automation expertise in streamlining data entry allows our team to focus on volume growth.

Director

MI Homes

Kirkendall Dwyer

PAteam makes complex solutions simple. They took my vision and turned it into an automated process that worked better than imagined.

Mason Johnson

Kirkendall

BPO Sector

If you want to avoid the pitfalls of building a scalable automation environment, PAteam are the masters at making that vision a reality.

Manager

Business Process Outsourcing

Unlock the Future of Work

One platform. Copilots that elevate people. Automation that scales everywhere. Let’s design a smarter, seamless operation for your customers, your teams, and your business.

unlock the future work
Scroll to Top