AI workflow automation means putting an AI step — reading unstructured input, classifying or routing it, drafting an output — inside a business process that previously required a person to do that reading, deciding or writing. For a Singapore SME the practical version is not a company-wide AI strategy: it is picking one high-volume, pattern-heavy workflow (supplier invoice capture, inbound enquiry triage, onboarding document checks), automating the mechanical parts, and keeping a human approval gate on anything that touches money, personal data or a customer’s outcome. Done that way, a first workflow is a 90-day project run by existing staff, not a transformation programme.

This guide covers what the term actually means, which workflows pay back first, how to choose tools, how to test accuracy before you cut over, what it costs beyond the licence fee, how Singapore’s grants and data protection rules shape the decision, and a 90-day plan with the artefacts you should have at the end of each phase.

What AI workflow automation actually means, and the one-page brief it produces

AI workflow automation combines two older ideas — business process automation and applied AI — into systems that handle messy inputs rather than only clean, structured data. Traditional automation needs a field in a form. AI automation can start from an email, a scanned delivery order, a WhatsApp photo of a signed job sheet or a supplier’s PDF invoice in their own peculiar layout.

That distinction matters because most SME bottlenecks are not in the tidy parts of a process. They sit in the handoffs, where somebody has to read something, work out what it is, and decide where it goes next. If you have ever watched an admin colleague open twenty emails, retype the contents into an accounting system and forward three of them to operations, you have watched the exact step AI automation is good at.

The output of this section should be a one-page brief per candidate workflow: what comes in, in what format, from whom; what a person currently does with it; what goes out; and which of those steps is reading, deciding or drafting. If you cannot fill that page from memory, you do not yet know the workflow well enough to automate it.

Three layers: rules, AI-assisted steps, agents

Separate three layers, because they have very different cost, risk and maintenance profiles.

Rules-based automation. If X happens, do Y. No judgement involved. Cheap, predictable, entirely deterministic, and it breaks silently when an input format changes. Most SMEs already have some of this — an email rule, an auto-generated reminder, a spreadsheet macro.

AI-assisted steps. A model reads, classifies, extracts or drafts inside an otherwise rules-based flow, and a person approves the output before it has effect. This is where most SME value currently sits, because the AI handles the ambiguity and the surrounding rules handle the plumbing.

Agents. A model plans a multi-step sequence itself — look up the order, check stock, draft the reply, send it — with limited review between steps. Genuinely powerful, and also where failures are quietest, because nobody sees the intermediate reasoning that went wrong.

Most SMEs should start in the middle layer and only move towards agents once the underlying workflow is well understood and instrumented. Which workflows to tackle in what order is covered in 5 enterprise workflows that should be redesigned first.

What AI adds to a workflow: reading, deciding, drafting

Strip away the marketing language and AI does one of three jobs inside a workflow:

  • Reading. Turning unstructured input into structured data — invoice number, vendor, amount, GST, due date; or name, masked identification reference, start date, department.
  • Deciding. Classifying, scoring or routing that structured data against rules or learned patterns — this is a quote request, this is a complaint, this invoice does not match any purchase order.
  • Drafting. Producing a first version of an output — a reply, a quotation, a summary, a shift roster — for a person to check, edit and send.

Almost every AI use case pitched to SMEs is a combination of these three. When a vendor demo confuses you, ask which of the three their product is doing, and what happens to the other two.

Where a human must stay in the loop

Keep a human approval step on anything that meets one of these conditions:

  • Money leaves the business (payments, refunds, credit notes).
  • Personal data is disclosed outside the company.
  • The output affects a person’s employment, pay, credit or access to a service.
  • The action is irreversible or visible to a customer without recall.

That is partly quality control and partly a Singapore compliance posture, which the data protection section below sets out in more detail.

Where it pays off first: examples by function

The workflows that automate well share three traits: they happen often, they follow a recognisable pattern most of the time, and getting one wrong is annoying rather than catastrophic. Here is where that shows up by function, and what the first slice of each looks like.

Finance and admin: invoice capture, reconciliation, claims

Supplier invoice and staff claim processing is the most common starting point because the documents repeat and the correct answer is usually unambiguous. The AI extracts fields, matches them against a purchase order or ledger entry, and pushes only the exceptions to a person.

Singapore gives you an unusual advantage here: a national e-invoicing network based on the Peppol standard, branded locally as InvoiceNow and run by IMDA, which transmits invoices as structured data directly between finance systems rather than as documents to be read. Where a supplier is on the network, there is no PDF to interpret at all — a far better foundation than optical character recognition plus guesswork. IRAS has also set out a phased GST InvoiceNow requirement/basics-of-gst/invoicenow-requirement-for-gst-registered-businesses) for GST-registered businesses, beginning with newly incorporated voluntary registrants, so the direction of travel for finance document flows in Singapore is towards structured data rather than scanned paper. Check the current phase dates on the IRAS page before you plan around them.

Practical sequence for a finance pilot:

  1. Ask your highest-volume suppliers whether they are on InvoiceNow, and move those who are, so their invoices bypass extraction entirely.
  2. Apply AI extraction only to the remainder — the PDF and email attachments.
  3. Auto-match against purchase orders or expected amounts within a tolerance you set deliberately, not by default.
  4. Route every mismatch to an exceptions queue with a named owner.
  5. Keep payment approval manual.

Staff claims are a close second and often easier, because the volume is predictable and the amounts are small enough that an error is recoverable. Receipts photographed on phones are the hard part; expect a residue that stays manual.

Sales and customer service: enquiry triage, quoting, follow-up

Inbound enquiry handling is high volume and low risk per item, which makes it a strong early candidate. The AI reads the enquiry, classifies intent, extracts the details needed to respond, checks them against a rate card or product list, and drafts a reply for a person to send.

The common mistake is jumping straight to a customer-facing chatbot. Internal drafting — where the AI writes and your staff send — gets you most of the time saving with almost none of the brand risk, and it produces a labelled record of corrections that tells you exactly where the model is weak before you ever expose it to a customer.

Follow-up sequencing is the quieter win: summarising a call, updating the CRM record, drafting the next touchpoint. Nobody enjoys CRM hygiene, which is why it does not happen.

Handling WhatsApp as an input channel

In Singapore a large share of SME enquiries, job instructions and delivery confirmations arrive on WhatsApp, often on a staff member’s personal device. That creates two problems before any AI question arises: the message is not in a system of record, and the personal data in it sits on a phone you do not control.

The sensible sequence is to fix the channel before automating it:

  • Move business traffic to a business account or a shared inbox tool so messages land somewhere exportable and auditable.
  • Establish who owns the number if the staff member leaves.
  • Only then add an AI step that reads incoming messages, extracts the job or enquiry details, and creates a record for a coordinator to confirm.

Photographs are the dominant WhatsApp input — signed job sheets, delivery dockets, meter readings, damaged goods. Extraction quality on a phone photo taken in a dim corridor is materially worse than on a scanned document, so plan for a manual residue and measure it rather than assuming it away.

Operations and HR: scheduling, document checks, onboarding

HR and operations in Singapore are document-heavy by nature: work pass applications and renewals, CPF contribution records, tax filings, onboarding packs, safety briefings, training records. AI is useful for completeness checking — has this new hire submitted everything required, does this renewal have the right supporting documents — and for drafting routine correspondence.

What it should not do is submit on your behalf. The official CPF employer and MOM work pass channels remain the system of record, so your automation should sit around those systems — preparing, checking, chasing, reminding — rather than pretending to replace them. A missed renewal deadline is an expensive way to learn that a model hallucinated a submission confirmation.

Workflows to leave alone for now

Be equally clear about what not to automate in year one:

  • Bespoke pricing or contract negotiation for major accounts. High variance, high value, and the cost of being subtly wrong is a lost margin you never see.
  • Anything requiring a regulated professional judgement.
  • Disciplinary, hiring or termination decisions. Drafting the paperwork is fine; deciding is not.
  • Processes that change every quarter. You will spend more time reconfiguring than the automation saves.
  • Anything where the input exists only on paper and nobody scans it. Solve the scanning first.

Choosing your first workflow

The workflows that look most impressive in a vendor demo are rarely the ones worth automating first. Pick on volume, variance and value.

The volume, variance and value test

  1. Volume. Does it happen daily or weekly, not quarterly? Setup effort is fixed; only frequency pays it back.
  2. Variance. Does it follow a recognisable pattern most of the time, or is every instance genuinely different? Variance is what kills pilots, and it is almost always higher than people estimate from memory.
  3. Value. Is the cost of an individual error low enough that a human check catches it in time?

High volume, low-to-moderate variance, moderate value is the sweet spot.

Map the workflow as it actually runs

Before looking at any software, write down the workflow as it actually happens, not as the process manual describes it. Sit with the person who does it and record: every system touched, every handoff, every point where someone makes a judgement call rather than following a rule, and every workaround that exists because a system does not do what it should.

This exercise frequently reveals that the real fix is redesigning the process rather than bolting AI onto a broken one. That trade-off is covered in AI workflow redesign vs buying more software.

A useful test while mapping: count the number of times the same piece of information is typed by a human. Anything typed three times is an automation candidate regardless of what else you conclude.

A second test: ask what happens when the person who runs the process is on leave. If the answer is “it piles up until they return”, you have found both a bottleneck and a business continuity risk, and the case for automating it is stronger than the time saving alone suggests.

A scoring checklist you can run in an afternoon

Score each candidate workflow out of 5 on each criterion below. This is a prioritisation aid, not a validated instrument — the value is in forcing a comparison, not in the arithmetic.

CriterionQuestion to askScore (1–5)
VolumeDoes this happen daily or weekly?
Pattern consistencyDo most cases follow the same steps, or does each differ?
Data availabilityIs the input already digital — email, PDF, form submission?
Error toleranceIs a wrong output inconvenient rather than costly?
Bandwidth freedWould this free meaningful hours, not minutes?
System fitDoes it connect to software you already run?
Owner identifiedIs there a named person who will own and tune it?
Continuity riskDoes the process stall when one person is away?
Personal data exposureDoes it avoid handling sensitive personal data in the first slice?

If your top-scoring workflow is not clearly ahead of the second, you have probably not mapped them in enough detail yet. Go back to the floor and watch the process again.

Tools and platforms: how to choose

Embedded AI, workflow builders, or custom build

There are three routes in, and most SMEs are best served by the first two.

ApproachBest whenWatch out for
Embedded AI in software you already ownThe workflow already lives inside your accounting, CRM or HR platformThe feature may sit on a higher pricing tier, or not exist yet; you move at the vendor’s pace
Configure a no-code workflow builderYou need to connect two to four existing systems with an AI step between themMaintenance falls on whoever built it; per-execution pricing scales with success
Custom buildThe workflow is genuinely core to how you compete, or no tool fitsNeeds technical ownership you may not have in-house; budget for maintenance, not just build

Start by auditing what you already pay for. Accounting, CRM, helpdesk and HR platforms have all shipped AI features recently, and an activated feature in a system your staff already use beats a new tool nobody logs into.

Integration reality: check before you commit

Most pilot delays are integration delays, not AI problems. Before signing, confirm in writing:

  • Does the tool have a native connector to your accounting, CRM or HR system, or does it need middleware?
  • Which version or tier of your existing system does that connector require?
  • Does the connector write back, or only read?
  • What happens to a record if the connection drops mid-run — is it retried, queued or lost?
  • Who holds the credentials, and what happens when that person leaves?

A tool that requires a data export, a spreadsheet and a manual import is not automation. It is a different kind of typing.

Pricing traps in per-document and per-execution models

Free tiers are genuinely useful for testing whether a workflow is worth automating. The trap is that most platforms price on volume — per document, per seat, per execution, per token — so a pilot that costs nothing in week one can cost real money once it runs your whole invoice queue. Before committing, model the cost at the volume you expect in month six, including retries and failed runs, and ask the vendor directly how failed executions are billed.

Also check what happens to your data and your configuration if you leave. Can you export the workflow logic? Can you export extracted data in a usable format? A workflow you cannot move is a dependency, not an asset.

Vendor questions with a pass mark

Treat this as a scored gate, not a conversation starter. Mark each answer pass, partial or fail, and refuse to proceed with more than one fail in the data rows.

QuestionWhat a pass looks like
Where is our data stored and processed, and in which jurisdictions?Named regions, in writing, in the contract or DPA
Is our data used to train your models?No by default, or a documented opt-out you have exercised
What is the retention period and deletion process?A stated period and a self-service or contractual deletion route
What accuracy should we expect on documents like ours?An offer to test on your documents, not a generic benchmark
How are exceptions surfaced?A named queue, dashboard or notification with an owner
What does the audit log capture, and can we export it?Field-level log, exportable without a support ticket
What is the pricing at three times current volume?A quotable figure or published tier table
How are failed or retried executions billed?A clear, unambiguous answer
Is the product on the Productivity Solutions Grant pre-approved list?Yes with category, or a straight no
What is the rollback plan if accuracy drops after a model update?Version pinning, advance notice, or a documented fallback

If a vendor cannot answer the data questions in writing, that is your answer.

Testing accuracy before you trust it

This is the step most SME pilots skip, and skipping it is why so many are abandoned in month three on the basis of a vague sense that “it kept getting things wrong”.

Build a gold set

Collect a sample of real inputs from a normal period — not your cleanest documents, and not only the awkward ones. Include the formats you actually receive: the tidy PDF, the scanned copy, the phone photograph, the email with the details in the body rather than an attachment. Have a person record the correct output for each by hand. That set is your yardstick for every configuration change you make afterwards.

Keep it small enough to re-run in an afternoon and large enough to cover your real variety. Store it with the correct answers alongside; the temptation to fix the gold set when the model disagrees with it must be resisted.

Decide what accuracy you need, field by field

Accuracy is not one number. On an invoice, a misread supplier name is a nuisance; a misread amount or GST figure is a financial error. Set a target per field and decide which fields must always be reviewed by a person regardless of confidence score.

Field typeFailure consequenceReview rule
Monetary amounts and taxFinancial error, possible filing impactAlways human-checked before posting
Bank or payment detailsFraud exposureNever auto-updated from a document
Dates and reference numbersRework, delayed matchingAuto-accept above threshold, sample-audit weekly
Descriptive textCosmeticAuto-accept
Classification or routingDelay, wrong queueAuto-accept with weekly misroute review

Run the same test after every change

Model updates, prompt changes, a new document template from a supplier, a switch of pricing tier — any of these can move accuracy. Re-run the gold set and compare. Without that habit you are relying on somebody noticing a pattern of errors in live work, which usually happens after a customer does.

AI agents: what works and what does not yet

Patterns that hold up in an SME context

The agent patterns that survive contact with a real business are narrow and reversible:

  • An agent that drafts a reply but does not send it.
  • An agent that checks stock levels and flags a reorder but does not place the order.
  • An agent that reconciles transactions and routes mismatches to a person.
  • An agent that prepares a submission pack and leaves the submission to a human.

The common thread is that the output is checked before it has an irreversible effect.

Write a scope statement before you build one

The artefact for this section is a short scope statement you can hand to whoever configures the agent. It should state, in plain sentences: what the agent may read; what it may write to and where; what it must never do without approval; what it should do when it is unsure; who is notified when it stops; and what the maximum blast radius of a single wrong run is. If the last answer is “we do not know”, the scope is too wide.

Guardrails, logging and approval gates

Every step that touches customer data, money or personal information should have an approval gate and a log recording what was read, what was decided, what a person approved, and what they changed. The changes are the valuable part: a log of human corrections tells you precisely where the automation is weak and when it is safe to widen the scope.

Minimum logging set:

  • Input received (source, timestamp, document reference)
  • Extracted or classified output
  • Confidence or exception flag
  • Human action taken: approved, edited, rejected
  • What was changed, if edited
  • Final output and destination system

Failure modes to plan for

Plan explicitly for the model being confidently wrong: a transposed invoice amount, a misclassified complaint routed to sales, a drafted reply with the wrong tone for a long-standing customer, a currency misread on an overseas supplier invoice, a duplicate posted because the same document arrived twice by different routes. Design so that a wrong output is caught at your checkpoint rather than by your customer.

Then decide, in advance, what triggers a pause: a defined spike in exception rate, a model version change, a supplier changing document format, or any error that reaches a customer. Write down who has the authority to switch it off and how they do it. A kill switch nobody knows how to operate is decoration.

Costing it out: budget, payback and hidden work

What drives cost beyond licence fees

Cost driverTypically appearsHow to control it
Workflow mapping and testingBefore go-liveTimebox it; map one workflow properly rather than five loosely
Integration and data cleanupBefore go-livePrefer tools that already connect to your systems
Building the gold setBefore go-liveKeep it small; reuse it for every later change
Exception handlingOngoing, every weekTrack exception rate weekly; tune or narrow scope
ReconfigurationWhen processes or formats changeName an owner; schedule a quarterly review
Training and change managementFirst two monthsInvolve the person doing the work from week one
Volume-based feesFrom month three onwardsModel costs at expected month-six volume

The licence is rarely the largest line. Exception handling and reconfiguration are.

A worked example: a 40-person logistics firm

The following is an invented scenario used to show the shape of a pilot. None of the figures are survey data; they are illustrative of a firm of this size.

Imagine a 40-person freight forwarder handling a few dozen supplier invoices a week, currently keyed into accounting software by one admin staff member who also handles customer billing queries and filing.

The pilot: AI invoice capture connected to the accounting platform, with the largest suppliers migrated to InvoiceNow where they support it, and bank feed reconciliation switched on. Everything that matches a purchase order within tolerance posts automatically. Everything that does not — wrong quantity, GST that does not reconcile, no matching PO — lands in an exceptions queue.

The shape of the change: the admin colleague stops typing and starts adjudicating. Their week shifts from keying every document to reviewing a handful of exceptions and chasing the suppliers responsible. Month-end closes earlier because the ledger is current rather than a week behind. When the firm next wins a contract that adds volume, it does not need another pair of hands to absorb it — which matters in Singapore, where adding headcount for routine data entry runs into work pass quotas and levies as well as a tight administrative hiring market.

The risks that surfaced in this scenario: two suppliers send invoices as photographs taken on a phone, which extract poorly and stay manual; and the tolerance threshold for auto-posting was initially set too loose, letting through small mismatches in the first fortnight before it was tightened.

A second example: a 120-person facilities services firm

Another invented scenario, with a different bottleneck. A 120-person cleaning and facilities maintenance company receives job requests by email and WhatsApp from building managers, dispatches crews, and closes jobs with a signed job sheet photographed on site. Its bottleneck is not invoicing; it is triage and closure.

The pilot reads inbound requests from a shared mailbox and a business WhatsApp account, extracts site, requested work, urgency and contact, creates a job record, and drafts an acknowledgement for the coordinator to send. Photographed job sheets are read at closure, matched to the job record, and flagged when a signature or completion time is missing. Crew dispatch stays human, because it depends on knowledge of who works well at which site.

The lesson from this shape of business: automate the ends of the process — intake and closure — and leave the judgement-heavy middle alone.

A third example: a 25-person engineering consultancy

A third invented scenario, this time in professional services. A 25-person engineering consultancy loses time to two things: assembling tender submissions from past project documents, and chasing timesheets so invoices can go out.

The tender work is tempting but wrong for a first pilot — every submission is different, the value of each is high, and a subtly wrong technical claim is expensive. The timesheet chase is the better first slice: a scheduled step that checks which timesheets are incomplete, drafts a personalised reminder to each engineer, and escalates to the project lead after a set period. No judgement, no personal data leaving the firm, trivial error cost, and it runs every week whether anyone remembers or not.

The broader point: in professional services the first automation is often administrative rather than intellectual, and that is fine. Cash collected a week earlier is a real result.

How to measure whether it worked

Capture baselines before you build anything, because you cannot reconstruct them afterwards:

  • Time per item. Timed, not estimated. Sample twenty items across a normal week.
  • Exception rate. How often a person must intervene beyond a quick approval.
  • Downstream error rate. Mistakes discovered later — in month-end reconciliation, a customer complaint, a rework.
  • Cycle time. Elapsed time from receipt to resolution, which is often what customers actually notice.
  • Cost per item. Total running cost divided by volume, recalculated at month six.
  • Staff experience. Ask the person doing the work whether their week improved. If it did not, adoption will decay.

If the exception rate stays stubbornly high after a month of tuning, the workflow had more variance than your scoring suggested. Narrow the scope — one document type, one supplier group, one enquiry category — rather than pushing forward.

What to do if the pilot fails

A failed pilot is only wasted if you learn nothing from it. Before abandoning, work through the causes in order, because they have different fixes:

SymptomLikely causeFix
High exception rate on a subset of inputsVariance underestimatedNarrow scope to the consistent subset
Errors concentrated in one fieldExtraction or prompt issueRe-test that field against the gold set; add a validation rule
Staff bypass the toolNo involvement, or it is slower than the old wayWatch them use it; the friction is usually one extra click or a missing shortcut
Nobody tunes itOwnership unfundedAllocate hours formally or stop
Cost overranVolume pricing at scaleRenegotiate, cap volume, or move the high-volume portion to structured input

If none of these apply and the process still resists automation, the honest conclusion is that the process needs redesigning first.

The Singapore context: grants, rules and skills

Funding routes worth checking first

Check these before paying full price.

The Productivity Solutions Grant, administered through GoBusiness, maintains a list of pre-approved digital solutions across categories including finance, HR, customer management and sales, with co-funding for qualifying SMEs. Search the list for your shortlisted tool before assuming a bespoke application is required — buying a listed solution is administratively far lighter than applying for a custom project. Check the current support level and eligibility conditions on the GoBusiness page itself, as both change.

Enterprise Singapore’s Enterprise Development Grant supports larger, more customised transformation projects rather than off-the-shelf software. It is the relevant route if you are building something specific to your business rather than configuring a listed product, and it expects a defined project scope and outcomes.

IMDA’s SMEs Go Digital programme is the vendor-neutral starting point, with sector-specific digital roadmaps and advisory support for smaller businesses working out what is appropriate for their industry.

Sequence matters: identify the workflow first, shortlist tools second, then check funding. Choosing a tool because it is funded, rather than because it fits the workflow, is how SMEs end up with subsidised shelfware.

PDPA, governance and data handling

Any workflow that processes customer or employee personal data through an AI tool must be checked against Singapore’s data protection rules, not just the vendor’s marketing claims. The Personal Data Protection Commission has issued advisory guidelines on the use of personal data in AI recommendation and decision systems, which address when consent or business improvement exceptions apply to using personal data with AI systems, and the accountability documentation organisations are expected to maintain.

Separately, IMDA and the AI Verify Foundation published a Model AI Governance Framework for Generative AI setting out dimensions including accountability, data provenance, incident reporting and testing. It is aimed broadly, but it is worth reading as a smaller business, because it shapes what larger customers and partners will start asking of their suppliers. Expect these questions to arrive in procurement questionnaires before they arrive in regulation.

A practical minimum for an SME pilot:

  • Keep a register of which AI tools process which categories of personal data, and where that data is stored.
  • Confirm in writing whether your data is used for vendor model training, and opt out if you can.
  • Check your collection notices still cover the purpose you are now using the data for.
  • Keep a human approval step on anything affecting a customer or employee outcome.
  • Retain logs showing what the system produced and what a person approved.
  • Name a person accountable for the tool, not just the process.
  • Decide what staff may and may not paste into consumer AI tools, and write it down in one page.

That last point matters more than most SMEs assume. The uncontrolled risk is usually not the pilot you governed carefully; it is the salesperson pasting a customer list into a free chatbot because nobody told them otherwise.

Skills, ownership and who runs it

You do not need to hire a data scientist to run a first pilot, but somebody must own it — mapping the workflow, reviewing outputs weekly, tuning the configuration and deciding when to widen scope. In most SMEs the right owner is the operations or finance person who already knows where the process breaks, given time and training rather than a new job title.

SkillsFuture’s enterprise-facing support, including the SkillsFuture Enterprise Credit, is designed to help fund exactly this kind of internal capability building. Training an existing employee to configure and supervise AI tools is usually more realistic than competing for technical hires in Singapore’s market, and the resulting knowledge stays attached to people who understand your processes.

Budget the owner’s time explicitly — a defined few hours a week during the pilot, agreed with their manager and visible in their objectives. Unfunded ownership is the most common quiet cause of pilot failure.

A 90-day rollout plan you can run internally

A 90-day structure keeps the project small enough to actually finish. Each phase has an artefact; if the artefact does not exist, the phase is not done.

Weeks 1–4: map, score, pick one

Map four or five candidate workflows using the checklist above, score them, and pick exactly one. Resist piloting two at once; you will learn less from each. Capture baseline metrics now — time per item, exception volume, cycle time — because after go-live you cannot recover them. Build the gold set while you are already handling live documents. Check the PSG pre-approved list and SMEs Go Digital resources so you are not paying full price for something co-fundable. Name the owner and agree their weekly time commitment with their manager.

Artefact: one workflow brief, one scored comparison table, one baseline sheet, one gold set.

Weeks 5–8: build the thin slice and shadow-run it

Build the narrowest useful version — one input type, one output, one approval gate — and run it in shadow mode alongside the existing manual process, comparing outputs without switching anyone over. This is where you discover whether your variance estimate was optimistic. Log every disagreement between the AI output and the human output, and categorise the causes weekly; the categories tell you what to fix.

If personal data is involved, set up consent positions, retention settings and logging now, not after go-live.

Artefact: a shadow-run comparison log, a categorised list of disagreement causes, and a written accuracy threshold you are willing to accept.

Weeks 9–12: cut over, measure, decide what is next

Once shadow accuracy is acceptable for your risk tolerance, cut over, keep the approval gate, and track your baseline metrics for at least four weeks before declaring success. Write a one-page retrospective covering what the AI handled well, what stayed manual and why, what the true running cost is at current volume, and what you would do differently. Then pick the second workflow using what you learned about your own process rather than the vendor’s pitch. For a broader view of which processes to tackle in what order, see our resources section.

Artefact: four weeks of post-cutover metrics against baseline, and the one-page retrospective.

Months 4–6: consolidate before expanding

The period after a successful pilot is where most of the value is either compounded or lost. Review the exception log monthly and fix the top recurring cause. Re-check costs against actual volume. Confirm the owner still has the time allocated. Re-run the gold set after any model or vendor update. Document the configuration well enough that a second person could take it over. Only then widen scope — another document type, another enquiry category, another site — one increment at a time.

Common mistakes to avoid

  • Starting with the most complex workflow because it is the most painful. Pain is not the same as suitability.
  • Buying before mapping. The tool shortlist should fall out of the workflow map, not the other way round.
  • No baseline. Without before-numbers you cannot prove value, and the project becomes a matter of opinion.
  • No gold set. You will argue about accuracy for months without one.
  • Skipping the shadow run. It is the cheapest risk control available.
  • Automating a broken process. You get errors faster.
  • No named owner. Shared ownership means nobody tunes it, and accuracy drifts.
  • Ignoring the person whose job changes. They know the exceptions, and they can quietly ensure the automation fails.
  • Letting the kill switch go undocumented. Somebody must know how to stop it at 6pm on a Friday.
  • Treating go-live as the finish. The first month after cutover is where the real configuration happens.

Want help applying this? See what Fetch builds.