AI for Business
How to Choose an AI Pilot Project: A Risk-Payoff Matrix
A two-axis risk and payoff matrix for picking your first AI pilot project: which process to start with, how many weeks to run it, how to write the success threshold, and why pilots stall.

Your first AI pilot project succeeds on how boring it is, not on how impressive it looks. We state that as a thesis because we have watched the opposite play out too many times: the pilot that starts with a customer-facing chatbot built to impress the board gets quietly shut down six months later, while the pilot that starts with invoice matching, a job nobody would ever present a slide about, goes into production in month three.
This piece answers the "which process should we start with" question with a two-axis matrix: risk and speed of payoff. We cover which quadrant to start in, how many weeks a pilot should run, how to write the success criterion, and why so many pilots never reach production, with concrete examples throughout. For the broader picture, see our end-to-end guide to AI for small business; here we focus only on the first step.
Why does the first step carry so much weight? Because the three barriers small companies report most often, lack of expertise, high cost and legal uncertainty, are all shrunk by a well-chosen first pilot. It needs little specialist knowledge, it turns on a small budget, and it never touches sensitive data. A badly chosen one amplifies all three.
Turkey offers a clean data point on how wide that gap is. According to the national statistics office's 2025 AI survey, only 6.6 percent of businesses with 10 to 49 employees use at least one AI technology, against 24.1 percent of those with 250 or more. Among companies that considered AI and held back, 74.2 percent cited missing expertise, 67.4 percent cost and 62.4 percent legal uncertainty. The numbers differ by country; the ranking of the barriers rarely does.
Why do AI pilot projects never reach production?
The main reason pilots stall is that they are chosen badly and measured badly; model quality sits far down the list. According to MIT NANDA's 2025 report, 95 percent of enterprise generative AI pilots produce no measurable profit-and-loss impact. That figure is routinely distorted into "95 percent of pilots fail". What the report actually says is narrower: the pilots run, but nobody can show the financial effect.
The report's own explanation is a "learning gap": the tools do not learn from the workflow, and the organizations do not manage the adaptation. It covers more than 300 enterprise-scale deployments, 52 case studies and 153 executive interviews, so it does not map one-to-one onto a 30-person company, but the lesson travels. Secondary reporting of McKinsey's 2025 survey adds that roughly two thirds of organizations remain stuck in pilot mode. That figure rests on a single relay; the direction is right, read the number with care.
The pattern we see in the field comes down to three points. The pilot is chosen because it is "technically interesting" or "a competitor announced one". Success is reported with a technical metric such as accuracy percentage rather than a business outcome a finance director can use. And the path to production is never defined before the pilot starts, so even a successful pilot ends with "so what happens now" hanging in the air. We covered the wider failure modes in six lessons on why AI projects fail; this piece is about preventing the first one.
How do you build the risk-payoff matrix?
The horizontal axis is speed of payoff, the vertical axis is risk. Score each candidate process from 1 to 5 on both and place it on the grid. The first pilot comes from the low-risk, fast-payoff quadrant. The high-risk, high-payoff quadrant is where the second or third pilot lives, never the first.
Three questions set the risk score. What do you lose if it makes a mistake: an email, a customer, or a court file? What data does the process touch: a product description, a customer's identity details, or a health record? Does the output reach an external customer without a human seeing it first? Each "yes" pushes the risk score up; three yeses disqualify the process as a first pilot.
Three questions set the payoff-speed score. How often is the task done: fifty times a day or once a month? How many minutes does each instance take? How many weeks until the result becomes measurable? A five-minute task done fifty times a day is a far better candidate than a one-hour task done once a month. This is where the common "impact versus effort" matrix falls short: it measures effort but not risk or adoption. In our matrix, effort is folded into payoff speed, and risk gets its own axis.
Add one simple design principle on top: AI produces the draft, a human approves it. No process that sends money automatically, changes records or communicates externally without review qualifies as a first pilot. That single rule cuts the risk axis roughly in half.
Which processes land in the low-risk, fast-payoff quadrant?
The low-risk, fast-payoff quadrant usually fills up with back-office work: classifying incoming email, reading and matching invoice and delivery-note data, drafting proposals and contracts, meeting summaries, product description writing. What they share is high volume, low stakes, and output whose usability can be judged in ten seconds.
- Incoming email and request classification: hundreds of messages, low error cost (a message in the wrong queue is moved by a person), results measurable in two weeks.
- Invoice, delivery note and order form reading: the AI extracts the fields, accounting approves; errors are caught at the approval step.
- Proposal and email drafting: the sales team edits and sends; writing time is measurable.
- Internal knowledge assistant: for procedures, product catalog and frequently asked questions; used only by staff, never opened to external customers.
The high-risk quadrant is where the most-requested projects sit: an autonomous chatbot talking to external customers, a model that sets prices or makes credit decisions, automatic payments or refunds. These make fine second projects and poor first ones. Which jobs call for a chatbot and which for plain automation is something we worked through in our decision tree for chatbots, agents and automation.
How long should an AI pilot project run, and what should it cost?
A pilot should run four to twelve weeks and have a written end date before it starts. Past ninety days a pilot enters what practitioners call "pilot purgatory": neither shut down nor scaled, while budget and management patience drain away. The most common mistake is extending a pilot without a clear go or no-go criterion.
Budget depends on scope, and the main driver is the number of integrations. One US practitioner's range runs from $12,000 to $40,000 for a narrow, small-team, 60-to-90-day pilot, up to $200,000 to $500,000 for three to five integrations, a hundred users and a regulated industry. Those are single-source US consulting prices; in other markets, carry over the ratio rather than the amount. Our budget-tier breakdown shows what a $2,500 first-pilot scope actually buys, line by line.
Where the money goes matters too. In a first pilot most of it goes into gathering the data and connecting to the existing system; the model line stays small. A single-integration pilot often comes in at a fifth of the cost of a three-integration one; you would expect a third, it turns out lower. That is the argument for a one-system, one-data-source rule in the first pilot.
How do you know the pilot succeeded?
If a single primary business metric, a baseline measurement and a written go or no-go threshold were set before the pilot started, success is beyond dispute. Without them, everyone sees something different when the pilot ends and the decision drifts.
The primary metric is written in business language: "processing time per email from 9 minutes to 3" or "invoice entry error rate from 4 percent to 1 percent." Accuracy, F1 score and other technical metrics can support the case but cannot lead it; no finance director approves a budget on an F1 score. The baseline is measured for one week before the pilot starts; otherwise "what was it before" becomes a guess.
Changing the criterion during the pilot or after the results arrive invalidates the decision. The threshold is written at the start, signed by the technical and business sides, and the decision follows it whatever the outcome. A stop decision is also a success: shutting down a $2,000 pilot prevents a $20,000 production system that would have failed.
How a 40-person packaging manufacturer used the matrix
An example from the field, company anonymized, figures rounded. The management of a 40-person packaging manufacturer came to us with three candidate projects: a chatbot that quotes prices to customers, entering dealer orders that arrive by email into the ERP, and predictive maintenance from machine fault logs.
We placed them on the matrix. The chatbot: direct contact with external customers, price statements, commercial loss on error; risk 5, payoff speed 3. Predictive maintenance: almost no data, results visible only after months; risk 2, payoff speed 1. Order emails: 30 to 40 a day, 6 to 8 minutes of manual entry each, a shipping mix-up on error but with a human approval step in between; risk 2, payoff speed 5.
The order emails won. The AI extracted product code, quantity and delivery date from each email and created a draft order in the ERP; a sales support employee approved the draft. The eight-week pilot's primary metric was entry time per order: it fell from 7 minutes to 1.5, and the faulty-entry rate dropped from 3 percent to 0.5. A one-page table replaced the board presentation, and the system went into production in month three. The chatbot is still on the agenda, but the company now has a working data flow and a measurement habit to build it on.
Five things not to do in a first pilot
Even with the matrix set up correctly, five habits sink pilots, and every one of them starts with good intentions. The list below was compiled from the common traits of pilots we have seen fail to reach production over the past two years.
- Growing the scope mid-flight. "Since it reads email, let it read WhatsApp too" turns an eight-week pilot into a six-month project. The second channel is the second pilot's subject.
- Mistaking the vendor demo for a pilot. A demonstration on the vendor's own data is not a pilot on yours. A pilot runs on your invoices and your emails.
- Bringing the user in last. If the employee who does the process is not at the table from day one, the tool that comes out will not be used. Adoption belongs among the metrics.
- Waiting for the data infrastructure. The company that says "let's build the big data platform first" never starts a pilot. A pilot on a single data source demonstrates the infrastructure need with a real example.
- Saving the result for the presentation. A short weekly table produces more decisions than a big end-of-pilot presentation; a surprise result is not a good result.
Frequently asked questions
Which process should a first AI pilot project target?
One that is high volume, low stakes, whose output can be checked in ten seconds and where a human approval step sits in the loop. Back-office work fits that description in most companies.
How long should the pilot run?
Four to twelve weeks. Past ninety days a pilot rarely produces a decision; write the end date on day one.
What does the pilot budget actually pay for?
Mostly data preparation and integration; the model is a small line. Because integration count drives the budget, keep the first pilot to a single system.
Why does a technically successful pilot get shelved?
Because the finance side never sees a business outcome and the path to production was never defined up front. An accuracy percentage is not a budget decision.
Are there government grants for a pilot?
It depends on the country, and AI-specific pilot funding is rarer than general digitalization or R&D support. If you operate in Turkey, we mapped the available programs in our guide to Turkish AI investment grants; note that we could not confirm a pilot-specific AI scheme there, so check current calls directly.
What should we do if the pilot fails?
Write down the lesson and move to the next candidate on the matrix. Keeping a failed pilot alive with "let's extend it a little" leads to pilots stacked on pilots and a melting budget.
So what should you actually do?
- Write down five candidate processes and score each from 1 to 5 on risk and on payoff speed. Do the scoring with the employee who performs the process, not from a manager's estimate.
- Pick from the low-risk, fast-payoff quadrant. If the project everyone wants is not in that quadrant, save it for the second pilot.
- Write the end date, the primary metric and the threshold on day one. All three fit on one page; have the technical and business sides sign it.
- Measure the baseline for a week. When the pilot ends, "what was it before" should not be a guess.
- Draw the path to production before the pilot starts. If "who takes it live, with what budget, in how many weeks, if it succeeds" has no answer, the pilot should not start.
Choosing an AI pilot project means choosing the most measurable job; the brightest idea waits its turn. Once five candidates are scored and placed on the grid, the first pilot usually announces itself. If yours does not, send us the scoring table and we will tell you which square we would start in.

Written by
Muhammet Fatih Batman
Founder & Editor
Founder of YZ Uzman, with 20+ years of experience in web design and software development.
Comments
No comments yet. Be the first to comment!