Industry Guides
AI for Consulting Firms: Where Report Writing Breaks
A field experiment with 758 consultants found AI sped up some tasks and reduced accuracy on others. Where exactly does the line fall in report production?

758 consultants. That was the sample in the field experiment Harvard Business School researchers ran with BCG. The result pointed two ways at once. On tasks where AI is strong, consultants completed 12.2% more work and worked 25.1% faster. On a task requiring them to interpret data alongside interview notes and arrive at the right business recommendation, consultants using AI were 19% less likely to reach the correct answer.
The second number matters more than the first, because that second task type is the core of a consulting report. The researchers called it a "jagged frontier": you cannot tell which side of the line a task falls on by looking at how hard it is. Of two tasks of similar difficulty, one speeds up and the other degrades.
So the right question for a consulting firm is not whether to use AI. It is which layer of the report to use it on. This piece separates those layers.
Where does AI genuinely help in report production?
AI is reliable on format and language work, and unreliable on judgement and attribution. Turning scattered field notes into structured findings, summarising a long regulatory text, drafting the first pass of an executive summary, reformatting the same finding for three different client templates: all of that sits in the safe zone.
Make it concrete. A consultant walking out of an audit has thirty fragmentary notes on their phone. "Labelling inconsistent in the warehouse", "training records stop in 2024", "corrective action form left blank". Turning those into a standard finding format, meaning observation, evidence and recommendation, takes hours and contributes very little intellectually. AI does it in minutes, and generally does it well.
The places it breaks:
- Regulatory and standard clause attribution. The model will write a clause number that does not exist, in an extremely confident tone. In compliance, safety and financial work this is fatal.
- Numerical analysis and table consistency. The percentage in the narrative and the figure in the table drift apart silently.
- Client-specific context. The same standard is applied differently at two sites. The model writes the general truth and does not know the exception on the ground.
- The link between finding and evidence. The model can produce a highly plausible finding that was never actually observed during the engagement. A reader cannot tell the difference.
Compressed into one sentence: you can delegate how the report is written. You cannot delegate what it says.
How serious is the fabricated source problem?
The cost of this one has been measured. In October 2025, a report Deloitte Australia prepared for a government department, on a contract worth roughly 440,000 Australian dollars, was found to contain citations to academic sources that did not exist and a fabricated quote from a federal court judgment. The firm repaid the final instalment of the contract.
The instructive part is who caught it. An academic reading the report noticed that publications he had never written were being attributed to him. Neither the consulting firm nor the client caught the error. An outside reader did.
The problem was not that AI wrote the report. The problem was that nobody checked the sources.
One more detail matters: the corrected version disclosed that an enterprise deployment of a commercial model had been used. So the assumption that "we bought the enterprise tier, the hallucination problem is handled" is wrong. Enterprise deployments improve data confidentiality. They do not guarantee accuracy.
The practical rule is simple: every number, every clause reference and every external citation in the report gets verified against its source by the person signing it, whether AI produced it or not. Skipping that step converts your time savings into reputational risk.
How do you catch a source the model invented?
Fabricated citations are not random. They follow a recognisable pattern: the model combines real author names, real journal titles and plausible dates into a source that does not exist. That is why you cannot spot them by reading. The only way to catch them is to go to the source.
A few signals speed up the review:
- The suspiciously perfect source. A citation whose title is almost a restatement of your claim deserves a second look.
- Unverifiable bibliographic detail. Author and year are present but there is no direct link. Models have no trouble inventing a citation line; working URLs are harder.
- Round, striking figures. Something like "30% time saved" appears in genuine studies and fabrications alike. Only the source itself separates them.
- Clause number mismatch. The standard reference may be correctly formatted but point at different content. Check the text of the clause, not the number.
The way to make this scale is to move verification into the writing rather than leaving it to the end. If the source link is captured alongside the finding as it is written, sign-off becomes a scan. Verification deferred to the end is, in practice, verification that never happens.
From field note to report: how does the workflow actually run?
Deployments that work share one trait: they place AI at a specific step rather than across the whole process. Break report production into steps and the delegable ones become obvious.
Take a safety consultancy. The steps are: observation and note-taking on site, converting notes into finding format, mapping findings to the relevant standard clauses, writing corrective action recommendations, and assembling the report in the client's format.
Steps two and five are fully delegable. The consultant records voice notes on site, those are transcribed and structured according to the firm's own finding template. At assembly, the client's required format is applied automatically. Together these represent a substantial share of the hours in a typical report and the smallest share of the thinking.
Step three, mapping to standards, is delegable only in a controlled setup: the model must quote from the firm's own verified reference archive rather than from memory, and show which document it took the text from. Step four, writing recommendations, can be delegated at draft level but should not reach the report without passing the eyes of someone who knows the client. Step one is not delegable at all. What was observed on site is the consultant's own observation, and the entire value of the report comes from it.
This split also changes how teams are staffed. Junior consultants used to spend most of their time on steps two and five. When that work shrinks, juniors start getting site experience earlier. There is a risk to watch, though: automating the entry-level work should not close off the channel through which people learn the job.
What are templates and RAG, and can a small firm build one?
The internal assistants large consultancies have been building since 2023 have one thing in common, and it is not the strength of the underlying model. It is the order of the archive they connect to. McKinsey's internal tool, Lilli, runs against the firm's document base built up over decades and answers with sources attached. PwC, Deloitte, EY, Bain and KPMG all built comparable closed assistants in the same period.
The pattern is usually called RAG: instead of producing an answer from memory, the model first retrieves relevant passages from your own document store and then writes using only those passages. Fabrication risk drops because the answer has a document behind it that can be shown.
Can a small consultancy build this? Technically yes, but the sequence matters. The value comes from the archive, not the model. If your past reports sit in unnamed folders, and three versions of the same standard are in circulation, the assistant you build will read that mess back to you in well-formed sentences.
So the correct order is: standardise your report templates first, move past reports into an organised archive second, add the AI layer last. In our experience the first two steps alone save a few hours per consultant per week, and they carry no licence fee at all.
Can you upload client documents to an AI tool?
There are two separate limits here, and most firms only think about one. The first is data protection law. The second is the confidentiality obligation in your engagement contract, which is frequently the more binding of the two and can be breached by documents containing no personal data whatsoever.
On the legal side the position is straightforward. Data protection rules are technology-neutral. If the text you paste into an AI tool contains personal data, that is a processing activity, and the principle of data minimisation says the unnecessary parts should never have been sent. Regulators in several jurisdictions have now published specific guidance on generative AI; read the current version from your own regulator when you set your internal policy, since this guidance is being revised frequently.
On the contractual side: uploading a client's financial statements, employee list or audit findings to a third-party service typically counts as "disclosure to a third party" under a standard confidentiality clause. The fact that the vendor does not train on your data does not remove that clause.
A workable framework looks like this:
- Use an enterprise or API-based deployment, and confirm the training exclusion in the contract text rather than the marketing page.
- Mask before you send: names become "Employee A", the company becomes "Client X". Analysis quality is barely affected.
- Write a usage policy instead of a ban. A banned tool gets used from personal accounts, and that usage cannot be audited.
- Raise it early with clients whose confidentiality terms are strict. Usage discovered later is a far bigger problem than usage disclosed upfront.
We covered the limits of putting contract text itself through a model in our guide to AI contract review.
Should you tell clients you used AI?
Yes, and volunteering it is incomparably better than having it discovered. In the Deloitte case the disclosure of the tool came after the error surfaced. From that point on, every statement reads as a defence; the same statement made upfront reads as methodological transparency.
It does not need to be framed as a confession. A few lines in the methodology section will do: at which stage AI support was used, how data was protected, and who performed verification. What clients actually want to know is not whether you used the tool but who checked the output.
Some enterprise clients and public procurement contracts have started including explicit clauses on this. Looking for that clause at signing costs less than arguing about it after delivery.
If the work got faster, should your fee drop?
That depends on your pricing model, and it is the question the sector is actually wrestling with. If you bill by days spent, making the work faster directly reduces your revenue. If you bill for the value delivered, the speed gain stays with you as margin.
Large firms are already living through this shift. Trade press reports describe McKinsey moving a meaningful portion of revenue toward outcome-based pricing, and the future of hourly billing in consulting is being openly debated. These accounts come more from sector journalism than from firm statements, so it is wiser to read the direction than the specific percentages.
For a small or mid-sized consultancy there are three ways out:
- Grow capacity: serve more clients with the same team. Easiest route, useless if your market is saturated.
- Fixed-price packages: "this scope of report, for this fee". The speed gain stays with you.
- Widen the deliverable: stop at the report no longer. Add implementation follow-up, periodic monitoring or training. Convert saved time into a new service line.
The bad route is passing the speed gain straight through as a discount and chasing more work for the same profit. That hands the productivity gain to your client rather than your firm, and turns competition into a price race.
Where to start
- Split your report process into layers. Which part is format work and which part is judgement? Delegate only the first.
- Write the verification step into policy. Every number, clause reference and external citation is checked at source before sign-off. Put it on the checklist.
- Fix the archive first. An AI layer built without template standards and an organised document store just accelerates the existing mess.
- Use enterprise deployments and mask data. These are complements, not alternatives.
- Review your pricing model now. If you price by days, the productivity gain is flowing to your clients rather than to you.
- Be transparent with clients. Saying it upfront beats having it noticed later by a wide margin.
What consulting sells is not the report. It is the judgement behind it. AI is making everything around that judgement, the compiling, the formatting, the writing, meaningfully cheaper. It is not making the judgement cheaper. Firms that write that distinction clearly into their process will work faster and keep the value of their signature intact.

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.
Comments
No comments yet. Be the first to comment!