Industry Guides
AI Audio Tour Guide: Multilingual Scripts and Narration
With 25,000 audio tours on every phone, who still pays a guide? What it really costs to script a tour with AI and voice it in five languages, where the law draws the line, and the place-name pronunciation trap.

With 25,000 audio tours sitting on every visitor's phone, who still pays for a guide? We have heard that question from a lot of licensed guides over the past two years. izi.TRAVEL, VoiceMap and SmartGuide grow a little every season, and the newest versions generate an AI narration on the spot for any landmark missing from their catalog. The answer is less dramatic than you might fear. But it starts with a different question: if an app can do that, why shouldn't the guide?
This article is about producing an AI audio tour guide of your own, and everything that feeds it: getting a tour script out of an LLM without inventing history, adapting it to five languages, voicing it, slicing it up for social media, and doing all of that inside the legal boundary of a guiding license. It is written for licensed guides, small tour operators and agencies alike. We covered the agency side of the business, the inquiry-to-quote pipeline, in our AI for travel agents guide; the wider picture of the sector sits in our AI by industry map.
One caveat up front. The prices and app figures below come mostly from vendor pages and comparison blogs; the legal points come from law-firm summaries, using Turkey, where we work with tour businesses, as the concrete case. Check the current price list before you sign anything, and your own regulator before you make a decision about your license.
If audio guide apps exist, why does anyone need a guide?
Because an app plays a recording, while a guide answers the question a visitor just asked, reroutes the group around weather and crowds, and is the person legally responsible for the tour. Most countries that license guides reserve that second job for licensed people. Audio guide apps replace the museum placard and the rented headset, not the guided-tour product.
Lawyers are not the only ones saying this. A peer-reviewed study in a Turkish university journal looked at an AI and augmented-reality digital guide built for Hagia Sophia. More than half of the guides surveyed did not believe AI could replace them at answering questions, managing a group, reading emotion, interpreting, or interacting. The same study flags apps that market "tours without a guide" as legally and ethically contested.
So here is the real shape of the question. Apps growing does not put guides out of work, but it does ask something of them: present the tour better than the three-language recording the visitor already has in their pocket. That is the actual job AI has in this profession.
Can AI write a tour script, and can you trust it?
It can write one; you cannot trust it. Large language models (the technology behind ChatGPT and Claude) produce a fluent tour script and, with equal confidence, invent dates, attributions and dynasties. That is called hallucination. Asking a model for the stratigraphy of an Anatolian mound, the architect of a mosque or the reading of an inscription and then repeating it to a group without checking is a risk no licensed guide can afford.
With that said, the practical benefit is real. If a guide has a thirty-year Ephesus narration scattered across notebooks and phone memos, a model is genuinely good at turning it into a structured script: an opening, stop-by-stop narration, transitions, timing estimates, likely questions. The knowledge comes from the guide, the form from the model. "Write me a 45-minute Ephesus script" and "here are my thirty years of notes, turn them into a 12-stop script and add no new dates or names" produce two very different results.
There is a second reason to stay skeptical. Research on these models shows they are weaker at generating text in non-Latin scripts than at understanding it. A Russian, Farsi, Arabic or Chinese script comes out rougher than English or German. When we get to language priorities below, this matters more than it looks.
What does it cost to voice one tour in five languages?
On the audio side, almost nothing: a 45-minute tour script is roughly 40,000 characters, which costs under a dollar per language on Azure or Google neural voices and about four dollars per language on ElevenLabs' top model. Five languages together cost less than lunch. The real cost is writing, fact-checking and adapting the script, and that is human hours.
The numbers, as they appear in 2026 comparison blogs (vendors change list prices often, so verify before you commit):
- Google Cloud TTS: standard voices at $4 per million characters, Neural2 at $16, Chirp 3 HD at $30, Studio voices at $160.
- Azure Neural TTS: $16 per million characters; HD voices at $22 (one blog says this was cut from $30 in March 2026, single source); a free tier of 500,000 characters a month.
- ElevenLabs: the Flash model around $50 per million characters, V3 around $100. Widely rated the most realistic, also the most expensive and slowest.
Let us add our own arithmetic. Say you run three tours in a region like Cappadocia: a valley hike, an underground city and an open-air museum, 45 minutes each, in English, German, Russian, Turkish and Spanish. Three tours x five languages x 40,000 characters = 600,000 characters. On Azure that is under $10; on ElevenLabs V3 about $60. Paid once, and the files serve you for years.
Now the human side, from our own experience: turning a guide's raw notes into a clean 12-stop script in the source language takes the guide three to four hours even with the model's help, because every sentence gets read with the question "did I say this or did the model add it?" Each foreign language then needs one to two hours from a native reader. The true cost of three five-language tours hides in those 30-40 hours; the audio files are the smallest line in the budget.
Which languages should you prioritize?
The ones your own guests actually speak, in the order your booking book shows them. Turkey is a useful illustration of how this plays out: about 64 million foreign visitors in 2025, and the three largest source markets were Russia (6.9 million), Germany (6.75 million) and the United Kingdom (4.27 million). Nationally that makes Russian, German and English the first three languages. But an operator on the Mediterranean coast and a guide in a small heritage town do not share a guest profile.
This is where the hallucination warning comes back. Russian, the largest market, is one of the languages models generate worst, because of the Cyrillic script. English and German scripts can go to the model and then to a proofreader; Russian needs something closer to a rewrite budget. "We translated it into five languages" is easy to say; being equally good in five languages is a separate project.
Then there are the secondary markets. In the first quarter of 2025 Turkey received 733,000 visitors from Iran and 560,000 from Bulgaria. A Farsi or Bulgarian audio narration would probably make a guide the only one in their region offering it, because nobody invests in those languages. The language with the least competition is usually the cheapest way to stand out. The same logic applies wherever you work: look for the third and fourth nationality in your own numbers.
Do AI voices pronounce local place names correctly?
Mostly no, and we found no neutral study measuring it; what we see in the field is an English or German neural voice reaching a name like "Göbeklitepe" or "Çavuşin" and reading it with its own language's rules, which sounds comic to anyone who knows the place. There is a fix, but it is manual: phonetic respelling or SSML (a markup language that steers pronunciation line by line) applied to every proper name.
The practical move: list every proper noun in your tour (places, people, periods, dishes), define the pronunciation you want for each one once, and reuse that dictionary in every language and every tour. A region's dictionary runs to 30-40 words; build it once, apply it everywhere. Skip this step and the audio guide loses its value in the first second a listener says "a machine made this."
Voices in the source language are usually in better shape. ElevenLabs, Azure and Google all offer Turkish voices, for instance; ElevenLabs' own page claims it adapts to regional Turkish accents, which is a vendor landing-page claim you should test rather than trust. Never pick a voice without listening to a two or three minute sample with your own ears; the free tiers exist for exactly that.
Can you clone your own voice and speak five languages?
Technically yes: tools like ElevenLabs copy your voice from a few minutes of recording and make it speak other languages. The appeal for a guide is obvious; the visitor hears you on the tour and hears you in the headset. Two cautions. The first is legal: a voice is personal data. Cloning your own voice is fine, but using a colleague's or a former employee's voice without their explicit consent is a data protection problem under GDPR-style regimes and Turkey's KVKK alike. The second is quality: cloned voices in a foreign language sound like "a Turk who speaks the language well" rather than a native speaker; some visitors like that, others find it off-putting.
Our recommendation: use your own voice for the source language and the established neural voices for foreign languages. And do not hesitate to tell visitors "this narration was written by your guide and voiced by AI"; transparency also buys you tolerance for the occasional pronunciation slip.
Where will you use your AI audio tour guide content?
In four places: before the tour (selling it), during the tour (headset narration and the guide's own notes), after the tour (a digital keepsake the visitor takes home) and on social media (a steady stream of posts). One script feeds all four; what AI does here is break a piece of content written once into different formats. According to a survey by the tourism research and events organization Arival, three in five tour operators already use AI to write product descriptions, with website copy and blog and social content next in line.
Make it concrete. Picture a small operator running boat tours out of a resort town on the Aegean. The guide's 40-minute narration on the castle and the ancient mausoleum has been cleaned up once and adapted into five languages. Out of that come:
- A published audio tour on an izi.TRAVEL or VoiceMap-style app: travelers who never book the tour still listen and learn the operator's name. (Those apps' "25,000 tours, 137 countries" figures are app-store marketing copy; read them as claims about their own size.)
- Booking page and OTA descriptions: five languages, one voice, same facts.
- A post-tour email: a three-minute "what you saw today" audio in the guest's language, plus a review link.
- Twenty social posts a month: every stop is a post, every post exists in five languages. The content calendar never runs dry because the source script already exists.
The order of work matters. Source-language script first (from the guide, formatted by the model), then fact-check (by the guide), then language adaptation (model plus native proofreader), then the pronunciation dictionary, and voicing last. People who reverse the order and generate audio first redo everything at the first correction.
Are izi.TRAVEL and its rivals competitors or channels?
For most guides, channels. These apps sell recorded narration for museums and city walks; they do not compete directly with your guided tour, but putting your own narration on them turns them into your shop window. It is no accident that VoiceMap emphasizes that its content is written by local storytellers and licensed guides: the apps need guides for content too.
Keep one distinction in mind, though. A guide uploading their own narration to an app is legally closer to publishing a book; recorded content is generally not treated as live guiding. An agency handing those audio files to an unlicensed person and sending them out with a group is unlicensed guiding, even if AI wrote every word. The line the law draws runs through the person leading the group; who produced the audio file does not move it. In Turkey, the guides' union and the travel agencies' association publish a model guide-agency contract that describes exactly how that relationship is supposed to be set up; most licensing regimes have an equivalent.
Frequently asked questions
Which tool should I pick: ElevenLabs, Azure or Google?
Tight budget and many languages: Azure or Google (broad language coverage, low prices, free tiers). Naturalness and emotion in your source language: ElevenLabs. The sensible mix for most guides is ElevenLabs for the home language and Azure for the rest. Before deciding, voice the same paragraph in all three and listen.
Does having ChatGPT write the script create a copyright problem?
Text generated from your own notes is yours. The risk is the model reproducing another guide's published narration or a guidebook from memory. Before you publish a script, make sure it rests on your own sources; the instruction "add no new information" also reduces this risk.
Does this change site entry rules or license checks?
No. Where licenses are checked at heritage sites, they are still checked; owning an audio file does not give a group the right to enter without a guide.
Should a small operator outsource this?
Do the first tour yourself; the free tiers are enough. Once you reach five languages and ten tours, a one-time professional setup for the pronunciation dictionary, file management and app integration comes out cheaper than doing it by hand.
So what should you do?
- Pick the one tour you tell best and hand your raw notes to the model with the instruction "structure this, add nothing new"; the first script takes an afternoon.
- Set the language order from your own guest book; national arrival statistics are only a starting point. Budget a rewrite, not a proofread, for Cyrillic and other non-Latin scripts.
- Build the pronunciation dictionary on the first tour; every later tour draws on it.
- Voice the same paragraph on all three tools' free tiers and let your ears make the purchase decision.
- Write the license boundary into your contracts: decide up front which agencies may use your audio files and under what conditions.
Back to the opening question: with 25,000 audio tours on every phone, who pays a guide? The answer has to do with whether the best of those 25,000 is yours. AI does not produce this profession's knowledge; it carries the voice of the guide who has it into five languages and four channels. If you want to talk through which tool stack and which language order make sense for your own tours, bring your notes and come see us.

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.
Comments
No comments yet. Be the first to comment!