Webinar: discover the new localization workspace for marketers.
Save your seat
AI Translation

The real cost of building in-house AI translation

Sasho Savkov,Updated on October 8, 2026·10 min read
Three stacks of coins, each taller than the last, illustrating the rising cost of building in-house AI translation

Someone probably sent you this because you're weighing a build-versus-buy decision for AI translation. It's a common dilemma: the model API is right there, and the initial coding is fast. But here is the reality check: the code is the cheapest part of the equation. This is what the reality of building a DIY translation system actually looks like.

It's Friday. Fifty strings, a deadline, a chat window. You paste in You have {count} open orders, ask for German, and what comes back is good — genuinely good. A native speaker shrugs: ship it.

First: if you ship a handful of locales, a few thousand words a month, and someone on your team reads the languages you ship — build the script and close this tab. You'll be right.

Second, the numbers underneath this article: we took four customers' content, about twenty language pairs, and translated it two ways — a bare model call, and the same call wrapped in the customer's context — then measured how much editing each output needed to match translations the customer had already approved. Context cut the editing by 20 to 57%. Same strings, same model. The difference isn't the model. It's everything around it.

A comparison of a bare model call against the same call wrapped in customer context: four customers, about twenty language pairs, and 20 to 57 percent less editing

Act I: You automate it

TL;DR. The script works, but then the chaos starts. Turns out the real problem isn't the model API but the fragmented “truth” scattered across your repos, CMS, and design tools. You aren't just writing code; you're trying to organize your company's data.

The demo was a success, so on Monday, you're responsible for five thousand strings instead of fifty.

The engine works. From here on, your biggest problem isn't the system breaking, but the workflow itself. All of the content needs to go into the system, then back again into production.

Product copy might live in the app, while other text sits on the website, in the help center, in emails, or in design files. Each source works slightly differently, and the translated text has to go back into the right place without changing formatting, breaking personalization, or overwriting anything that is already live.

Some of that is surprisingly difficult to automate. Translating text inside a design file, for example, is one thing. Putting every translation back into exactly the right place without changing the layout is another.

Other problems are easier to miss. Imagine the original string is Active Members: {0}, where {0} tells the app where to insert a number. The AI produces perfectly good French, but drops {0}. The translation sounds right, but the feature breaks.

When this happened to a customer of ours, it broke their app; their developers filed it as an outage, not a typo. So now you need checks to make sure important elements like this are preserved.

And there is a bigger issue. When you were translating fifty strings in a chat window, you read every result yourself. You were checking the quality. Once you automate thousands of translations, you stop reading every line. You can build checks for obvious problems — missing text, broken formatting, or an incorrect number of translations — but those checks still can't tell you whether the translation itself is actually any good.

Lesson 01. A script sends the app's string files to an LLM API in batches, checks placeholders, and writes translations back. CMS, help center, and Figma stay export-only.
Act I ledger: a script, a retry queue, and a bucket of results. Owned by a third of an engineer, costing low hundreds a month, with no way to tell whether anything is right.

Act II: You feed it your company context

TL;DR. The model speaks the language, but it doesn't speak you. To sound like your brand, you have to build four support systems — glossary, style, translation memory, and context — just to keep the AI in check. The code goes up in a week; the curation is forever.

The first real quality report doesn't come from your pipeline — it can't. It comes from the French country manager, in Slack: “This is correct, but it doesn't sound like us.”

The model's French was right. What it missed was something no model ships with: your terminology, your tone, your translation history. A human flagged it, weeks late — and humans are still your only reliable quality instrument. So review spreadsheets appear, and bilingual colleagues start checking the output.

Each feedback cluster becomes a system, and with an AI agent, those systems go up fast — a glossary matcher, a style-guide injector, and a retrieval index can all exist by Friday.

What is even harder: the job attached to each system. The calendar from here on is set by feedback loops — and quality depends on how long it takes a human to notice something is off — not by how fast the code gets written.

Let's take a closer look at the four support systems and what each requires.

Terminology

Reviewers notice this first. A wrong term isn't just a typo; it's a brand identity crisis when “order” appears three different ways on a single screen. The technical solution — a glossary — seems simple, but don't get distracted by the file format. The engineering trap isn't storing the terms; it's using them at scale. The AI needs to match your glossary terms to the specific grammar of every sentence in real time, which simple search-and-replace logic can't handle. You have to build a robust system that manages this context for every call. The payoff is real — it's one of the biggest quality levers we've measured in our A/B tests — but the maintenance is brutal. If your glossary isn't actively curated by a human, it's a liability, automatically injecting its own errors into your entire product.

Voice

Then come the smaller, vaguer complaints: the translation sounds too formal, or the tone doesn't quite match the brand. Those preferences become style guides per locale, living documents, each needing an author. But voice prompts aren't code: you can tell AI to sound more conversational, but you can't guarantee how it will interpret it.

One customer's style rule against question-form headings made the model answer the questions instead of translating them. Deleting the rule fixed it. That's the maintenance model for every instruction you'll ever add: you can't debug rules by reading them, only by watching what the model does with them — and more instructions don't necessarily equal more control. Sometimes, it creates new problems.

Consistency

This quarter's translations contradict work you approved last quarter. What do you do?

Well, you need to make previous translations available to the model. In localization, this is called translation memory: a database of source text and its previously approved translation, along with information about where it came from and who approved it.

The system can then find exact or similar past translations and feed the most relevant examples back into the prompt.

This improves consistency, but the quality of the translation memory matters. Feed it outdated or incorrect translations, and it can repeat those mistakes at scale.

Context

Quality plateaus, for a reason that sounds too simple: an eight-word string doesn't provide enough context for the model to translate it well. Is “mole” the animal, a spy, or a mark on someone's skin? You've been pasting hints into prompts since week two; now you build the machine version — descriptions, screenshots, surrounding strings, assembled per request, automatically.

Step back. You now have a terminology system that someone needs to maintain. A voice system with authors. A memory system with hygiene obligations. A context system with retrieval.

And it shows up on the invoice: our production pipeline sends roughly fifty input tokens for every one it gets back. In dollars, that lands at 8–15× the per-word number your stakeholder's spreadsheet started with. Nobody lied. The spreadsheet priced the model.

Lesson 02. Context assembly builds each model request from a glossary, style guides, translation memory, and descriptions and screenshots pulled from the app. Reviewer corrections stop at the spreadsheet.
Act II ledger: about one and a half people, an API bill plus reviewer invoices that exceed it, and tone drifting apart across eight locales.

Act III: You learn to trust it

TL;DR. Generating translations is the easy, cheap part. Knowing if they're actually right is where the budget explodes. Once you stop manually checking every line, you're flying blind until you build an entire internal system just to grade your own AI.

A placeholder ships broken despite the checker. The check fired. What didn't exist was the answer to “now what?”

When a check fires, the open questions are whether to retry with better context, route the string to a human, or block the release.

Spotting an error is easy. But every check needs a rule for what happens when something goes wrong, and someone has to make that call.

Mechanical checks are great for catching structure errors, but they're blind to meaning. When a translation is simply wrong in a language nobody on your team reads, you need a Judge — a second AI model to grade the first. It sounds logical, but the Judge has its own failure modes. It flags correct, approved terms as errors or gives inconsistent scores. You aren't debugging a translation anymore; you're debugging the Judge.

Sadly, judging the judge is not easy. You can test structural behaviors automatically, but the real quality checks are a manual human effort.

That simple request-capture tool you skipped in week one just became the only way to answer any question you have. You've stopped building a pipeline and started building an eval suite.

Now you're running a quality operation. You set a threshold: above this score, ship; below it, a human looks. It looks like a simple quality slider, but it's actually two conflicting dials welded together. Raise the bar, and you've effectively hired a team of reviewers without a budget request. Lower it, and you're shipping risk into languages you can't verify. There's no perfect setting — only the budget line for the sampling rate. How much of your output are you willing to let ship unexamined? That is the question every localization manager needs to answer.

Lesson 03. Output passes checks and a judge, then a threshold either publishes it or sends it to reviewers. An eval harness replays captured requests.
Act III ledger: two engineers, the curator from Act II, and a reviewer budget. The stakeholder asks why grading costs more than translating.

Act IV: It becomes an institution

Nothing in this act is caused by something breaking. Your headache now is that you've built a system that needs constant attention.

Your evals will inevitably report an inconvenient fact: no single model wins everywhere. One might excel at Turkish, while another dominates your marketing tone, and they'll leapfrog each other every quarter. The solution is a multi-provider setup. You can buy the gateway (LiteLLM); that part is solved and free. But the routing policy — deciding which model handles which task — cannot be bought. That's proprietary logic based entirely on your content, and only your evals can tell you which model is actually the best fit.

Once you've built a system that smart, you have a platform. And that platform needs governance, with permissions and clear ownership, the more teams use it. Who owns terminology? Who can change style guides? Who approves updates? All these decisions need rules, shared across the company.

Once you've got this in place, you will likely experience a period of bliss. But then a model update you never asked for shifts tone in two languages and none of the others; the order string from that first Friday changes, and nobody notices for a month. Prompts start to rot against new model behavior, glossaries age, eval sets go stale.

Decay has a flipside: AI models keep getting better, and you want to benefit from those improvements. But you can't just switch to a new model and hope for the best. You need to test it against your own content, in each language, before putting it into use. Quality testing is no longer something you do once. It becomes an ongoing part of running the system.

The final ledger. People: engineers, a curator, an eval owner, and reviewers. Money: models, the judge, infrastructure, and human review. Process: the system is operated, not finished.

A fair question at this point: doesn't all of this cost the vendor exactly the same?

Yes — but they pay it once and split it across thousands of customers. That's all “buy” means: splitting the bill.

Lesson 04. The platform adds a model gateway, a routing table driven by evals, shared terminology with team overrides, and a support channel that lands on the sprint board.

The engineer who built it — and bought anyway

One line that never makes the ledger: the right to not care.

Part of what a subscription buys is being able to say “it's broken” and keep working — someone else's week absorbs it. Nobody on your team has that right anymore. When the judge misbehaves the week its engineer is away, the other one learns it live, against a deadline that doesn't know anyone's on holiday.

The best-qualified DIY builder we've met — a staff engineer at a customer — didn't ask whether it could be built, but if he was ready to own the responsibility. He'd already built a localization system in the past that worked. But knowing everything that goes into it, he would never do it again. Nothing in this article was beyond him. He just priced the job before taking it.

The build, for him, was the good part — bounded, interesting, half of it agent-written, results he was proud of. What he declined was the job: the dials, the regression rituals, the support channel, the decay that nothing announces — indefinitely. That's what buying is: the job is the product. A vendor is a team whose whole product is holding those dials.

It also comes with an apprenticeship: none of this arrives as a ready-made package the month you decide to build. You discover each system by needing it, learn to operate it by operating it, and pay tuition in production. The map you've just read took years to draw — the know-how is part of what a vendor's price includes.

The company that was right to build

The engineer is one pole of the answer. Here's the other. There's a kind of company — big enough that entering a country is a regulatory event, not just a launch — for whom localization is market entry: the app, the compliance texts, the support content, every market, forever. At that scale the math flips. The platform bill dwarfs a team's payroll, and their team delivers something tailor-made that improves at their priorities, not a vendor's roadmap. Companies like that sometimes run on a platform for years and then build their own. They're right to, and nobody should pretend otherwise — including us.

Both stories answer the same question: how far is what you need from what's on offer? If you need far less — a script and a bilingual colleague cover it — build; it's cheaper. If you need far more — tailor-made, at a scale where owning the whole loop pays for itself — build; you'll close the loop faster at your own priorities. Everyone in between buys. Including the teams that need something special: they build their special piece on top of a bought foundation. That's what the engineer did — he didn't stop building. He stopped owning.

Lesson 05. Building it is a project. Owning it is a job. The only question is whose job it should be.

Five lessons

  1. Automation scales translation. It also makes mistakes harder to see.
  2. The model knows the language. It doesn't know you.
  3. Generating translations is cheap. Trusting them is expensive.
  4. Nothing has to break for the work to continue. Maintenance is the price of building.
  5. Building it is a project. Owning it is a job. The only question is whose job it should be.

So, build or buy?

You already know for yourself.

AI Translation

Author

Sasho.webp

Engineering Manager, AI

I'm Sasho, Engineering Manager at Lokalise where I manage the AI team, which mostly means I get to argue about evaluation metrics for a living and occasionally ship something that translates the internet slightly better.

Before this I spent a bunch of years in clinical NLP, including a PhD at Sussex teaching computers to read doctors' notes — which back then took years of research and is now roughly a weekend and a good prompt. Somewhere in there I picked up a handful of papers and patents, mostly on things like evaluating medical note generation and measuring how similar two sentences really are.

These days at Lokalise I help build the AI behind translation quality — RAG-based translation, evaluation pipelines, the unglamorous plumbing that keeps the AI-harness in check. It's going fine, mostly.

I studied Computational Linguistics at the University of Tübingen before my PhD at the University of Sussex, and I've spent 14 years across NLP research and AI engineering, with stops in academia, clinical AI, and now localization.

You can also find me at sasho.io and on Google Scholar.

AI translation quality.webp

How to get data-backed proof that AI translation works — on your own content

Here's a stat that should make every localization leader uncomfortable: 57% of localization teams say their number one barrier to using AI more is that they don't trust the quality. Not that the quality is bad. That they don't trust it. The distinction matters. AI translation has improved dramatically. Models are better, context-aware systems like

Read more How to get data-backed proof that AI translation works — on your own content

Context is king

How to give AI translation tools more context: A developer's guide to deterministic l10n

You run a clean, well-structured string file through AI translation. The output is linguistically correct, but completely wrong for the UI. This is the context deficit problem. Without structured input about what a string is and where it lives, models default to general-purpose language patterns. The output looks coherent on its own, but breaks when placed in the interface. Most dev teams try to fix this problem by switching models or adjusting prompts. The translation

Read more How to give AI translation tools more context: A developer's guide to deterministic l10n

AI translation quality evaluation: Looking beyond human review

AI translation quality evaluation: Looking beyond human review

The meeting is going well. Translation costs are down, turnaround times are shorter, and AI is taking on more of the work. Then your VP asks a question: “How do you know the quality of AI translations is good enough?” The expensive part usually starts once you try to answer that question inside a system you built: the real cost of buildi

Read more AI translation quality evaluation: Looking beyond human review

Stop wasting time with manual localization tasks.

Launch global products days from now.

  • Lokalise_Arduino_logo_28732514bb (1).svg
  • mastercard_logo2.svg
  • 1273-Starbucks_logo.svg
  • 1277_Withings_logo_826d84320d (1).svg
  • Revolut_logo2.svg
  • hyuindai_logo2.svg