AI translation quality, measured: Prove it on your content
Learn how to measure AI translation quality on your own content — in 15 minutes, no external tools needed.
Date: 📅 Available on-demand from August 26th, 2026
Key takeaways
Why teams don't trust AI yet
57% of localization teams cite a lack of confidence in AI quality as their biggest barrier to using AI more.
How Proof of Value works
A live product demo comparing Custom AI Profiles, base AI, and standard machine translation side by side.
Reading the metrics
What the quality metrics actually mean — and how to interpret them without a linguistics background.
A real result: 49% vs. 13%
How a European fintech team achieved a 49% perfect match with a Custom AI Profile, versus 13% with generic AI.
No new data required
How to use your existing reviewed translations as a starting point for evaluation — no additional setup needed.
Why teams don't trust AI yet
57% of localization teams cite a lack of confidence in AI quality as their biggest barrier to using AI more.
How Proof of Value works
A live product demo comparing Custom AI Profiles, base AI, and standard machine translation side by side.
Reading the metrics
What the quality metrics actually mean — and how to interpret them without a linguistics background.
A real result: 49% vs. 13%
How a European fintech team achieved a 49% perfect match with a Custom AI Profile, versus 13% with generic AI.
No new data required
How to use your existing reviewed translations as a starting point for evaluation — no additional setup needed.
Speakers

Alesia Nikalaichyk, Product Manager, Lokalise
Alesia is a Product Manager at Lokalise, with 5+ years of localization expertise. Thanks to previously having worked as a Customer Success Manager, she has a deep understanding of Enterprise customer needs, which she now applies to shaping impactful product strategies and results.

Sasho Savkov, Engineering Manager, Lokalise
Sasho is an Engineering Manager at Lokalise, where he leads the AI translation effort. He brings more than 15 years of experience in AI across academia, business sustainability, and healthtech, and now applies his expertise to the localization industry. Sasho combines deep technical knowledge with a passion for innovation.

Alesia Nikalaichyk, Product Manager, Lokalise
Alesia is a Product Manager at Lokalise, with 5+ years of localization expertise. Thanks to previously having worked as a Customer Success Manager, she has a deep understanding of Enterprise customer needs, which she now applies to shaping impactful product strategies and results.

Sasho Savkov, Engineering Manager, Lokalise
Sasho is an Engineering Manager at Lokalise, where he leads the AI translation effort. He brings more than 15 years of experience in AI across academia, business sustainability, and healthtech, and now applies his expertise to the localization industry. Sasho combines deep technical knowledge with a passion for innovation.
About this topic
AI translation quality evaluation is the process of measuring how closely machine-generated translations match human-reviewed quality benchmarks, using metrics run against a company's own content and language pairs rather than generic test sets. It gives localization teams objective, data-backed evidence for whether AI translation is reliable enough to use at scale, without requiring in-house linguistic expertise to interpret the results.
Full transcript
Marta Puerto (host): Hi everyone, and thanks for watching. I'm Marta, Senior Product Marketing Manager at Lokalise, and I'm excited to host this session on how to get data-backed proof of AI translation quality on your own content. Today we'll look at the AI trust gap we're seeing across the market and our own customers — for enterprise teams deciding whether to expand AI, justify the investment, or move away from human-review-only workflows, gut instinct isn't enough anymore. We'll show how Lokalise is closing that trust gap with a live demo of our latest release, then go into a technical deep dive on how it works under the hood, plus some real results and use cases. Joining me are two excellent speakers: Alesia from our product team and Sasho from our engineering team, who'll show you how to get that data-backed proof so you can trust AI translations with confidence.
The AI trust gap
57% of localization teams say their number one barrier to using AI more is that they don't trust the quality. AI has raised the ambition for global companies — more content, more content types, more markets, and faster shipping — to the point where the question has shifted from “can AI translate?” to “can we trust this enough to ship?” Right now, most companies are left with three imperfect options: trust AI blindly and hope for the best, pay for an external linguistic quality audit that can cost thousands of dollars and take two to three weeks, or avoid AI and fall behind.
The AI trust ladder
Lokalise's answer is what we call the AI trust ladder — four steps that build confidence in AI translation quality.
Context: Does the AI actually know your brand, or is it starting from zero every time? Without a strong context layer — a glossary, a style guide, past approved translations — AI output will always sound flat. Readiness starts with feeding AI your own accumulated knowledge instead of treating every translation as a fresh guess.
Proof: Can you verify that AI works on your own content, instead of relying on a generic benchmark? This is where most companies get stuck — they either skip the proof entirely, roll out on faith, or lean on a benchmark that doesn't reflect their reality. Real readiness means testing AI against your own historical, approved standard and getting a measurable answer.
Transparency: When you get a quality result, can you understand why? A good score alone doesn't build trust — it just relocates the doubt. Teams trust systems they can inspect: the reasoning, the comparisons, the edge cases.
Ownership: Is someone watching for drift and re-proving value as usage grows? This is the step most organizations skip, and it's exactly where governance tends to lack capability.
Introducing Custom AI Profiles
Alesia Nikalaichyk: Hi, I'm Alesia from the product team, and I'll walk you through how we put these steps into practice — starting with context. Our Custom AI Profiles feature lets you customize your translations so they no longer sound flat, but sound more like you. Think of a generic LLM as a junior translator starting from scratch, with no knowledge of your product, style, or preferences. Custom AI translations, powered by RAG (Retrieval Augmented Generation), work more like a senior translator who's worked with you for years — the AI leverages past translation examples and follows your prescribed terminology and tone.
In a Polish example, a generic LLM translated “Create a support ticket” literally as “Utwórz bilet wsparcia” — technically correct, but not the customer's preferred phrasing. If the customer's past translations consistently used “zgłoszenie do pomocy technicznej” instead, the AI picks up on that pattern and reuses it going forward, the way a senior translator would apply your known preferences.
RAG retrieves past translations and glossary terms from memory, and it stays current automatically: the more you translate, the more content becomes available as context, without the costly, slow retraining cycle that fine-tuning would require.
Introducing Proof of Value
Once context is in place, the natural next question is: can you prove it's working better, and how? That's why we built Proof of Value. Customers told us they didn't believe AI works for their own content — creative or marketing copy, transcreation versus literal translation, or content they'd invested heavily in reviewing — and weren't sure how AI compares to human translation, even when they'd tried customization elsewhere. Proof of Value addresses that directly, showing the impact AI is making on your content and giving you a way to bring that transparency to your stakeholders.
Live demo: setting up and running an evaluation
In the AI Profiles section, you can see the available profiles — starting with the base profile, a generic LLM translation that can use glossaries and style guides but no past translation examples, and an AI profile based on your translation memory (TM). You configure a profile by naming it, choosing LLM routing (or a specific model if your company requires one), and selecting which languages' translation examples the AI should see — since TM quality can vary, and low-quality examples can do more harm than good. You can also filter by date, to exclude older translations from a previous vendor, and choose which project or projects the profile applies to.
Once a profile is customized, its impact is easy to see but hard to quantify — which is what AI Profile Evaluation solves. Evaluation runs on your own data: we take a high-quality, human-approved set of your translations, retranslate the source text three ways — with your custom profile (using your existing translations as context), a base profile (glossary and style guide only, no translation examples), and Google Translate — and compare each result back to your original translation. You select which translations to trust for this comparison, either by a specific tag or by review status, so translation memory noise doesn't skew the result. The whole evaluation takes just a few minutes, with a notification once it's done.
In the demo, the custom AI profile produced at least 13% more edit-free (perfect match) translations than the base profile, with 34% of its translations exactly matching the human reference. Looking deeper, the custom profile also showed up to 31% less editing effort — meaning fewer changes would be needed to bring the AI translation in line with the reviewed reference. For some translations, especially longer or more nuanced ones, the custom profile matched the reference exactly, with a translation edit rate of zero, while the other profiles required more edits to reach the same result.
If your translation memory mixes very different content types — for example, creative marketing copy alongside stricter software strings — you can create additional profiles that use only translations with a specific tag or review status as context, instead of the full TM. This gives you more granular control over exactly what the AI learns from for each use case; the evaluation flow for these profiles is the same, minus the step of selecting a separate high-quality set, since the tagged or reviewed content already serves that purpose.
Technical deep dive: how the evaluation works
Sasho Savkov: Hello, I'm Sasho, and I'll walk through the technical side of this experiment. Customers often ask why our AI performs the way it does, and why a general-purpose LLM wouldn't do the same job — so we built a way to compare an AI profile with a lot of context, like our Custom AI Profiles, against less context-aware options like Google Translate and a standard LLM, using your own data rather than a public dataset. That matters because it shows whether RAG actually helps on your project, not just in a research paper.
To run the experiment, we split your high-quality data into a reference set and a context set. The reference set is the “exam” every model is tested against; the context set is what the RAG model can draw on as relevant past translations. The two sets never share the same key, though duplicate translations across them are allowed — and expected, since that mirrors how RAG works in production: if a translation exists in your translation memory and a similar one comes up for translation, RAG retrieves it as context. A translation with no duplicate never has its answer available to the RAG model as context.
Splitting works differently depending on the data source. Tag-based and review-status-based sets are split directly and need a minimum of 500 keys, since smaller sets risk too much noise to be meaningful. Translation memory works differently, since it spans multiple projects: the reference set comes from a single project's TM (minimum 200 keys), while the context set draws from the full translation memory (minimum 500 keys).
The four quality metrics
We use four metrics to compare candidate translations against your reference translations, because each tells a slightly different story. Perfect match is the simplest — the percentage of candidate translations that exactly match the reference. It's a strict metric: even a single differing word means it doesn't count, so scores won't always be very high, especially on longer sentences. BLEU and chrF are more forgiving, industry-standard metrics that compare surface form using sub-word units rather than whole words. Translation edit rate measures the effort a translator would need to turn the candidate into the reference — zero for a perfect match, low for a one-word difference, and it can exceed 1 for a very poor candidate, which is expected, not a bug.
Reading the results — and what this analysis can't tell you
The goal of this evaluation is to rank the models against each other, not to grade quality in absolute terms — there's no single right way to interpret a score, since it depends on the language, the metric, and the data involved. What matters is which model performs best on your content, and by how much.
Data quality matters more than anything else here, in two ways. First, low-quality or inconsistent context data teaches the RAG model bad examples, which narrows the gap between the custom and base profiles for the wrong reasons. Second, a flawed reference set acts as an unreliable judge: if your “trusted” translations contain errors, the evaluation will penalize models for producing the correct translation instead of the wrong one your data implies.
There are also three things this analysis can't do. It has a cold-start problem — you need historical, good-quality translations before you can run it at all. It can't tell you how much you'll save by choosing one model over another in the real world. And it's a snapshot: results reflect your data at the moment of evaluation, and can shift as your translation memory grows or changes.
Ownership: Proof of Value vs. Analytics
The last step on the trust ladder is ownership — proving the impact to your internal stakeholders on an ongoing basis. Lokalise has two features for this, and it's important to know when to use which. Proof of Value gives you a snapshot comparison between AI types at a fixed point in time, using the exact same source content, context, and models for each — it's the only way to get an apples-to-apples answer on which AI approach works best for you, but it's not something you'd run continuously, since generating three sets of AI translations isn't free or scalable to do all the time. Analytics, by contrast, shows the live performance of whichever AI you've deployed in production — useful for tracking cost and effort over time, but not a reliable way to compare AI types, since production conditions (glossaries, available TM, linguists, instructions) change from one point in time to the next.
A real customer result
A European fintech company ran an evaluation on German and Italian engagement communication content and saw one of the largest deltas we've observed between the custom and base profiles. When we asked what made the difference so large, the customer explained that this was their most personalized, culturally nuanced content — content that went through a three-cycle human review process. The more personalized and heavily edited your content history is, the more nuance the AI has to learn from, and the bigger the impact of customization tends to be. Customers who've run the evaluation have told us the results were definitive enough that it “would be silly not to use” the custom profile, that it noticeably reduced repeat edits linguists were making, and that AI responses got visibly closer to their preferred style over time.
Closing and next steps
Marta Puerto: Thank you, Alesia, and thank you, Sasho, for the deep dive. As a reminder, this feature is available from August onwards on our Enterprise and Elite plans, and on the Advanced plan with some limitations. Go ahead and test it on your own AI profiles — click Evaluate and see how it works — and reach out to your CSM any time with questions. See you at the next webinar!
Ready to see Lokalise in action?
Start your free trial or talk to our team today.
Case studies

Behind the scenes of localization with one of Europe’s leading digital health providers
Read more Case studiesSupport
Company
Localization workflow for your web and mobile apps, games and digital content.
©2017-2026
All Rights Reserved.