Translation Scoring: AI-Powered Translation Quality at Scale

Translation Scoring is the new AI feature that’s changing how teams approach translation quality β€” from instant per-string scoring right in the editor to automated workflows that scale.

 

Date: πŸ“… October 30th, 2025

Key takeaways

Takeaway 1 icon

How Translation Scoring works

A live demo of automated quality scoring in action, built on industry-standard Multidimensional Quality Metrics (MQM).

Takeaway 2 icon

Real-world use cases

How teams skip up to 80% of human reviews with AI Scoring, while maintaining quality.

Takeaway 3 icon

Implementation best practices

How to tailor quality thresholds to your own standards and build automated review workflows that scale.

Takeaway 4 icon

Tips to maximize impact

Leverage smart filtering to instantly surface only the strings that need human attention β€” and confidently publish the rest.

Speaker

Alesia_LP_new.webp

Alesia Nikalaichyk, Product Manager on AIML, Lokalise

Alesia is a Product Manager at Lokalise, with 5+ years of localization expertise. Thanks to previously having worked as a Customer Success Manager, she has a deep understanding of Enterprise customer needs, which she now applies to shaping impactful product strategies and results.

About this topic

 

Translation quality scoring uses AI to evaluate translated strings against the industry-standard Multidimensional Quality Metrics (MQM) framework, assigning each string an objective quality score. Manual translation review is one of the most time-consuming and costly parts of localization β€” automated scoring lets teams set their own quality thresholds, route only low-scoring strings to human reviewers, and confidently publish the rest at scale.

Full transcript

 

Introduction

 

Marta: Welcome, everyone. I'm Marta and I'll be moderating the webinar today. I'm joined by Alesia from the product team, and today we're going to be talking about translation quality and Translation Scoring. Here's the agenda: we'll start with an introduction to why translation quality assurance is important in your workflow. Then we'll talk about the challenge that usually comes with it β€” cost, quality, and speed. We'll show some examples of how it works, plus a demo of what we're doing at Lokalise, and some key takeaways at the end. And at the very end, we'll have time for Q&A as well.

Alesia Nikalaichyk is a Product Manager at Lokalise, building AI-based capabilities to improve translation quality, reduce manual review, and accelerate the continuous localization cycle. She has been in the localization industry for over five years, and before moving into product management she was a Customer Success Manager guiding enterprise teams on best practices β€” which has really helped her translate practical customer requirements into product strategy by partnering closely with customers. Without further ado, the mic is yours, Alesia.

Why translation quality assurance matters

 

Alesia Nikalaichyk: Thank you very much, Marta β€” it's a pleasure to have everyone here, and I'm happy to see some familiar faces among our customers. I want to start with the importance of translation quality assurance, or TQA, and there's a really good example that illustrates it. Most of you know Mercedes-Benz. They wanted to enter China, the world's second-largest luxury market, in 2009, and all they had to do was come up with the right spelling and representation of their brand name. They landed on a name that sounds phonetically like "Benz" β€” but later discovered it roughly translates as "rush to die." Not a great name for a company launching a car business in China. They fixed the mistake and changed the branding to a spelling that roughly translates as "run faster" β€” a much nicer motto. One localization mistake cost them millions of dollars.

At the same time, not all cases are equal. We're not always localizing brand names, and not every case is this risky. So today we'll talk about the importance of quality assurance, but also about the different strategies you can apply in different cases.

The challenge: cost, quality, and speed

 

Before we jump in, I asked the audience: what is your biggest challenge when it comes to translation quality assurance at scale? Looking at the poll responses, different themes pop up. Some are centered around time β€” review cycles taking too long. Some relate to quality β€” it's hard to assess or maintain consistent quality. And some are about cost. The responses are fairly balanced across all three.

It can be very hard to balance these three β€” maybe even impossible. You can have really high quality by investing a lot of money in human review, but that impacts speed. Or you can optimize for fast, affordable translations without much human review, which can impact quality. You can usually pick two out of three. This is where scoring comes in: it tackles all three areas β€” cost, quality, and speed β€” one by one.

Cost: from reviewing everything to reviewing what matters

 

Our second poll asked how teams currently handle translation quality control, and the answers ranged from reviewing everything to reviewing nothing. Both can be good strategies depending on the case. So which strategy works best?

Think of it as a graph. On the vertical axis is the risk of not reviewing content β€” which depends on how many people will see it, whether you're in a regulated industry, how visible the content is, and how easy a mistake is to revert. On the horizontal axis is content visibility, from low-visibility content that only a few users will see β€” internal documentation, FAQs β€” to highly visible content that most customers see. The Mercedes brand-name case sits in the top-right corner: highly visible, very high risk of not reviewing. But even within one company, different assets carry different degrees of risk. At the other extreme, content that's almost invisible and very low risk β€” like internal documentation β€” might be perfectly reasonable not to review. There are also highly visible but low-risk cases, like a funny social media post, and moderate-risk but low-visibility cases, like disclaimers or emails. Scoring fits best in that middle ground: it gives you assurance that translations are checked and the risk of issues is low, while optimizing human review. Depending on your use case β€” highly regulated finance versus entertainment content β€” the sizes of these boxes will differ, but in both cases scoring helps you figure out the right strategy and balance effort with risk.

Now look at it from a different perspective: acceptance rate, which shows how often people accept translations done by others. Quality is a tricky, subjective topic β€” we're all human and have different preferences. Some prefer short, concise copy; some prefer a more formal or informal tone. From past research, humans tend to agree with each other roughly 90% of the time, so that's what we treat as human-level quality: translations performed by humans or AI that are accepted by humans in 90% of cases. Depending on the approach β€” translating with LLMs without context, with context, or with AI customized on your past translation examples β€” you reach different acceptance rates. In successful scenarios of 75%+, around 80% of your content is never edited by humans, meaning it's already good.

Combine the two graphs and the message is: first, not all content carries the same risk, and for some content it's okay not to review everything. Second, for the content you do review β€” if you approach it properly with context β€” edits are only needed for about 20% of it, because the rest of the translations are already good. This is where scoring saves money and effort: if a human reviews 100% of content, you pay the full cost. Scoring surfaces the lower-scoring translations β€” the ones with some degree of doubt, roughly 20% β€” because in the rest of the cases AI failed to find any issue or ambiguity. By focusing human effort on the 20% of the most ambiguous and challenging segments that are likely to have issues, you can skip review for the rest β€” up to 80% less to process β€” while AI still gives you assurance.

Quality: measurable confidence with MQM

 

The second dimension of the triangle is quality. In our UI you can see the scoring in action: a score of 95, shown in green, with five points deducted for a minor stylistic issue, and guidance explaining that the phrase is a little more formal and less energetic than the source.

We base our scoring on MQM β€” Multidimensional Quality Metrics. If you're shown two translations, you might say one sounds a bit odd, but how do you actually quantify that and put it on a scale? That's what MQM does: it provides the rules for categorizing issues and building that scale.

First, it gives us interpretability, because it defines issue types and penalties β€” and alongside the score you also see the reasoning, the comment AI gives on what might be wrong. We found that if you show users just translations and scores, it's very hard to start believing the scoring really works. As soon as you see the reasoning, it becomes much more obvious why the AI thinks this way.

Second, MQM gives us severity levels that help quantify issues. There are three: critical, major, and minor. Critical issues are where words are missing or the meaning is changed β€” the source and target don't convey the same meaning. Major issues are also significant: they can impact grammar and readability. Minor issues are nuanced, preferential issues β€” small quirks like unnatural phrasing, or preferences for a more or less formal tone of voice.

Third, the MQM framework defines an error typology, which helps align the team: you immediately see what type of issue there is, and you can decide which types matter to you. If you know you don't care about certain categories β€” say, style or fluency β€” that guidance tells you which issues are important for you and which can be skipped during review.

We put all those categorized issues together and rank quality on a scale from 0 to 100. One hundred means AI failed to find any issue β€” in the AI's world, it's perfect. Then we deduct points depending on the type of issue found: 75 for critical, 25 for major, and 5 for minor. We also apply color coding: if there are only minor, highly preferential issues, scores usually land between 80 and 100 and show green. Any major or critical issue drives the score to 75 or below β€” below the 80 threshold β€” and that's where we suggest customers route content to human review.

Scoring in practice: three examples

 

Problem one: design breaks. If you're localizing UI strings, string overflow can be a really bad issue β€” a string that doesn't fit the screen leads to poor UI, upset teams, wasted time, and rework. Why do LLMs make these mistakes? Imagine being asked to describe yourself in ten to fifteen words β€” easy. Now imagine you must do it in exactly 170 characters. You can approximate how talkative to be, but you can catch yourself mid-sentence realizing you either have to cut a word short or overshoot the limit. That's what can happen to AI when translating. A human fixes it by rereading, counting characters, and finding a shorter translation β€” and that's exactly what scoring does. If AI overshoots the limit during translation, scoring catches it at the scoring stage: there's a simple check against the character limit. If the translation complies, it can be automatically approved; if the limit is exceeded, it's flagged as a critical design issue with a score below 80 and routed for review.

Problem two: not respecting the glossary. Companies create glossary terms to make translations sound exactly the way they want, to meet legal requirements, or to keep feature names consistent. Ignoring glossary terms can lead to legal issues, inconsistency, and customers losing the meaning. As a concrete example, I added the term "deposit" to the glossary with a required Latvian translation. Then we translated the source "Open a new deposit" into Latvian. I inserted a translation myself that ignored the glossary β€” it's actually quite tricky to force AI to make mistakes β€” and asked AI to score it. Scoring flagged an error saying the word "deposit" must be translated according to the glossary. So AI scoring not only learns from your glossary assets, it actively catches these issues. When I revised the translation using the glossary term β€” adjusted to the correct accusative case ending β€” and rescored it, we got a perfect score.

Problem three: style guides and tone of voice. Some companies invest a lot of effort in making their content sound the way they want β€” super formal, or warm and friendly for a young audience. An inconsistent voice leads to a confusing, off-brand experience. For this example we created a very simple style guide saying we're translating for a fintech app targeting a young audience, using a playful, conversational tone. The source was: "Oops! The unicorn tripped. Check your internet and retry." For those who don't speak Russian, there are two forms of "you" in Russian β€” formal and informal β€” with different verb endings. Scoring flagged an issue and gave the translation a 75, explaining that it used the formal form, which doesn't match the casual, conversational tone required by the style guide. So it actually reads your context and adapts to your preferences. Once I changed the endings to the informal form, the score improved to a green 95.

Why quality is subjective β€” and when it isn't

 

I also want you to feel how challenging this task is from the other side. We showed the audience three translations of the same friendly delivery-delay message β€” "Thanks for your order, it's running a bit late, more details will follow" β€” and asked which they preferred. Votes split mostly between options one and three, with some preferring option two. That's perfectly normal: all three convey the same meaning but differ slightly. The first is reassuring β€” no action needed. The second is more specific β€” a new estimate of Friday, October 24. The third is empowering β€” you can check your delivery options here. There are preference-level nuances each of us has, and that's healthy variability.

Then we asked the audience to score a translation pair where the meaning actually changed: the source promised free returns within fourteen days, while the target said there are simply no returns. Almost everyone gave it a thumbs down. We ran one more, even more obvious example, and again saw strong agreement that it was a bad translation β€” and that kind of mistake is serious, because misguiding users can create compliance or safety risks.

The point of the exercise: with friendly-tone preferences, there's variance in responses β€” a tone can be friendly but reassuring, or trade off against brevity β€” and the same is true for AI. Minor issues are about preferences, and it's okay to have healthy variability there; a 90 or 95 score can be a perfectly good translation depending on preferences. Don't strive for perfection and always aim for scores of 100. But when there are major issues that alter meaning, there's much less variance β€” both humans and LLMs spot them consistently.

Speed: automating review with workflows (demo)

 

The last dimension is speed, and this is where I'd like to do a quick demo. This is my Lokalise project, and it's empty β€” we'll do everything from scratch. I go to Workflows, click Create workflow, and define exactly how I want content translated. I pick the template that says I want my content translated by AI and later reviewed based on the score.

In the workflow, we define where it starts and the cadence: starting tomorrow, every day at 9 a.m., group all my newly added content and translate it this way. I'll pick Russian as the target language and set it to look for new and updated content to be picked up and retranslated. There's a step that wipes out any previous translations, then the new and updated content is translated with AI β€” and we can provide additional instructions and context to the AI. Scoring is automatically enabled for this task, so the translation stage returns both translations and scores.

In the review stage, we select the threshold β€” that's the main setting, defining which content gets human review. If you're very cautious, you can review everything, or everything below 90 or 95. But based on what we shared earlier, if you're okay with minor preferential issues, the suggested threshold is 80: any translation with at least one major or critical issue β€” anything that alters meaning or violates character limits β€” will be spotted and checked. There are extra settings for reviewer due dates, and we pick the reviewers β€” in this case, just myself. Then I activate the workflow. It's live, waiting for content and for its scheduled run time.

Next I upload a file with thirteen keys, and shortly we see them populated in the editor β€” it's also a fintech app. Rather than wait until tomorrow 9 a.m., I trigger the workflow manually β€” but normally it runs automatically without human involvement. The workflow starts with the AI translation step: under Tasks we can see a running AI translation task translating thirteen keys into Russian. Translations coming through AI translation tasks are scored automatically β€” and for any translation without a score, there's an icon in the editor to score it on demand.

It works like a charm: the translations populate along with their scores, and we can click on the scores to see which translations are perfect and which have issues. Checking Tasks again, a review task has been created. Compare the two: the AI translation task had thirteen keys and 77 words, but the human review task for Russian has much less β€” 48 words. Clicking through, all of those translations have scores below 80, exactly as configured in the workflow. That's how you save effort by reviewing only what matters. Reviewers can accept a translation as is, or fix it β€” and after changing it, they can rescore and see the new score (it can take a moment, since we also ask the AI to provide reasoning). Once the human confirms each segment β€” or confirms in bulk β€” the workflow completes and translations are ready for publishing.

People can also agree or disagree with scoring. As we showed earlier, it's okay to disagree, especially on minor issues. There are also cases where humans have context the AI doesn't β€” maybe the character limit was exceeded, but any shorter translation would harm quality, so the team decides to keep the longer translation and expand the UI instead.

Key takeaways

 

We started with translation quality assurance and its challenges, and looked at it from three perspectives. On the cost side, scoring helps you move from reviewing everything to reviewing what matters. On the quality side, it's context-aware β€” of your terminology and style β€” and gives you measurable confidence in quality. LLMs try really hard to find issues, so if they fail to find any, it's usually a good translation, or at least there's no ambiguity or change in meaning. And in terms of speed, it improves productivity: it can cut review cycles, or if you still review everything, it gives your reviewers insights into what might need fixing.

These features work together, so we suggest trying the holistic AI experience: AI translations powered by AI scoring, quality improved with style guides and glossaries, and workflows bringing the speed.

Marta: Thank you, Alesia. Before we move to Q&A: last month we had a webinar about custom AI profiles and customizing AI with our engineering manager, Sasha. And next week we're launching a new series called Shipped by Lokalise, where we'll be showing this Translation Scoring feature and others in action. Any questions we don't get to now, we'll answer by email after the session.

Q&A

 

Q: Can I empower my freelance translators to run the QA on their own and make adjustments?
If the question is about scoring translations performed by humans or freelancers: that's not part of the automated workflow we showed, where translations are automatically generated, scored, and routed for review. But you can ask your linguists β€” or do it yourself β€” to trigger scoring in the editor one by one. Next to each translation there's a scoring icon; if there's no score yet, you can score that translation manually and see the scoring guidance for inspiration.

Q: Is it possible to adjust the severity of issues ourselves for branded content, where preferential changes are important?
It's not possible right now to include or exclude specific categories or change their severity. What you can change is the threshold: if you want to check translations that have any minor issues, raise the threshold in your workflow and review everything below a score of 100. We're closely monitoring customer feedback on this topic, so please do share your feedback after trying the feature.

Q: We're not yet using Lokalise β€” I tried it yesterday but got a bit lost in all the features. Can I get a personal demo or consulting?
We suggest reaching out to our support team in the chat β€” they can help directly or connect you with one of our sales people who can deep-dive into your use case and suggest best practices. We'll also be sending follow-up emails after this webinar, so we'll get in touch with you for sure.

Q: How do we identify where the 20% of content for review is?
We ask the AI to check each translation and look for issues against the MQM framework and your own context β€” does the meaning match, does the style match the tone of voice defined in your style guide, are the glossary terms respected β€” so it's really personalized. After many months of testing and iteration, we've found the AI tries very hard to find issues; it may even flag a potential risk or note that the source itself is unclear and worth double-checking. So if AI genuinely fails to find any issue, it's usually a good translation, because the models are specifically directed to hunt for issues. From our internal tests and customer feedback, roughly 20–30% of content ends up with scores below 80 and gets routed for review. It also depends on your assets: the more specific and demanding your style guides and glossaries are, the higher the chance that not every requirement can be satisfied β€” for example, asking for a very friendly tone while inserting a very formal glossary term, or asking for creativity within very short character limits. Loosen your requirements and you'll see less content sent for human review; tighten them β€” especially with conflicting requirements β€” and you'll see slightly more issues flagged.

Q: Does the scoring learn over time based on our actions?
Very good question. By default, the AI takes context from your style guide and glossary. But scoring also works in combination with our feature for customizing AI on your past translations β€” the one my colleague Sasha covered in the previous webinar. If you use them together, scoring becomes aware of your past translations and preferences. The more you translate and the more edits you make, the more our system recognizes them, and scoring becomes more forgiving of those patterns because it understands how you prefer content to be translated. I suggest checking out the previous webinar and signing up for the custom AI profiles beta to try the two features together.

Q: I was told that with custom MT models, style guides won't play a role anymore. So where does the style guide used by the scoring feature come from?
There are two cases. If you use scoring without customized translations and there's no other context, the style guide is very useful: it defines the rules for how content should be translated β€” friendly versus formal tone, and so on β€” and scoring uses that same context. The AI uses the style guide both during translation and during the scoring stage. But if you customize the AI on your past translations, that customization can be superior to the style guide: instead of just following written rules, the AI sees precise examples of how you want things translated. In that case, scoring uses exactly the same context β€” it sees the past examples and takes those preferences into account. So it's either/or: use translations with a style guide and scoring will check against the style guide, or customize AI on your past translations and scoring will digest and apply the same style from those translations.

Marta: Perfect. Thank you very much, Alesia, and thank you to everyone who joined. We have a few questions that weren't answered, but due to time we're going to wrap up here and send you the answers by email. Thank you, and see you in the next webinar!

Ready to see Lokalise in action?

Start your free trial or talk to our team today.