AI Beyond Translation: Trends Shaping the Future of Localization

How AI is reshaping localization workflows and roles β€” context-aware translation, post-editing, and the rise of custom AI models β€” as teams face 5x content growth over three years with flat headcount.

 

Date: πŸ“… April 29th, 2025 πŸ• 11am EDT | 5pm CEST

Key takeaways

Takeaway 1 icon

Beyond generic AI models

The rise of custom AI models built for localization, not generic translation.

Takeaway 2 icon

β€œTranslator 2.0”

How the translator role is evolving into reviewer, prompt engineer, and quality strategist.

Takeaway 3 icon

Context-aware AI

Why context-aware AI matters for output that stays true to your brand.

Takeaway 4 icon

Cutting costs, not quality

Strategies to reduce localization costs without compromising on quality.

Takeaway 5 icon

Post-editing as a scaling skill

Why post-editing is becoming a critical skill for scaling localization output.

Speaker

Adam_LP.webp

Adam Ε oltys, Senior Lead Product Manager, Lokalise

Adam is senior lead product manager at Lokalise, focused on AI translation and helping companies expand into new markets with AI. He's spoken extensively with customers about AI over the past several years.

About this topic

AI in localization is expanding beyond raw machine translation to include custom, brand-tuned AI models, automated quality scoring, and post-editing workflows, where a human reviews and refines AI output rather than translating from scratch. This shift is partly driven by localization teams facing sharply rising content volumes without proportional increases in headcount, pushing teams to rely more on AI-assisted workflows to keep pace.

Full transcript

Host: Today with me is Adam. We are here to discuss the future of localization β€” how AI is transforming translation, and some things that aren't widely discussed yet but are shaping the future of AI in translation. Our agenda: why this topic matters today, where we are with AI translations, five trends reshaping translation, and finally how Lokalise is responding to these trends. Adam is senior lead product manager here at Lokalise. He's joined the team to work on AI, on next developments, and on helping companies expand into new markets with AI β€” he's spoken deeply with lots of customers about AI over the last few years.

Adam: Hello, everyone, it's really nice to see you all here. Before we go to the trends, I want to reflect on where we are today. Just a couple of years β€” maybe a decade β€” ago, relying on AI for really high-quality localization felt nearly impossible. Today it's a different story, because of the evolution of large language models. Their output is real-time β€” like traditional and neural machine translation, you can get a translation from an LLM immediately, at scale, and cheaply. But LLMs can also work with a wide range of context, capturing tone, intent, specific vocabulary, and learning from historical translations, which lets them mimic human translation to a real degree. So LLMs are positioned to disrupt how localization is done β€” and it's partially happening, though we're not yet seeing full-scale adoption. I think there are two blockers: providing the right context to the LLM, and setting up AI-first, AI-native workflows.

Trend 1: the translator's role is changing

Before the first trend, a quick poll: what percentage of Lokalise AI translations do you think are ready to be published without any post-editing? Answers ranged from zero to eighty percent in the chat. The real number is eighty-one percent β€” based on hundreds of thousands of AI suggestions Lokalise AI has provided over the last several months, across all customers and languages, where a human reviewer accepted the AI's suggestion without any edit.

The first trend is how the translator's role is changing. We don't see this as a battle between humans and machines, but as a symbiosis, and we see the translator's role developing in three areas. First, the translator shifts toward being a reviewer or post-editor focused on complex, high-value cases β€” long-form content, gaming UI, hero elements in marketing, critical parts of legal documents, or translations an AI has flagged as at risk of a critical quality issue (a feature like AI/LQA scoring, which we're developing). Second, there's a rise of annotators and language trainers: human translators prepare a high-quality dataset of a few hundred to a few thousand translations, which is then used as context or training data so an LLM can scale that quality knowledge to hundreds of thousands of translations. Third, there's a shift toward transcreation and creative localization β€” mainly for marketing and branding, where it's not just about fluent, accurate translation but about keeping it creative and unique; this is an area where translators can add a lot of additional value, and we already see this trend happening.

Trend 2: context is king β€” but quality of context matters most

Context is the main differentiator for LLMs, and it comes in three types: historical translations done for the same customer and similar purpose; localization assets like glossary and style guide, so the right terminology and style are respected; and context specific to a particular translation β€” for example, a description or screenshot showing where a short string like a CTA button actually appears. All three matter, but what I really want to stress is not the quantity of context, but its quality β€” feeding in accurate past translations and correct screenshots is critical to the resulting translation quality, and that's sometimes overlooked when thinking about AI translation at scale.

Another poll: what percentage of human translations get accepted, on average, without edits when reviewed by another human? Estimates in the chat ranged from twenty to eighty percent. From our tests, we estimate it's around ninety percent β€” high, but not a hundred. This varies by customer, language, context, and content type (we've seen a range of roughly eighty to ninety-nine percent), but on average, human translation quality sits around ninety percent. This tells us translations are never perfect, so AI translation shouldn't be benchmarked against a hundred percent acceptance rate, but against average human translation quality. It also shows that quality evaluation is subjective β€” the remaining ten percent aren't necessarily bad translations, just points where one reviewer sees an issue and another doesn't. This is a known phenomenon called low annotator agreement: three human translators asked to score the same translation might give three different scores.

Trend 3: custom large language models are key to human-like translation

This is the trend I believe is most critical, because custom LLMs are the key to human-like translations. An LLM can learn your past translations and preferred style, and adapt new AI-generated translations accordingly β€” like a translator aware of all your best past work. This achieves a high level of personalization, improving both quality and consistency with past translations. From our pilots with customers over the last few months, the impact on translation quality and consistency has been significant.

There are two mainstream methods for customizing an LLM: fine-tuning a mainstream model directly, or advanced retrieval-augmented generation (RAG) that's smart enough to pick the right past translations to use as context. For both β€” and for most other methods β€” what's critical is the quality of the input data. High-quality historical translations fed in for fine-tuning or as RAG context lead to high-quality AI translations. Conversely, lower-quality data can degrade performance and even make translations worse than just using a generic LLM with no training or context at all.

One more poll: what percentage of Lokalise AI translations done by custom models do you think are ready to publish without post-editing? Estimates ranged from eighty to ninety-five percent. The real number is also around ninety percent β€” from our pilot program with several enterprise customers who have quite picky QA processes, mirroring average human translation quality.

Trend 4: AI configuration is as critical as AI capability

Today you'll find a lot of AI-related features on the market β€” different models, context types, workflows. But what actually matters isn't the number of features, it's how you configure them and what data you feed them. What we're seeing is that customers, and platforms themselves, often try to fit AI into a workflow designed for a pre-AI (or MT-era) process β€” and that usually doesn't work well, because a human in that old workflow could work around flaws the AI can't. For example, a human translator who sees an obviously wrong or unrelated screenshot can recognize the mistake and go find a better image; an AI can't do that yet. Same with a mistranslated translation-memory match β€” a human translator might work around it, but an AI, told to match the past translation, won't realize it shouldn't. So AI adoption often exposes flaws that were already present in a workflow, and the resulting underwhelming quality isn't really about the technology being bad β€” it's about feeding AI the wrong data, in the wrong kind of workflow.

The real drivers of translation quality are selecting the right context, setting up the right style guide for the content type, curating high-quality data, and optimizing the workflow to balance automation with human review β€” putting AI and humans each in the right step, for example using scoring to flag which translations need review. Who drives this correct configuration? Two parties: localization platforms like Lokalise, which should guide customers in setting things up correctly for their use case, and AI-savvy localization managers who combine linguistic experience with at least basic tech/data fluency. Companies that nail this configuration β€” not just the raw technical capability, having AI-native workflows rather than treating AI as an afterthought β€” will be the ones that successfully adopt AI translation.

Trend 5: cutting cost without sacrificing quality

This trend is a culmination of the others, and it comes from a remaining challenge: over eighty percent of a high-maturity localization budget still goes to human translation and review. AI can cut those costs without sacrificing quality β€” there's been a lot of promise around this in recent years, but we're now actually at the point of delivering on it, if we nail context and configuration. There are two main levers for cost reduction. First, improving the AI translation quality itself: generic LLMs already provide roughly a seventy-five percent acceptance rate by default; using dynamic engine selection (Lokalise picking the best engine per content type and language) pushes that over eighty percent; and custom LLMs, tailored to brand, tone, and domain, push it to roughly ninety percent, matching human-like quality. Second, helping with quality evaluation at scale β€” automated scoring (AI/LQA) that assesses translation quality in real time, flagging at-risk or lower-quality translations for human review while accepting the rest as-is.

How Lokalise is responding

We're responding to these trends with several features. Dynamic engine selection: Lokalise AI already automatically selects the best-fitting LLM per language pair β€” currently the newest GPT and Claude model versions, with more LLMs coming soon. Custom large language models: we've piloted both fine-tuned and advanced-RAG models and proved we can reach human-like translation quality; we're now launching a beta program for a limited number of enterprise customers to use this at scale in production. Full context awareness: we already leverage glossary, style guide, and key descriptions for AI translations, with screenshot context coming soon. And workflow automation: we recently launched the ability to set up your own workflows and orchestrate translations, including AI, and we're soon integrating translation quality scoring into those workflows so the review step itself can be automated β€” translations can be auto-accepted based on score, or routed to human review.

Key takeaways and Q&A

Three numbers to remember: eighty-one percent of Lokalise AI translations are already accepted today without edits when reviewed; average human translation quality is around ninety percent; and our pilots proved custom LLMs can reach quality similar to that average human quality β€” human-like quality delivered at AI speed and AI cost.

Q: What domains are your custom LLMs trained on β€” are results good for SaaS?
Each custom model is trained specifically on a particular customer's data, so results are good for SaaS specifically (part of our pilot customers) as well as other industries we piloted with.

Q: How will Lokalise AI use screenshots to improve translations?
Once a screenshot is available for a translation β€” either uploaded manually or automatically pulled via an integration like Figma β€” we'll use it as context in the prompt for the LLM, telling it that the translation appears on that screen. In initial tests, this especially improves difficult, very short translations like buttons and CTAs, where context is otherwise hard to infer.

Q: How do you approach custom models if translation memory quality is unknown?
Today we work individually with customers to help them select which past translations are reliable and likely high quality, rather than assuming the whole translation memory is safe to use β€” we'd rather use a smaller set of translations with higher confidence than all available translations of unknown quality. We're looking into more automated ways to detect and curate quality at scale in the future.

Q: Is it possible to start building a custom LLM by feeding it existing translation memory for a specific project?
Yes β€” that's already part of the custom-models beta; translation memory can be used as one of the sources for training a custom model, or as context for RAG.

Q: What would an ideal AI-native workflow look like?
It starts with AI-ready translation assets β€” a style guide that's concise and focused on things AI tends to get wrong (rather than a 20-page document covering things like grammar, which LLMs already handle well), and glossaries without contradictory entries for the same term. Then the workflow itself starts with translation, followed by AI scoring; if scoring flags potential critical issues, it routes to human review to reach near-perfect quality; for high-volume, lower-criticality content using a custom model, the workflow may not need human review at all, accepting that some translations may have minor nuances rather than critical issues.

Q: How does cost compare between machine translation, human translation, and AI translation?
Human translation is significantly more expensive than either machine or AI translation. For AI translation, it depends: a vanilla AI translation without much context is close in cost to machine translation, but cost increases with the amount of context provided β€” so it varies.

The session closed with a reminder that attendees would receive a follow-up email with a trends report to download and share, and any unanswered questions would be answered directly by email.

Ready to see Lokalise in action?

Start your free trial or talk to our team today.