Language technology & AI
AI-ready terminology: why machine translation and GenAI need a termbase
Machine translation and large language models produce remarkably fluent text. But fluency is not accuracy: an engine will happily pick a plausible synonym for your product feature, translate a brand name, or use a term your legal team banned years ago.
Why AI amplifies terminology problems
General-purpose engines learn from general-purpose language. They do not know that your company says “sign in” and never “log on”, or that a regulated term must appear exactly as approved. And because AI produces content at scale, one wrong term is repeated across thousands of segments before anyone notices.
There is a second effect: if your translation memories and style guides already contain competing variants, AI systems trained or prompted with that data reproduce the inconsistency. Garbage in, fluent garbage out.
What “AI-ready” terminology means
An AI-ready termbase is one that machines can use as reliably as people:
- Concept-oriented, with one entry per concept across all languages.
- Clear usage status: approved terms, and forbidden variants that tools can flag or block.
- Definitions and context, so models and humans can disambiguate terms.
- Governed metadata such as domain and product line, to apply the right terms in the right content.
- Machine-readable export in standard formats such as TBX or CSV.
- Ownership, so the data stays current as products and messaging change.
If that sounds like the definition of a well-run termbase, it is. AI does not change what good terminology looks like; it raises the cost of not having it. (New to the topic? Start with what terminology management is.)
Where terminology plugs into the AI stack
- Machine translation glossaries. Many MT engines and translation management systems accept glossaries that force approved target terms.
- LLM prompts and retrieval. Supplying approved terms and definitions in the model’s context (directly in prompts or through retrieval) steers generated and translated content toward your terminology.
- GenAI assistants. Internal linguistic resources can be turned into an assistant, for example with Microsoft Copilot. Ask “What is the difference between click-through rate and click-to-open rate?” and it answers from the company termbase; ask “Should I translate job titles into French?” and it quotes the French style guide.
- Controlled authoring. Tools such as Acrolinx and Congree check terminology at the source, before content reaches any engine.
- Translation tools and QA. CAT tools and TMS term recognition and QA checks catch remaining deviations after translation.
How to get started
- Audit the terminology sources your MT and AI tools already use, intentionally or not.
- Prioritize high-impact terms: product names, UI labels, regulated and brand terms.
- Clean up duplicates and mark forbidden variants (see building a termbase).
- Connect the termbase to your MT engine glossaries and AI workflows.
- Measure how often post-editors still correct terminology, and feed the findings back into the termbase.
With AI reshaping the language industry, staying technology-agnostic matters more than ever. Engines will change; a clean, governed termbase is the asset you keep.
Frequently asked questions
How do you make machine translation use the right terminology?
Export approved terms from your termbase into the glossary feature of your MT engine or TMS, mark forbidden variants, and check terminology in QA. Keep the glossary synchronized with the termbase rather than maintaining it separately.
Can large language models follow a glossary?
LLMs follow terminology much more reliably when approved terms and definitions are supplied in their context, through prompts or retrieval, and when outputs are checked against the termbase.
Is terminology management still needed with AI translation?
More than before. AI produces content faster and in more places, so inconsistent terminology spreads faster too. A governed termbase is what keeps AI output on brand and compliant.
