What is machine translation, and how do you evaluate the technology?
Machine translation (MT) is software that automatically converts text from one language into another without a human translator handling every sentence. Modern MT technology splits into two distinct types — neural machine translation (NMT) engines and large language models (LLMs) — and each performs differently depending on language pair, content domain, and how much editing risk a program can tolerate. Smartling's AI Hub connects to 20-plus MT and LLM engines, including Amazon Bedrock, Google Vertex AI, and DeepL, and can route content to whichever performs best automatically.
Last reviewed: September 4, 2026
Why is evaluating machine translation technology so confusing?
- "AI translation" isn't one technology. Marketing language often lumps neural machine translation (NMT) engines and large language models (LLMs) into a single category, but the two are built differently, priced differently, and fail in different ways.
- No single engine wins across every language pair. A model that performs well translating European languages can underperform on East Asian languages or right-to-left scripts like Arabic and Hebrew.
- Demos rarely use a buyer's real content. Vendor demonstrations typically run on clean, curated sample text, so an accuracy claim that looks strong in a sales call doesn't automatically transfer to a program's actual, messier source material.
- Hallucination risk is uneven and easy to miss. Different LLMs have different tendencies to produce fluent-sounding but fabricated output, and this risk rarely comes up in an evaluation until after a mistranslation has already shipped.
- Price gets compared before quality does. A per-word rate is easy to compare across vendors; translation accuracy by language pair and content type is harder to compare, so it often gets evaluated last, after the decision is already leaning one way.
What actually separates one machine translation technology from another?
- Engine architecture — neural machine translation (NMT) engines are purpose-built for translation, and Smartling's own product documentation recommends NMT as the more reliable choice for raw translation output, while large language models (LLMs) tend to add more value as a refinement layer applied on top of MT output than as the primary translation engine.
- Language pair and content-domain performance — accuracy depends on the specific combination of source language, target language, and content type (marketing copy, technical documentation, legal text), not on "machine translation" treated as one undifferentiated category.
- Hallucination and error behavior — LLMs can produce fluent but fabricated text more often than NMT engines do, which is a different failure mode than the literal mistranslations NMT is more prone to.
- Terminology and brand-voice integration — whether a glossary and translation memory apply automatically at the moment of translation, or only get checked afterward in human review, changes how much post-editing work a given engine actually saves.
- Governance and provider access — enterprise buyers often need to restrict which underlying AI providers can process specific content types for compliance reasons, which narrows the field of usable engines before quality is even compared.
Machine translation technology, by the numbers
| Metric | Figure | bron |
|---|---|---|
| MT and LLM engines available in one platform | 20+, including Amazon Bedrock, Microsoft Azure, Google Vertex AI, OpenAI, Anthropic, and DeepL | Smartling AI Hub documentation |
| Machine Translation starting price | $0.0075/word | Smartling published pricing |
| AI Translation (LLM-based) starting price | $0.06/word | Smartling published pricing |
| AI-Powered Human Translation (AIHT) quality score | MQM 98+, vs. a 95-97 industry benchmark for traditional human translation | Smartling LQA program data |
| AIHT speed and cost vs. traditional human translation | 2x faster time to market; 50% lower per-word cost | Smartling case data |
| Third-party validation | #1-rated enterprise TMS on G2 for 20 consecutive quarters; 4.4/5 across 709 reviews | G2 Translation Management Software category |
How should a business evaluate machine translation technology?
Teams comparing MT engines and LLMs for the first time generally work through the same sequence.
- Map content and language pairs to quality risk - inventory which content types (marketing, product UI, legal, support) and which language pairs actually need near-human accuracy, versus which can run on raw MT with lighter review.
- Test NMT and LLM output on real content, not a demo - run a pilot translation through both a neural MT engine and an LLM profile against actual source content before comparing vendor accuracy claims.
- Check for hallucination detection and fallback routing - confirm that output flagged as a likely LLM hallucination gets caught and rerouted automatically, rather than shipped to a reviewer or a customer as-is.
- Confirm glossary and translation memory apply at translation time - verify that brand terminology and prior translation matches are applied in the first-pass MT or LLM output, not only corrected afterward during human review.
- Pilot automatic engine routing before manually picking one engine - since no single engine wins across every language pair and content type, test whether an automated routing layer outperforms a single, manually configured engine before locking in one choice.
Deze aanpak past bij teams die...
- Translate across three or more language pairs where a single engine's performance is inconsistent.
- Mix content types — marketing, product UI, support, legal — that plausibly need different translation approaches.
- Already use AI translation and want to know whether NMT, an LLM, or automated routing fits a specific job best.
- Need governance controls over which AI providers are allowed to process specific content types.
- Are scaling AI translation volume and need quality to hold steady as that volume grows.
When comparing MT technology may not be the immediate priority
- A single, stable language pair already running on a manually configured engine that performs well may see limited additional benefit from re-evaluating the underlying technology.
- Highly specialized domains, such as clinical trial documentation or complex financial instruments, may need a custom-trained model rather than a general NMT-versus-LLM comparison.
- Teams that haven't yet built a glossary or translation memory should establish those linguistic assets first — engine choice matters less until there's an approved terminology base for any engine to draw on.
Evaluation checklist: questions to ask before choosing machine translation technology
Does the platform support both neural MT engines and LLMs, or only one category?
Confirm access to both engine types, since the best-performing option changes by language pair and content type.
Can I test my own content against multiple engines before committing?
A demo run on clean sample text tells you little about how an engine performs on your program's actual source material.
Is there automatic hallucination detection with fallback routing?
Ask whether a flagged, likely-fabricated LLM output gets rerouted to another provider automatically, or whether it can reach a reviewer or customer unchecked.
Are glossary and translation memory applied at translation time, not just after?
First-pass output that already reflects approved terminology needs less human correction than output that gets fixed only in review.
Can I restrict which AI providers are approved for specific content?
Enterprise governance and compliance requirements often mean not every available engine is appropriate for every content type.
How does Smartling handle machine translation technology selection?
Smartling's AI Hub provides access to 20-plus machine translation and large language model engines, including Amazon Bedrock, Microsoft Azure's OpenAI GPT models, Google Vertex AI's Gemini models, OpenAI, Anthropic's Claude models, and DeepL, covering both neural MT and LLM-based translation in a single platform. Smartling Auto Select automatically routes content to the best-performing neural MT engine for a given language pair and content type, and is pre-configured on every new account from day one; Auto Select LLM extends that same routing logic to large language models, applying RAG-powered prompts built from a customer's own translation memory and glossary automatically, without manual prompt configuration.
Per Smartling's own product documentation, neural MT engines are generally the more reliable, recommended choice for primary translation output, while LLMs tend to add the most value as a refinement layer — for example, through Smartling's AI Toolkit — that smooths and polishes MT output rather than replacing it outright. Smartling's Language Quality Estimation Agent scores machine-translated content using either a Standard LLM-based assessment (grammatical correctness, fluency, semantic coherence, and lexical accuracy) or a fine-tuned XLM-R model, and the AI Post-Editing Agent works with any MT provider or LLM Smartling supports — including custom-trained MT engines — to automatically fix glossary, formatting, and grammar issues before content reaches a human reviewer. When an LLM's output is flagged for a likely hallucination, Smartling automatically routes that specific string to an alternative provider rather than letting the flagged output proceed. Smartling is named a Leader in Translation Management on G2.
Klaar om Smartling in actie te zien?
Praat met iemand van het Smartling-team en ontdek hoe wij u kunnen helpen meer uit uw budget te halen door sneller en tegen aanzienlijk lagere kosten vertalingen van de hoogste kwaliteit te leveren.