Guide
How to evaluate translation quality automatically
There are three main ways to evaluate translation quality automatically: reference-based metrics such as BLEU and COMET, reference-free quality estimation such as COMETKiwi, and LLM-as-a-judge with MQM error categories. They differ in what they need as input and what they can explain, and all of them evaluate text rather than the product the text appears in.
Updated October 2026
What do reference-based metrics measure?
Reference-based metrics compare a machine translation with one or more human reference translations of the same source.
BLEU scores how many word sequences (n-grams) the translation shares with the reference. It is fast and widely reported, but it rewards surface overlap. A correct paraphrase can score low, and it is most meaningful across a whole test set rather than for single sentences.
chrF works the same way at the character level, which suits languages with rich word forms.
COMET is a neural metric trained to predict human quality judgments. It reads the source, the translation and the reference, and it tracks human judgments more closely than overlap metrics.
MetricX, from Google, is another learned metric, with both reference-based and reference-free (QE) versions.
The catch is the input. These metrics need a reference translation, which you usually don’t have for new content in production.
What is reference-free quality estimation?
Quality estimation (QE) predicts translation quality from the source and the translation alone, with no reference. COMETKiwi is a widely used example. Because it needs no reference, QE can score new machine translation as it is produced, which makes it useful for triage: deciding which segments need human review.
QE returns a score, not always a reason. It still evaluates text segments, usually one at a time.
What is LLM-as-a-judge with MQM?
LLM-as-a-judge asks a large language model to evaluate a translation the way a reviewer would. In MQM-style setups such as GEMBA-MQM, the model marks error spans and assigns each one an MQM category (for example accuracy, fluency or terminology) and a severity, without needing a reference.
It gives explanations a team can act on. It also costs more per segment than a metric, its output can vary between runs, and research has found its error spans don’t always match human annotators. Like the other approaches, it usually sees only the text you send it.
How do the approaches compare?
Approach | Needs a reference? | Output | Good for | Misses |
|---|---|---|---|---|
BLEU, chrF | Yes | Overlap score | Comparing MT systems on a test set | Meaning, valid paraphrases, context |
COMET, MetricX | Yes | Learned quality score | Ranking systems and outputs against references | Content without references; UI context |
COMETKiwi (QE) | No | Quality score | Triaging new MT in production | Reasons for the score; UI context |
LLM-as-a-judge (MQM) | No | Error spans with MQM category and severity | Explainable review of text | UI context, unless the model sees the page |
LLM-as-a-judge on the rendered page (Loqalit) | No | MQM-scored findings with screenshots | Live websites and web apps | Bilingual files before delivery |
What does text-only evaluation miss?
Every approach above evaluates text: a segment, or a segment plus its source. None of them sees the user interface the text ends up in. So they can’t catch text that overflows its button, truncation at a specific screen width, layout breaks, RTL mirroring errors, or a translation that is correct as a string but wrong for the element it labels.
Loqalit applies LLM-as-a-judge to the rendered page. Its agents open your live or staging pages in a real browser, in every language, and evaluate each string where it renders. Every finding comes with a screenshot, a reason and an MQM severity. Localization QA for live websites and web apps
How do you validate machine translation quality in production?
In production you rarely have reference translations, so combine reference-free methods:
Score everything with QE to find the segments most likely to be wrong.
Review flagged and high-stakes segments with LLM-as-a-judge or human reviewers, using MQM categories so results are comparable.
Check the rendered product in every language, because some errors only exist on the page.
Track results over time, so a drop in quality after a release or an engine change is visible.
Frequently asked questions
Is BLEU still useful?
For comparing machine translation systems on a fixed test set, yes. For judging a single translation, or new content without a reference, it is a weak signal.
Does COMET need a reference translation?
The standard COMET models do. COMETKiwi is the reference-free variant used for quality estimation.
Is LLM-as-a-judge reliable enough to replace human review?
Not entirely. It is useful for finding and explaining likely errors at scale, but its judgments can vary and don’t always match human annotators. Keep human review for high-stakes content.
Which methods work without reference translations?
Quality estimation (such as COMETKiwi) and LLM-as-a-judge. Both can run on new machine translation in production.
