A company compares machine translations produced by its tool with human translations on the same set of documents. Which evaluation strategy should they use to compare the tool’s translation quality relative to human translations?
Choose an answer
Tap an option to check your answer.
Correct answer: Use the Bilingual Evaluation Understudy (BLEU) score to estimate the relative translation quality of the two methods..
Why this is the answer
The Bilingual Evaluation Understudy (BLEU) score is a widely used metric for evaluating the quality of machine-translated text by comparing it to one or more human-translated reference texts. It measures the n-gram overlap between the machine translation and the reference translations, providing a numerical score that indicates how closely the machine translation resembles the human translation. This makes it suitable for estimating the relative translation quality between different machine translation systems or, as in this case, between a machine translation and a human translation baseline. While it doesn't measure absolute quality in terms of perfect fluency or meaning, it effectively quantifies how much a machine translation "looks like" a human translation. BERTScore is another metric, but BLEU is the more traditional and commonly accepted metric for this type of direct comparison against human references.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed