A company uses Amazon SageMaker to produce article summaries in several languages and needs a metric to assess the quality of translated summaries across languages. Which evaluation metric should they use?
Choose an answer
Tap an option to check your answer.
Correct answer: Bilingual Evaluation Understudy (BLEU).
Why this is the answer
BLEU (Bilingual Evaluation Understudy) is the most appropriate metric because it specifically measures the similarity between a machine-translated text and a set of high-quality reference translations. This makes it ideal for evaluating the quality of translated summaries. ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is primarily used for evaluating summarization quality by comparing candidate summaries to reference summaries, but it doesn't inherently handle cross-language comparisons for translation accuracy. AUC (Area Under the ROC Curve) is used for evaluating the performance of binary classification models, not text generation or translation quality. Precision measures the proportion of true positive predictions among all positive predictions, which is not suitable for assessing the overall quality of translated text.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed