A company is comparing large language models for a text summarization task and needs a metric to evaluate summary quality. Which metric is appropriate?
Choose an answer
Tap an option to check your answer.
Correct answer: Recall-Oriented Understudy for Gisting Evaluation (ROUGE).
Why this is the answer
ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is the appropriate metric for evaluating summary quality. It compares an automatically generated summary against human-written reference summaries by counting overlapping units such as n-grams, word sequences, or word pairs. Higher ROUGE scores indicate better summary quality. Recall is a general classification metric and doesn't specifically assess summary content overlap. AUC is used for evaluating the performance of binary classification models. MSE is a regression metric, used for quantifying the difference between predicted and actual continuous values.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed