Which evaluation metric is commonly used to assess foundation models (FMs) on text summarization tasks?
Choose an answer
Tap an option to check your answer.
Correct answer: Bilingual Evaluation Understudy (BLEU) score.
Why this is the answer
The Bilingual Evaluation Understudy (BLEU) score is a standard metric for evaluating the quality of text generated by machine translation and text summarization models. It compares the generated text to one or more reference texts, calculating a score based on the overlap of n-grams (sequences of words). A higher BLEU score indicates greater similarity to the reference summary, suggesting better quality. F1 score is primarily used for classification tasks, measuring precision and recall, not text generation quality. Accuracy is also for classification, indicating the proportion of correct predictions. Mean squared error (MSE) is a regression metric, used to quantify the difference between predicted and actual numerical values. These metrics are unsuitable for evaluating the linguistic quality of generated text like summaries.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed