A company built a generative text summarization model with Amazon Bedrock and wants to use Bedrock’s automatic model evaluation to measure the model’s accuracy. Which evaluation metric is most appropriate for assessing the quality of generated summaries?
Choose an answer
Tap an option to check your answer.
Correct answer: BERTScore.
Why this is the answer
BERTScore is the most appropriate metric because it leverages pre-trained BERT embeddings to compare the semantic similarity between the generated summary and the reference summary. This approach captures the meaning and contextual relevance, which is crucial for evaluating summarization quality beyond simple word overlap. AUC is used for classification tasks to measure the separability of classes. F1 score is a harmonic mean of precision and recall, typically used for classification or information retrieval, and while it can be adapted for text generation, it often focuses on exact word matches rather than semantic meaning. RWK score is not a standard or widely recognized metric for generative text summarization evaluation.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed