A company is training a custom large language model for a chatbot aimed at teenagers. The chatbot should use the audience’s informal style, including creative spellings and abbreviations. Which metric is appropriate for evaluating the model’s performance?
Choose an answer
Tap an option to check your answer.
Correct answer: BERTScore.
Why this is the answer
BERTScore is appropriate because it measures semantic similarity between generated and reference text using contextual embeddings from BERT. This allows it to effectively evaluate the chatbot's ability to capture the informal style, creative spellings, and abbreviations common in teenage language, even if they deviate from standard grammar or spelling. F1 score is used for classification tasks, not for evaluating text generation quality. ROUGE and BLEU scores primarily focus on n-gram overlap, which would penalize the model for creative spellings and abbreviations that don't exactly match a reference, even if semantically correct and stylistically appropriate for the target audience.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed