After fine-tuning a large language model to answer help-desk queries, the company wants to measure whether accuracy improved. Which evaluation metric should they use?
Choose an answer
Tap an option to check your answer.
Correct answer: F1 score.
Why this is the answer
The F1 score is the most appropriate metric because it balances precision and recall, which are both crucial for evaluating help-desk query responses. Precision measures the proportion of correct answers among all answers given, ensuring the model doesn't provide incorrect information. Recall measures the proportion of correct answers identified among all actual correct answers, ensuring the model doesn't miss relevant information. A high F1 score indicates that the model is both accurate and comprehensive. Precision alone might show a high score if the model answers very few questions but gets them all right, missing many valid queries. Time to first token measures latency, not accuracy. Word error rate (WER) is typically used for speech recognition or machine translation to assess how well transcribed text matches a reference, not for the semantic accuracy of generated answers.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed