A social media firm wants to compare the toxicity of outputs from several LLMs available in SageMaker JumpStart while minimizing operational work. Which evaluation approach has the least operational overhead?
Choose an answer
Tap an option to check your answer.
Correct answer: Automatic model evaluation.
Why this is the answer
Automatic model evaluation, especially with built-in metrics for toxicity, offers the least operational overhead. SageMaker JumpStart provides pre-trained models and evaluation capabilities, allowing direct comparison of LLM outputs against predefined toxicity scores or benchmarks without extensive manual intervention. Crowd-sourced evaluation and model evaluation using human reviewers both require significant effort in setting up tasks, recruiting and managing evaluators, and analyzing subjective feedback. RLHF is a training technique, not purely an evaluation method, and involves iterative model fine-tuning based on human preferences, which is a much more complex and resource-intensive process than simply comparing model outputs.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed