A company intends to use a generative AI model to give users real-time service quotes. Which selection criterion is most important for this application?
Choose an answer
Tap an option to check your answer.
Correct answer: Low latency and inference performance optimized for fast responses.
Why this is the answer
For real-time service quotes, low latency and optimized inference performance are critical because users expect immediate responses. Delays directly impact user experience and the utility of the service. While the quality of training data is always important for model accuracy, and model size can affect performance, these are secondary to the immediate need for speed in a real-time interaction. Choosing a general-purpose model and having access to high-powered GPUs are considerations for model development and deployment, but they don't directly address the core requirement of fast user-facing responses as effectively as focusing on low latency and inference performance.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed