An AI practitioner is building a search application that needs to handle queries containing both text and images. Which type of foundation model is the best fit to power this multi-modal search?
Choose an answer
Tap an option to check your answer.
Correct answer: Multi-modal embedding model.
Why this is the answer
A multi-modal embedding model is the best fit because it can process and understand information from multiple modalities, such as text and images, and embed them into a shared vector space. This allows for similarity searches across different data types, meaning a text query can find relevant images, and an image query can find relevant text. A text embedding model is incorrect because it only processes text, making it unsuitable for image queries or finding images based on text. A multi-modal generation model is incorrect because its primary function is to generate new content (e.g., text from an image, or an image from text), not to embed and search existing multi-modal data. An image generation model is incorrect as it focuses solely on creating new images and does not handle text or multi-modal search.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed