Which prompt-based attack specifically reveals the model's configured behavior or the underlying prompt template?
Choose an answer
Tap an option to check your answer.
Correct answer: Extracting the prompt template.
Why this is the answer
Extracting the prompt template is a prompt-based attack where an attacker crafts inputs designed to reveal the hidden instructions or system prompts that guide the model's behavior. This attack aims to understand the model's internal configuration, which can then be used for further exploitation or to bypass safety mechanisms. Prompted persona switches involve tricking the model into adopting a different persona or role, but this doesn't directly reveal the underlying template. Exploiting friendliness and trust focuses on manipulating the model's helpfulness to extract sensitive information or perform harmful actions, not on exposing the template itself. Ignoring the prompt template is a general category of attacks where the model deviates from its instructions, but it's not a specific method for template extraction.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed