select TWO: LexaLegal, a legal-tech startup, needs an Amazon Comprehend custom entity recognizer to identify party names, legal clauses, and statute references. Their labeling team can either produce annotated documents with entity spans or curate lists of known statute identifiers. Which two training-data approaches can they use when creating a Comprehend custom entity recognizer?
Choose an answer
Tap an option to check your answer.
Correct answer: Provide annotated training documents that mark entity spans (labeled text) in the format Comprehend expects (document-level annotations)., Supply an entity list (a lexicon) containing known statute identifiers that Comprehend can use as part of training..
Why this is the answer
Amazon Comprehend custom entity recognizers can be trained using two primary methods. First, you can provide annotated training documents (labeled text) where specific entity spans are marked. This allows Comprehend to learn the patterns and context associated with each entity type, such as party names or legal clauses, directly from examples. Second, for entities with a finite and known set of values, like statute identifiers, you can supply an entity list (a lexicon). Comprehend uses this list to recognize exact matches or variations of these known entities. Incorrect options: Uploading a CSV of S3 object tags is not a supported training data format for custom entity recognizers. Sending raw PDFs to Textract and using its output directly without annotation is insufficient; the output still needs to be labeled for entity recognition training. Training using only Comprehend's built-in pre-trained entities would not allow for the recognition of custom entities like legal clauses or specific statute references.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed