A company processes hundreds of printed PDF and JPG documents daily and needs an automated, low-maintenance workflow to extract specific text fields and classify documents. Which solution requires the least operational effort?
Choose an answer
Tap an option to check your answer.
Correct answer: Use Amazon Textract to extract fields and Amazon Comprehend to classify documents..
Why this is the answer
The correct solution uses Amazon Textract for field extraction and Amazon Comprehend for document classification. Amazon Textract is a fully managed service designed for OCR and data extraction from documents, including forms and tables, making it ideal for the PDF and JPG input. Amazon Comprehend is a fully managed natural language processing (NLP) service that can classify documents based on their content, requiring minimal operational overhead. Using PaddleOCR in SageMaker would involve managing SageMaker endpoints and potentially training/fine-tuning models, which increases operational effort compared to fully managed services. Amazon Rekognition is primarily for image and video analysis, not document text classification, making it unsuitable for the classification task.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed