A pharmaceutical company audits clinical trial site documents (text) and needs to discover the top 10 topics across the corpus so auditors can prioritize reviews. Documents mentioning adverse events must be prioritized. A data scientist will use statistical topic modeling to find abstract topics and list top words per topic. Which algorithms are best suited? (Choose two.)
Choose an answer
Tap an option to check your answer.
Correct answer: Latent Dirichlet Allocation (LDA), Neural topic modeling (NTM).
Why this is the answer
Latent Dirichlet Allocation (LDA) is a widely used statistical topic modeling algorithm that identifies abstract topics within a collection of documents. It assumes documents are a mixture of topics, and topics are a mixture of words, making it suitable for discovering the top 10 topics and their associated words. Neural topic modeling (NTM) is another effective approach that leverages neural networks to learn latent topic representations, often outperforming traditional methods like LDA in capturing complex semantic relationships. Random forest classifier and linear support vector machine are supervised learning algorithms used for classification, not unsupervised topic discovery. Linear regression is a supervised algorithm for predicting continuous values.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed