Abstract
The analysis of short text documents has become a vital and challenging task. Topic models are utilized to extract topics from a large amount of text data. However, these topic models typically suffer from data sparsity problems when applied to short texts because of relatively lower word co-occurrence patterns. As a result, they tend to provide repetitive or trivial topics of poor quality. Therefore, we presented a DistilBERTopic model to remove the sparsity problem and discover quality topics more accurately from short texts. DistilBERTopic model utilized the pre-trained transformer-based language models, reduced the dimensionality effect on embedding, clustered these embeddings, and discovered the topics from short text documents. Experimental results demonstrate that the DistilBERTopic model achieves better classification and topic coherence than other state-of-the-art topic models for real-world datasets.
| Original language | English |
|---|---|
| Journal | Proceedings of the Australasian Language Technology Workshop |
| Volume | 20 |
| State | Published - 2022 |
| Externally published | Yes |
| Event | 20th Annual Workshop of the Australasian Language Technology Association, ALTA 2022 - Adelaide, Australia Duration: 14 Dec 2022 → 16 Dec 2022 |
Bibliographical note
Publisher Copyright:© 2020, Australasian Language Technology Association. All rights reserved.
ASJC Scopus subject areas
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'A DistilBERTopic Model for Short Text Documents'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver