Klasterisasi Judul Berita Online Isu Pemilu Prabowo Subianto dengan Kombinasi LLMS Embedding Dengan HDBSCAN
DOI:
https://doi.org/10.58344/locus.v4i10.4779Keywords:
Analisis Wacana Digital, HDBSCAN, Klasterisasi Teks Pendek, LLM, Pemilu 2024, Prabowo SubiantoAbstract
Penelitian ini bertujuan untuk mengelompokkan judul-judul berita politik daring yang berkaitan dengan Presiden Prabowo Subianto selama Pemilu 2024 menggunakan pendekatan berbasis embedding large language models (LLMS) dan algoritma klasterisasi HDBScan. Data yang digunakan dalam penelitian ini berjumlah 24.000 judul berita yang kemudian dianalisis mellaui beberapa tahap meliputi pra-pemrosesan teks, ekstraksi embedding menggunakan model OpenAI, reduksi dimensi menggunakan UMAP, serta klasterisasi berbasis densitas adaptif dengan HDBSCAN. Hasil penelitian menunjukkan terbentuknya 85 klaster tematik dan identifikasi sekitar 27,2% data sebagai noise. Hasil temuan pada penelitian ini mengindikasikan bahwa kombinasi embedding LLM dan HDBSCAN efektif dalam, mengungkap struktur semantik wacana politik digital dari data yang digunakan, serta mampu menangani karakteristik data teks pendek yang kompleks dan heterogen. Pendekatan ini memberikan kontribusi metodologis terhadap studi analisis media berbasis data besar dan menwarakan landasan bagi penelitian lanjutan dalam pemetaan itu publik di ruang digital. Hasil penelitian ini dapat digunakan sebagai sarana untuk penelitian lebih lanjut dengan studi kasus yang berbeda namun menggunakan algoritma yang sama.
References
Blanco?Portals, J., Peiró, F., & Estradé, S. (2021). Strategies for EELS data analysis: Introducing UMAP and HDBSCAN for dimensionality reduction and clustering. Microscopy and Microanalysis, 28(1), 109–122. https://doi.org/10.1017/s1431927621013696
Campello, R., Moulavi, D., Zimek, A., & Sander, J. (2015). Hierarchical density estimates for data clustering, visualization, and outlier detection. ACM Transactions on Knowledge Discovery from Data, 10(1), 1–51. https://doi.org/10.1145/2733381
Chajia, M., & Nfaoui, E. (2024). Customer churn prediction approach based on LLM embeddings and logistic regression. Future Internet, 16(12), 453. https://doi.org/10.3390/fi16120453
Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2019). Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT. https://doi.org/10.48550/arXiv.1810.04805
Fossheim, K. (2022). How can non-elected representatives secure democratic representation? Policy & Politics, 50(2), 243–260. https://doi.org/10.1332/030557321x16371011677734
He, H., Du, C., Fu, K., Wen, B., Sun, Y., Peng, J., & Chang, L. (2024). Human-like object concept representations emerge naturally in multimodal large language models. Research Square. https://doi.org/10.21203/rs.3.rs-4641719/v1
Hidayat, M. (2024). The 2024 general elections in Indonesia: Issues of political dynasties, electoral fraud, and the emergence of a national protest movement. IASJOL, 2(1), 33–51. https://doi.org/10.62033/iasjol.v2i1.51
Keraghel, I., Morbieu, S., & Nadif, M. (2024). Beyond words: A comparative analysis of LLM embeddings for effective clustering. In Proceedings (pp. 205–216). https://doi.org/10.1007/978-3-031-58547-0_17
Korade, N., Salunke, M., Bhosle, A., Asalkar, G., Lal, B., & Kumbharkar, P. (2025). Elevating intelligent voice assistant chatbots with natural language processing and OpenAI technologies. Indonesian Journal of Electrical Engineering and Computer Science, 37(1), 507–517. https://doi.org/10.11591/ijeecs.v37.i1.pp507-517
Lewis, P., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., & Zettlemoyer, L. (2020). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of ACL. https://doi.org/10.48550/arXiv.1910.13461
Motoki, F., Neto, V., & Rodrigues, V. (2023). More human than human: Measuring ChatGPT political bias. Public Choice, 198(1–2), 3–23. https://doi.org/10.1007/s11127-023-01097-2
Neto, A., Sander, J., Campello, R., & Nascimento, M. (2017). Efficient computation of multiple density-based clustering hierarchies. In Proceedings of the 2017 IEEE International Conference on Data Mining (ICDM). https://doi.org/10.1109/icdm.2017.127
Raffel, C., Shazeer, N., Roberts, A., et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21, 1–67. https://doi.org/10.48550/arXiv.1910.10683
Sadeghi, S., Bui, A., Forooghi, A., Lü, J., & Ngom, A. (2024). Can large language models understand molecules? BMC Bioinformatics, 25(1). https://doi.org/10.1186/s12859-024-05847-x
Suchanek, F., & Tuan, L. (2023). Knowledge bases and language models: Complementing forces. In Proceedings (pp. 3–15). https://doi.org/10.1007/978-3-031-45072-3_1
Weng, M., Wu, S., & Dyer, M. (2022). Identification and visualization of key topics in scientific publications with transformer-based language models and document clustering methods. Applied Sciences, 12(21), 11220. https://doi.org/10.3390/app122111220
Yang, C., Cao, B., & Fan, J. (2024). TEC: A novel method for text clustering with large language models guidance and weakly-supervised contrastive learning. Proceedings of the International AAAI Conference on Web and Social Media, 18, 1702–1712. https://doi.org/10.1609/icwsm.v18i1.31419
Ye, Z., Ai, Q., Liu, Y., de Rijke, M., Zhang, M., Lioma, C., & Ruotsalo, T. (2024). Generative language reconstruction from brain recordings. Research Square. https://doi.org/10.21203/rs.3.rs-4587150/v1
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Nanda Perdana, Handri Santoso

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-ShareAlike 4.0 International (CC-BY-SA). that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.




