Perbandingan Kinerja Model XGBoost dan IndoBERT untuk Deteksi Teks Hoax Terselubung Berbahasa Indonesia
DOI:
https://doi.org/10.55606/juisik.v6i2.2411Keywords:
Concealed Hoax, Generative AI, IndoBERT, Text Detection, XGBoostAbstract
The rapid advancement of Generative AI driven by Large Language Models (LLMs) has raised significant concerns regarding the integrity of digital information. Although text detection systems are evolving rapidly, the majority of current instruments predominantly support English, thereby leaving a substantial research gap in low-resource languages such as Indonesian. This study aims to conduct a comparative performance analysis between XGBoost and IndoBERT for detecting Indonesian text generated by Generative AI. Experiments were executed using a structured, hand-crafted dataset spanning three main domains: Science and Technology, Social and Culture, and Politics. The XGBoost model was evaluated utilizing Term Frequency-Inverse Document Frequency (TF-IDF) feature extraction, whereas IndoBERT relied on bidirectional contextual semantic understanding. The evaluation results on the controlled dataset demonstrated that both architectures achieved perfect accuracy (100%). However, under a stress test employing adversarial trap sentences designed with keyword stuffing techniques, the performance of the two models exhibited a sharp contrast. The global accuracy of XGBoost plummeted drastically to 50.00% because the TF-IDF representation failed to recognize the core logical essence of the sentences and was easily misled by the dominance of individual statistical keywords. Conversely, IndoBERT proved to be more robust by maintaining a global accuracy of 66.67% and recording maximum Precision, Recall, and F1-Score (1.000000) within the Social and Culture class. This demonstrates the capability of the Self-Attention Mechanism in capturing latent context and complex linguistic structures. The implications of these findings emphasize that in countering advanced text manipulation (covert hoaxes), the deployment of context-based models like IndoBERT is far more reliable. Nevertheless, architectural selection must remain dynamically aligned with local computing infrastructure capacity.
References
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
Chakraborty, S., Bedi, A. S., Zhu, S., An, B., Manocha, D., & Huang, F. (2023). On the possibilities of AI-generated text detection. arXiv. https://arxiv.org/abs/2304.04736
Chapman, P., Clinton, J., Kerber, R., Khabaza, T., Reinartz, T., Shearer, C., & Wirth, R. (2000). CRISP-DM 1.0: Step-by-step data mining guide. SPSS Inc.
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
Fraser, K. C., Dawkins, C., & Kiritchenko, S. (2024). Detection of AI-generated text: A survey. arXiv. https://arxiv.org/
Gehrmann, S., Strobelt, H., & Rush, A. M. (2019). GLTR: Statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 111–116). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-3019
Jurafsky, D., & Martin, J. H. (2023). Speech and language processing (3rd ed.). Stanford University. https://web.stanford.edu/~jurafsky/slp3/
Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. arXiv. https://arxiv.org/abs/2011.00677
Liu, Y., Li, X., & Li, Z. (2025). Machine learning approaches for AI-generated text detection using stylometric and TF-IDF features.
Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.
Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (2023). DetectGPT: Zero-shot machine-generated text detection using probability curvature. In Proceedings of the 40th International Conference on Machine Learning.
OpenAI. (2023). GPT-4 technical report. arXiv. https://arxiv.org/abs/2303.08774
Powers, D. M. W. (2020). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation.
Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427–437. https://doi.org/10.1016/j.ipm.2009.03.002
Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., & Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (pp. 843–857).
Wirth, R., & Hipp, J. (2000). CRISP-DM: Towards a standard process model for data mining. In Proceedings of the 4th International Conference on the Practical Applications of Knowledge Discovery and Data Mining (pp. 29–39).
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Jurnal ilmiah Sistem Informasi dan Ilmu Komputer

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.




