ANALISIS KOMPARATIF ALGORITMA NAIVE BAYES DAN SUPPORT VECTOR MACHINE DALAM KLASIFIKASI UJARAN KEBENCIAN DAN TEKS ABUSIF BERBAHASA INDONESIA
DOI:
https://doi.org/10.25134/fon.v22i1.471Kata Kunci:
klasifikasi teks; ujaran kebencian; bahasa kasar; Support Vector Machine; Naive Bayes; media sosial IndonesiaAbstrak
ABSTRAK: Penelitian ini mengkaji efektivitas algoritma Naive Bayes dan Support Vector Machine dalam mengklasifikasikan ujaran kebencian dan teks kasar berbahasa Indonesia, dengan tujuan memahami keunggulan dan pola kegagalan setiap model. Metode penelitian menggunakan dataset tweet beranotasi yang melalui pra-pemrosesan khusus dan ekstraksi fitur TF-IDF. Kinerja model Multinomial Naive Bayes (MNB) dan Linear Support Vector Classifier (LinearSVC) dievaluasi secara ketat menggunakan holdout test set, validasi silang, dan uji signifikansi statistik. Hasil penelitian mengonfirmasi superioritas signifikan LinearSVC atas MNB, dengan selisih akurasi 8,7%. Analisis mendalam mengungkap bahwa meskipun model sangat akurat mengidentifikasi konten bersih (clean), tantangan utama terletak pada klasifikasi teks kasar (abusive) yang sering tertukar dengan ujaran kebencian (hate speech). Pola kesalahan ini menunjukkan batas kabur antara kedua kategori, mengindikasikan spektrum kontinu ketimbang kelas diskrit. Diskusi menegaskan bahwa keunggulan LinearSVC berasal dari kemampuannya menangani data berdimensi tinggi dan kompleksitas linguistik khas media sosial Indonesia. Temuan ini menyoroti perlunya pendekatan pemodelan yang lebih peka terhadap nuansa bahasa dan gradasi keparahan konten. Implikasi studi mendorong penyempurnaan skema anotasi data, eksplorasi metode seperti klasifikasi multi-label atau regresi ordinal, serta pengembangan sistem moderasi hibrid yang memprioritaskan area abu-abu antara teks kasar dan ujaran kebencian untuk keadilan dan akurasi yang lebih baik.
KATA KUNCI: klasifikasi teks; ujaran kebencian; bahasa kasar; Support Vector Machine; Naive Bayes; media sosial Indonesia.
COMPARATIVE ANALYSIS OF NAIVE BAYES AND SUPPORT VECTOR MACHINE ALGORITHMS IN CLASSIFYING HATE SPEECH AND ABUSIVE INDONESIAN TEXT
ABSTRACT: This study examines the effectiveness of Naive Bayes and Support Vector Machine algorithms in classifying hate speech and abusive Indonesian text, aiming to understand the strengths and failure patterns of each model. The research method uses an annotated tweet dataset that undergoes specialized preprocessing and TF-IDF feature extraction. The performance of the Multinomial Naive Bayes (MNB) and Linear Support Vector Classifier (LinearSVC) models is rigorously evaluated using a holdout test set, cross-validation, and statistical significance testing. The results confirm the significant superiority of LinearSVC over MNB, with an accuracy difference of 8.7%. In-depth analysis reveals that while the model is highly accurate in identifying clean content, the main challenge lies in classifying abusive text, which is often confused with hate speech. This error pattern indicates a blurred boundary between the two categories, suggesting a continuous spectrum rather than discrete classes. The discussion affirms that LinearSVC's advantage stems from its ability to handle high-dimensional data and the linguistic complexity typical of Indonesian social media. These findings highlight the need for modeling approaches that are more sensitive to linguistic nuances and content severity gradations. The study's implications encourage the refinement of data annotation schemes, exploration of methods such as multi-label classification or ordinal regression, and the development of hybrid moderation systems that prioritize the gray area between abusive text and hate speech for greater fairness and accuracy.
KEYWORDS: text classification; hate speech; abusive language; Support Vector Machine; Naive Bayes; Indonesian social media.
Unduhan
Referensi
Alfina, I., Mulia, R., Fanany, M. I., & Ekanata, Y. (2017). Hate speech detection in the Indonesian language: A dataset and preliminary study. In 2017 International Conference on Advanced Computer Science and Information Systems (ICACSIS). IEEE. https://doi.org/10.1109/ICACSIS.2017.8355039
Amal, I., & Pamungkas, E. W. (2025). Enhancing hate speech detection in Indonesia code-mixed tweets: The role of oversampling and undersampling techniques. In 2025 International Conference on Smart Computing, IoT and Machine Learning (SIML 2025). IEEE.
Asti, A. D., Budi, I., & Ibrohim, M. O. (2021). Multi-label classification for hate speech and abusive language in Indonesian-local languages. In 2021 International Conference on Advanced Computer Science and Information Systems (ICACSIS). IEEE.
Fortuna, P., & Nunes, S. (2018). A survey on automatic detection of hate speech in text. ACM Computing Surveys, 51(4), 1–30. https://doi.org/10.1145/3232676
Hana, K. M., Adiwijaya, Al Faraby, S., & Bramantoro, A. (2020). Multi-label classification of Indonesian hate speech on Twitter using support vector machines. In 2020 International Conference on Data Science and Its Applications (ICoDSA). IEEE.
Ibrohim, M. O., & Budi, I. (2018). A dataset and preliminaries study for abusive language detection in Indonesian social media. Procedia Computer Science, 135, 222–229. https://doi.org/10.1016/j.procs.2018.08.169
Ibrohim, M. O., & Budi, I. (2019). Multi-label hate speech and abusive language detection in Indonesian Twitter. In Proceedings of the Third Workshop on Abusive Language Online (ALW3).
Ibrohim, M. O., & Budi, I. (2019). Translated vs non-translated method for multilingual hate speech identification in Twitter. International Journal on Advanced Science, Engineering and Information Technology, 9(1), 294–301. https://doi.org/10.18517/ijaseit.9.1.6826
Ibrohim, M. O., Setiadi, M. A., & Budi, I. (2019). Identification of hate speech and abusive language on Indonesian twitter using the word2vec, part of speech and emoji features. In Proceedings of the 2019 2nd International Conference on Algorithms, Computing and Artificial Intelligence (ACAI 2019). ACM. https://doi.org/10.1145/3377713.3377782
Joachims, T. (1998). Text categorization with support vector machines: Learning with many relevant features. In C. Nédellec & C. Rouveirol (Eds.), *Machine Learning: ECML-98* (pp. 137–142). Springer. https://doi.org/10.1007/BFb0026683
Kementerian Komunikasi dan Informatika Republik Indonesia. (2023). Laporan tahunan konten negatif 2023. Kementerian Kominfo RI.
Koto, F., & Rahmaningtyas, W. (2018). Exploring Indonesian abusive language detection in Twitter. In 2018 International Conference on Asian Language Processing (IALP). IEEE. https://doi.org/10.1109/IALP.2018.8629157
Kusumawati, R., D’Arofah, A., & Pramana, P. A. (2019). Comparison performance of Naive Bayes classifier and support vector machine algorithm for Twitter’s classification of Tokopedia services. Journal of Physics: Conference Series, 1196(1), Article 012043. https://doi.org/10.1088/1742-6596/1196/1/012043
Louisa, R. F., & Murwantara, I. M. (2025). Performance comparison of SBERT and FastText in social media analysis. In *ICoCSETI 2025 - International Conference on Computer Sciences, Engineering, and Technology Innovation, Proceeding*.
Prabowo, F. A., Ibrohim, M. O., & Budi, I. (2019). Hierarchical multi-label classification to identify hate speech and abusive language on Indonesian Twitter. In 2019 6th International Conference on Information Technology, Computer and Electrical Engineering (ICITACEE). IEEE.
Putri, K. R., & Cahyani, D. E. (2025). Application of the Naïve Bayes and Support Vector Machine in sentiment analysis for the 2024 Indonesian presidential candidates. AIP Conference Proceedings, 3045(1), Article 050003. https://doi.org/10.1063/5.0212345
Putri, S. D. A., Ibrohim, M. O., & Budi, I. (2021). Abusive language and hate speech detection for Indonesian-local language in social media text. In K. Arai (Ed.), Advances in Information and Communication (pp. 417–431). Springer. https://doi.org/10.1007/978-3-030-90016-8_30
Putri, S. D. A., Ibrohim, M. O., & Budi, I. (2021). Abusive language and hate speech detection for Javanese and Sundanese languages in tweets: Dataset and preliminary study. In 2021 11th International Workshop on Computer Science and Engineering (WCSE 2021).
Putri, T. T. A., Sriadhi, S., Sari, R. D., Sihombing, J. L., Sembiring, E. B., Sitorus, M. S., Saragih, M. H., Ramadhani, N., Irmayani, D., & Hutahaean, H. D. (2020). A comparison of classification algorithms for hate speech detection. IOP Conference Series: Materials Science and Engineering, 830(3), Article 032029. https://doi.org/10.1088/1757-899X/830/3/032029
Rohmawati, U. A. N., Sihwi, S. W., & Cahyani, D. E. (2018). SEMAR: An interface for Indonesian hate speech detection using machine learning. In 2018 International Seminar on Research of Information Technology and Intelligent Systems (ISRITI). IEEE. https://doi.org/10.1109/ISRITI.2018.8864478
Russell, S., & Norvig, P. (2010). Artificial intelligence: A modern approach (3rd ed.). Prentice Hall.
Sireesha, M., & Tamilselvan, S. (2024). Detecting fake political news on social media by support vector machine compared to Naive Bayes. In 2024 2nd International Conference on Computational and Characterization Techniques in Engineering and Sciences (IC3TES). IEEE.
We Are Social. (2024, January). Digital 2024: Indonesia. DataReportal. https://datareportal.com/reports/digital-2024-indonesia
Wulansari, R., Wibowo, H., Adriani, M., & Irawan, B. (2020). Indonesian abusive and hate speech Twitter dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/rwulansari/indonesian-abusive-and-hate-speech-twitter-dataset
Yovellia Londo, G. L., Kartawijaya, D. H., Ivariyani, H. T., Wibowo, A., Maharani, D. A., & Ariyandi, D. (2019). A study of text classification for Indonesian news article. In Proceeding of the 2019 International Conference of Artificial Intelligence and Information Technology (ICAIIT).
Unduhan
Diterbitkan
Terbitan
Bagian
Lisensi
Hak Cipta (c) 2026 Rolanda Difandana, Ian Imaduddin (Penulis)

Artikel ini berlisensiCreative Commons Attribution-ShareAlike 4.0 International License.