A Hybrid Rule-Based and Multinomial Naïve Bayes System for Sentiment and Intent Classification of Indonesian Public Reports with Sarcasm Detection

Authors

  • Suhendri Suhendri Universitas Majalengka
  • Sahal Ubaidillah Gunardo Universitas Majalengka
  • Kartika Dwi Mulyana Universitas Majalengka
  • Ruli Susanti Universitas Majalengka
  • Dika Alfaizal Akbar Universitas Majalengka
  • Amelia Putri Universitas Majalengka

DOI:

https://doi.org/10.24076/intechnojournal.2026v8i1.2765

Keywords:

Citizen Report Classification, Indonesian NLP, Multinomial Naive Bayes, Sentiment Analysis, TF-IDF

Abstract

Public service reporting systems in Indonesia face significant challenges in processing large volumes of unstructured citizen feedback efficiently. This study proposes Aspiralytica, a mobile-based citizen report classification system that integrates TF-IDF feature extraction with a Multinomial Naive Bayes (MNB) classifier within a hybrid rule-based and machine learning architecture. The system simultaneously performs three-class sentiment classification (positive, negative, neutral) and five-class intent classification (complaint, appreciation, request, emergency, suggestion), with automated priority level determination and a rule-based sarcasm detection module achieving F1 of 0.8980. Evaluated on an augmented dataset of 1,137 sentiment-labeled and 2,187 intent-labeled Indonesian-language citizen report texts using Stratified 10-Fold Cross-Validation, the proposed MNB model achieved sentiment classification accuracy of 96.59% (F1: 96.59%) and intent classification accuracy of 96.97% (F1: 96.96%). An ablation study confirmed TF-IDF with MNB as the dominant performance driver, and a computational efficiency benchmark empirically justified MNB selection with mean inference latency of 0.3955 ms and throughput of 52,312 requests per second. The system is deployed as a FastAPI backend integrated with a React Native mobile frontend, delivering real-time classification through a citizen-facing interface.

References

[1] S. Syafarudin, A. Haris, S. Tinggi, I. Administrasi, and Y. Makassar, “Digital Transformation in Public Services: A Study of E-Government Implementation in Indonesia,” International Journal of Law and Society, vol. 2, no. 4, pp. 169–179, Nov. 2025, doi: 10.62951/IJLS.V2I4.797.

[2] R. N. Bewinda, H. Prabowo, E. Indrayani, and G. Gatiningsih, “Revisiting E-Government Services in the Provincial Government of DKI Jakarta: A Case Study on the Management of Public Complaints,” Jurnal Komunikasi Ikatan Sarjana Komunikasi Indonesia, vol. 9, no. 2, pp. 406–420, Dec. 2024, doi: 10.25008/JKISKI.V9I2.1123.

[3] R. Kosasih and A. Alberto, “Sentiment analysis of game product on shopee using the TF-IDF method and naive bayes classifier,” ILKOM Jurnal Ilmiah, vol. 13, no. 2, pp. 101–109, Aug. 2021, doi: 10.33096/ilkom.v13i2.721.101-109.

[4] I. Verawati and S. N. Jaelani, “Analisis Sentimen Pengguna Twitter Terhadap Bus Listrik Menggunakan Naïve Bayes,” JURNAL MEDIA INFORMATIKA BUDIDARMA, vol. 8, no. 2, pp. 832–842, Apr. 2024, doi: 10.30865/MIB.V8I2.7030.

[5] M. L. F. Martanto and W. Istiono, “Sentiment Analysis of M-Paspor App Reviews Using Multinomial Naive Bayes,” Journal of Logistics, Informatics and Service Science, vol. 11, no. 10, pp. 311–326, 2024, doi: 10.33168/JLISS.2024.1017.

[6] Dhendra and V. Gayuh Utomo, “Benchmarking IndoBERT and Transformer Models for Sentiment Classification on Indonesian E-Government Service Reviews,” Jurnal Transformatika, vol. 23, no. 1, pp. 86–95, Jul. 2025, doi: 10.26623/transformatika.v23i1.12095.

[7] E. Di. Madyatmadja, B. N. Yahya, and C. Wijaya, “Contextual Text Analytics Framework for Citizen Report Classification: A Case Study Using the Indonesian Language,” IEEE Access, vol. 10, pp. 31432–31444, 2022, doi: 10.1109/ACCESS.2022.3158940.

[8] S. M. Intani, B. I. Nasution, M. E. Aminanto, Y. Nugraha, N. Muchtar, and J. I. Kanggrawan, “Automating Public Complaint Classification Through JakLapor Channel: A Case Study of Jakarta, Indonesia,” 2022 IEEE International Smart Cities Conference (ISC2), 2022, doi: 10.1109/ISC255366.2022.9922346.

[9] I. Alpiana, W. Yustanti, and Y. Yamasari, “Optimization and Evaluation of IndoBERT, BiLSTM–BiGRU, and Hybrid XGBoost for Predicting Complaint Subcategories in the E-Layanan System of Universitas Negeri Surabaya,” Journal of Education and Informatics Research, vol. 6, no. 2, p. 2025, Dec. 2025, Accessed: May 29, 2026. [Online]. Available: https://journal.trunojoyo.ac.id/jedumatic/article/view/31890

[10] D. Suhartono, W. Wongso, and A. Tri Handoyo, “IdSarcasm: Benchmarking and Evaluating Language Models for Indonesian Sarcasm Detection,” IEEE Access, vol. 12, pp. 87323–87332, 2024, doi: 10.1109/ACCESS.2024.3416955.

[11] M. Bayer, M. A. Kaufhold, B. Buchhold, M. Keller, J. Dallmeyer, and C. Reuter, “Data augmentation in natural language processing: a novel text generation approach for long and short text classifiers,” International Journal of Machine Learning and Cybernetics, vol. 14, no. 1, pp. 135–150, Jan. 2023, doi: 10.1007/s13042-022-01553-3.

[12] D. Pluscec and J. Snajder, “Data Augmentation for Neural NLP,” arXiv:2302.11412v1, Feb. 2023. Available: http://arxiv.org/abs/2302.11412

[13] L. Zhang, “Features extraction based on Naive Bayes algorithm and TF-IDF for news classification,” PLoS One, vol. 20, no. 7 July, Jul. 2025, doi: 10.1371/journal.pone.0327347.

[14] L. Xiang, “Application of an Improved TF-IDF Method in Literary Text Classification,” Advances in Multimedia, vol. 2022, 2022, doi: 10.1155/2022/9285324.

[15] M. Abbas, K. Ali Memon, and A. Aleem Jamali, “Multinomial Naive Bayes Classification Model for Sentiment Analysis,” 2019. doi: 10.13140/RG.2.2.30021.40169.

[16] V. W. Lumumba, D. Kiprotich, M. L. Mpaine, N. G. Makena, and M. D. Kavita, “Comparative Analysis of Cross-Validation Techniques: LOOCV, K-folds Cross-Validation, and Repeated K-folds Cross-Validation in Machine Learning Models,” American Journal of Theoretical and Applied Statistics 2024, Volume 13, Page 127, vol. 13, no. 5, pp. 127–137, Oct. 2024, doi: 10.11648/J.AJTAS.20241305.13.

[17] M. Chen, K. Ubul, X. Xu, A. Aysa, and M. Muhammat, “Connecting Text Classification with Image Classification: A New Preprocessing Method for Implicit Sentiment Text Classification,” Sensors (Basel), vol. 22, no. 5, Feb. 2022, doi: 10.3390/s22051899.

[18] M. Jin and N. Aletras, “Complaint Identification in Social Media with Transformer Networks,” Online, 2020. doi: 10.18653/v1/2020.coling-main.157.

[19] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” Online, 2020. doi: 10.18653/v1/2020.coling-main.66.

[20] A. Romadhony, S. Al Faraby, R. Rismala, U. N. Wisesti, and A. Arifianto, “Sentiment Analysis on a Large Indonesian Product Review Dataset,” Journal of Information Systems Engineering and Business Intelligence, vol. 10, no. 1, pp. 167–178, 2024, doi: 10.20473/jisebi.10.1.167-178.

Downloads

Published

2026-07-31

Issue

Section

Articles