Klasifikasi Diabetes Menggunakan Algoritma K-Nearest Neighbor dengan Preprocessing dan Min-Max Scaling pada Pima Indians Diabetes Dataset

Authors

  • Nurdilla Institut Teknologi dan Bisnis Indonesia
  • Roberto Kaban Institut Teknologi dan Bisnis Indonesia

Keywords:

Classification, Diabetes Mellitus, K-Nearest Neighbor, Machine Learning, Python

Abstract

Diabetes mellitus is a chronic metabolic disease characterized by high blood glucose levels and has the potential to cause various complications if not detected early. The use of machine learning technology is increasingly developing in the health sector because it can help the process of analyzing and classifying diseases based on patient data. This study aims to apply the K-Nearest Neighbor (KNN) algorithm to classify diabetes using the Pima Indians Diabetes Dataset. The dataset used consists of 768 patient data with 8 predictor attributes and 1 target attribute. The research stages include data cleaning and improvement through preprocessing, data normalization using the Min-Max Scaling method, dividing the dataset into training data and testing data with a ratio of 80:20, and the application of the KNN algorithm with a K value of 5. Model performance evaluation was carried out using a Confusion Matrix which produces Accuracy, Precision, Recall, and F1-Score values. Based on the test results, the model obtained Accuracy of 74.68%, Precision of 66.00%, Recall of 60.00%, and F1-Score of 62.86%. These results demonstrate that the KNN algorithm is capable of classifying diabetes data with fairly good performance based on available health attributes. This research is expected to serve as a reference in the development of machine learning-based decision support systems to aid in the identification of diabetes.

Downloads

Download data is not yet available.

References

[1] M. A. Hama Saeed, “Diabetes type 2 classification using machine learning algorithms with up-sampling technique,” Journal of Electrical Systems and Inf Technol, vol. 10, no. 1, p. 8, Feb. 2023, doi: 10.1186/s43067-023-00074-5.

[2] M. Y. Shams, Z. Tarek, and A. M. Elshewey, “A novel RFE-GRU model for diabetes classification using PIMA Indian dataset,” Sci Rep, vol. 15, no. 1, p. 982, Jan. 2025, doi: 10.1038/s41598-024-82420-9.

[3] H. H. Rashidi, S. Albahra, S. Robertson, N. K. Tran, and B. Hu, “Common statistical concepts in the supervised Machine Learning arena,” Front. Oncol., vol. 13, p. 1130229, Feb. 2023, doi: 10.3389/fonc.2023.1130229.

[4] R. G. Wardhana, G. Wang, and F. Sibuea, “PENERAPAN MACHINE LEARNING DALAM PREDIKSI TINGKAT KASUS PENYAKIT DI INDONESIA,” JOISM, vol. 5, no. 1, pp. 40–45, Jul. 2023, doi: 10.24076/joism.2023v5i1.1136.

[5] A. K. Ermy Pily, Oktavianda, F. Aprilia, Rahmaddeni, and L. Efrizoni, “Komparasi Algoritma K-Nearest Neighbors dan Naïve Bayes dalam Klasifikasi Penyakit Diabetes Gestasional,” ijcs, vol. 13, no. 1, Feb. 2024, doi: 10.33022/ijcs.v13i1.3714.

[6] L. N. Mukarromah, Z. Fatah, and I. Yunita, “KLASIFIKASI PENYAKIT DIABETES MENGGUNAKAN METODE K-NEAREST NEIGHBORS (KNN),” vol. 1, no. 4, 2024.

[7] F. D. Musa, D. Purwanto, S. Amri, A. Fadlurohman, and A. Fitriyanan, “Klasifikasi Dataset Diabetes menggunakan Algoritma K-Nearest Neighbors,” 2024.

[8] H. M. Saleh, “A Comprehensive Review of Data Mining Techniques for Diabetes Diagnosis Using the Pima Indian Diabetes Dataset,” EDRAAK, vol. 2024, pp. 39–42, Apr. 2024, doi: 10.70470/EDRAAK/2024/006.

[9] Novian Ikhsan, “Klasifikasi penyakit diabetes menggunakan metode K-Nearest Neighbors (KNN) berdasarkan data klinis,” infotech, vol. 7, no. 1, pp. 19–26, Jun. 2026, doi: 10.37373/infotech.v7i1.2110.

[10] “Pima Indians Diabetes Database.” Accessed: Jun. 02, 2026. [Online]. Available: https://www.kaggle.com/datasets/uciml/pima-indians-diabetes-database

[11] A. T. Akbar, H. Prapcoyo, and R. Husaini, “SMOTE and K-Means Preprocessing for Classification by Logistic Regression on Pima Indian Diabetes Dataset,” vol. 20, no. 2, 2023.

[12] D. S. F. Azzahrah and A. Alamsyah, “Comparison of Probabilistic Neural Network (PNN) and k-Nearest Neighbor (k-NN) Algorithms for Diabetes Classification,” Recursive J. of Informatics, vol. 1, no. 2, pp. 73–82, Sep. 2023, doi: 10.15294/rji.v1i2.66078.

[13] Muhammad Randy Fachrezi, Hafiz Aryanda, Alwi Syahputra, and Risma Riansyah, “Sistem Klasifikasi Diabetes Mellitus Menggunakan Algoritma K-Nearest Neighbor (KNN) Berbasis Web,” JUIKTI, vol. 2, no. 1, pp. 119–130, Jan. 2026, doi: 10.64803/juikti.v2i1.116.

[14] T. Emmanuel, T. Maupong, D. Mpoeleng, T. Semong, B. Mphago, and O. Tabona, “A survey on missing data in machine learning,” J Big Data, vol. 8, no. 1, p. 140, Oct. 2021, doi: 10.1186/s40537-021-00516-9.

[15] M. Phongying and S. Hiriote, “Diabetes Classification Using Machine Learning Techniques,” Computation, vol. 11, no. 5, p. 96, May 2023, doi: 10.3390/computation11050096.

[16] A. I. ElSeddawy, F. K. Karim, A. M. Hussein, and D. S. Khafaga, “Predictive Analysis of Diabetes-Risk with Class Imbalance,” Computational Intelligence and Neuroscience, vol. 2022, pp. 1–16, Oct. 2022, doi: 10.1155/2022/3078025.

[17] M. Sandeep Kumar, M. Zubair Khan, S. Rajendran, A. Noor, A. Stephen Dass, and J. Prabhu, “Imbalanced Classification in Diabetics Using Ensembled Machine Learning,” Computers, Materials & Continua, vol. 72, no. 3, pp. 4397–4409, 2022, doi: 10.32604/cmc.2022.025865.

Published

2026-07-31

How to Cite

Nurdilla , N., & Kaban, R. (2026). Klasifikasi Diabetes Menggunakan Algoritma K-Nearest Neighbor dengan Preprocessing dan Min-Max Scaling pada Pima Indians Diabetes Dataset. System Information and Computer Technology (SYNCTECH), 2(2), 54–65. Retrieved from https://librarium.id/index.php/synctech/article/view/54

Issue

Section

Articles