Abstract
The goal of this paper is to compare the performance of the statistical and machine learning classification methods in diagnosing the death of liver cancer patients based on demographic characteristics, risk factors, and medical interventions. For this purpose, five methods; include random tree, C4.5, random forest, support vector machine (SVM) and logistic regression, all of which are supervised methods, were selected. The data used in this research are the real data of 165 patients diagnosed with liver cancer in a hospital in Portugal. The aim variable is the patient\\\\'s death during the trial period, and the aforementioned group was monitored for a year. There are twenty-six qualitative and twenty-three quantitative variables in this diverse dataset. In total, 10.22% of the dataset is missing data, and just eight patients have full information in every field (4.85%). Additionally, there is some class disparity (63 cases classified as \\\\"Dead\\\\" and 102 as \\\\"Alive\\\\"). With 73.33% accurate detection, the SVM approach was found to be the most effective approach. After that, the random forest method with 71.52% had a more correct identification ratio than the others. The treesC4.5 method had the lowest correct diagnosis with 58.18%. Although, based on the ROC Area index, the random forest method performed better (with an area under the curve = 0.789) than the SVM method (with an area under the curve = 0.711). In total, SVM and random forest methods worked with a large difference compared to others in diagnosing the death of patients with liver cancer
Keywords
C4.5
liver Cancer
logistic regression
Random Forest
Random tree
SVM
Abstract
تهدف هذه الورقة البحثية إلى مقارنة أداء أساليب التصنيف الإحصائية وأساليب التعلم الآلي في تشخيص وفاة مرضى سرطان الكبد بناءً على الخصائص الديموغرافية وعوامل الخطر والتدخلات الطبية. ولتحقيق هذا الهدف، تم اختيار خمسة أساليب، هي: الشجرة العشوائية، وخوارزمية C4.5، والغابة العشوائية، وآلة المتجهات الداعمة (SVM)، والانحدار اللوجستي، وجميعها أساليب خاضعة للإشراف. وتتكون البيانات المستخدمة في هذا البحث من بيانات حقيقية لـ 165 مريضًا تم تشخيص إصابتهم بسرطان الكبد في أحد مستشفيات البرتغال. المتغير الرئيسي هو وفاة المريض خلال فترة الدراسة، وقد تمت متابعة المجموعة المذكورة لمدة عام. تحتوي مجموعة البيانات المتنوعة هذه على 26 متغيرًا نوعيًا و23 متغيرًا كميًا. إجمالًا، 10.22% من مجموعة البيانات مفقودة، وثمانية مرضى فقط لديهم معلومات كاملة في كل حقل (4.85%). بالإضافة إلى ذلك، يوجد تباين في التصنيف (63 حالة مصنفة على أنها \\\\"متوفاة\\\\" و102 حالة على أنها \\\\"على قيد الحياة\\\\"). بنسبة دقة كشف بلغت 73.33%، تبين أن أسلوب SVM هو الأكثر فعالية. يليه أسلوب الغابة العشوائية بنسبة دقة 71.52%، متفوقًا على الأساليب الأخرى. أما أسلوب treesC4.5 فقد سجل أدنى نسبة تشخيص صحيح بلغت 58.18%. مع ذلك، وبناءً على مؤشر مساحة منحنى ROC، تفوق أسلوب الغابة العشوائية (بمساحة تحت المنحنى 0.789) على أسلوب SVM (بمساحة تحت المنحنى 0.711). إجمالًا، أظهر أسلوبا SVM والغابة العشوائية تباينًا كبيرًا في تشخيص وفيات مرضى سرطان الكبد مقارنةً بالأساليب الأخرى.
Keywords
الشجرة العشوائية، C4.5، الغابة العشوائية، آلة المتجهات الداعمة، الانحدار اللوجستي، سرطان الكبد