Journal Article

·2021 OPEN ACCESS

Performance analysis of set partitioning formulations on the rule extraction from random forests

Mert Edali YTU

Pamukkale University Journal of Engineering Sciences

Abstract

Random Forests is a widely used machine learning algorithm for classification and regression problems from different domains. Although they are generally accurate, their interpretability is low compared to their building blocks: single decision trees. Using the fact that each member of a Random Forest is a decision tree, we propose different set partitioning formulations to extract interpretable if-then rules from Random Forests. Our experiments on well-known classification and regression datasets show that the original set partitioning model formulation significantly reduces the number of rules while keeping the accuracy at acceptable levels. We also propose a modification to the problem's objective function, which aims to reduce the number of extracted rules further. We observe a further reduction in the number of extracted rules while the accuracy values stay nearly the same. Although the set partitioning problem is NP-hard, we obtain optimal results for most datasets within twenty minutes.

Keywords

Extraction (chemistry) Computer science Random forest Set (abstract data type) Data mining Artificial intelligence Chromatography Chemistry Programming language

Subject Areas

Advanced Clustering Algorithms Research ·Artificial Intelligence ·Physical Sciences
Data Management and Algorithms ·Signal Processing ·Physical Sciences
Rough Sets and Fuzzy Logic ·Computational Theory and Mathematics ·Physical Sciences

Citations by Year

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Life in Land 58%