Abstract
This study focuses on detecting a specific type of financial statement fraud in which companies artificially inflate their profitability by underreporting the cost of sales and adding the difference to inventories. The dataset used in this study comprises real financial statements from 500 publicly listed companies in Turkey, obtained from the Public Disclosure Platform, all of which received unqualified audit opinions. Of these, 300 samples were used as-is (non-fraudulent), while 200 were manually manipulated by domain experts based on a defined fraud scenario, creating synthetic fraudulent instances. The dataset includes 23 financial ratios derived from balance sheet and income statement items. Several machine learning algorithms—including Random Forest, XGBoost, Gradient Boosting, Support Vector Machine, and Multilayer Perceptron—were evaluated under two scenarios: a baseline using all features and a feature selection scenario using a correlation-based method. Model performance was assessed via 5-fold stratified cross-validation and test evaluation using accuracy, precision, recall, and F1-score. The results indicate that correlation-based feature selection improved model performance, particularly in ensemble models like Random Forest and Bagged Trees. This study demonstrates that machine learning can effectively detect manipulations related to the cost of sales in financial reporting.
Keywords
Subject Areas
OpenAlex SDG Match
SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).