This study evaluates multiple machine-learning models for detecting highly imbalanced credit card fraud, finding that the configured XGBoost model achieved the best overall performance with an AP of 0.9084 and F1 score of 0.8534. However, the results are specific to the synthetic dataset and experimental setup, with several limitations preventing broader conclusions about XGBoost’s superiority.
Credit card fraud detection is a highly imbalanced classification problem in which headline accuracy can be misleading. In the held-out Kaggle test file used in this study, only 2,145 of 555,719 transactions (0.386%) were fraudulent; a classifier that labeled every transaction legitimate would therefore achieve approximately 99.61% accuracy while detecting no fraud. The analysis evaluated four configured machine-learning pipelines - Decision Tree, Random Forest, XGBoost, and a Multilayer Perceptron (MLP) - together with equal and validation-AP-weighted soft-voting ensembles. Average Precision (AP) was the primary ranking metric, along with precision, recall, F1 score, PR AUC, false positives (FP), and false negatives (FN) used to characterize operating tradeoffs. The selected XGBoost configuration produced the highest observed test AP (0.9084) and F1 score (0.8534) on this synthetic dataset. The soft-voting ensembles reduced some false positives but did not exceed XGBoost in AP. These findings apply to the specific configured pipelines, feature engineering choices, and recurring synthetic customer/merchant population studied here; they do not establish that XGBoost is generally the best fraud-detection method. The study also identifies limitations related to repeated validation use, unequal imbalance treatments, synthetic-data artifacts, probability calibration, and the absence of uncertainty estimates or entity-disjoint evaluation.
Related Projects