Explainable AI for delinquency risk monitoring in U.S. auto lease securitizations
-
DOIhttp://dx.doi.org/10.21511/imfi.23(3).2026.16
-
Article InfoVolume 23 2026, Issue #3, pp. 215–234
- 12 Views
-
3 Downloads
This work is licensed under a
Creative Commons Attribution 4.0 International License
Type of the article: Research Article
Abstract
This study evaluates whether explainable machine-learning models can provide reliable and operationally interpretable short-horizon delinquency monitoring in U.S. auto lease securitization panels. Six public SEC ABS-EE trust-family panels were harmonized at the contract-month level. The primary outcome is one-month-ahead incident escalation to 30 or more days past due among contracts below 30 days past due at the feature month. Fully tuned penalized logistic regression (M1), unconstrained gradient boosting (M2), and governance-constrained gradient boosting (M3/X-LEASE) were assessed through chronological development, calibration, locked out-of-time testing, contract-held-out validation, six leave-one-issuer-out experiments, and later temporal evaluation. The locked test comprised 198,301 observations and 768 events. M2 produced the strongest discrimination (AUC-ROC 0.8151; PR-AUC 0.0269), while M3 exceeded the logistic benchmark by PR-AUC but did not outperform M2. At an exact 1% review capacity, each model generated 1,984 alerts; M2 detected 104 events, compared with 61 for M3, so the hypothesized recall advantage of X-LEASE was not supported. Full-sample TreeSHAP analysis showed stable feature rankings across tested issuers, later periods, and contract-level resampling. M3 relied more heavily on credit score and issuer controls and had a more concentrated explanation structure, but this does not establish superior auditability or fairness. The results support human-reviewed early-warning monitoring within the tested securitization panels, not lifetime default prediction, automated adverse decisions, or universal transferability.
Acknowledgments
The authors acknowledge the public availability of SEC EDGAR Form ABS-EE, Exhibit 102 asset-level data. No individuals or institutions outside the author team provided paid analytical, editorial, or funding support for this manuscript.
- Keywords
-
JEL Classification (Paper profile tab)C53, C55, G23, G28
-
References29
-
Tables16
-
Figures4
-
- Figure 1. Temporal model-development and evaluation workflow
- Figure 2. Locked-test precision-recall and calibration curves
- Figure 3. Held-out-issuer and temporal robustness
- Figure 4. Full locked-test normalized raw-feature SHAP importance
-
- Table 1. Sample composition and selected descriptive statistics
- Table 2. Model specifications and implemented X-LEASE constraints
- Table 3. Locked-test discrimination, calibration, and probability performance
- Table 4. Exact matched top-1% confusion matrices and operational alert burden
- Table 5. Contract-held-out, LOIO, temporal, and fundamentals-only robustness summary
- Table 6. Global SHAP importance and explanation-stability summary
- Table A1. SEC source coverage and grouped analytical panels
- Table A2. Analytical sample chronology
- Table B1. Descriptive statistics and missingness summary
- Table C1. Model specifications and selected settings
- Table C2. Locked-test discrimination, calibration, and probability performance
- Table C3. Exact matched top-1% confusion matrices
- Table D1. Robustness summary
- Table D2. Explainability and stability diagnostics
- Table E1. Deterministically selected local cases
- Table F1. Layered governance interpretation of frozen X-LEASE outputs
-
- Aas, K., Jullum, M., & Løland, A. (2021). Explaining individual predictions when features are dependent: More accurate approximations to Shapley values. Artificial Intelligence, 298, 103502.
- Board of Governors of the Federal Reserve System. (2026, April 17). Revised guidance on model risk management (SR 26-2). Board of Governors of the Federal Reserve System, Office of the Comptroller of the Currency, & Federal Deposit Insurance Corporation.
- Brown, I., & Mues, C. (2012). An experimental comparison of classification algorithms for imbalanced credit scoring data sets. Expert Systems with Applications, 39(3), 3446-3453.
- Bücker, M., Szepannek, G., Gosiewska, A., & Biecek, P. (2022). Transparency, auditability, and explainability of machine learning models in credit scoring. Journal of the Operational Research Society, 73(1), 70-90.
- Bussmann, N., Giudici, P., Marinelli, D., & Papenbrock, J. (2021). Explainable machine learning in credit risk management. Computational Economics, 57(1), 203-216.
- Černevičienė, J., & Kabašinskas, A. (2024). Explainable artificial intelligence (XAI) in finance: A systematic literature review. Artificial Intelligence Review, 57, Article 216.
- Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794). Association for Computing Machinery.
- Crisanto, J. C., Leuterio, C. B., Prenio, J., & Yong, J. (2024). Regulating AI in the financial sector: Recent developments and main challenges (FSI Insights No. 63). Bank for International Settlements.
- Dryha, Z., Levchenko, O., Sergiienko, L., Kochubei, L., Domashenko, O., & Shcheglova, K. (2026). Reproducibility package for “Explainable AI for Delinquency Risk Monitoring in U.S. Auto Lease Securitizations” (Version 1.0.0) [Data set]. Zenodo.
- European Banking Authority (EBA). (2023). Follow-up report on the use of machine learning for internal ratings-based models (EBA/REP/2023/28).
- European Parliament and Council. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence. Official Journal of the European Union.
- Financial Stability Board (FSB). (2017). Artificial intelligence and machine learning in financial services: Market developments and financial stability implications.
- Financial Stability Board (FSB). (2024). The financial stability implications of artificial intelligence.
- Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189-1232.
- Hernes, M., Kozierkiewicz, A., Maleszka, M., Rot, A., Kozina, A., Matenczuk, K., Janus, J., & Wróbel, E. (2021). Deep learning for repayment prediction in leasing companies. European Research Studies Journal, 24(2), 1134-1148.
- Hong Kong Institute for Monetary and Financial Research (HKIMR). (2021). Artificial intelligence and big data in the financial services industry: A regional perspective and strategies for talent development (HKIMR Applied Research Report No. 2/2021).
- Kozina, A., Kuźmiński, Ł., Nadolny, M., Miałkowska, K., Tutak, P., Janus, J., Płotnicki, F., Walaszczyk, E., Rot, A., Dziembek, D., & Król, R. (2023). The default of leasing contracts prediction using machine learning. Procedia Computer Science, 225, 424-433.
- Kuiper, O., van den Berg, M., van der Burgt, J., & Leijnen, S. (2021). Exploring explainable AI in the financial sector: Perspectives of banks and supervisory authorities. arXiv.
- Kumar, I. E., Venkatasubramanian, S., Scheidegger, C., & Friedler, S. A. (2020). Problems with Shapley-value-based explanations as feature importance measures. In Proceedings of the 37th International Conference on Machine Learning (pp. 5491-5500). PMLR.
- Lessmann, S., Baesens, B., Seow, H.-V., & Thomas, L. C. (2015). Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research, 247(1), 124-136.
- Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765-4774.
- Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220-229). Association for Computing Machinery.
- National Institute of Standards and Technology (NIST). (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1).
- Pérez-Cruz, F., Prenio, J., Restoy, F., & Yong, J. (2025). Managing explanations: How regulators can address AI explainability (FSI Occasional Papers No. 24). Bank for International Settlements.
- Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1, 206-215.
- Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432.
- U.S. Securities and Exchange Commission (SEC). (n.d.). EDGAR full text search.
- Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., & Steyerberg, E. W. (2019). Calibration: The Achilles heel of predictive analytics. BMC Medicine, 17, Article 230.
- Wynants, L., van Smeden, M., McLernon, D. J., Timmerman, D., Steyerberg, E. W., & Van Calster, B. (2019). Three myths about risk thresholds for prediction models. BMC Medicine, 17, Article 192.


