[1] Bochkovskiy, A., Wang, C. Y., & Liao, H. Y. M. (2020). YOLOv4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934.
[2] Dosovitskiy, A., Beyer, L., Kolesnikov, A., et al. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR).
[3] Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A. (2010). The Pascal Visual Object Classes Challenge. International Journal of Computer Vision, 88(2), 303–338.
[4] He, K., Gkioxari, G., Dollar, P., & Girshick, R. (2020). Mask R-CNN. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(2), 386–397.
[5] Howard, A. G., Zhu, M., Chen, B., et al. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861.
[6] Kaggle. (2025). Carvana Image Masking Challenge Dataset. Available at:
https://www.kaggle.com/c/ carvana-image-masking-challenge.
[7] Lin, T. Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017). Feature Pyramid Networks for Object Detection. CVPR.
[8] Lin, T. Y., Goyal, P., Girshick, R., He, K., and Dollar, P. (2017). Focal Loss for Dense Object Detection. ICCV.
[9] Lundberg, S. M., and Lee, S. I. (2017). A Unified Approach to Interpreting Model Predictions. In NeurIPS.
[10] Redmon, J., & Farhadi, A. (2018). YOLOv3: An incremental improvement. arXiv preprint arXiv:1804.02767.
[11] Ren, S., He, K., Girshick, R., & Sun, J. (2017). Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6), 1137–1149.
[12] Samek, W., Wiegand, T., & Muller, K. R. (2019). Explainable artificial intelligence: Understanding, visualizing and interpreting deep learning models. IEEE Signal Processing Magazine, 36(6), 56–65.
[13] Selvaraju, R. R., Cogswell, M., Das, A., et al. (2020). Grad-CAM: Visual explanations from deep networks via gradient-based localization. International Journal of Computer Vision, 128(2), 336–359.
[14] Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning (ICML).
[15] The Chinese University of Hong Kong. (2025). CompCars Dataset. Available at: http://mmlab.ie.cuhk.edu.hk/datasets/comp_cars.
[16] Ultralytics. YOLOv8 Documentation. GitHub Repository. Available at: https://github.com/ultralytics/ultralytics.