Application of the XGBoost Method with CNN Feature Extraction for AI-Generated Video Detection
DOI:
https://doi.org/10.54783/influencejournal.v8i3.396Keywords:
AI video Detection, Convolutional Neural Network (CNN), Extreme Gradient Boosting (XGBoost), GenVidBench, Hybrid Approach.Abstract
The rapid advancement of generative artificial intelligence has substantially improved the visual quality of synthetic video, making it increasingly difficult to distinguish from authentic footage and creating new challenges for digital content verification. This study develops and evaluates an AI-generated video detection system using a hybrid approach that integrates a Convolutional Neural Network (CNN) as a feature extractor with Extreme Gradient Boosting (XGBoost) as the classifier. The research uses a 70,000-video subset of the GenVidBench-143k benchmark, selected through stratified sampling for computational efficiency. Feature extraction is performed in parallel using ResNet50, which captures high-level semantic features, and VGG16, which captures local texture detail; the resulting 2,048- and 512-dimensional vectors are concatenated into a single 2,560-dimensional representation per frame. A temporal aggregation stage transforms frame-level features into video-level features using Average Pooling and Max Pooling, and a Cost-Sensitive Learning strategy is applied through the scale_pos_weight, sample_weight, and max_delta_step parameters of XGBoost to address the marked class imbalance between real and AI-generated videos. The experimental results show that the hybrid ResNet50-VGG16 architecture combined with Average Pooling aggregation and scale_pos_weight optimization yields the best performance, achieving 92.93% accuracy, with 0.89 precision and 0.85 recall for the real-video class, and 0.94 precision and 0.96 recall for the AI-video class. A robustness analysis further shows that the model performs strongly on low-resolution video (<720p), reaching 94.00% accuracy, and maintains 88.00% accuracy on high-resolution video (>=720p); the modest decline at higher resolution reflects the increasingly seamless visual quality of modern generative models. These findings indicate that the hybrid CNN-XGBoost approach is a promising and computationally efficient strategy for synthetic video detection, although broader training-data variation is still required to reduce detection errors on high-quality video in real-world scenarios.
References
Abbas, F., & Taeihagh, A. (2024). Unmasking deepfakes: A systematic review of deepfake detection and generation techniques using artificial intelligence. Expert Systems with Applications, 252(PB), 124260. https://doi.org/10.1016/j.eswa.2024.124260
Bieder, F., Sandkühler, R., & Cattin, P. C. (n.d.). Comparison of methods generalizing max- and average-pooling.
Chen, H., & Wang, X. (2024). VideoCrafter2: Overcoming data limitations for high-quality video diffusion models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7310-7320.
Ding, G., Sener, F., & Yao, A. (2023). Temporal action segmentation: An analysis of modern techniques. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(2), 1011-1030.
Fernandes, Y. A., & Fatma, Y. (2025). Metode deep learning dalam teknologi deepfake: Systematic literature review. JATI (Jurnal Mahasiswa Teknik Informatika), 9(2), 3403-3410.
Francis, N. (2025). Deepfake detection and defense: An analysis of techniques and robustness. Graduate Thesis and Dissertation post-2024.
Heidari, A., Jafari Navimipour, N., Dag, H., & Unal, M. (2024). Deepfake detection using deep learning methods: A systematic and comprehensive review. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 14(2), e1520. https://doi.org/10.1002/widm.1520
Hong, W., Ding, M., Zheng, W., Liu, X., & Tang, J. (2022). Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868.
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., ... & Liu, Z. (2024, June). Vbench: Comprehensive benchmark suite for video generative models. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 21807-21818). IEEE.
Humidan, A. S., Abdullah, L. N., & Halin, A. A. (2022, December). Detection of compressed deepfake video drawbacks and technical developments. In 2022 5th International conference on signal processing and information security (ICSPIS) (pp. 11-16). IEEE.
Ismail, A., Elpeltagy, M., S. Zaki, M., & Eldahshan, K. (2021). A new deep learning-based methodology for video deepfake detection using XGBoost. Sensors, 21(16), 5413.
Jolliffe, I. T., & Cadima, J. (2016). Principal component analysis: A review and recent developments. Philosophical Transactions of the Royal Society A, 374(2065), 20150202. https://doi.org/10.1098/rsta.2015.0202
Karathanasis, A., Violos, J., & Kompatsiaris, I. (2025). A comparative analysis of compression and transfer learning techniques in deepfake detection models. Mathematics, 13(5), 887.
Khachatryan, L., Movsisyan, A., Tadevosyan, V., Henschel, R., Wang, Z., Navasardyan, S., & Shi, H. (2023, October). Text2video-zero: Text-to-image diffusion models are zero-shot video generators. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 15908-15918). IEEE.
Khalid, M., Raza, A., Younas, F., Rustam, F., Villar, M. G., Ashraf, I., & Akhtar, A. (2024). Novel sentiment majority voting classifier and transfer learning-based feature engineering for sentiment analysis of deepfake tweets. IEEE Access, 12, 67117-67129.
Kuang, L., Wang, Y., Hang, T., Chen, B., & Zhao, G. (2022). A dual-branch neural network for DeepFake video detection by detecting spatial and temporal inconsistencies. Multimedia Tools and Applications, 81(29), 42591-42606.
Li, J., Zhang, C., Zhu, W., & Ren, Y. (2025). A comprehensive survey of image generation models based on deep learning. Annals of Data Science, 12(1), 141-170. https://doi.org/10.1007/s40745-024-00544-1
Mancy, H., Elpeltagy, M., Eldahshan, K., & Ismail, A. (2025). Hybrid-Optimized Model for Deepfake Detection. International Journal of Advanced Computer Science & Applications, 16(4), 148.
Ni, Z. L., Qiangyu, Y. A. N., Yuan, T., Huang, M., Hu, H., Chen, X., & Wang, Y. (2025). Genvidbench: A challenging benchmark for detecting ai-generated video.
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., ... & Rombach, R. (2024, May). Sdxl: Improving latent diffusion models for high-resolution image synthesis. In International Conference on Learning Representations (Vol. 2024, pp. 1862-1874).
Pontorno, O., Guarnera, L., & Battiato, S. (2025). DeepFeatureX-SN: Generalization of deepfake detection via contrastive learning. Multimedia Tools and Applications, 84(39), 47721-47740.
Saikia, P., Dholaria, D., Yadav, P., Patel, V., & Roy, M. (2022, July). A hybrid CNN-LSTM model for video deepfake detection by leveraging optical flow features. In 2022 international joint conference on neural networks (IJCNN) (pp. 1-7). IEEE.
Sholeh, M., Lestari, U., & Andayati, D. (2025). Hyperparameter Optimization Using Grid Search and Random Search to Improve the Performance of Prediction Models with Decision Trees. Jurnal Riset Multidisiplin dan Inovasi Teknologi, 3(03), 453-464.
Trivedi, A. K., Mahajan, T., Maheshwari, T., Mehta, R., & Tiwari, S. (2025). Leveraging feature fusion ensemble of VGG16 and ResNet-50 for automated potato leaf abnormality detection in precision agriculture: AK Trivedi et al. Soft Computing, 29(4), 2263-2277.
Wang, C., Deng, C., & Wang, S. (2020). Imbalance-XGBoost: leveraging weighted and focal losses for binary label-imbalanced classification with XGBoost. Pattern recognition letters, 136, 190-197.
Wang, J., Yuan, H., Chen, D., Zhang, Y., Wang, X., & Zhang, S. (2023). Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571.
Zohra, F., Zhao, C., Liu, S., & Ghanem, B. (2025, June). Effectiveness of max-pooling for fine-tuning clip on videos. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 3282-3291). IEEE.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 INFLUENCE: INTERNATIONAL JOURNAL OF SCIENCE REVIEW

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.














