A Lightweight Hybrid CNN-BiLSTM Model with Adaptive Temporal Attention for Violence Detection in Surveillance Videos
DOI:
https://doi.org/10.47839/ijc.25.2.4652Keywords:
violence detection, edge devices, adaptive temporal attention, MobileNetV3-S, Bi-LSTMAbstract
Detecting violent actions in surveillance videos is a critical task for ensuring public safety in smart city environments. Although numerous deep learning approaches have been proposed, most of them favor either accuracy or computational efficiency, but rarely achieve both simultaneously. This limitation restricts their deployment on resource-constrained edge devices. In this paper, we propose a lightweight yet effective hybrid architecture that combines MobileNetV3-S for spatial feature extraction with a BiLSTM enhanced by an Adaptive Temporal Attention mechanism for temporal modeling. Despite its compact design (1.00 GFLOPs and 3.66 million parameters), the proposed model achieves competitive performance on four public benchmark datasets. The experimental results demonstrate that the proposed approach provides an excellent trade-off between accuracy and efficiency, making it suitable for real-time smart surveillance applications.
References
H. M. B. Jahlan and L. A. Elrefaei, “Detecting violence in video based on deep features fusion technique,” arXiv preprint arXiv:2204.07443, 2022.
A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan et al., “Searching for mobilenetv3,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1314–1324.
A. Traoré and M. A. Akhloufi, “Violence detection in videos using deep recurrent and convolutional neural networks,” in 2020 IEEE international conference on systems, man, and cybernetics (SMC). IEEE, 2020, pp. 154–159.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), vol. 1. Ieee, 2005, pp. 886–893.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proceedings of the IEEE international conference on computer vision, 2015, pp. 4489– 4497.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” Advances in neural information processing systems, vol. 27, 2014.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 2625–2634.
S. Sharma, B. Sudharsan, S. Naraharisetti, V. Trehan, and K. Jayavel, “A fully integrated violence detection system using cnn and lstm.” International Journal of Electrical & Computer Engineering (2088-8708), vol. 11, no. 4, 2021.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803.
M.-S. Kang, R.-H. Park, and H.-M. Park, “Efficient spatio-temporal modeling methods for real-time violence recognition,” IEEE Access, vol. 9, pp. 76 270–76 285, 2021.
H. Chen, X. Mei, Z. Ma, X. Wu, and Y. Wei, “Spatial–temporal graph attention network for video anomaly detection,” Image and Vision Computing, vol. 131, p. 104629, 2023.
S. Hochreiter, “The vanishing gradient problem during learning recurrent neural nets and problem solutions,” International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, vol. 6, no. 02, pp. 107–116, 1998.
Z. Liu, L. Wang, W. Wu, C. Qian, and T. L. Tam, “Temporal adaptive module for video recognition. in 2021 ieee,” in CVF International Conference on Computer Vision (ICCV), 2021, pp. 13 688–13 698.
¸S. Aktı, G. A. Tataroglu, and H. K. Ekenel, “Vision-based fight detection ˘ from surveillance cameras,” in 2019 Ninth International Conference on Image Processing Theory, Tools and Applications (IPTA). IEEE, 2019, pp. 1–6.
E. Bermejo Nievas, O. Deniz Suarez, G. Bueno García, and R. Sukthankar, “Violence detection in video using computer vision techniques,” in Computer Analysis of Images and Patterns: 14th International Conference, CAIP 2011, Seville, Spain, August 29-31, 2011, Proceedings, Part II 14. Springer, 2011, pp. 332–339.
M. Cheng, K. Cai, and M. Li, “Rwf-2000: An open large scale video database for violence detection,” in 2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021, pp. 4183–4190.
S. Sudhakaran and O. Lanz, “Learning to detect violent videos using convolutional long short-term memory,” in 2017 14th IEEE international conference on advanced video and signal based surveillance (AVSS). IEEE, 2017, pp. 1–6.
R. Halder and R. Chatterjee, “Cnn-bilstm model for violence detection in smart surveillance,” SN Computer science, vol. 1, no. 4, p. 201, 2020.
L. Ciampi, C. Santiago, J. P. Costeira, F. Falchi, C. Gennaro, and G. Amato, “Unsupervised domain adaptation for video violence detection in the wild.” in IMPROVE, 2023, pp. 37–46.
Q. Liang, Y. Li, B. Chen, and K. Yang, “Violence behavior recognition of two-cascade temporal shift module with attention mechanism,” Journal of Electronic Imaging, vol. 30, no. 4, pp. 043 009–043 009, 2021.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 6450–6459.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
M. Khan, A. El Saddik, W. Gueaieb, G. De Masi, and F. Karray, “Vd-net: An edge vision-based surveillance system for violence detection,” IEEE Access, vol. 12, pp. 43 796–43 808, 2024.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6202–6211.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6299–6308.
Downloads
Published
How to Cite
Issue
Section
License
International Journal of Computing is an open access journal. Authors who publish with this journal agree to the following terms:• Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
• Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
• Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.