📄 Sciences Methods and Technologies
International Journal (SciMeTech)

Volume 2 · Issue 2 · 2026
ISSN: 3085-5284
Hybrid Multi-Person Tracking Framework for Dense and Dynamic Environments
Oussama Lachihab, My Ahmed El Kiram, Latifa Er-ray
Pages 186–191 · Laboratory of Computer Science and Smart Systems, Faculty of Science Semlalia, Cadi Ayyad University, Marrakech MOROCCO
Abstract
Multi-object tracking in crowded scenes remains challenging due to occlusions, appearance ambiguity, and unreliable detections. We propose a tracking framework that leverages both body and face cues to improve identity consistency. Our method uses YOLOv8s to detect persons and faces, followed by a spatial association step to link them into person-level observations. We then employ DINOv2 to extract feature embeddings from detected regions, which are used within a tracking pipeline that combines motion and appearance information. The tracker adopts a cascade matching strategy to associate detections across frames and handle challenging cases such as occlusions and missed detections. Experimental results on a sequence of MOT17 dataset demonstrate that incorporating complementary cues can help maintain identity consistency, while also revealing limitations in dense scenarios.
Keywords: Multi-object tracking, Crowded scenes, DINOv2, YOLOv8, Data association, Re-identification

References

  1. Sharma, H., & Kanwal, N. (2025). Video surveillance in smart cities: current status, challenges & future directions. Multimedia Tools and Applications, 84(16), 15787-15832.
  2. Quintana-Ramirez, I., Sequeira, L., & Ruiz-Mas, J. (2021). An edge-cloud approach for video surveillance in public transport vehicles. IEEE Latin America Transactions, 19(10), 1763-1771.
  3. Sujkowski, M., Kozuba, J., Uchroński, P., Banaś, A., Pulit, P., & Gryżewska, L. (2023). Artificial intelligence systems for supporting video surveillance operators at international airport. Transportation Research Procedia, 74, 1284-1291.
  4. Arroyo, R., Yebes, J. J., Bergasa, L. M., Daza, I. G., & Almazán, J. (2015). Expert video-surveillance system for realtime detection of suspicious behaviors in shopping malls. Expert systems with Applications, 42(21), 7991-8005.
  5. Song, X., Sun, L., Lei, J., Tao, D., Yuan, G., & Song, M. (2016). Event-based large scale surveillance video summarization. Neurocomputing, 187, 66-74.
  6. Stadler, D., & Beyerer, J. (2023). An improved association pipeline for multi-person tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 3170-3179).
  7. Fu, T., Chen, Y., Chen, Z., Zhao, M., Li, B., & Xue, X. (2025). CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios. arXiv preprint arXiv:2507.02479.
  8. Sun, P., Cao, J., Jiang, Y., Yuan, Z., Bai, S., Kitani, K., & Luo, P. (2022). Dancetrack: Multi-object tracking in uniform appearance and diverse motion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 2099321002).
  9. Luo, W., Zhao, X., & Kim, T. K. (2014). Multiple object tracking: A review. arXiv preprint arXiv:1409.7618, 1(1), 1.
  10. Veeramani, B., Raymond, J. W., & Chanda, P. (2018). DeepSort: deep convolutional networks for sorting haploid maize seeds. BMC bioinformatics, 19(Suppl 9), 289.
  11. Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., ... & Wang, X. (2022, October). Bytetrack: Multi-object tracking by associating every detection box. In European conference on computer vision (pp. 1-21). Cham: Springer Nature Switzerland.
  12. Zeng, F., Dong, B., Zhang, Y., Wang, T., Zhang, X., & Wei, Y. (2022, October). Mort: End-to-end multiple-object tracking with transformer. In European conference on computer vision (pp. 659-675). Cham: Springer Nature Switzerland.
  13. Meinhardt, T., Kirillov, A., Leal-Taixe, L., & Feichtenhofer, C. (2022). Trackformer: Multi-object tracking with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 8844-8854).
  14. Dendorfer, P., Rezatofighi, H., Milan, A., Shi, J., Cremers, D., Reid, I., ... & Leal-Taixe, L. (2020). Mot20: A benchmark for multi object tracking in crowded scenes. arXiv preprint arXiv:2003.09003.
  15. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., ... & Bojanowski, P. (2023). Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193.
  16. Lachihab, O., El Kiram, A., & Errayj, L. (2026). Fine-Tuning YOLOv8s for Unified Human and Face Detection in Crowded Environments. Engineering, Technology & Applied Science Research, 16(1), 32348-32356.