Comparing Segmentation Models for Object Goal Navigation and Semantic Mapping Nesne Hedefli Navigasyon ve Semantik Haritalama için Segmentasyon Modellerinin Karşilaştirmasi


Aykor K., ATA B.

34th Signal Processing and Communications Applications Conference, SIU 2026, İstanbul, Türkiye, 7 - 10 Temmuz 2026, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/siu71813.2026.11636787
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Anahtar Kelimeler: Mask R-CNN, Object Goal Navigation, Semantic Mapping, Semantic Segmentation, YOLO
  • Çukurova Üniversitesi Adresli: Evet

Özet

This study compares five segmentation models for semantic mapping in indoor robot navigation: COCO-pretrained YOLOv8x-seg, YOLOv11x-seg, YOLOv26x-seg, and both pretrained and fine-tuned Mask R-CNN. The models are evaluated through 1,000-episode Object Goal Navigation experiments on the HM3D dataset and semantic map quality analysis across 7 scenes from the MP3D dataset. Fine-tuned Mask R-CNN achieves the highest navigation performance with 59.7% success rate and 28.8% SPL, yet produces the sparsest maps with an average of 4.0 categories and 657 cells. In contrast, pretrained YOLO models generate significantly richer maps with 8.1 - 8.6 categories and 3,862 - 4,997 cells on average, despite lower success rates of 40.8% - 43.8%. The results demonstrate that fine-tuning improves task performance on the target dataset while reducing general semantic mapping capability on different datasets.