A visual questioning answering approach to enhance robot localization in indoor environments

Peña-Narvaez, Juan Diego; Martín, Francisco; Guerrero Hernández, José Miguel; Pérez-Rodríguez, Rodrigo

A visual questioning answering approach to enhance robot localization in indoor environments

Archivos

fnbot-17-1290584 (1).pdf (2.77 MB)

Fecha

2023-11-27

Autores

Peña-Narvaez, Juan Diego

Martín, Francisco

Guerrero Hernández, José Miguel

Pérez-Rodríguez, Rodrigo

Editor

Frontiers in Neurorobotics

URI

https://hdl.handle.net/10115/27453

DOI

https://doi.org/10.3389/fnbot.2023.1290584

Citas

0 citas en

Resumen

Navigating robots with precision in complex environments remains a significant challenge. In this article, we present an innovative approach to enhance robot localization in dynamic and intricate spaces like homes and offices. We leverage Visual Question Answering (VQA) techniques to integrate semantic insights into traditional mapping methods, formulating a novel position hypothesis generation to assist localization methods, while also addressing challenges related to mapping accuracy and localization reliability. Our methodology combines a probabilistic approach with the latest advances in Monte Carlo Localization methods and Visual Language models. The integration of our hypothesis generation mechanism results in more robust robot localization compared to existing approaches. Experimental validation demonstrates the effectiveness of our approach, surpassing state-of-the-art multi-hypothesis algorithms in both position estimation and particle quality. This highlights the potential for accurate self-localization, even in symmetric environments with large corridor spaces. Furthermore, our approach exhibits a high recovery rate from deliberate position alterations, showcasing its robustness. By merging visual sensing, semantic mapping, and advanced localization techniques, we open new horizons for robot navigation. Our work bridges the gap between visual perception, semantic understanding, and traditional mapping, enabling robots to interact with their environment through questions and enrich their map with valuable insights. The code for this project is available on GitHub "https://github.com/juandpenan/topology_nav_ros2"

Descripción

The usage of a visual large language model to localize a robot in an indoor environment

Palabras clave

visual question answering , robot localization , robot navigation , semantic map , robot mapping

Citación

Peña-Narvaez JD, Martín F, Guerrero JM and Pérez-Rodríguez R (2023) A visual questioning answering approach to enhance robot localization in indoor environments. Front. Neurorobot. 17:1290584. doi: 10.3389/fnbot.2023.1290584

Colecciones

Artículos de Revista

Página completa del ítem

Excepto si se señala otra cosa, la licencia del ítem se describe como Atribución 4.0 Internacional

A visual questioning answering approach to enhance robot localization in indoor environments

Archivos

Fecha

Autores

Título de la revista

ISSN de la revista

Título del volumen

Editor

URI

DOI

Citas

Resumen

Descripción

Palabras clave

Citación

Colecciones