Table of Contents
Fetching ...

RaycastGrasp: Eye-Gaze Interaction with Wearable Devices for Robotic Manipulation

Zitiantao Lin, Yongpeng Sang, Yang Ye

TL;DR

RaycastGrasp addresses joystick limitations in assistive robotics by enabling hands-free, gaze-driven manipulation in an egocentric MR environment. It integrates first-person gaze raycasting with YOLOv8-based object recognition and semantic label mapping to align user intent with robot perception, enabling a Franka Emika arm to grasp real-world objects. Experimental results show high reliability, including a spatial precision of $0.05~\mathrm{m}$ and single-pass intention/object recognition accuracy above $>88\%$, with object detection confidence also exceeding $>88\%$. This approach enhances intuitiveness and accessibility for assistive robotics and lays groundwork for more robust multi-object disambiguation in real-world settings.

Abstract

Robotic manipulators are increasingly used to assist individuals with mobility impairments in object retrieval. However, the predominant joystick-based control interfaces can be challenging due to high precision requirements and unintuitive reference frames. Recent advances in human-robot interaction have explored alternative modalities, yet many solutions still rely on external screens or restrictive control schemes, limiting their intuitiveness and accessibility. To address these challenges, we present an egocentric, gaze-guided robotic manipulation interface that leverages a wearable Mixed Reality (MR) headset. Our system enables users to interact seamlessly with real-world objects using natural gaze fixation from a first-person perspective, while providing augmented visual cues to confirm intent and leveraging a pretrained vision model and robotic arm for intent recognition and object manipulation. Experimental results demonstrate that our approach significantly improves manipulation accuracy, reduces system latency, and achieves single-pass intention and object recognition accuracy greater than 88% across multiple real-world scenarios. These results demonstrate the system's effectiveness in enhancing intuitiveness and accessibility, underscoring its practical significance for assistive robotics applications.

RaycastGrasp: Eye-Gaze Interaction with Wearable Devices for Robotic Manipulation

TL;DR

RaycastGrasp addresses joystick limitations in assistive robotics by enabling hands-free, gaze-driven manipulation in an egocentric MR environment. It integrates first-person gaze raycasting with YOLOv8-based object recognition and semantic label mapping to align user intent with robot perception, enabling a Franka Emika arm to grasp real-world objects. Experimental results show high reliability, including a spatial precision of and single-pass intention/object recognition accuracy above , with object detection confidence also exceeding . This approach enhances intuitiveness and accessibility for assistive robotics and lays groundwork for more robust multi-object disambiguation in real-world settings.

Abstract

Robotic manipulators are increasingly used to assist individuals with mobility impairments in object retrieval. However, the predominant joystick-based control interfaces can be challenging due to high precision requirements and unintuitive reference frames. Recent advances in human-robot interaction have explored alternative modalities, yet many solutions still rely on external screens or restrictive control schemes, limiting their intuitiveness and accessibility. To address these challenges, we present an egocentric, gaze-guided robotic manipulation interface that leverages a wearable Mixed Reality (MR) headset. Our system enables users to interact seamlessly with real-world objects using natural gaze fixation from a first-person perspective, while providing augmented visual cues to confirm intent and leveraging a pretrained vision model and robotic arm for intent recognition and object manipulation. Experimental results demonstrate that our approach significantly improves manipulation accuracy, reduces system latency, and achieves single-pass intention and object recognition accuracy greater than 88% across multiple real-world scenarios. These results demonstrate the system's effectiveness in enhancing intuitiveness and accessibility, underscoring its practical significance for assistive robotics applications.
Paper Structure (15 sections, 3 equations, 5 figures)

This paper contains 15 sections, 3 equations, 5 figures.

Figures (5)

  • Figure 1: Overview of the RaycastGrasp system
  • Figure 2: Alignment of virtual passthrough layer and physical layer
  • Figure 3: Class label mapping between user gaze detection and robot camera recognition
  • Figure 4: User observing real objects through VR headset in a mixed reality environment
  • Figure 5: Object view from the robot-mounted camera perspective