An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction

Guanting Shen; Zi Tian

An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction

Guanting Shen, Zi Tian

TL;DR

This work presents a novel multimodal HRI framework that combines advanced vision-language models, speech processing, and fuzzy logic to enable precise and adaptive control of a Dobot Magician robotic arm.

Abstract

Interpreting human intent accurately is a central challenge in human-robot interaction (HRI) and a key requirement for achieving more natural and intuitive collaboration between humans and machines. This work presents a novel multimodal HRI framework that combines advanced vision-language models, speech processing, and fuzzy logic to enable precise and adaptive control of a Dobot Magician robotic arm. The proposed system integrates Florence-2 for object detection, Llama 3.1 for natural language understanding, and Whisper for speech recognition, providing users with a seamless and intuitive interface for object manipulation through spoken commands. By jointly addressing scene perception and action planning, the approach enhances the reliability of command interpretation and execution. Experimental evaluations conducted on consumer-grade hardware demonstrate a command execution accuracy of 75\%, highlighting both the robustness and adaptability of the system. Beyond its current performance, the proposed architecture serves as a flexible and extensible foundation for future HRI research, offering a practical pathway toward more sophisticated and natural human-robot collaboration through tightly coupled speech and vision-language processing.

An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction

TL;DR

Abstract

Paper Structure (22 sections, 7 equations, 5 figures, 1 table)

This paper contains 22 sections, 7 equations, 5 figures, 1 table.

Introduction
System Architecture
Hardware Components
Processing Distribution
System and Data Workflow
User Interaction
Methods
Robot End-Effector Detection
Speech Processing
Wake-Up Command Recognition
Speech-to-Text Conversion
Action Extraction from Text
Object Detection with Florence-2
Fuzzy Logic Control
Controller Structure
...and 7 more sections

Figures (5)

Figure 1: Information flow. Hardware elements and data flow of the robotic manipulation system.
Figure 2: UI Handler Process.
Figure 3: Robot controller.
Figure 4: Percentage contribution of different stages to total time taken and aggregated error.
Figure 5: Lemon and hand detection.Left - The system identifies the user’s hand and the lemon object. Right - The robot positions itself to pick up the selected object.

An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction

TL;DR

Abstract

An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction

Authors

TL;DR

Abstract

Table of Contents

Figures (5)