ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Ruiping Liu; Jiaming Zhang; Angela Schön; Karin Müller; Junwei Zheng; Kailun Yang; Anhong Guo; Kathrin Gerling; Rainer Stiefelhagen

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Ruiping Liu, Jiaming Zhang, Angela Schön, Karin Müller, Junwei Zheng, Kailun Yang, Anhong Guo, Kathrin Gerling, Rainer Stiefelhagen

TL;DR

This paper introduces ObjectFinder, a wearable system that unifies open-vocabulary object detection (YOLO-World) with a multimodal language model (GPT-4) to support open-ended, interactive search for blind users in unfamiliar environments. The design emphasizes flexible target queries, real-time egocentric localization, and intent-driven feedback branches (navigation and scene description/open questions), implemented across three modules and validated by an exploratory study with eight blind participants against BeMyAI and Google Lookout. Findings show that ObjectFinder provides essential localization and scene-context information, enhances independence, and supports discovery of incidental targets, though users desire customization of information load and smoother hardware interaction. The work highlights practical opportunities and tensions in integrating vision- and AI-based assistive tech, offering concrete directions for future development, including richer descriptions, reliability improvements, and portable, socially acceptable hardware.

Abstract

Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing description- and detection-based assistive technologies do not sufficiently support the multifaceted nature of interactive object search tasks. We present ObjectFinder, an open-vocabulary wearable assistive system for interactive object search by blind people. ObjectFinder allows users to query target objects using flexible wording. Once the target object is detected, it provides egocentric localization information in real-time, including distance and direction. Users can then initiate different branches to gather detailed information based on their intent towards the target object, such as navigating to it or perceiving its surroundings. ObjectFinder is powered by a seamless combination of open-vocabulary models, namely an open-vocabulary object detector and a multimodal large language model. The ObjectFinder design concept and its development were carried out in collaboration with a blind co-designer. To evaluate ObjectFinder, we conducted an exploratory user study with eight blind participants. We compared ObjectFinder to BeMyAI and Google Lookout, popular description- and detection-based assistive applications. Our findings indicate that most participants felt more independent with ObjectFinder and preferred it for object search, as it enhanced scene context gathering and navigation, and allowed for active target identification. Finally, we discuss the implications for future assistive systems to support interactive object search.

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

TL;DR

Abstract

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (9)