Table of Contents
Fetching ...

Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda

André Torneiro, Diogo Monteiro, Paulo Novais, Pedro Rangel Henriques, Nuno F. Rodrigues

TL;DR

This systematic review addresses how zero-/few-shot Vision-Language Models can transform urban infrastructure monitoring by surveying 32 studies (2021–2025) and organizing them into seven application domains. It reveals a landscape dominated by street-view datasets and modular backbones, with growing use of multimodal VLMs but persistent challenges in cross-city generalization, deployment feasibility, and ethical considerations. The authors propose a five-pillar research agenda—hybrid lightweight architectures, unified benchmarks, deployment-first design, embedded ethics, and reproducible ecosystems—to drive toward scalable, inclusive, and real-world urban AI. The work provides a structured map and concrete directions for researchers and practitioners aiming to deploy robust, transparent, and context-aware urban perception systems at city scale.

Abstract

Urban monitoring of public infrastructure (such as waste bins, road signs, vegetation, sidewalks, and construction sites) poses significant challenges due to the diversity of objects, environments, and contextual conditions involved. Current state-of-the-art approaches typically rely on a combination of IoT sensors and manual inspections, which are costly, difficult to scale, and often misaligned with citizens' perception formed through direct visual observation. This raises a critical question: Can machines now "see" like citizens and infer informed opinions about the condition of urban infrastructure? Vision-Language Models (VLMs), which integrate visual understanding with natural language reasoning, have recently demonstrated impressive capabilities in processing complex visual information, turning them into a promising technology to address this challenge. This systematic review investigates the role of VLMs in urban monitoring, with particular emphasis on zero-shot applications. Following the PRISMA methodology, we analyzed 32 peer-reviewed studies published between 2021 and 2025 to address four core research questions: (1) What urban monitoring tasks have been effectively addressed using VLMs? (2) Which VLM architectures and frameworks are most commonly used and demonstrate superior performance? (3) What datasets and resources support this emerging field? (4) How are VLM-based applications evaluated, and what performance levels have been reported?

Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda

TL;DR

This systematic review addresses how zero-/few-shot Vision-Language Models can transform urban infrastructure monitoring by surveying 32 studies (2021–2025) and organizing them into seven application domains. It reveals a landscape dominated by street-view datasets and modular backbones, with growing use of multimodal VLMs but persistent challenges in cross-city generalization, deployment feasibility, and ethical considerations. The authors propose a five-pillar research agenda—hybrid lightweight architectures, unified benchmarks, deployment-first design, embedded ethics, and reproducible ecosystems—to drive toward scalable, inclusive, and real-world urban AI. The work provides a structured map and concrete directions for researchers and practitioners aiming to deploy robust, transparent, and context-aware urban perception systems at city scale.

Abstract

Urban monitoring of public infrastructure (such as waste bins, road signs, vegetation, sidewalks, and construction sites) poses significant challenges due to the diversity of objects, environments, and contextual conditions involved. Current state-of-the-art approaches typically rely on a combination of IoT sensors and manual inspections, which are costly, difficult to scale, and often misaligned with citizens' perception formed through direct visual observation. This raises a critical question: Can machines now "see" like citizens and infer informed opinions about the condition of urban infrastructure? Vision-Language Models (VLMs), which integrate visual understanding with natural language reasoning, have recently demonstrated impressive capabilities in processing complex visual information, turning them into a promising technology to address this challenge. This systematic review investigates the role of VLMs in urban monitoring, with particular emphasis on zero-shot applications. Following the PRISMA methodology, we analyzed 32 peer-reviewed studies published between 2021 and 2025 to address four core research questions: (1) What urban monitoring tasks have been effectively addressed using VLMs? (2) Which VLM architectures and frameworks are most commonly used and demonstrate superior performance? (3) What datasets and resources support this emerging field? (4) How are VLM-based applications evaluated, and what performance levels have been reported?
Paper Structure (27 sections, 6 figures, 14 tables)

This paper contains 27 sections, 6 figures, 14 tables.

Figures (6)

  • Figure 1: Flowchart of the review protocol
  • Figure 2: A functional taxonomy of VLM applications in urban contexts. The framework categorizes the 32 reviewed studies into seven key domains based on their primary research goal. The references for studies associated with a node are listed beneath the respective node.
  • Figure 3: Functional roles of models across reviewed studies. Vision-only architectures are dominant, but VLMs and hybrid integrations are gaining traction.
  • Figure 4: Most cited models across the reviewed corpus. CLIP, Grounding DINO, and GPT-3.5 are among the most prevalent.
  • Figure 5: CityBench: Overview of the eight urban task families used to evaluate LLM/VLM capabilities across cities. Source from CityBench
  • ...and 1 more figures