Table of Contents
Fetching ...

Semantic4Safety: Causal Insights from Zero-shot Street View Imagery Segmentation for Urban Road Safety

Huan Chen, Ting Han, Siyu Chen, Zhihao Guo, Yiping Chen, Meiliu Wu

TL;DR

The paper tackles the challenge of deriving street-level accident risk indicators from street-view imagery and quantifying their causal effects across accident types. It introduces Semantic4Safety, which uses zero-shot semantic segmentation to extract eleven streetscape indicators from multiview street-view images at about thirty thousand accident sites in Austin, models five accident types with an XGBoost classifier, and interprets predictions with SHAP. A causal pipeline based on Generalized Propensity Score weighting and Average Treatment Effect estimation reveals heterogeneous, accident-type–specific effects, with scene complexity, exposure, and road geometry emerging as dominant drivers. The framework yields interpretable risk assessments and corridor-level diagnostics, offering a scalable, data-informed tool to guide urban planning and safety interventions.

Abstract

Street-view imagery (SVI) offers a fine-grained lens on traffic risk, yet two fundamental challenges persist: (1) how to construct street-level indicators that capture accident-related features, and (2) how to quantify their causal impacts across different accident types. To address these challenges, we propose Semantic4Safety, a framework that applies zero-shot semantic segmentation to SVIs to derive 11 interpretable streetscape indicators, and integrates road type as contextual information to analyze approximately 30,000 accident records in Austin. Specifically, we train an eXtreme Gradient Boosting (XGBoost) multi-class classifier and use Shapley Additive Explanations (SHAP) to interpret both global and local feature contributions, and then apply Generalized Propensity Score (GPS) weighting and Average Treatment Effect (ATE) estimation to control confounding and quantify causal effects. Results uncover heterogeneous, accident-type-specific causal patterns: features capturing scene complexity, exposure, and roadway geometry dominate predictive power; larger drivable area and emergency space reduce risk, whereas excessive visual openness can increase it. By bridging predictive modeling with causal inference, Semantic4Safety supports targeted interventions and high-risk corridor diagnosis, offering a scalable, data-informed tool for urban road safety planning.

Semantic4Safety: Causal Insights from Zero-shot Street View Imagery Segmentation for Urban Road Safety

TL;DR

The paper tackles the challenge of deriving street-level accident risk indicators from street-view imagery and quantifying their causal effects across accident types. It introduces Semantic4Safety, which uses zero-shot semantic segmentation to extract eleven streetscape indicators from multiview street-view images at about thirty thousand accident sites in Austin, models five accident types with an XGBoost classifier, and interprets predictions with SHAP. A causal pipeline based on Generalized Propensity Score weighting and Average Treatment Effect estimation reveals heterogeneous, accident-type–specific effects, with scene complexity, exposure, and road geometry emerging as dominant drivers. The framework yields interpretable risk assessments and corridor-level diagnostics, offering a scalable, data-informed tool to guide urban planning and safety interventions.

Abstract

Street-view imagery (SVI) offers a fine-grained lens on traffic risk, yet two fundamental challenges persist: (1) how to construct street-level indicators that capture accident-related features, and (2) how to quantify their causal impacts across different accident types. To address these challenges, we propose Semantic4Safety, a framework that applies zero-shot semantic segmentation to SVIs to derive 11 interpretable streetscape indicators, and integrates road type as contextual information to analyze approximately 30,000 accident records in Austin. Specifically, we train an eXtreme Gradient Boosting (XGBoost) multi-class classifier and use Shapley Additive Explanations (SHAP) to interpret both global and local feature contributions, and then apply Generalized Propensity Score (GPS) weighting and Average Treatment Effect (ATE) estimation to control confounding and quantify causal effects. Results uncover heterogeneous, accident-type-specific causal patterns: features capturing scene complexity, exposure, and roadway geometry dominate predictive power; larger drivable area and emergency space reduce risk, whereas excessive visual openness can increase it. By bridging predictive modeling with causal inference, Semantic4Safety supports targeted interventions and high-risk corridor diagnosis, offering a scalable, data-informed tool for urban road safety planning.
Paper Structure (23 sections, 8 equations, 9 figures, 3 tables)

This paper contains 23 sections, 8 equations, 9 figures, 3 tables.

Figures (9)

  • Figure 1: Overview of the proposed Semantic4Safety framework. We first process SVI using zero-shot segmentation to extract 11 indicators across four categories. These indicators are then used to predict five accident types through XGBoost, with SHAP providing both local and global-level interpretability. Finally, Generalized Propensity Score (GPS) weighting and ATE estimation are applied to quantify the causal effects of indicators across accident categories.
  • Figure 2: The study area is located in Austin, Texas, United States. The main panel provides a detailed view of the study area boundaries within Austin, where SVI and traffic accident data were collected and analyzed.
  • Figure 3: Illustration of input SVIs and corresponding zero-shot semantic segmentation outputs.
  • Figure 4: Spatial distribution of five accident categories across the study area.
  • Figure 5: Spatial distribution of street‑view indicators and road types across the study area.
  • ...and 4 more figures