Multi-View Learning with Multimodal Fusion for Urban Walkability Prediction
Main Article Content
Abstract
Urban walkability is central to public health, sustainability, and quality of life, yet traditional assessments based on surveys and field audits are costly and difficult to scale. Recent computational approaches have used either satellite imagery, street view imagery, or population indicators, but each captures only a partial view: satellites provide spatial layout, street views capture pedestrian conditions, and population data reflect activity patterns. We present WalkCLIP: a multimodal framework that fuses these complementary perspectives to predict urban walkability. Our approach adapts vision–language representations with walkability-focused captions, enhances them with a spatial aggregation module to capture neighborhood context, and integrates non-visual population dynamics embeddings. Evaluated on 4,660 locations in Minneapolis–Saint Paul, the model outperforms unimodal and multimodal baselines in both predictive accuracy and spatial consistency. These findings demonstrate that combining visual and behavioral signals provides a scalable and reliable method for assessing urban walkability.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

All work in MURAJ is licensed under a Creative Commons Attribution-Noncommercial 4.0 License
Copyright remains with the individual authors.