We’re excited to share that our paper “Challenges in the Evaluation of Machine Learning Techniques in Generative Urban Design” has been published in the International Journal of Architectural Computing (IJAC), as part of the eCAADe Special Issue 2026 on Digital Design and Data-Driven Intelligence.
The paper, co-authored by Haya Brama, Tal Grinshpoun, Agata Dalach, and Jonathan Dortheimer, tackles a question that has been largely overlooked as machine learning tools for generative urban design (GUD) have proliferated: how do we know if these tools are actually any good?
Traditional evaluation of urban design outputs relies on expert judgment, which is time-consuming, subjective, and struggles to keep pace with the sheer volume of designs that generative models can produce. Existing computational metrics, borrowed from other domains, often fail to capture what actually matters for urban design quality — our experiments show they can even favor disrupted, incoherent layouts over well-designed ones.
To address this gap, the paper poses three research questions: what evaluation methods have been used in GUD research, how suitable are they, and how can the evaluation process be improved? Combining a critical literature review with experimental testing, we propose two new strategies:
- Domain-Guided FID — an adaptation of the Fréchet Inception Distance metric, tailored to align with urban design principles rather than generic image similarity
- Visual Language Model assessment — using VLMs to emulate expert-style evaluation of generated design outputs
Together, these approaches offer a more robust and comprehensive way to assess generative urban design tools, and point toward the kind of standardized evaluation framework the field needs to build trust among designers, planners, and policymakers.
Read the full paper at the International Journal of Architectural Computing.