ARTICLE
31 August 2026

Evaluating Large Language Models in Green Building Design Decisions: A Case-Based Analysis of Energy, Water, and Material Trade-Offs

Reem AL-qawas1 Dan Shao1*
Show Less
1 School of Art and Design, Dalian Polytechnic University, 116034, Dalian, China
JWA 2026 , 10(4), 14–18; https://doi.org/10.26689/JWA.v10i4.15268
© 2026 by the Author. Licensee: Bio-Byword Scientific Publishing Pty Ltd, Australia. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution 4.0 International License ( https://creativecommons.org/licenses/by/4.0/ )
Abstract

Background: AI has shown potential across technical professions, yet its application to green building design—navigating trade-offs between energy efficiency, water conservation, and materials lifecycle assessment (LCA)—remains unevaluated. No prior study has benchmarked frontier large language models (LLMs) using a professor-validated instrument combining multiple-choice question (MCQ) and True/False (T/F) items probing context-dependent sustainability reasoning. Objective: To comparatively benchmark four frontier LLMs—Claude, ChatGPT-5.2, DeepSeek-R1, and Gemini 3.1—on a validated 80-item instrument (40 MCQs and 40 True/False statements) derived from 20 green building design scenarios spanning materials, energy, water, and cross-domain lifecycle trade-off reasoning. Methods: 80 items from 20 original scenarios were validated by two professors. MCQs required selection of the optimal design alternative under competing sustainability constraints; True/False items used context-dependent partial truths (Statements 1–20: False via contextual override; 21–40: True) to evaluate resistance to overgeneralization. Items were grounded in LEED v4.1, ASHRAE 90.1, and ISO 14044. Each model was evaluated under identical zero-temperature conditions using binary scoring. Pearson Chi-square tests were applied for inter-model and inter-format comparisons. Results: Claude achieved the highest overall accuracy (95.00%; 95% CI: 90.22–99.78%), followed by DeepSeek-R1 (93.75%), ChatGPT-5.2 (88.75%), and Gemini 3.1 (83.75%). No statistically significant difference across chatbots was found (x2 = 7.251, df = 3, P = 0.064). No significant MCQ-vs-T/F differential was observed for any model (all P > 0.05); pooled T/F accuracy (91.3%) marginally exceeded pooled MCQ accuracy (89.4%). Conclusions: This is the first study to benchmark multiple frontier LLMs in green building design using a purpose-built, professor-validated instrument evaluating both design-alternative selection and resistance to absolute sustainability heuristics. Consistent performance declines in cross-domain lifecycle scenarios reveal residual multi-framework integration limitations across all models. Findings support LLM integration into green building education and decision-support, while underscoring the need for expert oversight in complex design contexts.

Keywords
Large language models
Green building design
Sustainable architecture
Benchmark evaluation
Energy efficiency
Water conservation
LCA
LEED
Claude
ChatGPT
DeepSeek
Gemini
References

[1] Biswas SS, 2023, Role of ChatGPT in public health. Annals of Biomedical Engineering, 2023(51): 868–869.

[2] Ali SR, Dobbs TD, Hutchings HA, et al., 2023, Using ChatGPT to Write Patient Clinic Letters. The Lancet Digital Health, 2023(5): e179–181.

[3] Wang Y, Liang L, Li R, et al., 2024, Comparison of ChatGPT, Claude, and Bard in Support of Technical Decision-making. Journal of Multidisciplinary Healthcare, 2024(17): 3917–3929.

[4] DeepSeek-AI, 2025, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv: 2501.12948.

[5] Wu J, Jiang M, Fan J, et al., 2025, Arch-Eval Benchmark for Assessing Chinese Architectural Domain Knowledge in LLMs. Scientific Reports, 2025(15): 1–12.

[6] He Q, 2024, How Well Do LLMs Perform in Green Building Assessment? iCCPMCE-2024.

[7] Nguyen HC, Dang HP, Nguyen TL, et al., 2025, Accuracy of Latest LLMs in Answering MCQs in Technical Domains. PLoS ONE, 2025(20): e0317423.

[8] U.S. Green Building Council, 2021, LEED v4.1 Building Design and Construction Reference Guide. USGBC; 2021.

[9] Lu C, Li S, Lu Z, 2021, Building Energy Prediction Using Artificial Neural Networks: A Literature Survey. Energy and Buildings, 2021(233): 110919.

[10] Luo X, Zhang Y, Lu J, et al., 2024, Multi-objective Optimization of Office Park Building Envelope for Nearly Zero Energy. Journal of Building Engineering, 2024(86): 108713.

[11] Yu L, Sun Y, Xu Z, et al., 2021, Multi-agent Deep Reinforcement Learning for HVAC Control. IEEE Transactions on Smart Grid, 12(1): 407–419.

[12] Zhong S, Aseniero BA, Groom AI, et al., 2025, Towards Interactive AI-assisted Material selection for Sustainable Building Design. Companion Publication of the 2025 ACM Designing Interactive Systems Conference, 567–573.

[13] Raeissi MM, Knapen R, 2025, Applications of Generative LLMs in Environmental Science: A Systematic Review. Advances in Environmental and Engineering Research, 6(2): 1–20.

[14] Gao Y, Yiu TW, Shen X, et al., 2026, LLMs in Smart Construction: A Systematic Review. Engineering, Construction and Architectural Management, 33(15): 159–181. https://doi.org/10.1108/ECAM-12-2024-1402

[15] Center for AI Safety, Scale AI, HLE Contributors Consortium, 2026, A Benchmark of Expert-level Academic Questions to Assess AI Capabilities. Nature, 2026(649): 1139–1146.

[16] Shen Y, Heacock L, Elias J, et al., 2023, ChatGPT and other LLMs are Double-Edged Swords. Radiology, 307(2): e230163. https://doi.org/10.1148/RADIOL.230163

[17] Kambhampati S, 2024, Can LLMs Plan? The Science and Nonsense of LLM Reasoning Claims. Communications of the ACM, 67(4): 34–41.

[18] Turpin M, Michael J, Perez E, et al., 2023, Language Models Don’t Always Say What They Think. Advances in Neural Information Processing Systems, 2023(36): 74952–74965.

Share
Back to top