Volume 10,Issue 4
Evaluating Large Language Models in Green Building Design Decisions: A Case-Based Analysis of Energy, Water, and Material Trade-Offs
Background: AI has shown potential across technical professions, yet its application to green building design—navigating trade-offs between energy efficiency, water conservation, and materials lifecycle assessment (LCA)—remains unevaluated. No prior study has benchmarked frontier large language models (LLMs) using a professor-validated instrument combining multiple-choice question (MCQ) and True/False (T/F) items probing context-dependent sustainability reasoning. Objective: To comparatively benchmark four frontier LLMs—Claude, ChatGPT-5.2, DeepSeek-R1, and Gemini 3.1—on a validated 80-item instrument (40 MCQs and 40 True/False statements) derived from 20 green building design scenarios spanning materials, energy, water, and cross-domain lifecycle trade-off reasoning. Methods: 80 items from 20 original scenarios were validated by two professors. MCQs required selection of the optimal design alternative under competing sustainability constraints; True/False items used context-dependent partial truths (Statements 1–20: False via contextual override; 21–40: True) to evaluate resistance to overgeneralization. Items were grounded in LEED v4.1, ASHRAE 90.1, and ISO 14044. Each model was evaluated under identical zero-temperature conditions using binary scoring. Pearson Chi-square tests were applied for inter-model and inter-format comparisons. Results: Claude achieved the highest overall accuracy (95.00%; 95% CI: 90.22–99.78%), followed by DeepSeek-R1 (93.75%), ChatGPT-5.2 (88.75%), and Gemini 3.1 (83.75%). No statistically significant difference across chatbots was found (x2 = 7.251, df = 3, P = 0.064). No significant MCQ-vs-T/F differential was observed for any model (all P > 0.05); pooled T/F accuracy (91.3%) marginally exceeded pooled MCQ accuracy (89.4%). Conclusions: This is the first study to benchmark multiple frontier LLMs in green building design using a purpose-built, professor-validated instrument evaluating both design-alternative selection and resistance to absolute sustainability heuristics. Consistent performance declines in cross-domain lifecycle scenarios reveal residual multi-framework integration limitations across all models. Findings support LLM integration into green building education and decision-support, while underscoring the need for expert oversight in complex design contexts.
[1] Biswas SS, 2023, Role of ChatGPT in public health. Annals of Biomedical Engineering, 2023(51): 868–869.
[2] Ali SR, Dobbs TD, Hutchings HA, et al., 2023, Using ChatGPT to Write Patient Clinic Letters. The Lancet Digital Health, 2023(5): e179–181.
[3] Wang Y, Liang L, Li R, et al., 2024, Comparison of ChatGPT, Claude, and Bard in Support of Technical Decision-making. Journal of Multidisciplinary Healthcare, 2024(17): 3917–3929.
[4] DeepSeek-AI, 2025, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv: 2501.12948.
[5] Wu J, Jiang M, Fan J, et al., 2025, Arch-Eval Benchmark for Assessing Chinese Architectural Domain Knowledge in LLMs. Scientific Reports, 2025(15): 1–12.
[6] He Q, 2024, How Well Do LLMs Perform in Green Building Assessment? iCCPMCE-2024.
[7] Nguyen HC, Dang HP, Nguyen TL, et al., 2025, Accuracy of Latest LLMs in Answering MCQs in Technical Domains. PLoS ONE, 2025(20): e0317423.
[8] U.S. Green Building Council, 2021, LEED v4.1 Building Design and Construction Reference Guide. USGBC; 2021.
[9] Lu C, Li S, Lu Z, 2021, Building Energy Prediction Using Artificial Neural Networks: A Literature Survey. Energy and Buildings, 2021(233): 110919.
[10] Luo X, Zhang Y, Lu J, et al., 2024, Multi-objective Optimization of Office Park Building Envelope for Nearly Zero Energy. Journal of Building Engineering, 2024(86): 108713.
[11] Yu L, Sun Y, Xu Z, et al., 2021, Multi-agent Deep Reinforcement Learning for HVAC Control. IEEE Transactions on Smart Grid, 12(1): 407–419.
[12] Zhong S, Aseniero BA, Groom AI, et al., 2025, Towards Interactive AI-assisted Material selection for Sustainable Building Design. Companion Publication of the 2025 ACM Designing Interactive Systems Conference, 567–573.
[13] Raeissi MM, Knapen R, 2025, Applications of Generative LLMs in Environmental Science: A Systematic Review. Advances in Environmental and Engineering Research, 6(2): 1–20.
[14] Gao Y, Yiu TW, Shen X, et al., 2026, LLMs in Smart Construction: A Systematic Review. Engineering, Construction and Architectural Management, 33(15): 159–181. https://doi.org/10.1108/ECAM-12-2024-1402
[15] Center for AI Safety, Scale AI, HLE Contributors Consortium, 2026, A Benchmark of Expert-level Academic Questions to Assess AI Capabilities. Nature, 2026(649): 1139–1146.
[16] Shen Y, Heacock L, Elias J, et al., 2023, ChatGPT and other LLMs are Double-Edged Swords. Radiology, 307(2): e230163. https://doi.org/10.1148/RADIOL.230163
[17] Kambhampati S, 2024, Can LLMs Plan? The Science and Nonsense of LLM Reasoning Claims. Communications of the ACM, 67(4): 34–41.
[18] Turpin M, Michael J, Perez E, et al., 2023, Language Models Don’t Always Say What They Think. Advances in Neural Information Processing Systems, 2023(36): 74952–74965.