Abstract
Recent advancements in remote sensing technologies have enabled the development of innovative methods to improve search and rescue efforts, assess structural damage, and streamline disaster management following earthquakes. Various imaging modalities—such as satellite imagery, UAVs, airborne platforms, and terrestrial systems—are utilized for damage evaluation and debris detection. Distinguishing post-earthquake imagery from pre-event data captured by UAVs is essential for effective disaster response. In this context, deep learning models have been increasingly employed to rapidly classify images as "damaged" or "undamaged." This study offers a comprehensive comparison of multiple deep learning architectures for image classification, including six convolutional neural networks (CNNs) and three variants of Vision Transformers (ViTs). The results show that Vision Transformer models, especially the ViT Large variant, consistently outperformed CNN counterparts, achieving an accuracy of 96.12%, sensitivity of 99.18%, and MCC of 0.9187. The superior performance of ViTs highlights their enhanced ability to discriminate between damaged and undamaged classes, providing balanced and reliable results. Furthermore, the ViT Large model surpassed results from previous earthquake image classification studies, demonstrating strong generalization on domain-specific datasets. These outcomes underscore the promise of transformer-based architectures as leading-edge solutions for complex and imbalanced image classification challenges in disaster assessment.
Keywords
Subject Areas
Related YTU Activities
Sustainability activities and projects connected to this article.