Back To Projects

A Comprehensive Review on Deep Learning Architectures for Image Segmentation

Madhurima M.

This paper provides a comprehensive review of recent advancements in deep learning-based models for image segmentation, including U-Net, Mask R-CNN, and transformer-based models, analyzing their performance, challenges like data shortage, and potential future directions for improved real-time deployment and interpretability.


At a micro level, image segmentation is an important computer vision task that assists in applications in various fields such as medical diagnostics, self-driving cars, and robotics. The introduction of deep learning in the scene led to great improvements in the performance of segmentation models in terms of accuracy, efficiency, and size. This paper presents a comprehensive literature review of recent advancements pertaining to the deep learning-based architectures for the image segmentation task. Such models include the well-known U-Net and Mask R-CNN models, as well as more recent models including Attention U-Net, segmentation using Generative Adversarial Networks (GANs) and transformer-based models including Vision Transformer (ViT) for segmentation tasks. We also include the analysis of hybrid architectures that utilize both CNN and transformer components for better performance in segmentation tasks. Other important issues such as data shortage and over-difficulty of annotation tasks are also considered, as well as the influence of data augmentation and few apologies on these problems. In addition to that, we also perform a comparative analysis of these architectures and evaluate and compare their performance on several benchmark datasets including their merits and demerits and usage. Finally, we identify current challenges and suggest promising future research directions which will include the development of more efficient models for real-time deployment and improvements in model interpretability and explainability.

Explore More!

Madhurima M.
Victoria Lloyd

Related Projects

Drone Obstacle Detection Using YOLOv5

This project was built with the goal of performing real-time obstacle detection on a drone.
Vihaan B.
Mentored by
DeepSolar Bangladesh: A Novel Convolutional Neural Network (CNN) Architecture for the Detection of Solar Panels from Low Resolution Satellite Imagery in Developing Countries

Due to its environmental benefits and decreasing costs, the supply of solar energy is growing at an accelerating pace globally. However, the decentralised nature of solar makes it difficult to keep track of the different photovoltaic (PV) systems deployed across a country. There is a critical need for highly accurate, comprehensive national databases of solar systems, which would allow policymakers, researchers, and the government to study socioeconomic trends in solar deployment. Manual surveys have shown to be inaccurate. The 2018 DeepSolar study by Yang et. al developed a deep-learning framework and national solar deployment database for the US using high-quality satellite imagery, which proved to be a much more efficient and accurate approach. However, satellite imagery in developing countries such as Bangladesh is of much lower resolution and quality, and performed poorly with the original DeepSolar model by Yang et. al. Our study highlights the implementation of a novel convolutional neural network (CNN) in detecting solar panels through low resolution Google Static Maps API satellite imagery data.
Khondoker F.
Mentored by Barbie Duckworth
Allez Go: AI Fencing Referee

Technology in fencing is generally an underdeveloped field and automated referees present potentially significant benefits to the sport. Automated referees will offer a more consistent call compared to a group of human referees with slightly different interpretations of the fencing rules.
Jason M.
Mentored by Anna Orosz