Abstract
Diabetic retinopathy is one of the most common microvascular complications of diabetes and a major cause of blindness, so timely automated grading is critical. This work builds a hybrid model with an EfficientNet-B3 convolutional backbone and a lightweight Vision Transformer encoder to capture both local lesions and global retinal patterns. The pipeline covers image preprocessing, hierarchical token projection, positional encoding, self-attention-based aggregation, and classification, with Grad-CAM for interpretability. Trained with stratified five-fold cross-validation and ensembling on the public APTOS 2019 dataset and externally validated on IDRiD, it reached 96.75% accuracy and a 0.9455 macro F1-score, with Grad-CAM highlighting microaneurysms, hemorrhages, hard exudates, and neovascularization.
Read on publisher