1. Home
  2. Medical News
  3. Ophthalmology
advertisement

Retinal AI Benchmark Finds Compact Pretrained Models Outperform

Retinal AI benchmark finds compact pretrained models outperform
08/24/2026

Key Takeaways

  • Across retinal imaging classification tasks in OCT and color fundus photography, pretrained models delivered better accuracy on every task than scratch-trained counterparts, and each comparison was statistically significant.
  • Compact models in the roughly 27 to 29 million parameter range matched the best overall result on OCT, diabetic macular edema, and glaucoma classification.
  • The paper reports that RETFound-DINOv2-CFP's clearest advantage was on DR severity grading, where it exceeded the best compact model by 1.54 percentage points.
  • SwinV2-tiny was the most consistent architecture across tasks, while DINOv2-small showed the strongest parameter efficiency.
Retinal image classification increasingly forces a practical choice between large retinal-specific foundation models and smaller general-purpose backbones for common optical coherence tomography and fundus tasks. Balanced disease classification may reward broad visual feature learning differently than imbalanced severity grading, where harder class distinctions can expose advantages that only appear on more challenging tasks. To test whether added scale materially changes benchmark performance across those settings, investigators compared compact and large models under shared training conditions.

Investigators in the Isztl et al benchmark of compact and retinal foundation models for retinal image classification compared nine retinal imaging model configurations spanning 22.8M to 303M parameters across optical coherence tomography (OCT) 8-class disease classification, diabetic macular edema (DME) severity, glaucoma (GL) classification, and diabetic retinopathy (DR) severity grading from color fundus photography. Four architectures were tested with both pretrained and random initialization, whereas DINOv2 variants and RETFound models were assessed with pretrained weights only. Training used unified end-to-end fine-tuning with minimal augmentation and unweighted cross-entropy loss, and performance was aggregated with accuracy, macro-averaged area under the receiver operating characteristic curve, macro-averaged F1-score, and Cohen’s kappa after statistical testing of pretrained-versus-scratch differences and examination of size-performance tradeoffs.

Pretraining was the clearest signal, improving accuracy by 5.18 to 18.41 percentage points versus random initialization, with all pretrained-versus-scratch comparisons significant at p < 0.05. In the SwinV2-tiny and RETFound task-level comparison, RETFound-DINOv2-CFP reached 71% validation accuracy on DR versus 70% for SwinV2-tiny, a 1.54 percentage-point lead. Otherwise, SwinV2-tiny ranked first on OCT, DME, and GL, scratch-trained models were more variable, and parameter count alone did not reliably track accuracy across tasks.

Each task was evaluated on a single dataset or dataset combination, so external generalizability to other devices, populations, and acquisition settings was not tested. The benchmark did not report calibration, uncertainty, failure-mode, mechanistic interpretability, or demographic subgroup analyses. Computational efficiency was inferred from parameter count rather than direct latency, memory, or energy measurements, and the use of public benchmark datasets rather than a North American clinical deployment cohort means the findings describe comparative model behavior under benchmark conditions rather than real-world U.S. screening performance.

According to the authors, pretraining was the most consistent performance driver in this benchmark, and compact hierarchical pretrained models remained competitive on most tasks tested. The clearest absolute advantage for the largest retinal-specific model appeared on the imbalanced DR grading task under these datasets and training conditions.

Clinician Questions

Which retinal imaging tasks favored the largest retinal-specific foundation model in this benchmark?

In this benchmark, the largest retinal-specific model led only on diabetic retinopathy severity grading; OCT, diabetic macular edema, and glaucoma classification did not show the same scale advantage. The pattern suggests any benefit from larger retinal-specific backbones was task dependent rather than a general property of retinal artificial intelligence classification.

How was fairness maintained when compact and large retinal imaging models were compared?

The comparison used largely standardized training conditions across tasks, including end-to-end fine-tuning, a unified optimization setup, a fixed training schedule, minimal augmentation, and common aggregate metrics. Four architectures were evaluated both with pretrained and random initialization, whereas DINOv2 and RETFound models were assessed with pretrained weights only.

What limits how far these retinal AI benchmark results can be generalized?

Each retinal imaging task was tested on a single dataset or dataset combination without an external validation cohort, and the benchmark did not report calibration, uncertainty, failure-mode, interpretability, or demographic subgroup analyses. The findings therefore compare model behavior on benchmark datasets rather than proving real-world clinical deployment performance.

Recommended Reading

Register

We’re glad to see you’re enjoying ReachMD…
but how about a more personalized experience?

Register for free