- Sep 2026: This repository is created. The code will be released soon.
Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing conditions. We introduce GeoSET, the first generalist model for SET, built around a single pretrained parent that is adapted to downstream datasets under a common protocol. We curate over 3 million high-quality SAR–EO pairs from a collection of more than 10 million SAR observations, spanning diverse sensors, spatial resolutions, and ground sampling distances. To bridge the modality gap between SAR observations and a pretrained image generator, we develop a speckle-robust SAR encoder and pretrain the conditional generator on this heterogeneous corpus. The resulting parent supports efficient adaptation across downstream datasets through low-rank adaptation (LoRA), updating only 0.60% of the generator parameters and requiring approximately one hour per dataset. Across six downstream benchmarks, GeoSET achieves state-of-the-art results in FID and DISTS with full fine-tuning or LoRA, demonstrating effective transfer across heterogeneous SAR–EO domains.
Across full fine-tuning and LoRA, GeoSET achieves the best reported FID on all six downstream benchmarks and the best DISTS on five.
Columns (g)–(h): GeoSET with LoRA and full fine-tuning; (i): ground-truth EO.
FID and DISTS on six benchmarks, normalized for each dataset–metric pair as 100 × best / value (outer ring = best).
All competing methods are retrained and evaluated on the same splits. FID and DISTS are the primary metrics; LPIPS, PSNR and SSIM are retained as complementary fidelity measures.
Please visit our project page for the interactive gallery and more results.
- Stage 1 · Speckle-robust SAR encoder: reconstructs the original SAR observation from a speckle-perturbed copy through a frozen decoder.
- Stage 2 · Generalist pretraining: a SAR-conditioned FLUX.2 flow transformer is trained on 3,204,744 curated SAR–EO pairs.
- Stage 3 · Downstream adaptation: the same parent is adapted to each benchmark by LoRA (0.60% of generator parameters, about one hour per dataset) or full fine-tuning.
The code and pretrained models will be released soon.
- Inference code
- Pretrained models
- Training scripts
- Evaluation scripts
If you find GeoSET useful, please consider citing:
@article{do2026geoset,
title={GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation},
author={Do, Jeonghyeok and Kim, Munchurl},
journal={arXiv preprint arXiv:2609.37496},
year={2026}
}Our prior work on SAR-to-EO image translation, C-DiffSET (project page):
@article{do2026cdiffset,
title={C-diffset: Leveraging latent diffusion for sar-to-eo image translation with confidence-guided reliable object generation},
author={Do, Jeonghyeok and Lee, Jaehyup and Lee, Seungchul and Kim, Munchurl},
journal={IEEE Transactions on Circuits and Systems for Video Technology},
year={2026},
publisher={IEEE}
}



