Automating Underwater Search and Rescue Under Different Levels of Visibility with Narrow and General Purpose AI Models
Autonomous underwater vehicles (AUVs) depend on reliable vision in low visibility, yet we lack clear evidence on when specialized detectors (YOLO-World) outperform multimodal vision-language models (Gemini 2.5 Pro.) This gap limits the informed model selection for underwater tasks. To address it, we examine how these two paradigms behave when applied to the same operational task focused on object presence determination. The dataset was generated using a synthetic underwater simulator spanning Low, Medium, and High visibility. Both models are evaluated on accuracy and latency. YOLO-World performs better in Low visibility and runs faster, while Gemini improves in clearer scenes but requires more computation. These findings indicate that model choice should align with visibility and runtime constraints, with detectors suited for real-time use and a multimodal model for offline tasks.
@inproceedings{Solorzano_2026,
title = {Automating Underwater Search and Rescue Under Different Levels of Visibility with Narrow and General Purpose AI Models},
author = {Solorzano, Cecilia Delgado and Guynup, Chase and Arnold, Emma and Weng, Nan and McNeese, Nathan and Bertrand, Jeff and Madathil, Kapil Chalil and Flathmann, Christopher and Gramopadhye, Anand and Toxtli-Hernandez, Carlos},
booktitle = {2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)},
address = {Rome, Italy},
publisher = {IEEE},
pages = {1--7},
year = {2026},
month = July,
isbn = {979-8-3195-0598-9},
doi = {10.1109/ICECET65726.2026.11632779},
url = {https://ieeexplore.ieee.org/document/11632779}
}