Publication

Bloom or Bluff? Benchmarking Vision–Language Models Against Classical Machine Learning for Harmful Algal Bloom Detection from Satellite Imagery

Jul 2, 2026 · 1 author · 3 topics

Abstract

What are the main findings? Commercial vision-language models (VLMs; GPT-4o, GPT-5.5, Claude Sonnet 4.6) flagged harmful algal blooms with 73-93% false positive rates on bloom-absent satellite images, below the best RGB classifier and a trivial always-present baseline. The VLMs collapsed 60-70% of severity predictions into "moderate" regardless of actual conditions and identified at most one of the two severe blooms. A multi-spectral SVM (10 Sentinel-2 bands) achieved the strongest overall detection (F1 = 0.833, 79.5% accuracy, 27% false positive rate), outperforming every other method tested. Multi-spectral feature importances ranked the chlorophyll indices NDCI and Floating Algae Index among the top predictors, the same spectral signatures used in NOAA's operational algorithms. What are the implications of the main findings? The 12-percentage-point accuracy gain from RGB to multi-spectral features within the same SVM quantifies the spectral information gap, carried by the red-edge and shortwave infrared bands, that constrains RGB-only models and VLM vision encoders. For aquatic monitoring that requires spectral discrimination, domain-specific classifiers on multi-spectral data remain the right tool, and VLMs cannot substitute for them in spectral analysis. VLMs may serve a complementary role as narrative generators for non-specialist audiences, but their inability to access spectral bands is a hard boundary for current architectures.

Showing the abstract — retrieve the full paper via the Exa API.

Authors

Harsh Deep Singh Narula

Topics

Marine and coastal ecosystemsOil Spill Detection and MitigationMarine and coastal plant biology

About

PublishedJul 2, 2026
TypeArticle
Citations0

Powered by the Exa API