Abstract
What are the main findings? Commercial vision-language models (VLMs; GPT-4o, GPT-5.5, Claude Sonnet 4.6) flagged harmful algal blooms with 73-93% false positive rates on bloom-absent satellite images, below the best RGB classifier and a trivial always-present baseline. The VLMs collapsed 60-70% of severity predictions into "moderate" regardless of actual conditions and identified at most one of the two severe blooms. A multi-spectral SVM (10 Sentinel-2 bands) achieved the strongest overall detection (F1 = 0.833, 79.5% accuracy, 27% false positive rate), outperforming every other method tested. Multi-spectral feature importances ranked the chlorophyll indices NDCI and Floating Algae Index among the top predictors, the same spectral signatures used in NOAA's operational algorithms. What are the implications of the main findings? The 12-percentage-point accuracy gain from RGB to multi-spectral features within the same SVM quantifies the spectral information gap, carried by the red-edge and shortwave infrared bands, that constrains RGB-only models and VLM vision encoders. For aquatic monitoring that requires spectral discrimination, domain-specific classifiers on multi-spectral data remain the right tool, and VLMs cannot substitute for them in spectral analysis. VLMs may serve a complementary role as narrative generators for non-specialist audiences, but their inability to access spectral bands is a hard boundary for current architectures.
Showing the abstract — retrieve the full paper via the Exa API.