LLM Evaluations: A Survey of Programmatic, Human, and LLM-as-Judge Approaches
Senior Solutions Architect, Amazon Web Services, New York, USA; Sarath Babu Poovassery Krishnan
The rapid progress of large language models (LLMs) has created an urgent need for robust, multifaceted evaluation methods that address both objective performance and subjective qualities. This survey reviews stateof-the-art developments from 2023–2025 in LLM evaluation, focusing on three paradigms: (1) programmatic (automated) evaluation, (2) human evaluation, and (3) LLM-as-a-judge approaches. We analyze the advantages, limitations, and practical trade-offs of each method, and highlight recent innovations in benchmarking, human feedback, and scalable AI-assisted evaluation. Our review concludes that comprehensive LLM assessment demands integrated strategies balancing automation, human judgment, and model-based augmentation, supported by dynamic, transparent, and community-driven practices.