Publication

LLM Evaluations: A Survey of Programmatic, Human, and LLM-as-Judge Approaches

Senior Solutions Architect, Amazon Web Services, New York, USA; Sarath Babu Poovassery Krishnan

International Journal of Innovative Research in Science Engineering and TechnologyJun 30, 2025
Abstract

The rapid progress of large language models (LLMs) has created an urgent need for robust, multifaceted evaluation methods that address both objective performance and subjective qualities. This survey reviews stateof-the-art developments from 2023–2025 in LLM evaluation, focusing on three paradigms: (1) programmatic (automated) evaluation, (2) human evaluation, and (3) LLM-as-a-judge approaches. We analyze the advantages, limitations, and practical trade-offs of each method, and highlight recent innovations in benchmarking, human feedback, and scalable AI-assisted evaluation. Our review concludes that comprehensive LLM assessment demands integrated strategies balancing automation, human judgment, and model-based augmentation, supported by dynamic, transparent, and community-driven practices.

1Authors

AuthorAffiliation
Senior Solutions Architect, Amazon Web Services, New York, USAAmazon (United States)
Sarath Babu Poovassery KrishnanAmazon (United States)

Showing the abstract — retrieve the full paper via the Exa API.

Powered by the Exa API