Evaluating large language models for complex supply chain optimisation: a focus on routing problems
Sadeque Hamdan, Nadjib Brahimi
This study evaluates the ability of general-purpose large language models (LLMs) to solve structured logistics and supply chain optimisation problems. We benchmark nine routing and procurement-routing problems, ranging from the travelling salesman problem and vehicle routing problem to richer variants with capacity limits and time windows, among others. We also introduce an additional inventory-routing extension. Using fixed instances, prompt templates, and external verification scripts, we evaluate eight GenAI model snapshots across three benchmark rounds, from early 2025 models to newer 2026 releases. Outputs are checked for feasibility, objective-value correctness, and solution quality relative to solver-based references or dedicated heuristic incumbents. The results show that LLMs can often generate feasible and high-quality solutions for unconstrained or moderately constrained routing problems, but their reliability decreases as operational constraints interact. Newer models improve feasibility and arithmetic consistency, especially on the core routing families, yet highly constrained procurement-routing and inventory-routing instances remain difficult. Comparisons with problem-tailored heuristics show that advanced GenAI can sometimes produce competitive candidate solutions, but dedicated heuristics remain more robust and computationally efficient. Overall, the study positions GenAI as a promising lightweight decision-support companion for logistics optimisation, while emphasising the continued need for external verification, structured prompting, and problem-specific optimisation methods.