Abstract
Product management demands continuous prioritization, triage, and roadmap planning. While large language models (LLMs) show increasing capability in software engineering, rigorous empirical comparison against human decision-makers in product management remains absent. This paper presents a controlled study benchmarking AI against humans across three tasks: requirement prioritization (T1), issue triage (T2), and roadmap grouping (T3). Using 2,847 labeled issues from 12 opensource projects, we evaluate three GPT-4 configurations (LLMonly, RAG-augmented, structured protocol) against individual and consensus human baselines. RAG-augmented AI achieves NDCG@10 of 0.847 (95% CI: [0.831, 0.863]), statistically comparable to human consensus (0.856) via Wilcoxon signed-rank tests, while significantly outperforming individuals$(0.792, p<0.001)$. Humans retain superior performance on novel edge cases and constraint satisfaction. We provide cross-model generalization analysis, qualitative failure cases, and complete reproducibility artifacts.
Showing the abstract — retrieve the full paper via the Exa API.