Enterprise LLM Router: Learning Quality–Capacity–Capability Trade-offs from 2026 Model Metadata
Abstract
Enterprise language-model selection is a constrained decision problem, not a single leaderboard lookup. This study integrated three 2026 metadata tables covering 22 models from eight providers and evaluated benchmark quality, capability requirements, paid rate-limit capacity, and discount support. The files joined exactly on provider and model without missing cells or duplicate keys. Quality was the mean percentile rank of MMLU, HumanEval, MATH, and Arena Elo; capacity was the mean paid-RPM and paid-TPM percentile for 19 models with numeric limits; discount support averaged batch and cached-input discounts. Nested leave-one-model-out and leave-one-provider-out tests compared ridge regression, k-nearest neighbors, random forest, and a training-mean baseline. The router was evaluated on 20 capability-gated candidate sets under five utility regimes, creating 100 scenarios, followed by a 0.1-resolution weight analysis. Ridge achieved leave-one-model-out MAE 13.09 and Spearman ρ .771, but provider-held-out performance weakened to MAE 22.35 and R² −.082. Capability breadth produced the lowest routing regret (5.60), ahead of ridge (8.46) and observed-quality routing (10.60). The eight-model Pareto frontier exposed specialization: o1 led quality, Gemini 2.5 Flash led numeric paid capacity, and DeepSeek V3 combined high capacity with maximum discount support. The evidence favors transparent capability filtering and abstention when price, latency, deployment license, or context evidence is absent.
Keywords
- Large language models; model routing; enterprise AI; multi-criteria selection; benchmark metadata; rate limits; capability constraints; provider generalization
How to Cite
Lily Peng. (2026). Enterprise LLM Router: Learning Quality–Capacity–Capability Trade-offs from 2026 Model Metadata. Journal of Computer Science and Information Technology, Vol. 3 No. 1 (2026): Journal of Computer Science and Information Technology, 104-120. https://doi.org/10.61424/jcsit.v3i1.953
