Accept familiar APIs
Applications and agents use OpenAI Chat, OpenAI Responses, or Anthropic Messages.
Apache-2.0 open-source GenAI gateway
Route applications, coding agents, subagents, and always-on automation across hosted and private models—based on capability, quality, cost, latency, and your policy. Developers keep one stable API while your platform team continuously improves the model mix behind it.
OpenAI Chat · OpenAI Responses · Anthropic Messages · Codex CLI · Claude Code · Private vLLM/SGLang
Agent economics changed the equation
Agents are not single prompts. Coding agents inspect repositories, call tools, replay context, retry steps, and launch subagents. Application, CI/CD, and monitoring agents run around the clock doing the same. Most of those steps do not require the most expensive model.
Turn model choice into platform policy instead of asking every developer, application, and subagent to select providers manually.
Conceptual illustration, not measured customer data.
How it works
The user requests a deployment-defined model group. The platform team controls the eligible providers, models, endpoints, and policy behind that contract.
Applications and agents use OpenAI Chat, OpenAI Responses, or Anthropic Messages.
Dialect, tools, images, reasoning, structured output, payload size, token caps, and access.
Choose among validated hosted or private targets with weighted, failover, scored, TypeScript, or external policy.
Preserve request-time cost, latency, throughput, attempts, errors, cache, and fallback behavior.
Enterprise hardware efficiency
Connect smaller models on older or lower-cost enterprise hardware, larger models on premium accelerators, and external frontier APIs behind the same endpoint.
Register validated vLLM, SGLang, or other compatible private inference services as routing targets.
Keep top accelerators and frontier APIs available for requests that need complex reasoning, multimodality, or specialized capability.
Developers and agents keep the same model-group name while the enterprise changes the infrastructure mix behind it.
Smart Router selects among configured inference endpoints. The enterprise serving platform remains responsible for GPU scheduling, model placement, and replicas.
Transparent savings scenario
Agents do not stop when the workday does. Between concurrent coding agents, delegated subagents, CI/CD automation, monitoring, and embedded applications, a developer's share of enterprise agent traffic runs to about 100 million tokens per week — and agent traffic lands in the most expensive corners of the price sheet.
Modeled annual savings per developer
Central 85% scenario. The 80–90% range equals approximately $94,000–$106,000 per developer per year.
Illustrative model, not a guarantee. Actual results depend on workload, cache behavior, context length, reasoning effort, model mix, internal hardware allocation, quality thresholds, and commercial pricing.
Product capabilities
Control policy, access, cost attribution, traffic, upstream capacity, and operational evidence from the individual caller key through the selected provider and model.
Quality contracts
Open source by design
Metrum GenAI Smart Router is Apache-2.0 open-source software. Review the behavior, run it in your environment, extend deployment-owned policy, and validate it against your workloads.
Go deeper
The public documentation covers configuration, compatibility, validation, security boundaries, reporting, deployment, and operations.
Metrum GenAI Smart Router
Start with the open-source software, explore the documentation, or get enterprise support, evaluation, and optional private-managed deployment from Metrum.