Smart news for curious minds.

Nerd News Network
Networking

Nvidia's latest solution to soaring enterprise AI costs is...a router?

NeMo Switchyard brings GPT-5-style model routing to the mainstream..

Lead image for “Nvidia's latest solution to soaring enterprise AI costs is...a router?”.
Image: The Register — Networks
Share

NeMo Switchyard brings GPT-5-style model routing to the mainstream.

The short version

  • NeMo Switchyard brings GPT-5-style model routing to the mainstream Soaring AI infrastructure costs and model pricing, combined with uncertain returns on investment, threaten to stall enterprise adoption.
  • Switchyard essentially functions as a proxy that sits between the inference server’s API endpoint and the models.
  • But rather than sending every request to the same model, Switchyard can be configured to route prompts to different models in order to optimize for cost, latency, or output quality.

What happened

By routing some requests to smaller, cheaper, and potentially locally hosted AI models, Nvidia claims Switchyard can cut job completion costs by 74 percent relative to using Claude Opus 4.8 alone, albeit with an approximately six-point accuracy tradeoff. The key metric in all of this is completion cost rather than price per token.

Why it matters

A model might cost one-tenth as much as OpenAI’s or Anthropic’s top model, but if it requires 10x the tokens to complete the request, it isn't actually cheaper.

Summary by Nerd News Network. Read the full article at The Register — Networks via the links above and below.

Share