NVIDIA-NeMo/SwitchyardPublic

Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.

AI summary: A fast Rust proxy and routing library for transparently translating and directing LLM traffic.

Stars
2.8K
+58 today
Forks
257
Watchers
9
Open issues
43
Open PRs
35
Contributors
~50
Commits
341
Branches
59

PythonApache-2.0Created May 19, 2026Last push todayLatest release v0.2.0+143 stars this week+1.6K this month

Star history

since Jun 28, 2026
01K2KJun 2026Jul 2026Aug 2026Sep 2026
2.8K stars as of Sep 10, 2026, tracked back to Jun 28, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 2 commits2026-06-30: 3 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 2 commits2026-07-07: 3 commits2026-07-08: 2 commits2026-07-09: 4 commits2026-07-10: 3 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 6 commits2026-07-15: 3 commits2026-07-16: 7 commits2026-07-17: 2 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 5 commits2026-07-21: 6 commits2026-07-22: 7 commits2026-07-23: 13 commits2026-07-24: 9 commits2026-07-25: 0 commits2026-07-26: 1 commit2026-07-27: 11 commits2026-07-28: 13 commits2026-07-29: 10 commits2026-07-30: 23 commits2026-07-31: 13 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 12 commits2026-08-04: 14 commits2026-08-05: 14 commits2026-08-06: 5 commits2026-08-07: 7 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 7 commits2026-08-11: 7 commits2026-08-12: 2 commits2026-08-13: 8 commits2026-08-14: 7 commits2026-08-15: 3 commits2026-08-16: 0 commits2026-08-17: 8 commits2026-08-18: 13 commits2026-08-19: 12 commits2026-08-20: 13 commits2026-08-21: 6 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 6 commits2026-08-25: 9 commits2026-08-26: 7 commits2026-08-27: 6 commits2026-08-28: 3 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 4 commits2026-09-01: 5 commits2026-09-02: 2 commits2026-09-03: 6 commits2026-09-04: 6 commits2026-09-05: 1 commit
341 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Rising fast

    +143 stars this week

  • Actively maintained

    Pushed within 48 hours

  • Well documented

    High community health score

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    12 trending appearances

What Switchyard does

Switchyard acts as an intelligent intermediary layer that routes and translates requests between various Large Language Model APIs, such as OpenAI and Anthropic formats. It allows applications, like coding agents or custom software, to communicate using their native protocol while seamlessly routing the traffic to entirely different backends like vLLM, Ollama, or NVIDIA NIM. Beyond simple translation, it provides sophisticated routing algorithms—including A/B benchmarking, signal-driven stage routing, and LLM-as-classifier routing—while actively recording operational metrics for performance monitoring. Built in Rust, it ensures minimal latency overhead whether run as a standalone server, embedded library, or command-line launcher.

Switchyard is built for AI platform engineers, backend developers, and AI researchers who need to manage, route, and monitor traffic across multiple LLM providers. It is specifically useful for teams building complex AI applications that require interoperability between different API standards or need to abstract away the underlying inference engines. Users should be comfortable configuring proxy servers or writing Rust code.

  • Protocol Translation: seamlessly converts requests and responses between OpenAI Chat, Anthropic Messages, and OpenAI Responses formats.
  • Advanced Traffic Routing: supports complex routing algorithms including random distribution, signal-driven stages, and custom user-defined logic.
  • Operational Metrics Export: natively exposes Prometheus metrics detailing request counts, error rates, token usage, and routing latency.
  • Flexible Deployment Modes: can be utilized as a standalone proxy server, a command-line launcher for coding agents, or an embedded Rust library.
  • Multi-Backend Support: abstracts the complexity of connecting to diverse LLM providers and self-hosted inference engines like vLLM or Ollama.
  • High-Performance Core: developed in Rust to ensure the proxy adds negligible latency to the critical path of LLM requests.

Where teams use it

Agent Protocol Adaptation

Used by developers to point coding agents expecting an Anthropic API towards open-source models running locally on Ollama.

A/B Model Benchmarking

Used by AI teams to silently split production traffic across multiple LLM providers to evaluate response quality and latency.

LLM Infrastructure Observability

Used by platform engineers to centralize token usage tracking and error monitoring across disparate AI backends via Prometheus.

Dynamic Request Routing

Used by application backends to route complex queries to expensive frontier models while directing simple queries to cheaper, local models based on classifiers.

Getting started: uv tool install --python 3.10 "nemo-switchyard[cli]"

README

main branch

Switchyard

Switchyard

Switchyard routes each LLM call to the cheapest model that can still do the job. Without changing a line of your agent.

Get started →

Accuracy versus total cost on Terminal-Bench 2.1. Switchyard's staged, escalation, and classifier routes reach 71-76% accuracy for 13-30% less than the Opus 4.8 baseline, while single fixed models stay below 56%.

*Total cost based on average ISP token cost

What is Switchyard

Switchyard picks which model serves each LLM call.

Use Switchyard

Switchyard runs inside gateways you may already have.

  • NeMo Relay — a native plugin. Load a routes.toml into a Relay deployment you already run. Setup →
  • LiteLLM — a routing plugin for LiteLLM's Router and proxy. examples/litellm
  • More integrations coming soon.
flowchart LR
    subgraph R["LiteLLM · NeMo Relay"]
        P["Switchyard"]
    end
    P--> M["Efficient model"]
    P--> N["Capable model"]
    P--> O[etc.]
    G[You] -->|"request"| P
    style P fill:#76B900,stroke:#5A8F00,color:#000
Loading

Integrate Switchyard into your gateway or harness

Embed the routing algorithms in your own. Switchyard picks the model; your harness makes the call, so your transport, retries, and credentials stay untouched.

  • Install: pip install nemo-switchyard
  • Then follow Path 2 — Embed the Library: construct an algorithm, drive its step stream, make the answer call.
  • Also available for Rust as switchyard-libsy; Path 2 has the Cargo.toml block.
flowchart LR
    subgraph R["Your LLM gateway / harness"]
        P["Switchyard"]
    end
    P--> M["Efficient model"]
    P--> N["Capable model"]
    P--> O[etc.]
    G["Your users"] -->|"request"| P
    style P fill:#76B900,stroke:#5A8F00,color:#000
Loading

Run Switchyard as a standalone proxy

A server in front of an agent, when you have no gateway to put Switchyard in:

cargo install --locked switchyard-server
switchyard-server --config routes.toml --port 4000

Point Claude Code, Codex CLI, or any OpenAI/Anthropic SDK client at the proxy. Switchyard decides per turn which model serves it.

flowchart LR
    P["Switchyard<br/>standalone proxy"]
    P--> M["Efficient model"]
    P--> N["Capable model"]
    P--> O[etc.]
    G[You] -->|"unchanged native API"| P
    style P fill:#76B900,stroke:#5A8F00,color:#000
Loading

Components

Pre-1.0 software. APIs, configuration, and routing behavior can change between releases — pin the version you integrate.

Component Stability Use it for Guidance
switchyard-libsy Beta Routing embedded in your own gateway or harness. You own model calls, credentials, and retries. Trial integrations. API will change before v1.0.
switchyard-llm-client Alpha HTTP model calls and protocol translation alongside libsy. Experiments and pilots.
switchyard-runner Alpha Running configured routes inside another runtime, such as NeMo Relay. Integration work and supervised pilots.
switchyard-server Demo A standalone OpenAI- and Anthropic-compatible proxy. Demos and evaluation only. Not for production.

Get Started

Three paths, in the same order as above. Each is self-contained: start at step 1, stop when you reach the result named under the heading.

Path 1 — Load the NeMo Relay Plugin

You finish with an existing NeMo Relay deployment routing through Switchyard. Requires NeMo Relay >=0.8.1,<0.9.0 and a Rust toolchain to build the plugin.

1. Build, package, and register the plugin. Follow steps 1–3 of the install guide in the plugin README. They build the shared library, package it into a bundle with a digest-bearing relay-plugin.toml, and register it with nemo-relay plugins add.

2. Write the Switchyard deployment to /etc/switchyard/routes.toml — the same version-1 TOML the proxy uses. Copy the file from step 2 of Path 3 below.

3. Point the plugin at the deployment. Add a config table to the [[plugins.dynamic]] entry that nemo-relay plugins add wrote, plus the policy override that lets Relay load the unsigned bundle. Use exactly one deployment source: a path, as here, or the config nested under switchyard_config.

[[plugins.dynamic]]
manifest = "./plugins/switchyard/relay-plugin.toml"

[plugins.dynamic.config]
priority = 0
switchyard_config_path = "/etc/switchyard/routes.toml"

[plugins.policy.overrides."nvidia.switchyard"]
attestation = "integrity_only"

4. Enable, validate, and restart Relay.

nemo-relay plugins enable nvidia.switchyard
nemo-relay plugins validate nvidia.switchyard

Relay now runs any algorithm switchyard-runner supports, while Switchyard owns provider HTTP dispatch.

Details: switchyard-nemo-relay-plugin and the TOML schema reference.

Path 2 — Embed the Library

You finish with your own harness picking a model per request and still making every model call itself. Shown in Python; the Rust API has the same shape.

1. Install. The Step and LlmResponse API below is newer than the nemo-switchyard 0.2.0 release on PyPI, which exposes an older LlmTarget based interface. Until the next release, build from source (requires a Rust toolchain):

pip install git+https://github.com/NVIDIA-NeMo/Switchyard.git

For Rust, the v0.2.0 tag has the older run_stream shape too, so depend on the repository's main branch and pin the rev you tested:

[dependencies]
async-trait = "0.1"
futures = "0.3"
switchyard-libsy = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", branch = "main" }
switchyard-protocol = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", branch = "main" }
tokio = { version = "1", features = ["macros", "rt"] }

2. Construct an algorithm. Target names are whatever your harness calls its models. This is the stage router from the benchmark; random, llm_task_classifier, and llm_classifier are built the same way.

from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

algorithm = stage_router(
    "capable",
    "efficient",
    picker="efficient_first",
    confidence_threshold=0.5,
)

3. Drive it. run_stream takes a normalized Switchyard request dict, not an OpenAI wire payload: the Request shape from switchyard-protocol, with messages whose content is a list of typed blocks. It yields steps. A CallModel step is a classifier or judge call — make it with your own client and hand back the normalized response wrapped in LlmResponse.Agg. Done carries the pick.

async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
    error: Exception | None = None
    for model in models:
        try:
            return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
        except Exception as exc:
            error = exc
    raise error or RuntimeError("no candidate models")


async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                try:
                    call.respond(await call_with_fallback(call.request, call.models, clients))
                except Exception as error:
                    call.fail(error)
            case Step.Done(outcome):
                if outcome.response is not None:
                    return outcome.response
                return await call_with_fallback(
                    outcome.request, outcome.selected_model_ids, clients
                )
    raise RuntimeError("algorithm ended without a decision")

clients maps each target name to your existing client; each call takes a normalized request dict and returns a normalized response dict. call.models and outcome.selected_model_ids list candidates in order, so the helper tries each one before giving up. outcome.request is the request to send, which may carry a rewrite the algorithm applied. When outcome.response is set, routing already produced the answer and no further call is needed.

4. Make the answer call with your own HTTP client, retries, and credentials, as call_with_fallback does above. A complete runnable version, including streaming responses, is in examples/libsy.py.

Type reference: switchyard-libsy and switchyard-protocol. In Rust the loop is Algorithm::run_stream yielding Step::CallModel and Step::Done, with switchyard-llm-client's run available to drive it for you.

Path 3 — Run the Standalone Proxy

You finish with a server on localhost:4000 that any OpenAI or Anthropic client can call. Needs Rust with Cargo.

1. Install the server.

cargo install --locked switchyard-server

2. Write routes.toml. A stage router over the same model pair as the benchmark above: how to reach a provider, which models to use, how to choose between them. --config takes any path; this writes it to the current directory.

cat > routes.toml <<'TOML'
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
TOML

Every key is documented in the TOML schema reference.

3. Start it. --dry-run loads the config, prints server OK: and the model IDs it exposes, then exits without starting the server.

export OPENROUTER_API_KEY="your-openrouter-key"  # pragma: allowlist secret
switchyard-server --config routes.toml --dry-run
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

4. Send a request. The route's id is the model name clients ask for.

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"hello"}]}'

The same route also answers on /v1/messages (Anthropic Messages) and /v1/responses (OpenAI Responses). /v1/stats reports which target served what, and /metrics exposes Prometheus counters for requests, errors, latency, tokens, and routing overhead.

5. Point a coding agent at it.

export ANTHROPIC_BASE_URL="http://localhost:4000"
export ANTHROPIC_MODEL="switchyard"
claude

Codex CLI and other OpenAI clients use the OpenAI variables instead:

export OPENAI_BASE_URL="http://localhost:4000/v1"

Routing Algorithms

Most use an LLM as a judge. All of them pick between an efficient model and a capable one; what differs is when the decision is made and how.

Algorithm How it decides Route type Benchmark
Capability The first request is judged by an LLM. llm_classifier 71.2% at $79.32
Stage Tool responses are judged by pattern matching or an LLM. stage_router 72.7% at $68.19
Capability + Stage Combines the two above. composite not yet benchmarked
Escalation Starts efficient. Responses are judged by an LLM for issues, then escalated. llm_classifier + mode = "escalation" 75.7% at $85.00
Advisor Gate One model serves every turn; a stronger advisor approves its plans and "done" claims, or sends it back. advisor lifts a weak executor 43.8% → 54.7%
Sub-Agent-Aware Delegated sub-agent traffic routes separately from the parent agent. subagents on passthrough or stage_router not yet benchmarked
Custom The first request is judged by an LLM against criteria you define, routing among 2+ of your own models. llm_classifier + target_selector policy not yet benchmarked
Random Each request is routed at random, uniform or weighted. random baseline mechanism

Benchmarks are Terminal-Bench 2.1 against a $98.06 Opus 4.8 baseline at 76.0%. A passthrough route registers one target under one model ID with no routing decision. See the Routing Overview for the common route shape and self-hosted targets.

Documentation

Benchmark Provenance

Configuration Accuracy Total cost vs. Opus 4.8 baseline
Opus 4.8 baseline 76.0% $98.06
Escalation 75.7% $85.00 99.6% of accuracy, 13.3% cheaper
Stage 72.7% $68.19 95.7% of accuracy, 30.5% cheaper
Capability 71.2% $79.32 93.7% of accuracy, 19.1% cheaper
Kimi K2.6 alone 55.8% $76.28
GLM 5.2 alone 52.4% $16.47
DeepSeek V4 Pro alone 48.7% $96.92
Ultra 3 alone 39.0% $29.66

These are the v0.2.0 Terminal-Bench 2.1 results from Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard. Those runs used NVIDIA-internal inference endpoints, so absolute solve rates may shift on another serving stack; the routing parameters are the ones that ran.

The escalation deployment is checked in at benchmark/routing-profiles/tb21-escalation-opus-glm-deepseek.toml, with OpenRouter targets substituted so it is publicly runnable. To run the harness, see benchmark/README.md; for latency and routing overhead rather than task success, see Soak Testing.

Community

License

Apache 2.0 License. Copyright NVIDIA Corporation.

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

3 total
  1. v0.2.0v0.2.0Aug 10, 2026

    # Switchyard v0.2.0 Switchyard v0.2.0 is a substantial redesign of the project around a native Rust server and the new libsy library architecture. Across [193 commits](https://github.com/NVIDIA-NeMo/Switchyard/compare/v0.1.0...v0.2.0), this release separates orchestration, provider-neutral protocols, translation, model transport, and serving into focused components that can be used together or embedded independently. Switchyard remains pre-alpha software. APIs and configuration may change before a stable release. ## libsy Redesign `switchyard-libsy` provides a provider-neutral framework for multi-LLM orchestration. Its `Algorithm` abstraction decides which semantic model targets to call, in what order, and how to combine their results. Algorithms do not own provider SDKs or an HTTP stack. They yield model-call steps back to their host, allowing the same algorithm to run inside the Switchyard server, another proxy, an agent runtime, or an application with custom model clients. The driver supports both buffered and streaming execution, concurrent model calls, routing decisions, request processors, classifiers, and session-scoped state. The redesign establishes clear crate bound

  2. v0.1.0v0.1.0Jun 30, 2026

    # Switchyard v0.1.0 Switchyard v0.1.0 was the first public release. It introduced the Python proxy, OpenAI and Anthropic format translation, YAML route bundles, initial routing strategies, coding-agent launchers, and request metrics. This release is retained for historical reference. New users should start with v0.2.0. - [nemo-switchyard 0.1.0 on PyPI](https://pypi.org/project/nemo-switchyard/0.1.0/) - [Full v0.1.0 changelog](https://github.com/NVIDIA-NeMo/Switchyard/blob/v0.1.0/CHANGELOG.md)

  3. v0.0.1v0.0.1Jun 30, 2026

Code frequency

additions and deletions
+147.3K-147.3KWeek of 2026-06-28: +147,292 linesWeek of 2026-06-28: -682 linesWeek of 2026-07-05: +11,491 linesWeek of 2026-07-05: -6,426 linesWeek of 2026-07-12: +12,102 linesWeek of 2026-07-12: -5,664 linesWeek of 2026-07-19: +22,125 linesWeek of 2026-07-19: -30,477 linesWeek of 2026-07-26: +31,116 linesWeek of 2026-07-26: -48,612 linesWeek of 2026-08-02: +14,374 linesWeek of 2026-08-02: -36,849 linesWeek of 2026-08-09: +9,049 linesWeek of 2026-08-09: -48,161 linesWeek of 2026-08-16: +18,412 linesWeek of 2026-08-16: -6,624 linesWeek of 2026-08-23: +84,941 linesWeek of 2026-08-23: -6,640 linesWeek of 2026-08-30: +6,462 linesWeek of 2026-08-30: -2,039 linesWeek of 2026-09-06: +0 linesWeek of 2026-09-06: -0 linesJun 28, 2026Sep 6, 2026
+357.4K lines added, -192.2K removed over the last year.

Commits per week

last 52 weeks
710Week of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 5 commitsWeek of 2026-07-05: 14 commitsWeek of 2026-07-12: 18 commitsWeek of 2026-07-19: 40 commitsWeek of 2026-07-26: 71 commitsWeek of 2026-08-02: 52 commitsWeek of 2026-08-09: 34 commitsWeek of 2026-08-16: 52 commitsWeek of 2026-08-23: 31 commitsWeek of 2026-08-30: 24 commitsSep 7, 2025Aug 30, 2026
341 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 1 commitsSun 15:00 — 0 commitsSun 16:00 — 0 commitsSun 17:00 — 0 commitsSun 18:00 — 0 commitsSun 19:00 — 0 commitsSun 20:00 — 0 commitsSun 21:00 — 0 commitsSun 22:00 — 0 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 6 commitsMon 10:00 — 6 commitsMon 11:00 — 6 commitsMon 12:00 — 5 commitsMon 13:00 — 7 commitsMon 14:00 — 3 commitsMon 15:00 — 5 commitsMon 16:00 — 3 commitsMon 17:00 — 3 commitsMon 18:00 — 7 commitsMon 19:00 — 1 commitsMon 20:00 — 2 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 3 commitsTue 0:00 — 4 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 1 commitsTue 4:00 — 1 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 4 commitsTue 9:00 — 6 commitsTue 10:00 — 9 commitsTue 11:00 — 7 commitsTue 12:00 — 8 commitsTue 13:00 — 7 commitsTue 14:00 — 11 commitsTue 15:00 — 9 commitsTue 16:00 — 1 commitsTue 17:00 — 4 commitsTue 18:00 — 1 commitsTue 19:00 — 2 commitsTue 20:00 — 0 commitsTue 21:00 — 1 commitsTue 22:00 — 0 commitsTue 23:00 — 3 commitsWed 0:00 — 0 commitsWed 1:00 — 3 commitsWed 2:00 — 2 commitsWed 3:00 — 0 commitsWed 4:00 — 1 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 2 commitsWed 8:00 — 4 commitsWed 9:00 — 2 commitsWed 10:00 — 4 commitsWed 11:00 — 5 commitsWed 12:00 — 2 commitsWed 13:00 — 8 commitsWed 14:00 — 5 commitsWed 15:00 — 5 commitsWed 16:00 — 5 commitsWed 17:00 — 5 commitsWed 18:00 — 2 commitsWed 19:00 — 1 commitsWed 20:00 — 1 commitsWed 21:00 — 0 commitsWed 22:00 — 1 commitsWed 23:00 — 1 commitsThu 0:00 — 2 commitsThu 1:00 — 0 commitsThu 2:00 — 1 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 1 commitsThu 6:00 — 1 commitsThu 7:00 — 1 commitsThu 8:00 — 2 commitsThu 9:00 — 3 commitsThu 10:00 — 4 commitsThu 11:00 — 9 commitsThu 12:00 — 7 commitsThu 13:00 — 6 commitsThu 14:00 — 12 commitsThu 15:00 — 5 commitsThu 16:00 — 13 commitsThu 17:00 — 8 commitsThu 18:00 — 4 commitsThu 19:00 — 1 commitsThu 20:00 — 1 commitsThu 21:00 — 1 commitsThu 22:00 — 2 commitsThu 23:00 — 1 commitsFri 0:00 — 2 commitsFri 1:00 — 3 commitsFri 2:00 — 1 commitsFri 3:00 — 0 commitsFri 4:00 — 1 commitsFri 5:00 — 2 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 1 commitsFri 9:00 — 2 commitsFri 10:00 — 4 commitsFri 11:00 — 9 commitsFri 12:00 — 5 commitsFri 13:00 — 5 commitsFri 14:00 — 6 commitsFri 15:00 — 3 commitsFri 16:00 — 4 commitsFri 17:00 — 2 commitsFri 18:00 — 4 commitsFri 19:00 — 1 commitsFri 20:00 — 1 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 0 commitsSat 0:00 — 2 commitsSat 1:00 — 0 commitsSat 2:00 — 1 commitsSat 3:00 — 1 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 0 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Sep 10, 2026monthly#9+2,617
Aug 24, 2026weekly#11+635
Aug 23, 2026weekly#11+635
Aug 22, 2026weekly#10+642
Aug 21, 2026weekly#5+932
Aug 18, 2026weekly#3+1,435
Aug 17, 2026weekly#3+1,435
Aug 16, 2026weekly#4+1,326
Aug 15, 2026daily#7+408
Aug 15, 2026weekly#5+1,195
Aug 14, 2026daily#7+408
Aug 13, 2026daily#8+421