diff --git a/assets/img/blog/cost-based-semantic-routing/semantic-routing-flow.svg b/assets/img/blog/cost-based-semantic-routing/semantic-routing-flow.svg new file mode 100644 index 000000000..bfc2381d7 --- /dev/null +++ b/assets/img/blog/cost-based-semantic-routing/semantic-routing-flow.svg @@ -0,0 +1,75 @@ + + Cost-based semantic routing with agentgateway and vLLM Semantic Router + A coding agent sends model auto to agentgateway. vLLM Semantic Router selects either GPT-5.4 nano for routine coding or GPT-5.5 for complex coding. Agentgateway records catalog-priced telemetry. + + + + + + + + + + + + + + + + + + + + + + + + + + + Coding agent + model: auto + + + + + + agentgateway + routes the selected model + records the routing outcome + + OpenAI-compatible + + + + + + vSR semantic policy + selects the model tier + + + + + + + Routine coding + gpt-5.4-nano + lower-cost model + + + + + + Complex coding + gpt-5.5 + higher-capability model + + + + + + + Catalog-priced telemetry + model selection, tokens, cost, latency + Prometheus, logs, traces, Grafana + diff --git a/content/blog/2026-07-17-semantic-routing-llm-costs.md b/content/blog/2026-07-17-semantic-routing-llm-costs.md new file mode 100644 index 000000000..7e88e9f42 --- /dev/null +++ b/content/blog/2026-07-17-semantic-routing-llm-costs.md @@ -0,0 +1,158 @@ +--- +title: "Reduce LLM Spend With Semantic Routing You Can Measure" +category: "Deep Dive" +toc: false +publishDate: 2026-07-17 +author: "Daneyon Hansen" +description: "Use agentgateway, vLLM Semantic Router, and catalog-priced telemetry to reduce routine coding-model spend without routing every request to the cheapest model." +--- + +I spend a lot of time in Go and Rust, where a quick "one more thing" in a long +coding chat can turn a tidy refactor into an hour of debugging. After the hard +part is over, routine tests and follow-up questions can still burn +premium-model credits. Sending every request to a cheaper model is not the +answer: distributed-systems design, correctness work, and difficult debugging +need more capable help. + +This example combines [vLLM Semantic Router (vSR)](https://vllm-semantic-router.com/) +with [agentgateway](https://agentgateway.dev/) to send routine Go and Rust +coding tasks to `gpt-5.4-nano` and escalate complex correctness, +distributed-systems, and deep-debugging work to `gpt-5.5`. The goal is a +reproducible cost reduction without quietly routing everything to the cheapest +tier. + +## Make the Decision Observable + +{{< reuse-image src="img/blog/cost-based-semantic-routing/semantic-routing-flow.svg" alt="Flow diagram: a coding agent sends model auto to agentgateway, vLLM Semantic Router selects either GPT-5.4 nano for routine coding or GPT-5.5 for complex coding, and agentgateway records catalog-priced telemetry" caption="One model name in, the appropriate model tier out - with the routing decision and realized cost visible at the gateway" >}} + +The client sends an OpenAI-compatible request with `model: "auto"`. An +agentgateway policy sends it to vSR, which evaluates semantic, complexity, +keyword, context, and structure signals before selecting a model. Agentgateway +forwards the request and records the outcome. Cost is not a classifier input; +it is measured after the request, keeping the policy explainable. + +`auto` opts into that automatic selection. A client can instead request +`gpt-5.4-nano` or `gpt-5.5` directly to force a tier and bypass automatic +routing; the request still passes through agentgateway. Organizations +that require the policy can [validate and reject unsupported model +values](https://agentgateway.dev/docs/kubernetes/latest/traffic-management/transformations/validate/), +including anything other than `auto`. + +## Price the Outcome at the Gateway + +My company needs prices, not just token counts, to make a budget decision. +Agentgateway loads a model cost catalog and exposes realized request cost in +metrics, logs, and traces. The [model cost catalog +guide](https://agentgateway.dev/docs/kubernetes/latest/llm/cost-controls/costs/) explains how +to generate and load the catalog with `agctl`. + +For this evaluation, I use the integrated [OpenTelemetry stack](https://agentgateway.dev/docs/kubernetes/main/observability/otel-stack/). +Prometheus, Loki, Tempo, the OpenTelemetry Collector, and Grafana let +engineering and finance inspect the same model-selection event. + +## A Small, Reproducible Evaluation + +I built the [cost-based semantic routing demo](https://github.com/danehans/agentgateway-demos/tree/main/cost-based-semantic-routing) +around a compact 50-prompt Go and Rust dataset. Half of the prompts are routine +implementation or test work; the other half involve correctness or +distributed-systems problems. I send every prompt through two lanes: + +| Lane | Purpose | +|---|---| +| `routed` | vSR selects `gpt-5.4-nano` or `gpt-5.5`. | +| `always_expensive` | Every request uses `gpt-5.5`; the cost baseline. | + +Each run has an ID on every request, so its cost query excludes unrelated +cluster traffic. I deliberately leave out a forced-cheap lane: the useful +comparison is whether routing saves money while still escalating complex work. + +In a clean run on July 16, 2026, all 100 primary requests completed with exact +model-catalog lookups. vSR selected `gpt-5.4-nano` for 38 requests and +`gpt-5.5` for 12. Catalog-priced agentgateway metrics reported `$0.375954` for +the routed lane versus `$0.916940` for the always-expensive lane, a **59.0% +cost reduction**. Routed p50 latency was **5.58 seconds**, compared with +**7.75 seconds** for the always-expensive lane. P95 was 17.18 seconds for +routed traffic and 19.18 seconds for the always-expensive lane. + +![Catalog-priced semantic-routing result](/img/blog/cost-based-semantic-routing/20260716T193923Z-chart.svg) + +The chart reached 74% tier agreement and sent 12 of 25 prompts labelled complex +to GPT-5.5. This is not a claim of perfect model selection, but it shows the +savings did not come from a blanket downgrade: 24% of routed requests still +used the expensive model. It is a policy sanity check, not an answer-quality +benchmark or a substitute for application task-success and user-feedback data. + +The default dataset uses independent requests. For long-running coding agents, +the optional cache-transition evaluation uses two ordered ten-turn Go and Rust +conversations with a stable prefix. It reports provider-observed cache reads by +model transition without mixing cache behavior into the primary evaluation. + +## Try It + +The demo automates the setup work I would otherwise repeat by hand. It uses the [agentgateway semantic-routing +example](https://github.com/agentgateway/agentgateway/tree/main/examples/llm-semantic-routing), +creates or reuses a kind cluster, and installs agentgateway, vSR, a model cost +catalog, and the OpenTelemetry stack. It checks readiness, routing, +catalog-priced metrics, logs, and traces before the primary requests run. + +```bash +git clone https://github.com/danehans/agentgateway-demos.git +cd agentgateway-demos/cost-based-semantic-routing + +export OPENAI_API_KEY='sk-...' +./demo.sh all --yes +``` + +The default run sends 106 billable requests: two routing probes, four smoke-test +requests, and 100 primary requests. It writes request-level JSONL, an +evaluation-scoped Prometheus report, and an SVG chart under `results/`. + +When I want to change the policy, I edit the fetched vSR values, redeploy with +`./demo.sh router`, and run `./demo.sh eval --yes` again. + +## Use It From Codex + +Codex lets me choose a model manually, but a user-level profile can instead +send every request to a gateway such as agentgateway with the stable `auto` +model name: + +```toml +# ~/.codex/agentgateway.config.toml +model = "auto" +model_provider = "agentgateway" + +[model_providers.agentgateway] +name = "Corporate agentgateway" +base_url = "https://my.corp.agentgateway.com/v1" +wire_api = "responses" +``` + +I start Codex with `codex --profile agentgateway`. It sends OpenAI Responses +API traffic to agentgateway; vSR selects the tier and agentgateway records the +decision and realized cost. I validated routine Go test traffic to nano and an +advanced leader-election request to GPT-5.5. + +The profile keeps credentials out of the configuration file. Codex documents +[custom model providers](https://learn.chatgpt.com/docs/config-file/config-advanced#custom-model-providers) +and user-level [configuration profiles](https://learn.chatgpt.com/docs/config-file/config-advanced#profiles). + +## From Example to Rollout + +Semantic routing is not a magic cost switch. Start with an explainable policy, +compare it with always using the expensive model, confirm it uses both tiers, +and tune it against task completion, feedback, retries, and escalation rates. +Agentgateway supplies the accounting and observability; vSR makes the semantic +decision. + +This example is intentionally static: it uses configured keywords, semantic +candidates, weights, and thresholds rather than learning from application +outcomes. Those signals need periodic review as traffic changes. The +agentgateway community and vSR project are continuing work to make semantic +routing easier to tune, operate, and measure against real application traffic, +so stay tuned. + +## What's Next + +After validating a tiering policy, use [agentgateway cost controls](https://agentgateway.dev/docs/kubernetes/main/llm/cost-controls/) +to attribute usage with virtual keys, observe realized spend by model and +consumer, and enforce budget or spend limits. diff --git a/static/img/blog/cost-based-semantic-routing/20260716T193923Z-chart.svg b/static/img/blog/cost-based-semantic-routing/20260716T193923Z-chart.svg new file mode 100644 index 000000000..cc6431b33 --- /dev/null +++ b/static/img/blog/cost-based-semantic-routing/20260716T193923Z-chart.svg @@ -0,0 +1,70 @@ + + Semantic routing evaluation results + Semantic routing spend, routing agreement, model mix, latency, and cache transitions for run 20260716T193923Z. + + + Semantic routing evaluation + Run 20260716T193923Z | Catalog-priced agentgateway metrics + + + SPEND REDUCTION + 59.0% + Routed versus always expensive + + ROUTING AGREEMENT + 74.0% + Expected tiers in this coding sample + + ROUTED P50 LATENCY + 5.58 s + -28.0% versus always expensive + + COST PER DEMO RUN (USD) + Semantic routing$0.3760Always expensive$0.9169 + + + ROUTED MODEL MIX + + + 76% lower-cost + 38 gpt-5.4-nano | 12 gpt-5.5 + + + COMPLEX-PROMPT ESCALATION + + + 48% to GPT-5.5 + 12/25 complex prompts sent to GPT-5.5 + + + END-TO-END LATENCY + 5.58 s p50 + 17.18 s p95 semantic routing + Always expensive: 7.75 s p50, 19.18 s p95 + + + CONVERSATION CACHE TRANSITIONS + 0 switches + 0 ordered continuation turns + + + PROVIDER CACHE READS + 0 tokens + + + CACHE READS BY TRANSITION + No ordered multi-turn transitions in this run + + Agreement uses expected tiers in the checked-in coding sample; cache tokens are reported by the upstream provider. +