Project 14
One front door for many language models. Applications call a single OpenAI-compatible endpoint, and Switchyard picks the model for each job, falls back when a provider fails, keeps spending in check, and only switches to a new model if it tests as good as the old one. It also serves a small model I fine-tuned on a GPU for one narrow job, and the experiments here decide when that is worth doing.
Each request is authenticated, checked against per-tenant rate limits and budgets, routed by task, stripped of personal data, and looked up in a semantic cache before it reaches a provider chain with retries, fallbacks and a circuit breaker. Providers are free OpenAI-compatible endpoints plus a self-hosted model; Anthropic and Bedrock adapters exist but were only tested against mocked HTTP. Usage is recorded in Postgres and shown in a React dashboard, and a rollout gate controls how new models take traffic.
What happens to a request?
The cache and the provider chain sit behind the guardrails, so a cached or redacted answer never skips the limits. Usage records feed both the dashboard and the live canary check.
I fine-tuned Qwen2.5 models with LoRA on 1,400 real complaints to route each one to one of five queues, then tested on 400 complaints the models had never seen. All three sizes ended up within noise of each other (roughly 1.5 points on 400 items), and the 7B model, which started highest, gained least.
Served behind the gateway, the 1.5B router scored 99% on a 100-item golden set at 173 ms median, against 90% at 355 ms for a 30B API model. Across eight API models and the router, most large models were within a few points of each other, so speed decides, and reasoning models took 5 to 10 times longer for no gain on classification.
A new model only takes traffic if it passes a gate: a paired bootstrap must show its accuracy is not worse than the baseline by more than 5 points. The three fine-tuned models passed, a seeded regression that broke 30% of correct answers was rejected, and Qwen3 235B tied the baseline exactly and still failed. That is deliberate: 100 items cannot establish equality within 5 points, and too little evidence should not promote a model. A live run with real backends promoted a healthy 20% canary, then rolled it back when its endpoint was removed, while all 200 of 200 user requests were still served.
No paid provider was called, so cost is reported as tokens and time, never dollars. The Anthropic and Bedrock adapters were never run against the real services. The routing labels come from a rule rather than people, both golden sets are only 100 items, and latencies are measured on a shared free endpoint, so they move with load. The filing-answer comparison was judged by two models that had to agree; a later check in JudgeLab found that pair accepts 0.8% of wrong numeric answers, though it did not cover free-text answers. The self-hosted server is a small transformers server that handles one request at a time, not vLLM. Rate limits and budgets live in process memory, so they are only correct for one replica, and streaming is emulated as a single chunk. One model on the first comparison timed out on most requests and is excluded. The AWS Terraform module validates but was never applied.
What I Learned
Tech Stack