Why Use FreeRouter Even With One LLM Gateway
The usual starting point for an LLM feature is a single gateway. You sign up with OpenRouter, Vercel AI Gateway, or Cloudflare AI Gateway, put the key in .env, point your SDK at its base URL, and build from there. That's a reasonable call — one account, one bill, one set of docs to read.
What it also does is make that gateway the shape of the integration. Its model ids live in your config. Its request and response format is the one your code is written against. Its error bodies are what your retry logic matches on, and its dashboard is the only place your traffic is visible. None of those is a problem by itself. Together they're the list of things that have to change if the backend ever does, and the list gets longer as the app grows.
The constraint that usually shows up first is model availability. New models are announced regularly, and whether you can use one this week depends on whether your gateway has listed it. If it hasn't, the choices are to wait, or to run a second integration next to the first — another client, another set of model ids, another set of failure cases to handle.
FreeRouter is built to be swapped in whenever that becomes worth solving. A key speaks one of several request shapes — openai, openrouter, anthropic, or google — so the client you already have keeps sending the bodies it already sends, and the migration is two changes: base URL and key. Waiting doesn't make that harder.
There's still a case for doing it early, while there's exactly one gateway behind it and nothing to migrate. That's the rest of this post, followed by the setup.
What it looks like with one gateway
Your app talks to a FreeRouter key. The key carries the shape, the routing rule, any model remaps, and whatever tools you've switched on. Today the rule has one target: the gateway you already pay. Everything else is a dashboard row you haven't added yet.
flowchart LR
App["Your app"] --> Key["FreeRouter key: shape, rule, remaps, tools"]
Key --> GW["Your gateway (your own key)"]
GW --> Model["model"]
Key -.->|"a row you add later"| GW2["second gateway"]
Key -.->|"a row you add later"| Remap["remapped model"]
One target today. The dotted paths are dashboard edits, not deploys.
The reasons, one at a time
1. The code change happens once
Swapping a base URL and a key is a small change, and it stays small only while the app is small. Every service, worker, test fixture, and environment added later is another place the gateway's name is written down. Doing it against today's codebase is the cheapest version of this change you'll get.
After that, routing lives in config that's evaluated per request. Change a rule in the dashboard and the next request already follows it — no redeploy, no rollout.
2. Your app keeps speaking its own dialect
The shape is a property of the key, not of the backend. If your code is written against OpenRouter — models fallback arrays, provider preferences, all of it — set the key to the openrouter shape and keep sending those bodies verbatim. Anthropic client hitting /v1/messages? Use the anthropic shape. FreeRouter translates between what your key speaks and what each gateway expects.
That matters when the gateway you want next speaks a different format than the one you built against: it's a translation on our side of the wire rather than a refactor on yours.
3. New models arrive as remaps
A model remap rewrites one incoming model id to a specific provider key plus a literal upstream model, before the routing rule runs. It works fine with a single gateway: when that gateway lists something new, point the id your app already sends at the new model, on the same provider key.
Each remap carries a percent (1–100), rolled per request, so percent: 20 sends a fifth of that traffic to the new model while the other four fifths follow the normal rule. Compare in the logs, then take it to 100. Your app never learns a new model id, and if the destination returns a 5xx, a 429, or a timeout, the request falls back to your normal targets.
4. Failover becomes a configuration change
One target means no failover — there's nothing to fall over to. What matters is how little it takes to change that: add a second provider key, drop it below the first in a priority rule, done. FreeRouter retries the next target on connection errors, 5xx, 429, a model the target doesn't offer, or no response headers inside the first-byte budget (4s by default). A 400 that would fail everywhere is returned immediately instead of being retried around the list.
When a failover happens, X-FreeRouter-Failover names every skipped attempt as gateway:reason — ttfb-timeout, rate-limited, no-model, server-error. That's retry behavior you configure rather than write.
5. One key per app, revocable on its own
Your gateway key gets stored once in the dashboard, encrypted and scoped to the workspace (BYOK — you still hold the account and the bill). Your app carries an fr_live_ key instead. If that leaks, revoke it and mint another; your other keys and the underlying provider key are untouched.
You can also pin a key to specific client addresses with an IP allowlist, at the workspace level, the key level, or both. A request from outside the list gets a 403 with code ip_not_allowed, and the message names the address we actually saw — which is the one you want when your servers sit behind NAT or a fixed egress.
6. Logs that don't change when the backend does
Every proxied request logs the gateway, model, status, attempt path, and a time-to-first-byte split into Router, Upstream, and Duration, plus an X-FreeRouter-Request-Id you can quote. The dashboard rolls that up into volume, error rate, latency, and breakdowns by model and provider.
With one gateway that's a second opinion on your vendor's numbers. With two it's one screen instead of two vendor dashboards and a spreadsheet. Response bodies are never stored; request bodies only if you switch on capture in Settings for debugging.
7. Capabilities attach to the key, not the gateway
Some things you'd otherwise wire into the app attach to the FreeRouter key instead: MCP tools like web search, weather, and finance, offered to the model with no change to your request body; Companion Ads and Monetizable Keyterms as optional siblings of the usual body; and a decisioner that lets Jev pick among a key's targets while your client still sends and receives an ordinary chat completion.
These don't care which gateway is behind the key. Turn one on now, change backends in a year, and it's still on.
8. The whole thing is scriptable
A management key (fr_mgmt_) mints inference keys, stores provider keys, and sets routing rules and remaps over /api/v1 — so per-tenant or per-environment keys can be provisioned from your own code instead of clicked. The same surface is exposed as an MCP server at https://api.freerouter.com/mcp if you'd rather have an agent do it. Read-only management keys serve the usage endpoints and 403 on writes, which is what you want in a reporting job.
What this costs: nothing during the preview, and your inference bill doesn't move — BYOK means the gateway still charges your account directly, with your key. The overhead is a network hop, and it's measured per request rather than asserted (see the limits below).
Setting it up
Ten minutes, most of it in the dashboard. You need your existing gateway's API key and the ability to change two values in your app config.
- Sign up. Create an account — Google or email plus an OTP. You land in a workspace. Getting Started has the screens.
- Add your gateway key as a provider. Providers → add the gateway you already use, paste its key. It's encrypted at rest and scoped to the workspace. Leave the same key in your app's
.envfor now; nothing stops you running both paths while you cut over. - Create a FreeRouter key with the shape your app already speaks. API Keys → new key. Leave it on openai if you're using the OpenAI SDK or anything compatible. Pick openrouter if your bodies are OpenRouter-flavored, anthropic if your client calls
/v1/messages, google forgenerateContent. Match the client you already have. - Point the routing rule at that one provider. Attach a rule with a single target: your provider key, 100%. A one-target rule is a legitimate rule. You're not configuring failover yet.
- Sanity-check it in the Playground. Pick the key, pick a model, send a prompt. You get the response, a per-attempt trace, and a curl command. Worth knowing: the Playground calls your provider from the dashboard rather than through
api.freerouter.com, and its runs are never written to Logs. - Change two values in your app. Base URL to
https://api.freerouter.com/v1, key to yourfr_live_key. Nothing else — same SDK, same model ids, same request bodies. - Send one real request and open Logs. Check
X-FreeRouter-Providernames your gateway, and look at the Router vs Upstream split on the log row. The first live request through the inference API swaps the dashboard's setup flow for charts.
The two changes, in curl
Export the key and send the request you already send. Same body you were sending to your gateway:
export FREEROUTER_API_KEY=fr_live_…
curl https://api.freerouter.com/v1/chat/completions \
-H "Authorization: Bearer $FREEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-i \
-d '{
"model": "openai/gpt-4o-mini",
"messages": [{ "role": "user", "content": "Ping" }],
"temperature": 0.2
}'
The body is an ordinary chat completion: { id, model, choices, usage }. The -i is there so you can read the routing headers — X-FreeRouter-Provider (who served it), X-FreeRouter-Attempts (how many targets were tried, 1 today), and X-FreeRouter-Request-Id (quote this when you ask us anything).
The same thing in your app
If you're on the OpenAI SDK, the diff is the constructor. Nothing downstream of it changes:
# fr_ping.py — the whole migration is the two kwargs below.
# Run: FREEROUTER_API_KEY=fr_live_… python fr_ping.py
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.freerouter.com/v1", # was your gateway's URL
api_key=os.environ["FREEROUTER_API_KEY"], # was your gateway's key
)
resp = client.chat.completions.with_raw_response.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Ping"}],
)
print("served by", resp.headers.get("x-freerouter-provider"))
print("attempts ", resp.headers.get("x-freerouter-attempts"))
print(resp.parse().choices[0].message.content)
Drop with_raw_response once you've seen the headers and it's a plain client.chat.completions.create(...) again. Node, Go, and anything else OpenAI-compatible work the same way — the migration quickstart has those snippets.
flowchart TD
S1["Add your gateway key under Providers"] --> S2["Create an fr_live_ key with your app's shape"]
S2 --> S3["Routing rule: one target, 100%"]
S3 --> S4["Playground sanity check"]
S4 --> S5["Swap base URL + key in the app"]
S5 --> S6["One real request"]
S6 --> S7{"Does the provider header name your gateway?"}
S7 -->|yes| S8["Done — routing is now config"]
S7 -->|no| S9["Check the rule and the key's shape"]
Four dashboard steps, two edited lines, one verification request.
Adding a second gateway later
When there's a reason to add one — a model your gateway doesn't list, a price difference worth testing, a reliability problem — the work is a dashboard edit. Store the new gateway's key under Providers, add it as a second target with a small weight, and compare the By provider breakdown and error rate for a day.
Over the management API the same move is one request body — your existing target, plus a 5% canary:
PUT /api/v1/keys/:id/routing
{
"strategy": "split",
"targets": [
{ "provider_key_id": "<your-current-key-id>", "gateway": "openrouter", "weight": 95 },
{ "provider_key_id": "<the-new-key-id>", "gateway": "vercel", "weight": 5 }
]
}
Raise the weight when you like the numbers, or flip the rule to priority so the new one is a backup rather than a share of traffic. A weight of 0 means failover-only: never chosen first, still tried when the primary fails. Either way there's no deploy, and the request your app sends is the one it sent last month. That's what Learny's move to Darkbloom looked like in practice.
Limits and failure modes
- It's one more hop. Routing decisions come from a pre-warmed in-memory cache, so simple routing is sub-1 ms; remaps, request translation, and failover drain push it toward 20–30 ms. You don't have to take that on faith — every log row splits Router and Upstream, so you can read your own number.
- One target means no failover. If your only gateway is down, you get an error, same as calling it directly. FreeRouter doesn't add availability until you add a second target; what it adds today is the ability to do that without a release.
- Playground and Test key don't exercise remaps. Both run the key's base targets with a fixed test model. A green test means connectivity, not remap correctness — verify remaps with live traffic and look for the remap badge in Logs.
- The google shape can't stream. MCP tools run either way — the loop works on streamed and buffered requests, merged into one response — but google is JSON-only for now. Call
:generateContentwithout streaming, or use an openai-shaped key for the same model. - Spend cards are OpenRouter-only and aggregate. Only OpenRouter exposes spend over an API, and it reports per provider key, not per call. FreeRouter doesn't store token counts or cost per request, so don't plan on billing off the Logs tab.
- The key is a server secret. An
fr_live_key can spend against every provider key on its routing rule. Keep it server-side, rotate it from the dashboard if it leaks, and use IP allowlists if you want a second lock. - Decisions keys are a different animal. A systemone key is not a chat key, and a TypeSafe provider key is Decisions-only — it can't be a chat routing target. If you want Jev in front of ordinary completions, that's a decisioner on a chat key, not this shape. See the systemone shape.
Next steps
Get a FreeRouter key, put your current gateway behind it as the only target, and run the curl above. If the header names your gateway, you're done, and routing decisions from then on are configuration rather than code.