← All resources

How to Use FreeRouter with Agent-Native

Agent-Native is Builder.io's open-source TypeScript framework for apps where the agent and the UI share one database, one set of actions, and one application state. You write a capability once as a defineAction() with a Zod schema, and the framework exposes it, in their words, across "UI, agent, HTTP, MCP, A2A, and CLI." The agent calls it as a tool, your React code calls it as a typed function, and both paths run the same validation and the same permission checks.

Application state lives in Postgres through Drizzle, so there's no sync layer between what the agent knows and what's on screen. The UI writes the current view into a navigation key, the agent reads it before it acts, and useDbSync() pushes the agent's writes back into the React Query cache as they land. "Archive this" works because both sides are reading the same row.

What people build with it

The repo ships about eighteen templates, and they're finished apps rather than starters. Each one scaffolds with a single command — npx @agent-native/core@latest create my-app --template <name> — and most of them run live on a subdomain of agent-native.com, so you can use one before you clone it.

  • Inbox and calendar. Mail is a keyboard-first Gmail client their README pitches against Superhuman: the agent reads, drafts, sends, and files, and it always knows which thread is open. Calendar syncs Google Calendar and adds Calendly-style public booking links, so finding an hour next week is one turn instead of six clicks.
  • Company knowledge. Brain is the Glean slot. It ingests approved Slack channels, meetings, transcripts, and GitHub issues and PRs, then answers with quoted evidence and links back to the source rather than a summary you have to take on faith.
  • Records and data. CRM runs a typed, bitemporal record grid — either as its own system of record on plain SQL or as a scoped companion to a connected HubSpot or Salesforce workspace. Analytics turns plain-English questions into charts and dashboards, with Amplitude and FullStory as the stated comparison. Forms is Typeform: describe the form, refine it visually, publish it, keep the submissions in your own database.
  • Documents and media. Content is a local-file MDX workspace in the Notion and Obsidian shape. Clips records and understands meetings, screens, and voice notes. Design generates and refines interactive designs, Slides builds on-brand decks, and Plan turns a Claude Code or Codex plan into a reviewable document with diagrams, wireframes, comments, and share links.
  • Orchestration. Dispatch is the control plane you run next to the others: a Slack and Telegram inbox, a secrets vault, scheduled jobs, approvals, and an orchestrator that hands work to the right specialist app over A2A. Factory takes work in one end — Slack feedback, pull-request evidence — and puts governed agent work and shipped changes out the other, with the human checkpoints you choose.

The list isn't the interesting part; the shapes are. These are the things teams keep rebuilding in-house — an assistant over the company's email, a question box over the warehouse, an intake form that routes its own submissions, an internal tool that does the boring half of a workflow — and the framework treats each one as an ordinary app that happens to have an agent wired into the same actions the UI already calls. Picking the closest template and deleting what you don't need is a legitimate way to start.

It also means the model workload is different in every one of them. Mail triage is short, constant, and cheap. A Factory run is long, tool-heavy, and expensive to lose two-thirds of the way through. Clips leans on transcription. Dispatch fires jobs on a schedule with nobody watching the screen. Wanting a different model behind different surfaces — or the same model from somewhere else when one provider is having a bad hour — comes up well before it feels like an architecture decision.

Where the model choice lives

Agent-Native selects its engine from whichever provider key you set. ANTHROPIC_API_KEY gets you Claude, OPENAI_API_KEY the GPT family, GOOGLE_GENERATIVE_AI_API_KEY Gemini, and there are entries for OpenRouter, Groq, Mistral, Cohere, and local Ollama; AGENT_ENGINE breaks the tie when several are present. The model itself is picked in Settings. Drop a key in and you're building the same afternoon, which is a reasonable way to start — and Claude is the path their docs say is most heavily tested.

What that leaves you with is a provider key, an endpoint, and a model name living in the app's config. Changing any of the three means editing that config and redeploying, once per environment you run. None of it is expensive on its own. It adds up because every model you want to try, every provider outage, and every new environment goes through the same loop.

The framework already has the seam for avoiding that: its OpenAI engine accepts OPENAI_BASE_URL, so it will talk to any endpoint with the same API shape. Their environment-variable docs say so directly — supported models span Claude, GPT, and Gemini "and any provider that speaks the OpenAI API shape via OPENAI_BASE_URL, including OpenAI-compatible gateways like LiteLLM." FreeRouter answers the same way. It's the ordinary base-URL-and-key migration — the app keeps sending the bodies it already sends, and the fr_live_ key decides which gateway serves each request. No fork of Agent-Native, no adapter, no extra code. Waiting doesn't make it harder either; the two changes are the same next year.

The rest of this post is the setup, why the seam works, and where it stops helping.

What it looks like when Agent-Native talks to FreeRouter

The app uses its OpenAI engine and sends ordinary chat completions to one base URL. The key carries the shape, the routing rule, and any model remaps; the gateways behind it hold your provider keys. Adding a gateway later is a row in the dashboard, and the app never hears about it.

flowchart LR
    App["Agent-Native app (openai engine)"] --> Key["FreeRouter key: openai shape, rule, remaps"]
    Key --> GW1["Gateway one (your key)"]
    Key -.->|"a row you add later"| GW2["Gateway two (your key)"]
    GW1 --> M["provider/model"]
    GW2 --> M2["provider/model"]

The app talks to one key. The dotted path is a dashboard edit, not a deploy.

Why this works

1. The engine already accepts a different endpoint

The ai-sdk:openai engine reads OPENAI_BASE_URL and sends model requests there instead of to https://api.openai.com/v1. Set it to https://api.freerouter.com/v1, put your FreeRouter key in OPENAI_API_KEY, and the app sends exactly what it sent yesterday. FreeRouter answers in the same chat completions format, documented under API shapes.

Two details make the fit closer than plain OpenAI compatibility. When the base URL isn't OpenAI's own, the engine talks Chat Completions rather than the Responses API — which is exactly what a key on the openai shape serves. It also stops matching your model id against its built-in list, so the provider/model id your key serves goes upstream as typed.

The engine to avoid is ai-sdk:openrouter. The framework exposes two endpoint overrides in total, OPENAI_BASE_URL and OLLAMA_BASE_URL, and there's no OpenRouter equivalent — so a FreeRouter key has nowhere to go on that engine. If your app runs on it today, move to ai-sdk:openai and keep OpenRouter as a gateway behind the key instead.

2. The model id becomes a dashboard entry

The model name still lives in the app's Settings, in the provider/model form the key serves — openai/gpt-4o-mini, say. To run something else behind that id, add a model remap on the key with a percent from 1 to 100: send a fifth of the traffic to the new model, compare in the logs, then take it to 100. If the destination errors or the mapping is wrong, requests fall back to the key's normal targets instead of failing.

That's the same arrangement as putting one gateway behind a key. The app holds a URL and a key; everything else is config that's evaluated per request.

3. A second gateway is where failover comes from

A priority rule tries your gateways in order. If the first one is rate-limited, down, or slow to the first byte, the request goes to the next, and the response says so: X-FreeRouter-Failover lists the skipped attempts as gateway:reason, and X-FreeRouter-Provider names who finally answered. Every request is logged with a Router, Upstream, and Duration split and an X-FreeRouter-Request-Id you can quote.

Agent runs are long and stateful, so a failure two-thirds of the way through a turn is the expensive kind. Failover won't rescue a bad prompt or a broken tool — it covers the routine outages, configured once instead of written as retry code around the engine.

Setting it up

  1. Create the key. In the dashboard, create an fr_live_ key on the openai shape and attach the gateway whose provider key it should use first. The shape is a property of the key, so this one always speaks chat completions.
  2. Prove the key before touching the app. Run the curl below against https://api.freerouter.com/v1/chat/completions with a model id your key serves. If X-FreeRouter-Provider names your gateway, the key is good, and anything that breaks after this is app config.
    export FREEROUTER_API_KEY=fr_live_…
    curl -i https://api.freerouter.com/v1/chat/completions \
      -H "Authorization: Bearer $FREEROUTER_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"Say ok."}]}'
  3. Point the app at the key. Set OPENAI_API_KEY to the fr_live_ key and OPENAI_BASE_URL to https://api.freerouter.com/v1 in the app's environment, then choose the provider/model id in Settings. It's the same pair of values the framework's own docs give for LiteLLM or a local proxy.
  4. Run an agent turn and open Logs. Send a prompt through agent chat, then read the log row: the Router vs Upstream split tells you what the hop cost, and the request id ties that row to the headers you got back.

The same call from Node, off the same environment variable:

const res = await fetch("https://api.freerouter.com/v1/chat/completions", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.FREEROUTER_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "openai/gpt-4o-mini",
    messages: [{ role: "user", content: "Say ok." }],
  }),
});
console.log(res.headers.get("x-freerouter-provider"));
console.log((await res.json()).choices[0].message.content);

The body is an ordinary chat completion: { id, model, choices, usage }. The headers worth reading are X-FreeRouter-Provider (who served it), X-FreeRouter-Attempts (how many targets were tried), and X-FreeRouter-Request-Id (quote this when you ask us anything).

flowchart TD
    A["Create fr_live_ key, openai shape"] --> B["curl chat completions, check provider header"]
    B --> C["Set OPENAI_API_KEY plus OPENAI_BASE_URL"]
    C --> D["Pick provider/model in Settings"]
    D --> E["Agent turn, check Logs"]

Prove the key on the wire before you change the app's environment.

What you get later

Three changes then cost no code. Add a second gateway as a backup target on the same rule. Move part of your traffic to another model with a remap percent, watch the logs for a day, then move the rest. When a provider has a bad hour, the traffic that would have failed lands on the other target instead.

The gateway you're on now stays in the picture. OpenRouter works fine as one target among several, which is the arrangement the Opencode post recommends for an editor. The key still sends vendor/model ids, and the dashboard decides where they land.

Limits and failure modes

  • One target means no failover. A key with a single gateway has nothing to retry against. The paper trail still works; resilience starts with the second row.
  • FreeRouter does not choose your model or write your prompts. Agent-Native's docs call Claude their most heavily tested path, and swapping in another model changes agent behavior in ways no router can see. Evaluate the model itself — identical plumbing doesn't mean identical results.
  • The same variables feed more than chat. In Agent-Native, OPENAI_API_KEY also backs realtime voice and transcription fallbacks, so those surfaces follow the endpoint override too. This post covers chat; test the others separately if you use them.
  • Overhead is small and measurable. Simple routing is sub-1 ms; remaps, translation, and failover drain push it toward 20–30 ms. Every log row splits Router and Upstream, so you can read your own number rather than trusting ours.
  • Logs hold routing metadata only. Request bodies are stored only with capture switched on, and response bodies never. Debug from the headers and the request id first.

Next steps

Create the key, prove it with the curl above, and point one environment at it. The routing rule you write today is the one that absorbs the next gateway without a release.

Route your first request today.

Bring your own keys and start routing in minutes.

Get your key