Fireworks AI through one OpenAI-compatible gateway
Call Fireworks’ fast open-model inference with your OpenAI SDK through Routeplane — with fallback, in-data-plane guardrails, residency routing, and cost attribution.
Fireworks AI serves open-weight models over an OpenAI-compatible API on its /inference/v1 base. Routeplane’s adapter is a faithful passthrough — Fireworks’ accounts/fireworks/models/... ids pass through verbatim — with the gateway’s fallback and governance around it.
Fireworks pairs well with Together as a second open-model host in a fallback chain.
Drop in your OpenAI client
Point your existing OpenAI SDK at the gateway and add two headers — your virtual key and the provider. Nothing else about your code changes.
curl https://api.routeplane.ai/v1/chat/completions \
-H "content-type: application/json" \
-H "x-routeplane-api-key: rp_your_gateway_key" \
-H "x-routeplane-provider: fireworks" \
-d '{"model":"accounts/fireworks/models/llama-v3p3-70b-instruct","messages":[{"role":"user","content":"Hello!"}]}'
import openai
client = openai.OpenAI(
api_key="rp_your_gateway_key",
base_url="https://api.routeplane.ai/v1",
default_headers={
"x-routeplane-api-key": "rp_your_gateway_key",
"x-routeplane-provider": "fireworks", # route to Fireworks AI
},
)
resp = client.chat.completions.create(
model="accounts/fireworks/models/llama-v3p3-70b-instruct",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: 'rp_your_gateway_key',
baseURL: 'https://api.routeplane.ai/v1',
defaultHeaders: {
'x-routeplane-api-key': 'rp_your_gateway_key',
'x-routeplane-provider': 'fireworks', // route to Fireworks AI
},
});
const completion = await client.chat.completions.create({
model: 'accounts/fireworks/models/llama-v3p3-70b-instruct',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(completion.choices[0]?.message.content);
Or use the Routeplane SDK
The Routeplane SDKs subclass the official OpenAI clients, wire up the x-routeplane-* headers for you, and add typed access to routing metadata and the non-OpenAI surfaces.
pip install routeplane
from routeplane import Routeplane
client = Routeplane(
api_key="rp_your_gateway_key",
provider="fireworks", # route to Fireworks AI
)
resp = client.chat.completions.create(
model="accounts/fireworks/models/llama-v3p3-70b-instruct",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
npm i @routeplane/sdk
import { Routeplane } from '@routeplane/sdk';
const client = new Routeplane({
apiKey: process.env.ROUTEPLANE_API_KEY!, // rp_...
provider: 'fireworks', // route to Fireworks AI
});
const completion = await client.chat.completions.create({
model: 'accounts/fireworks/models/llama-v3p3-70b-instruct',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(completion.choices[0]?.message.content);
Add a fallback chain
Try Fireworks first; fall back to Together — two open-model hosts. Make the provider header a comma-separated list and the gateway walks it in order, skipping any provider whose circuit is open.
curl https://api.routeplane.ai/v1/chat/completions \
-H "content-type: application/json" \
-H "x-routeplane-api-key: rp_your_gateway_key" \
-H "x-routeplane-provider: fireworks,together" \
-d '{"model":"accounts/fireworks/models/llama-v3p3-70b-instruct","messages":[{"role":"user","content":"Hello!"}]}'
Fireworks AI on the gateway
The adapter supports chat completions, native streaming, and embeddings.
Namespaced model ids
Fireworks model ids look like accounts/fireworks/models/llama-v3p3-70b-instruct. Send them verbatim in model; the adapter passes them through unchanged.
Fireworks AI model pricing
List prices for the Fireworks AI models you can call through the gateway — click any model for its per-token cost page.
Sovereign routing
Sovereign routing works with Fireworks AI the same way it works everywhere: add x-routeplane-residency to a request, and when the gateway classifies regulated personal data in it, only providers resident in the requested region stay eligible — a hard constraint that overrides the provider header. Declare where Fireworks AI runs by setting FIREWORKS_REGION.
Keep reading
- Quickstart — send your first request in a few minutes.
- Providers & routing strategy — how eligibility, fallback, and strategies work.
- Python SDK and TypeScript SDK — the full typed clients.
- All providers — every model provider behind the gateway.
Frequently asked questions
Is Routeplane a drop-in replacement for calling Fireworks AI directly?
Yes. The gateway is OpenAI-compatible, so switching is a base-URL change — point your existing OpenAI SDK at https://api.routeplane.ai/v1 and set the x-routeplane-provider header to fireworks. You keep your code and gain automatic fallback, in-data-plane PII guardrails, sovereign routing, and per-team cost attribution.
Which Fireworks AI models can I use through Routeplane?
Any Fireworks AI model the provider serves. See the list prices linked above for the ones we publish.
Can I enforce data residency on Fireworks AI requests?
Yes. Add the x-routeplane-residency header (for example IN) and, when a request carries regulated personal data, the gateway restricts routing to providers eligible in that region — overriding the provider header if it has to. The Fireworks AI adapter reads FIREWORKS_REGION to declare where it is resident.
Route your first request this week.
Point your existing OpenAI-compatible client at Routeplane, set one header, and get fallback, guardrails, and sovereign routing across every provider.