Private beta DPDP-ready sovereign routing is live, request access
Groq

gemma2-9b-it pricing

List price for the gemma2-9b-it API from Groq — input cost, output cost, and context window, plus what it actually costs to run across a few common workloads. All figures in USD per 1M tokens.

Last updated: July 2026
Input
$0.200 / 1M tokens
Output
$0.200 / 1M tokens
1M in + 1M out
$0.400 blended
Context window
8K tokens

What gemma2-9b-it costs for typical workloads

Estimated cost per request and per 1,000 requests at gemma2-9b-it's list price, for a few representative input/output token mixes. Your real usage will vary with prompt and response length.

Workload Input tokens Output tokens Cost / request Cost / 1,000
Short chat turn 1,000 300 $0.00026 $0.260
RAG answer 8,000 500 $0.00170 $1.70
Agent step (tool-heavy) 20,000 2,000 $0.00440 $4.40
Long-document summary 100,000 1,000 $0.020 $20.20

Provider-published list price, USD per 1M tokens, current as of July 2026. Batch, cached-input, and volume discounts are not applied. Verify with Groq before relying on these figures for billing.

Route to the cheapest model automatically.

Stop hand-tuning which model gets which request. Routeplane's difficulty router scores each prompt and picks the optimal model per request — downgrading easy calls to a cheaper model and reserving the expensive ones for the hard prompts, with automatic fallback if a provider is down. One OpenAI-compatible endpoint, every provider behind it.