gemma2-9b-it pricing
List price for the gemma2-9b-it API from Groq — input cost, output cost, and context window, plus what it actually costs to run across a few common workloads. All figures in USD per 1M tokens.
What gemma2-9b-it costs for typical workloads
Estimated cost per request and per 1,000 requests at gemma2-9b-it's list price, for a few representative input/output token mixes. Your real usage will vary with prompt and response length.
| Workload | Input tokens | Output tokens | Cost / request | Cost / 1,000 |
|---|---|---|---|---|
| Short chat turn | 1,000 | 300 | $0.00026 | $0.260 |
| RAG answer | 8,000 | 500 | $0.00170 | $1.70 |
| Agent step (tool-heavy) | 20,000 | 2,000 | $0.00440 | $4.40 |
| Long-document summary | 100,000 | 1,000 | $0.020 | $20.20 |
Provider-published list price, USD per 1M tokens, current as of July 2026. Batch, cached-input, and volume discounts are not applied. Verify with Groq before relying on these figures for billing.
Similar models to compare
Models at a comparable price point — click through for the same per-token and per-workload breakdown.
Route to the cheapest model automatically.
Stop hand-tuning which model gets which request. Routeplane's difficulty router scores each prompt and picks the optimal model per request — downgrading easy calls to a cheaper model and reserving the expensive ones for the hard prompts, with automatic fallback if a provider is down. One OpenAI-compatible endpoint, every provider behind it.