· inference api
Frontier and open models, at the lowest published price.
One OpenAI-compatible endpoint. Change your base URL, keep your code. Served on dedicated Azure AI Foundry capacity.
| Model | Context | Input / 1M | Output / 1M | Features |
|---|---|---|---|---|
gpt oss 120b openai/gpt-oss-120b | 131K | $0.036 | $0.167 | tools, structured |
DeepSeek V4 Flash 0731 deepseek-ai/DeepSeek-V4-Flash-0731 | 164K | $0.059 | $0.176 | tools, structured |
DeepSeek V4 Flash deepseek-ai/DeepSeek-V4-Flash | 164K | $0.088 | $0.176 | tools, structured |
Llama 3.3 70B Instruct meta-llama/Llama-3.3-70B-Instruct | 131K | $0.132 | $0.392 | tools, structured |
Kimi K2.7 Code moonshotai/Kimi-K2.7-Code | 262K | $0.666 | $3.33 | tools, structured |
Kimi K2.6 moonshotai/Kimi-K2.6 | 262K | $0.735 | $3.43 | tools, structured |
DeepSeek V4 Pro deepseek-ai/DeepSeek-V4-Pro | 164K | $1.27 | $2.55 | tools, structured, reasoning |
Prices are per million tokens and billed on actual usage. Cached input is billed at the cached rate where the model supports it. Fetched live from the API — this table and your invoice read the same source.
Sign in
The same account you use for gitgot. No separate signup, no sales call.
Create a key
One click in the dashboard. The key is shown once and stored only as a hash.
Change your base URL
Point any OpenAI SDK at our endpoint. Streaming, tool calling and structured output all work unchanged.
from openai import OpenAI
client = OpenAI(
base_url="https://inference.gitgot.ai/v1",
api_key="gg-...", # from the dashboard
)
client.chat.completions.create(
model="deepseek-ai/DeepSeek-V4-Pro",
messages=[{"role": "user", "content": "hello"}],
stream=True,
)Prepaid, never surprising
Top up what you want to spend. When the balance runs out, requests return 402 — there is no invoice at the end of the month and no way to overspend.
Priced per token, no cliffs
One rate per model, whatever your context length. No doubling above a threshold and no peak-hour pricing.
Dedicated capacity
Served on reserved Azure AI Foundry throughput across multiple deployments, with automatic failover between them.
Honest about retention
We do not store prompts or completions. Microsoft applies abuse monitoring to Azure-served models and may retain flagged content, so we cannot claim zero retention — see the privacy page.