Skip to content

· inference api

Frontier and open models, at the lowest published price.

One OpenAI-compatible endpoint. Change your base URL, keep your code. Served on dedicated Azure AI Foundry capacity.

ModelContextInput / 1MOutput / 1MFeatures
gpt oss 120b
openai/gpt-oss-120b
131K$0.036$0.167tools, structured
DeepSeek V4 Flash 0731
deepseek-ai/DeepSeek-V4-Flash-0731
164K$0.059$0.176tools, structured
DeepSeek V4 Flash
deepseek-ai/DeepSeek-V4-Flash
164K$0.088$0.176tools, structured
Llama 3.3 70B Instruct
meta-llama/Llama-3.3-70B-Instruct
131K$0.132$0.392tools, structured
Kimi K2.7 Code
moonshotai/Kimi-K2.7-Code
262K$0.666$3.33tools, structured
Kimi K2.6
moonshotai/Kimi-K2.6
262K$0.735$3.43tools, structured
DeepSeek V4 Pro
deepseek-ai/DeepSeek-V4-Pro
164K$1.27$2.55tools, structured, reasoning

Prices are per million tokens and billed on actual usage. Cached input is billed at the cached rate where the model supports it. Fetched live from the API — this table and your invoice read the same source.

01

Sign in

The same account you use for gitgot. No separate signup, no sales call.

02

Create a key

One click in the dashboard. The key is shown once and stored only as a hash.

03

Change your base URL

Point any OpenAI SDK at our endpoint. Streaming, tool calling and structured output all work unchanged.

drop-in
from openai import OpenAI

client = OpenAI(
    base_url="https://inference.gitgot.ai/v1",
    api_key="gg-...",           # from the dashboard
)

client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4-Pro",
    messages=[{"role": "user", "content": "hello"}],
    stream=True,
)

Prepaid, never surprising

Top up what you want to spend. When the balance runs out, requests return 402 — there is no invoice at the end of the month and no way to overspend.

Priced per token, no cliffs

One rate per model, whatever your context length. No doubling above a threshold and no peak-hour pricing.

Dedicated capacity

Served on reserved Azure AI Foundry throughput across multiple deployments, with automatic failover between them.

Honest about retention

We do not store prompts or completions. Microsoft applies abuse monitoring to Azure-served models and may retain flagged content, so we cannot claim zero retention — see the privacy page.