RESEARCH: converting an 8B transformer into an attention-free diffusion modelREAD THE RESEARCH →
OPEN-WEIGHT INFERENCE

Open models.
Lower inference costs.

Run Qwen, Gemma, and GPT-OSS well below the market’s same-model average. Pay for the input and output tokens you use.

Qwen 3.8 27B is available now in the dashboard.

Building on the API? Read the inference API reference for the base URL, the model id and a curl you can paste.

MODEL PRICING

Every token, accounted for.

Input and output rates, per million tokens.
Obit’s published list price, beside the market’s same-model average.

Qwen 3.8 27B is ready to use in the dashboard. Contact us to set up any of the other models.

QWENReady in the dashboard

Qwen 3.8 27B

USD / 1 million tokens

OBIT63–68% less
Input
$0.10
Output
$1.00

Cached input: $0.01 / 1M

Market average · same model
Input
$0.3146
Output
$2.6667

Save $0.2146 per 1M input tokens and $1.6667 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
Qwen 3.8 27Bxhigh · on Obit34$0.10$1.00
GPT-5.6 Lunamax38$0.20$1.20
Claude Sonnet 5max38$2.00$10.00

Qwen 3.8 27B on Obit costs less per token than GPT-5.6 Luna and Claude Sonnet 5, on input and on output.

Nearby scores do not imply identical performance. Sources & methodology ↓

Start generating
OPENAISetup by request

GPT-OSS 20B

USD / 1 million tokens

OBIT30% less
Input
$0.0293
Output
$0.1152

Cached input: $0.0029 / 1M

Market average · same model
Input
$0.0418
Output
$0.1646

Save $0.0125 per 1M input tokens and $0.0494 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
GPT-OSS 20Bhigh · on Obit9$0.0293$0.1152
GPT-4.1 mininon-reasoning10 est.$0.40$1.60

Nearby scores do not imply identical performance. “est.” marks an Artificial Analysis estimate. Sources & methodology ↓

Contact us to set up
QWENSetup by request

Qwen 3.6 35B-A3B

USD / 1 million tokens

OBIT30% less
Input
$0.0961
Output
$0.7143

Cached input: $0.0096 / 1M

Market average · same model
Input
$0.1373
Output
$1.0204

Save $0.0412 per 1M input tokens and $0.3061 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
Qwen 3.6 35B-A3Breasoning · on Obit19$0.0961$0.7143
GPT-5 minihigh17$0.25$2.00

Nearby scores do not imply identical performance. Sources & methodology ↓

Contact us to set up
GOOGLESetup by request

Gemma 4 31B

USD / 1 million tokens

OBIT30% less
Input
$0.1925
Output
$0.4282

Cached input: $0.0193 / 1M

Market average · same model
Input
$0.275
Output
$0.6117

Save $0.0825 per 1M input tokens and $0.1835 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
Gemma 4 31Breasoning · on Obit15$0.1925$0.4282
Claude Haiku 4.5non-reasoning15 est.$1.00$5.00

Nearby scores do not imply identical performance. “est.” marks an Artificial Analysis estimate. Sources & methodology ↓

Contact us to set up
GOOGLESetup by request

Gemma 4 26B-A4B

USD / 1 million tokens

OBIT30% less
Input
$0.0727
Output
$0.2565

Cached input: $0.0073 / 1M

Market average · same model
Input
$0.1038
Output
$0.3664

Save $0.0311 per 1M input tokens and $0.1099 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
Gemma 4 26B-A4Breasoning · on Obit17 est.$0.0727$0.2565
GPT-5 minihigh17$0.25$2.00

Nearby scores do not imply identical performance. “est.” marks an Artificial Analysis estimate. Sources & methodology ↓

Contact us to set up
Pricing sources & comparison methodology · market snapshot September 13, 2026

Obit’s rates are list prices, not a discount off anyone else’s. Every live model’s input, output, and cached-input price is declared in pricing.json — the machine-readable source for the OpenAI-compatible API at https://relay.obitmc.com/v1. This page renders that declaration and re-reads the same file in your browser, so a change to a list price shows here without a site deploy. Cached input is a prompt prefix served from a worker’s prefix cache rather than recomputed, so a repeated prefix costs less. It is priced 90% below that model’s input rate. The first send of a prefix pays the full rate.

The market average is the arithmetic mean of available paid providers for the same model on OpenRouter, captured September 13, 2026. We average each provider’s available endpoints first, then weight providers equally. Free and unavailable endpoints are excluded. Input and output are compared separately, and every percentage on this page is (market − Obit) ÷ market for that model, rounded to a whole percent.

Models marked “setup by request” are not live on Obit yet and have no declared list price. Their figures are indicative, 30% below that same market average, and their cached rate sits 90% below their input rate. We confirm the rate when we set one up.

Market comparisons use standard uncached text-token rates in USD per million tokens. Batch discounts, other modalities, and custom cloud deployments are excluded. Total workload cost depends on token usage, including reasoning tokens.

Comparisons use the same Artificial Analysis Intelligence Index v4.3, with the evaluated reasoning effort shown for each model. We selected GPT/Claude models within four rounded index points; scores describe the evaluated models, not a benchmark of Obit’s deployments. Some reference models are older generations.

LIVE AVAILABILITY

API & model status

View status history ↗

View current availability and uptime history. · Percentages reflect recorded checks.

BUILT AROUND YOUR WORKLOAD

Choose what you run.
And where you run it.

Make use of idle GPUs.

Obit brings together spare capacity in datacenters and idle GPUs from mining operations to serve open-weight models. Putting existing hardware to use helps us keep inference costs down. See where your prompt runs and what those machines can and can't see.

Keep it on AWS or GCP.

Need the security controls and infrastructure assurances of AWS or GCP? We can provide a dedicated endpoint that runs your workload only on your chosen cloud. Contact us to discuss your requirements and pricing.

Bring your own model.

Custom models are supported. Email jboesch@obitmc.com to discuss your model and setup.