Pricing comparison

Evidence for buyer-cost claims in the Litepaper. Last reviewed: 16 August 2026.
Savings are model-specific. The old 20–40% band is not a uniform guarantee. This page snapshots live Malibu rates against public OpenRouter list prices on the same family.

Live Malibu pricing

Same payload on https://malibu.tech/v1/rate-card. Full formula: Metering & billing. Canonical regimes: Economics.

Dated snapshot — 16 August 2026

Malibu card: generated_at 2026-07-29 (version 51d4eb9d…).
OpenRouter: public GET https://openrouter.ai/api/v1/models on 16 Aug 2026 (prompt/completion fields × 1e6 → USD per 1M tokens).
Blend column uses 1M prompt + 0.5M completion tokens (same mix as the worked example below). Malibu IDs: meta-llama/llama-3.2-3b-instruct, qwen3-8b, meta-llama/llama-3.1-8b-instruct, google-gemma-4-26b-a4b-it, qwen3-32b, openai/gpt-oss-20b, nemotron-3-nano-30b-a3b. OpenRouter IDs: meta-llama/llama-3.2-3b-instruct, qwen/qwen3-8b, meta-llama/llama-3.1-8b-instruct, google/gemma-4-26b-a4b-it, qwen/qwen3-32b, openai/gpt-oss-20b, nvidia/nemotron-3-nano-30b-a3b.
Only Llama 3.2 3B and Qwen3 8B were warm on 16 Aug 2026. A cheaper rate-card row does not mean the model routes today. Check GET /v1/models. Quantization, latency, and receipt semantics differ from OpenRouter even when the family name matches.

How to rerun the comparison

  1. Download https://api.malibu.tech/v1/rate-card and record generated_at.
  2. Download OpenRouter (or another aggregator) list prices with date.
  3. Convert Malibu credits: usd = credits_per_mtok / 1e6 * usd_per_million_credits.
  4. Compare the same token mix. Include prompt_cache_hit_rate_per_mtok if your agent replays context.
  5. Exclude closed flagship SKUs with no OSS equivalent.

Default row (no model match)

If a request hits the default rate-card row: 0.50/0.50 / 1.00 per 1M prompt/completion — usually more expensive than aggregator OSS list. Prefer an explicit priced model.

When Malibu is not cheaper

  • Cold catalog rows — price without supply is not a savings.
  • Some 20B–32B SKUs — this snapshot is flat or slightly above OpenRouter list.
  • Latency / availability — consumer Mac routing vs dedicated GPU clusters.
  • Confidential compute — cooperative-trust, not TEE. Economic comparison only.