BLACKBOX AI
BLACKBOX AI routes inference across 300+ frontier and open-weight models through one encrypted endpoint your agents, CLI, and IDE can share. Enterprise Inference deploys models like Nemotron 3 Ultra on single-tenant GPUs, while the Blackbox Router connects closed models with zero data retention enforced at the gateway. One annual token commit covers the API, CLI, VS Code extension, and Agents API.
Most inference providers charge per seat or hide routing markups. BLACKBOX bills per token at published list rates, with commit discounts that improve as spend grows. Artificial Analysis ranked BLACKBOX #1 on Nemotron 3 Ultra output speed at 454 tokens per second, 30% faster than the second provider and 2.7x cheaper per token.
Platform teams and security-conscious enterprises use it when they need customer-managed keys, PII stripping before closed models, SAML SSO, and contractual zero retention without running their own GPU fleet. A forward-deployed engineer configures routing, eval harnesses, and provider migrations at contract start.
300+ models through one endpoint with smart routing, failover, and prompt caching
Nemotron 3 Ultra ranked #1 at 454 tokens per second on Artificial Analysis benchmarks
Single-tenant Enterprise Inference with reserved GPU capacity isolated per customer
Agents API, CLI, and VS Code extension metered from the same token commit
PII removed before closed models on Enterprise with contractual zero data retention
Forward-deployed engineer included for migration, routing tuning, and production integrations
Per-token billing with no platform fees or per-seat charges on published list rates.
Ranked #1 Nemotron 3 Ultra provider at 454 t/s with 2.7x lower cost than the #2 provider.
One commit covers API, CLI, VS Code, and Agents API across the organization.
Forward-deployed engineer included for migration and production integration work.
All plans require talking to sales; there is no self-serve signup on the pricing page.
Unused token commitment expires at the end of the billing period.
Dedicated single-tenant deployment and data residency are Enterprise-only features.
How does BLACKBOX AI pricing work?
BLACKBOX AI uses annual token commits sized to your workload, billed at published per-model rates with no platform fees or per-seat charges. The commit balance decreases as teams consume tokens across the API, CLI, Agents API, and VS Code extension.
Is the BLACKBOX AI API OpenAI-compatible?
Yes. BLACKBOX AI exposes OpenAI-compatible endpoints at api.blackbox.ai, including /v1/chat/completions, /v1/embeddings, and /v1/images/generations. You can use the official OpenAI Python or Node.js SDK by changing only the base URL and API key.
What models does BLACKBOX AI support?
BLACKBOX AI routes 300+ models through one endpoint, including Claude Opus 5, Gemini 3.7 Flash, Kimi K3, Nemotron 3 Ultra, and GLM 5.3. Per-token input and output rates for each model are listed on the pricing page.
How does BLACKBOX AI keep prompts private?
BLACKBOX AI enforces end-to-end encryption on every connection and contractual zero data retention on routed traffic. Enterprise plans add PII removal before prompts reach closed models, SAML SSO, SCIM, RBAC, and dedicated single-tenant deployments.
Can I use BLACKBOX AI in CI/CD pipelines?
Yes. The BLACKBOX CLI supports headless mode for non-interactive runs in Docker, GitHub Actions, GitLab CI, and any environment with Node 20+. Usage meters against the same organizational token commit as the API.
What is included in a BLACKBOX AI Enterprise contract?
BLACKBOX AI Enterprise includes a dedicated forward-deployed engineer, implementation at no added cost, single-tenant deployment, data residency options, SAML SSO, audit logs, custom SLAs, and volume discounts on per-token rates.

