Cerebras
Cerebras 构建了晶圆级 AI 芯片,并运行当今最快的 LLM 推理平台之一。其 CS-2 和 CS-3 系统为需要在生产规模下实现低延迟响应的团队提供云 API、专用私有终端和本地部署。
晶圆级引擎是一个单芯片,比典型的 GPU 大 58 倍,专为训练和推理工作负载设计。在云端,开发者可以通过简单的 API 密钥调用开源模型,如 Llama、Qwen、GLM 和 GPT,其吞吐量在支持的模型中常常达到每秒数千个 Token。
Cerebras 服务于 AI 原生的创业公司、企业研究团队以及覆盖医疗保健、网络安全和药物发现的全球公司。客户包括 OpenAI、Meta、GSK、Notion 和 Mayo Clinic。您可以免费开始使用推理 API,通过按需付费的开发者计费进行扩展,或与销售团队联系获取专用容量和定制模型权重。
Wafer-Scale Engine芯片比标准GPU大58倍
云推理API,支持Llama、Qwen、GLM和GPT OSS模型
Gemma 4在Cerebras硬件上每秒运行1500+个tokens
GPT OSS 120B在开发者层每秒大约处理3000个tokens
使用CS-2或CS-3进行本地部署,实现完整数据和模型控制
通过AWS Marketplace、OpenRouter、Hugging Face和Vercel进行合作伙伴集成
免费账户在推理API上包含$5的额度,根据定价页面
免费账户包含$5额度,用于测试Cerebras支持的模型
已发布的推理速度在支持的开源模型上达到每秒数千个tokens
灵活的部署选项涵盖云API、专用端点和本地CS-2或CS-3系统
通过AWS Marketplace、OpenRouter、Hugging Face和Vercel的合作伙伴分发降低集成难度
Cerebras Code Pro 和 Max 订阅在定价页面上显示已售罄。
企业定价和专用容量需要联系销售团队。
预览模型如 GLM 4.7 标记为仅供评估,不得用于生产。
Does Cerebras offer a free inference API tier?
Yes. Cerebras provides a free trial with $5 in free credits after you create an account at cloud.cerebras.ai. It includes access to Cerebras-powered models, fast inference, and community support through Discord.
How much does the Cerebras Developer tier cost?
The Cerebras Developer tier uses self-serve pay-as-you-go billing starting at $10. It includes 10x higher rate limits than the free tier and higher priority processing, with per-token prices listed for models like GPT OSS 120B.
What is Cerebras Code and how much does it cost?
Cerebras Code is a coding-focused subscription with Pro at $50 per month for up to 24 million tokens per day and Max at $200 per month for up to 120 million tokens per day. Both tiers were listed as sold out on the pricing page at the time of research.
Which models can I run on Cerebras inference?
Cerebras serves open models including GLM, OpenAI-compatible OSS models, Qwen, and Llama through its cloud API. The pricing page lists developer-tier access to models like ZAI GLM 4.7 and GPT OSS 120B with published per-token rates.
Can I deploy Cerebras on my own infrastructure?
Yes. Cerebras offers on-prem deployment with CS-2 and CS-3 systems for teams that need full control over models, data, and infrastructure in their own data center or private cloud.
How do I access Cerebras inference through third-party platforms?
Cerebras inference is available through partner APIs including AWS Marketplace, OpenRouter, Hugging Face Hub, and Vercel AI Gateway. Each partner provides its own onboarding flow while routing requests to Cerebras hardware.

