Cerebras
World-class LLM inference speeds via wafer-scale AI chips, ideal for agentic workloads.
From $0.10/MTok
Get Started - Category
- LLM Platforms & APIs
- Pricing
- From $0.10/MTok
- Vendor
- Cerebras Systems
- Website
- cloud.cerebras.ai
About Cerebras
Cerebras delivers industry-leading inference throughput using its wafer-scale CS-3 AI chips, achieving speeds far exceeding GPU-based systems. Supports LLaMA 3.1 70B and 405B models, making it ideal for latency-sensitive agentic pipelines.
What you get
- Wafer-scale chip for extreme speed
- LLaMA 3.1 70B & 405B support
- Ideal for agentic low-latency workloads
- Simple REST API
- Competitive token pricing
- Streaming responses
Specifications
| Context Window | 128K tokens |
|---|---|
| Tool Use | No |
| Vision | No |
| Streaming | Yes |
| Open Source | Yes |
| Self-Host | No |
| Starting Price | $0.10/MTok input |