Groq
Ultra-fast LLM inference via custom LPU chips, running LLaMA, Mixtral, and Gemma.
Free / $0.05/MTok Free tier available
Try Free - Category
- LLM Platforms & APIs
- Pricing
- Free / $0.05/MTok
- Website
- console.groq.com
About Groq
Groq delivers the fastest available LLM inference using its proprietary Language Processing Unit (LPU) chips. Runs popular open-source models like LLaMA 3, Mixtral, and Gemma with a developer-friendly, OpenAI-compatible API and generous free tier.
What you get
- Fastest available inference speed
- LLaMA 3, Mixtral, Gemma models
- OpenAI-compatible API
- Free tier with generous rate limits
- Low-latency for agentic workflows
- Streaming support
Specifications
| Context Window | 128K tokens |
|---|---|
| Tool Use | Yes |
| Vision | No |
| Streaming | Yes |
| Open Source | Yes |
| Self-Host | No |
| Starting Price | Free tier available |