Groq
Groq runs open models on custom LPU chips for fast inference. 2026 pricing starts near $0.05 per million tokens, with Llama 4 and MCP support.
About Groq
What is Groq (2026 update)
Groq is an AI inference platform in the coding and development category that runs large language models on its own custom chip, called the LPU (Language Processing Unit), instead of standard GPUs. The pitch is speed: Groq serves open models like Llama and GPT-OSS through an API called GroqCloud, aimed at developers who need low-latency responses for chat, voice, and agentic applications. Because Groq builds its own chip rather than renting GPU capacity, it can price some of its smaller models well below typical GPU-based API pricing while still keeping response times fast.
What is new in 2026
As of mid-2026, GroqCloud added day-zero access to Meta’s Llama 4 Scout and Llama 4 Maverick models, plus newer enterprise-oriented models such as MiniMax M2.5 and Qwen3-VL 32B Instruct. Groq introduced a beta Remote MCP server integration that connects models to external tools through Anthropic’s Model Context Protocol, alongside support for the OpenAI Responses API. On the voice side, Groq’s Orpheus text-to-speech model now runs at roughly 100 characters per second and added new Arabic Saudi voices, bringing its total voice roster to six. Groq also expanded its Compound and Compound Mini systems, agentic wrappers that give a model built-in web search and code execution at around 450 tokens per second with a 131k context window.
Key features
- Custom LPU hardware built for fast, low-latency token generation
- GroqCloud API access to Llama 4 Scout/Maverick, GPT-OSS, and other open models
- Remote MCP integration in beta for connecting models to external tools
- Compound and Compound Mini agentic systems with built-in web search and code execution
- Speech-to-text and text-to-speech, including the Orpheus voice model
- Batch API and prompt caching, JSON/structured outputs, and LoRA inference support
Pricing in 2026
Groq bills per million tokens and prices vary by model. A free tier is available with no credit card required, limited to around 30 requests per minute for prototyping.
| Plan | Price | What you get |
|---|---|---|
| Free tier | Free | API access with rate limits around 30 requests per minute, no credit card needed |
| Llama 3.1 8B Instant | $0.05 per 1M input tokens, $0.08 per 1M output tokens | Fast, low-cost model for simple tasks |
| Llama 4 Scout | $0.11 per 1M input tokens, $0.34 per 1M output tokens | Newer Llama 4 model with day-zero GroqCloud access |
| Llama 3.3 70B Versatile | $0.59 per 1M input tokens, $0.79 per 1M output tokens | Larger model for more complex reasoning |
| Batch API / prompt caching | 50% off standard rate each, stackable to about 25% of on-demand cost | Discounted pricing for asynchronous or repeated-prompt workloads |
Who should use it
Groq suits developers building latency-sensitive applications, such as voice assistants, real-time chat, and agentic tools, where response speed matters as much as model quality. Startups testing multiple open models can use the free tier before committing to paid volume, and teams already using Llama or GPT-OSS models can switch to Groq mainly for the speed and lower per-token cost on smaller models. Companies building voice products can also lean on Groq’s Orpheus text-to-speech and its speech-to-text models to keep an entire voice pipeline on one provider instead of stitching together separate services.
Limitations and alternatives
Groq focuses on open and licensed third-party models rather than training its own frontier model, so it depends on releases from Meta, OpenAI’s open models, and others for its lineup, and older models get deprecated over time per its published schedule. Pricing on larger models is still higher than the cheapest small models, so cost-sensitive teams should check the model-by-model rate card before committing. Alternatives include Together AI and Fireworks AI for similar open-model hosting, and Cerebras for another specialized inference chip aimed at high-throughput workloads.
Frequently asked questions
What makes Groq different from other AI API providers?
Groq runs models on its own custom LPU chip rather than standard GPUs, which the company says gives it faster token generation speed, especially useful for real-time voice and chat applications.
How much does the Groq API cost?
Pricing varies by model, ranging from about $0.05 per million input tokens for small models like Llama 3.1 8B Instant up to $0.59 per million input tokens for larger models like Llama 3.3 70B.
Does Groq have a free tier?
Yes, Groq offers a free tier with no credit card required, though it is limited to around 30 requests per minute, which is enough for prototyping but not production traffic.
Does Groq support tool use and agents in 2026?
Yes, Groq added a beta Remote MCP integration for connecting models to external tools, plus Compound and Compound Mini agentic systems with built-in web search and code execution.
Information last verified: September 2026.
Related Tools