AI Hardware · On-Prem
Groq
Groq builds and operates its own on-premise inference infrastructure around custom Language Processing Units (LPUs) and GroqChip hardware, exposing this via stateless, multi-region HTTP/gRPC APIs used by customers’ gateways, queues, and orchestrators.[1][3][6][7] The production patterns described in public materials place Groq as a dedicated low-latency inference backend behind customer-managed proxies, Redis queues, and Kubernetes-based orchestrators, rather than as a full cloud stack provider.[1][3][4][6]
Foundation proximity: Native (own / self-hosted weights)
Foundational models & how they consume them
PrimaryLlama 3 70B (and related Llama-family variants)· Groq-hosted on Groq LPU / GroqChip hardware · Native (own / self-hosted weights)
Cloud & datastore
On-PremCompliance
—Others that build on Meta · Llama
BasetenCerebrasCloudflareDatabricksDellFireworks AIExplore
FAQ
Does Groq use a foundational AI model?
Groq uses Llama 3 70B (and related Llama-family variants) (Native (own / self-hosted weights)).
What cloud does Groq run on?
Groq runs on On-Prem.
How close does Groq run to the foundational model?
Groq is Native (own / self-hosted weights) — close to the metal, so cheaper tokens and more control over outputs.
Sources
groq.com ↗Also on the map
← Full orbit map (260 companies) · Score your own stack →
Architecture inferred from public sources · confidence low · verify before betting on a detail.