◆ Gentoo Logic · Modeling Warehouse

Who builds on Meta · Llama

15 companies build on Meta · Llama

15 companies in the Modeling Warehouse build on Meta · Llama, ranked by how directly they consume it — Native first, Middleware last. Avg Foundation Proximity Score 97/100.

Native 14Direct API 0Cloud-hosted 1Middleware 0

Baseten

AI-infra · Multi-Cloud · Native (own / self-hosted weights)

Model-inference platform serving open models (Llama, etc.) on autoscaling GPU infra.

Cerebras

AI-hardware · On-Prem · Native (own / self-hosted weights)

Designs wafer-scale AI chips and serves open models (Llama) at record speed on its own systems.

Cloudflare

Infrastructure · On-Prem · Native (own / self-hosted weights)

Cloudflare runs its own Linux-based servers and private backbone as a globally anycasted edge and core network, with most products (CDN, Zero Trust, R2, KV, Queues, Agents, Containers) built on an internal proxy chain and the Workers/Durable Objects developer platform.[3][9][10] Control planes and schedulers (for Containers, Workers AI, agents, etc.) are themselves implemented on Cloudflare’s stack, indicating a predominantly self-hosted, on-prem architecture rather than dependence on hyperscale public clouds.[7][8][10]

Databricks

Data Platform · Multi-Cloud · Native (own / self-hosted weights)

Databricks runs a two-layer architecture where a Databricks-managed control plane hosts the UI, APIs, metadata, and orchestration services, while customer workloads execute in a compute/data plane inside the customer’s AWS, Azure, or GCP account (or Databricks’ serverless account) against cloud object storage using Spark and Delta Lake.[3][4][9][11] Around this core, Databricks positions an open-core lakehouse stack (Delta Lake, Spark, MLflow) with proprietary governance, serverless, and AI platform services, integrated into medallion-style production patterns.[1][2][6][12]

Dell

Hardware/Enterprise · Hybrid · Native (own / self-hosted weights)

Dell AI Factory packages on-prem open models (Llama) + partners (Cohere) for enterprise/sovereign deployment.

Fireworks AI

AI Infra · Multi-Cloud · Native (own / self-hosted weights)

Fast inference for open models on its own serving stack.

Groq

AI Hardware · On-Prem · Native (own / self-hosted weights)

Groq builds and operates its own on-premise inference infrastructure around custom Language Processing Units (LPUs) and GroqChip hardware, exposing this via stateless, multi-region HTTP/gRPC APIs used by customers’ gateways, queues, and orchestrators.[1][3][6][7] The production patterns described in public materials place Groq as a dedicated low-latency inference backend behind customer-managed proxies, Redis queues, and Kubernetes-based orchestrators, rather than as a full cloud stack provider.[1][3][4][6]

Hugging Face

AI/LLM · AWS · Native (own / self-hosted weights)

The model/dataset hub on AWS, storing weights as Git LFS repos over S3 and running Inference Endpoints/Spaces. Stewards open libraries (Transformers, Diffusers) — the open-source center of gravity for ML.

Lightning AI

AI Infra · Multi-Cloud · Native (own / self-hosted weights)

Builds PyTorch Lightning + a studio to train/serve open models (Llama, etc.) on cloud GPUs.

Replicate

AI Infra · GCP · Native (own / self-hosted weights)

Runs open models as one-click APIs on its own GPU fleet.

SambaNova

AI-hardware · On-Prem · Native (own / self-hosted weights)

Own RDU AI chips plus its Samba-1 model and hosted open models for enterprise/government.

Scale AI

AI/Data · AWS · Native (own / self-hosted weights)

Beyond data-labeling, Scale trains its own Defense Llama (fine-tuned Llama 3) for U.S. national security on its Donovan federal platform; evaluates frontier models for the DoD CDAO.

Snowflake

Data Platform · Multi-Cloud · Native (own / self-hosted weights)

Snowflake runs a proprietary three-layer, multi-cluster shared-data architecture on AWS, Azure, and GCP, separating compressed columnar storage on cloud object stores from MPP compute warehouses and a distributed cloud services control plane.[1][16][19] Internally it is a service-oriented system with independently scalable storage, compute, and metadata/transaction services built for OLAP workloads.[16][19]

Together AI

AI Infra · Hybrid · Native (own / self-hosted weights)

Together AI operates its own GPU cloud and managed Kubernetes clusters as an AI-native infrastructure layer, while also integrating with external cloud storage and streaming services for data and analytics workloads.[2][11][12] The production stack is centered on open-weight models deployed as containerized workloads on Together-managed GPU clusters, exposed via serverless inference APIs and voice/agent pipelines.[3][4][9][11][16]

Oracle

Enterprise · On-Prem · Cloud-hosted (Bedrock/Vertex/Azure)

Runs its own OCI; OCI Generative AI hosts Cohere + Llama for enterprise.

← Back to the full orbit map (all 260 companies) · Score your own stack →