Baseten
AI-infra · Multi-Cloud · Native (own / self-hosted weights)
Model-inference platform serving open models (Llama, etc.) on autoscaling GPU infra.
Who builds on Meta · Llama
15 companies in the Modeling Warehouse build on Meta · Llama, ranked by how directly they consume it — Native first, Middleware last. Avg Foundation Proximity Score 97/100.
AI-infra · Multi-Cloud · Native (own / self-hosted weights)
Model-inference platform serving open models (Llama, etc.) on autoscaling GPU infra.
AI-hardware · On-Prem · Native (own / self-hosted weights)
Designs wafer-scale AI chips and serves open models (Llama) at record speed on its own systems.
Infrastructure · On-Prem · Native (own / self-hosted weights)
Cloudflare runs its own Linux-based servers and private backbone as a globally anycasted edge and core network, with most products (CDN, Zero Trust, R2, KV, Queues, Agents, Containers) built on an internal proxy chain and the Workers/Durable Objects developer platform.[3][9][10] Control planes and schedulers (for Containers, Workers AI, agents, etc.) are themselves implemented on Cloudflare’s stack, indicating a predominantly self-hosted, on-prem architecture rather than dependence on hyperscale public clouds.[7][8][10]
Data Platform · Multi-Cloud · Native (own / self-hosted weights)
Databricks runs a two-layer architecture where a Databricks-managed control plane hosts the UI, APIs, metadata, and orchestration services, while customer workloads execute in a compute/data plane inside the customer’s AWS, Azure, or GCP account (or Databricks’ serverless account) against cloud object storage using Spark and Delta Lake.[3][4][9][11] Around this core, Databricks positions an open-core lakehouse stack (Delta Lake, Spark, MLflow) with proprietary governance, serverless, and AI platform services, integrated into medallion-style production patterns.[1][2][6][12]
Hardware/Enterprise · Hybrid · Native (own / self-hosted weights)
Dell AI Factory packages on-prem open models (Llama) + partners (Cohere) for enterprise/sovereign deployment.
AI Infra · Multi-Cloud · Native (own / self-hosted weights)
Fast inference for open models on its own serving stack.
AI Hardware · On-Prem · Native (own / self-hosted weights)
Groq builds and operates its own on-premise inference infrastructure around custom Language Processing Units (LPUs) and GroqChip hardware, exposing this via stateless, multi-region HTTP/gRPC APIs used by customers’ gateways, queues, and orchestrators.[1][3][6][7] The production patterns described in public materials place Groq as a dedicated low-latency inference backend behind customer-managed proxies, Redis queues, and Kubernetes-based orchestrators, rather than as a full cloud stack provider.[1][3][4][6]
AI/LLM · AWS · Native (own / self-hosted weights)
The model/dataset hub on AWS, storing weights as Git LFS repos over S3 and running Inference Endpoints/Spaces. Stewards open libraries (Transformers, Diffusers) — the open-source center of gravity for ML.
AI Infra · Multi-Cloud · Native (own / self-hosted weights)
Builds PyTorch Lightning + a studio to train/serve open models (Llama, etc.) on cloud GPUs.
AI Infra · GCP · Native (own / self-hosted weights)
Runs open models as one-click APIs on its own GPU fleet.
AI-hardware · On-Prem · Native (own / self-hosted weights)
Own RDU AI chips plus its Samba-1 model and hosted open models for enterprise/government.
AI/Data · AWS · Native (own / self-hosted weights)
Beyond data-labeling, Scale trains its own Defense Llama (fine-tuned Llama 3) for U.S. national security on its Donovan federal platform; evaluates frontier models for the DoD CDAO.
Data Platform · Multi-Cloud · Native (own / self-hosted weights)
Snowflake runs a proprietary three-layer, multi-cluster shared-data architecture on AWS, Azure, and GCP, separating compressed columnar storage on cloud object stores from MPP compute warehouses and a distributed cloud services control plane.[1][16][19] Internally it is a service-oriented system with independently scalable storage, compute, and metadata/transaction services built for OLAP workloads.[16][19]
AI Infra · Hybrid · Native (own / self-hosted weights)
Together AI operates its own GPU cloud and managed Kubernetes clusters as an AI-native infrastructure layer, while also integrating with external cloud storage and streaming services for data and analytics workloads.[2][11][12] The production stack is centered on open-weight models deployed as containerized workloads on Together-managed GPU clusters, exposed via serverless inference APIs and voice/agent pipelines.[3][4][9][11][16]
Enterprise · On-Prem · Cloud-hosted (Bedrock/Vertex/Azure)
Runs its own OCI; OCI Generative AI hosts Cohere + Llama for enterprise.
← Back to the full orbit map (all 260 companies) · Score your own stack →