How AI/LLM companies build with AI — which foundational models they run and how directly, ranked by proximity to the model. Avg Foundation Proximity Score 100/100.
AI/LLM · Multi-Cloud · Native (own / self-hosted weights)
AI21 Labs develops and serves its own Jamba family of LLMs and related services via its AI21 Studio API and private deployments, while also distributing Jamba models through third‑party clouds such as Azure and upcoming NVIDIA APIs.[5][6][7] Public materials describe deployment options across public cloud and private/on‑prem environments but do not reveal a full production stack beyond this high‑level multi‑cloud posture.[3][5][6]
AI/LLM · Hybrid · Native (own / self-hosted weights)
Aleph Alpha’s current product layer appears to be PhariaAI, a sovereign enterprise stack that includes knowledge capture, development, operation, access control, and monitoring components. Public materials also show first-party API access and containerized deployment patterns, but the exact production datastore and hosting mix are not fully disclosed, so the infrastructure is best characterized as hybrid rather than purely on-prem or purely cloud.[2][5][10]
AI/LLM · Multi-Cloud · Native (own / self-hosted weights)
Anthropic runs **Claude** and related services on a safety-first **multi-cloud** compute fabric spanning AWS Trainium2, Google TPUv7 and NVIDIA GPUs, fronted by Kubernetes‑based microservices (API gateways, orchestration, rate limiting, caching, and safety filters), with state held in PostgreSQL, vector stores, Redis, and cloud object storage.[1][10][13] Production offerings like Claude API and Managed Agents expose a fully managed orchestration and agent runtime, while emerging self‑hosted sandboxes move tool execution into customer infrastructure but keep Claude inference, routing, and session state on Anthropic’s cloud.[2][3][11]
AI/LLM · Multi-Cloud · Native (own / self-hosted weights)
Enterprise-focused LLM lab training its own Command/Embed/Rerank models, served across clouds (Google Cloud, Oracle, AWS) and deployable in-VPC for data-sensitive customers. North platform targets RAG.
AI/LLM · Multi-Cloud · Native (own / self-hosted weights)
RAG-native enterprise platform training its own grounded language models.
AI/LLM · AWS · Native (own / self-hosted weights)
The model/dataset hub on AWS, storing weights as Git LFS repos over S3 and running Inference Endpoints/Spaces. Stewards open libraries (Transformers, Diffusers) — the open-source center of gravity for ML.
AI/LLM · Azure · Native (own / self-hosted weights)
Trains its own Inflection models (now enterprise-focused) on Azure supercompute.
AI/LLM · Hybrid · Native (own / self-hosted weights)
Liquid AI builds and serves its own Liquid Foundation Models (LFM and LFM2) with a custom hybrid liquid/convolution/attention architecture, optimized for both data-center and fully on-device deployment across CPUs, GPUs, and NPUs.[1][3][4][9][11] Public materials emphasize hardware-in-the-loop training and edge/PC deployment, but do not expose a full production cloud stack, suggesting a mix of self-hosted/model-serving infrastructure plus OEM/partner integrations rather than a single public hyperscaler.[9][11]
AI/LLM · Multi-Cloud · Native (own / self-hosted weights)
Mistral runs a multi-cloud architecture where its hosted La Plateforme and Studio offerings sit atop partner clouds (Google Cloud, AWS, Azure, SAP, IBM, Snowflake, NVIDIA, Outscale), while open-weight models are also deployable on‑prem via standard inference containers.[10][11] The production stack is therefore split between first‑party EU‑centric hosting for regulated workloads and distribution of models through major cloud marketplaces and managed services.
AI/LLM · Hybrid · Native (own / self-hosted weights)
OpenAI runs large GPU superclusters and Kubernetes-based orchestration across Azure and its own data centers in a hybrid setup, with Azure as the primary cloud and custom HPC clusters (e.g., Stargate) for training and serving frontier models.[15][16][19] Core application and API workloads use relational stores like PostgreSQL for accounts/settings and globally scalable databases such as Azure Cosmos DB plus Kafka streams for high-volume conversation, analytics, and event data.[1][8][12]
AI/LLM · Multi-Cloud · Native (own / self-hosted weights)
Reka trains and serves its own multimodal encoder–decoder frontier models (Core, Flash, Edge) on custom Kubernetes-based GPU clusters spanning multiple vendors, using PyTorch on large H100/A100 fleets and a separate A10/A100 inference stack.[2][5] Public deployment is via Reka’s own web app and API endpoints (chat.reka.ai, platform.reka.ai, showcase.reka.ai), with an OpenAI-compatible API server for Edge provided through vLLM and Hugging Face artifacts.[2][7][12]
AI/LLM · On-Prem · Native (own / self-hosted weights)
The provided sources do not expose xAI’s actual production infrastructure, and most search results are generic XAI architecture guidance rather than xAI company disclosures. The only xAI-specific material in the results points to an internal product architecture (Home Mixer, Thunder, Phoenix, Grox) but does not establish cloud, database, or compliance stack details for xAI’s core production environment.[9][5]