◆ Gentoo Logic · Modeling Warehouse

AI/LLM · AI architecture

12 AI/LLM companies — how they build with AI

How AI/LLM companies build with AI — which foundational models they run and how directly, ranked by proximity to the model. Avg Foundation Proximity Score 100/100.

Native 12Direct API 0Cloud-hosted 0Middleware 0

AI21 Labs

AI/LLM · Multi-Cloud · Native (own / self-hosted weights)

AI21 Labs develops and serves its own Jamba family of LLMs and related services via its AI21 Studio API and private deployments, while also distributing Jamba models through third‑party clouds such as Azure and upcoming NVIDIA APIs.[5][6][7] Public materials describe deployment options across public cloud and private/on‑prem environments but do not reveal a full production stack beyond this high‑level multi‑cloud posture.[3][5][6]

Aleph Alpha

AI/LLM · Hybrid · Native (own / self-hosted weights)

Aleph Alpha’s current product layer appears to be PhariaAI, a sovereign enterprise stack that includes knowledge capture, development, operation, access control, and monitoring components. Public materials also show first-party API access and containerized deployment patterns, but the exact production datastore and hosting mix are not fully disclosed, so the infrastructure is best characterized as hybrid rather than purely on-prem or purely cloud.[2][5][10]

Anthropic

AI/LLM · Multi-Cloud · Native (own / self-hosted weights)

Anthropic runs **Claude** and related services on a safety-first **multi-cloud** compute fabric spanning AWS Trainium2, Google TPUv7 and NVIDIA GPUs, fronted by Kubernetes‑based microservices (API gateways, orchestration, rate limiting, caching, and safety filters), with state held in PostgreSQL, vector stores, Redis, and cloud object storage.[1][10][13] Production offerings like Claude API and Managed Agents expose a fully managed orchestration and agent runtime, while emerging self‑hosted sandboxes move tool execution into customer infrastructure but keep Claude inference, routing, and session state on Anthropic’s cloud.[2][3][11]

Cohere

AI/LLM · Multi-Cloud · Native (own / self-hosted weights)

Enterprise-focused LLM lab training its own Command/Embed/Rerank models, served across clouds (Google Cloud, Oracle, AWS) and deployable in-VPC for data-sensitive customers. North platform targets RAG.

Contextual AI

AI/LLM · Multi-Cloud · Native (own / self-hosted weights)

RAG-native enterprise platform training its own grounded language models.

Hugging Face

AI/LLM · AWS · Native (own / self-hosted weights)

The model/dataset hub on AWS, storing weights as Git LFS repos over S3 and running Inference Endpoints/Spaces. Stewards open libraries (Transformers, Diffusers) — the open-source center of gravity for ML.

Inflection AI

AI/LLM · Azure · Native (own / self-hosted weights)

Trains its own Inflection models (now enterprise-focused) on Azure supercompute.

Liquid AI

AI/LLM · Hybrid · Native (own / self-hosted weights)

Liquid AI builds and serves its own Liquid Foundation Models (LFM and LFM2) with a custom hybrid liquid/convolution/attention architecture, optimized for both data-center and fully on-device deployment across CPUs, GPUs, and NPUs.[1][3][4][9][11] Public materials emphasize hardware-in-the-loop training and edge/PC deployment, but do not expose a full production cloud stack, suggesting a mix of self-hosted/model-serving infrastructure plus OEM/partner integrations rather than a single public hyperscaler.[9][11]

Mistral AI

AI/LLM · Multi-Cloud · Native (own / self-hosted weights)

Mistral runs a multi-cloud architecture where its hosted La Plateforme and Studio offerings sit atop partner clouds (Google Cloud, AWS, Azure, SAP, IBM, Snowflake, NVIDIA, Outscale), while open-weight models are also deployable on‑prem via standard inference containers.[10][11] The production stack is therefore split between first‑party EU‑centric hosting for regulated workloads and distribution of models through major cloud marketplaces and managed services.

OpenAI

AI/LLM · Hybrid · Native (own / self-hosted weights)

OpenAI runs large GPU superclusters and Kubernetes-based orchestration across Azure and its own data centers in a hybrid setup, with Azure as the primary cloud and custom HPC clusters (e.g., Stargate) for training and serving frontier models.[15][16][19] Core application and API workloads use relational stores like PostgreSQL for accounts/settings and globally scalable databases such as Azure Cosmos DB plus Kafka streams for high-volume conversation, analytics, and event data.[1][8][12]

Reka AI

AI/LLM · Multi-Cloud · Native (own / self-hosted weights)

Reka trains and serves its own multimodal encoder–decoder frontier models (Core, Flash, Edge) on custom Kubernetes-based GPU clusters spanning multiple vendors, using PyTorch on large H100/A100 fleets and a separate A10/A100 inference stack.[2][5] Public deployment is via Reka’s own web app and API endpoints (chat.reka.ai, platform.reka.ai, showcase.reka.ai), with an OpenAI-compatible API server for Edge provided through vLLM and Hugging Face artifacts.[2][7][12]

xAI

AI/LLM · On-Prem · Native (own / self-hosted weights)

The provided sources do not expose xAI’s actual production infrastructure, and most search results are generic XAI architecture guidance rather than xAI company disclosures. The only xAI-specific material in the results points to an internal product architecture (Home Mixer, Thunder, Phoenix, Grox) but does not establish cloud, database, or compliance stack details for xAI’s core production environment.[9][5]

← Back to the full orbit map (all 260 companies) · Score your own stack →