Data Platform · Multi-Cloud
Databricks
Databricks runs a two-layer architecture where a Databricks-managed control plane hosts the UI, APIs, metadata, and orchestration services, while customer workloads execute in a compute/data plane inside the customer’s AWS, Azure, or GCP account (or Databricks’ serverless account) against cloud object storage using Spark and Delta Lake.[3][4][9][11] Around this core, Databricks positions an open-core lakehouse stack (Delta Lake, Spark, MLflow) with proprietary governance, serverless, and AI platform services, integrated into medallion-style production patterns.[1][2][6][12]
Foundation proximity: Native (own / self-hosted weights)
Foundational models & how they consume them
PrimaryDBRX· Databricks (self-hosted in Databricks control/compute planes) · Native (own / self-hosted weights)
SecondaryLlama family (e.g., Llama 2 / Llama 3)· Self-hosted open models on Databricks clusters or serverless model serving · Native (own / self-hosted weights)
TertiaryOther open-source foundation models (e.g., MPT, Falcon, etc.) used via Databricks Model Serving· Self-hosted on Databricks or customer cloud compute · Native (own / self-hosted weights)
TertiaryProprietary frontier models (e.g., OpenAI GPT family) when customers integrate via external APIs from notebooks/jobs· Third-party API providers · Middleware / wrapper
Cloud & datastore
Multi-CloudDelta Lake (transactional lakehouse tables on cloud object storage)[6][10]Apache Spark (distributed compute engine for ETL/ML/streaming)[11][12]Unity Catalog (centralized governance and metadata layer)[1][5]Cloud object storage (Amazon S3, Azure Data Lake Storage, Google Cloud Storage) as primary data store[6][4][12]Medallion (Bronze/Silver/Gold) data modeling pattern on top of Delta tables[2][12]Compliance
Separation of control plane (Databricks-managed backend services) and compute/data plane (customer cloud account) for isolation and governance[3][4][7][11]Network isolation via VPC/VNet injection, CIDR planning, and private connectivity for data plane resources[1][5][7]Workspace- and catalog-level access control with Unity Catalog (role-based access, row/column-level policies)[1][5][10]Encryption of data in transit and at rest via underlying cloud provider primitives (KMS/Key Vault, TLS)[5][10]Governance features including lineage tracking, masking, and auditability built into lakehouse stack[1][10]Similar companies
SnowflakeConfluentOthers that build on proprietary / self-built models
Abnormal SecurityAbridgeAdobeAI21 LabsAleph AlphaAmbienceExplore
FAQ
Does Databricks use a foundational AI model?
Databricks uses DBRX, Llama family (e.g., Llama 2 / Llama 3), Other open-source foundation models (e.g., MPT, Falcon, etc.) used via Databricks Model Serving, Proprietary frontier models (e.g., OpenAI GPT family) when customers integrate via external APIs from notebooks/jobs (Native (own / self-hosted weights)).
What cloud does Databricks run on?
Databricks runs on Multi-Cloud.
What database does Databricks use?
Databricks uses Delta Lake (transactional lakehouse tables on cloud object storage)[6][10], Apache Spark (distributed compute engine for ETL/ML/streaming)[11][12], Unity Catalog (centralized governance and metadata layer)[1][5], Cloud object storage (Amazon S3, Azure Data Lake Storage, Google Cloud Storage) as primary data store[6][4][12], Medallion (Bronze/Silver/Gold) data modeling pattern on top of Delta tables[2][12].
How close does Databricks run to the foundational model?
Databricks is Native (own / self-hosted weights) — close to the metal, so cheaper tokens and more control over outputs.
Sources
databricks.com ↗Also on the map
← Full orbit map (260 companies) · Score your own stack →
Architecture inferred from public sources · confidence low · verify before betting on a detail.