Databricks Lakebase is a fully managed Postgres database that runs inside the Databricks platform, with compute separated from storage, autoscaling, scale to zero and instant branching (Databricks docs, Databricks blog). It matters because the operational database your application writes to can now sit beside the analytical tables your pipelines produce, under one governance model, instead of in a separate estate with its own pipelines moving rows back and forth.
It is generally available on AWS, and on Azure since 2 March 2026 (Azure release notes). Databricks' GA announcement describes the lakebase category as an operational database architecture that decouples compute from storage (Databricks blog). The engine is open source Postgres, not a fork and not a new SQL dialect; the change is in the cloud architecture underneath it (Databricks product page).
What the architecture actually gives you
Data lives in cloud object storage in open formats, and a serverless Postgres engine runs elastically on top. Databricks calls this LTAP, Lake Transactional/Analytical Processing, where transactions and analytics read from one governed copy (Databricks product page).
The object model is worth learning before you touch anything, because every limit and every permission hangs off it. A project is the top-level container. Inside a project are branches; each project starts with a root branch called production that can't be deleted. Each branch has one primary read/write compute, optional read replicas, its own Postgres roles, and its own databases. The default branch arrives with a database called databricks_postgres (Manage projects).
Project
└── Branches (production, development, staging, ...)
├── Computes (R/W compute, read replicas)
├── Roles (Postgres roles)
└── Databases (Postgres databases)
The published project limits are specific enough to design against. You get 500 branches per project, 500 roles and 500 databases per branch, 20 concurrently active computes, 6 read replicas per branch, 3 root branches, 30 days maximum history retention, and a scale to zero time between 60 seconds and 7 days (Manage projects). The default branch is exempt from the active compute cap, so production stays reachable when somebody spins up fifteen test branches at once.
Postgres 16, 17 and 18 are supported, with 17 as the default and 18 selectable at project creation (Postgres compatibility). Extensions include pgvector and PostGIS, and standard clients work, whether that is psql, pgAdmin, DBeaver or any Postgres driver (Databricks product page).
Branching and restore are the same mechanism
A branch is an independent database environment that shares storage with its parent through copy-on-write. Create one from the latest state of the parent, or from a point in time inside the restore window (Manage branches). Branches can expire: set a TTL and a background process deletes the branch when it reaches the timestamp, which is how you stop CI environments accumulating. Databricks suggests 2 to 4 hours for CI/CD pipelines and 1 to 7 days for feature development, with 30 days the maximum. Billing follows the same idea, an expiring branch is billed only for the data that changes on it, a permanent branch for its full data size (Manage branches).
Here is the branch creation call from the Python SDK, as the documentation shows it:
from databricks.sdk import WorkspaceClient
from databricks.sdk.service.postgres import Branch, BranchSpec, Duration
w = WorkspaceClient()
branch_spec = BranchSpec(
ttl=Duration(seconds=604800), # 7 days
source_branch="projects/my-project/branches/production",
)
result = w.postgres.create_branch(
parent="projects/my-project",
branch=Branch(spec=branch_spec),
branch_id="development",
).wait()
Point-in-time restore uses the same primitive, and this is the part teams misread. A restore doesn't modify your branch in place. It creates a new root branch holding the data as it existed at the chosen moment, while the original branch keeps running and existing connections keep working. To use the restored data you update your application's connection configuration to point at the new branch (Point-in-time restore). Two consequences follow. First, you can validate a restore before anyone depends on it, which is the right order of operations. Second, the three root branch limit means you may have to delete an old restore before you can take a new one, so clean up as part of the incident while it is still fresh.
The restore window is configurable from 2 to 30 days, default 7, and it applies to every branch in the project. A longer window costs more storage (Point-in-time restore). Restores also apply to all Postgres databases on the branch, beyond the one you were troubleshooting.
A database branch per coding agent is a use of this that deserves its own treatment, and we gave it one in Lakebase database branching: a database per coding agent.
Getting lakehouse data into Postgres, and Postgres changes back out
Synced tables are the first direction. A Unity Catalog table, managed or external Delta, managed or external Iceberg, or a view or materialized view, syncs into Lakebase as a read-only Postgres table that applications query with low-latency lookups. The Unity Catalog schema name becomes the Postgres schema name and the documentation's examples give the Postgres table a _synced suffix, a convention worth keeping so nobody mistakes it for a table the app owns (Synced tables).
The quickstart shows the payoff in one query. Build a gold table in Unity Catalog, sync it, then join it in Postgres against the table your application writes to:
SELECT p.id, p.name, p.value, s.segment, s.engagement_score
FROM playing_with_lakebase p
JOIN "default".user_segments_synced s ON p.id = s.user_id;
That is from the Databricks quickstart, including the quoting of default, which is a reserved keyword in Postgres and must be quoted every time if that is your schema name (Serve lakehouse data).
Pick the sync mode deliberately, because it's the main cost lever. Snapshot copies everything each run and is roughly 10x more efficient when more than 10% of the source rows change per cycle. Triggered runs on demand or on a schedule and propagates inserts, updates and deletes each refresh, but gets expensive below 5 minute intervals. Continuous streams with seconds of lag and a 15 second minimum interval, lowest lag and highest cost (Synced tables). Triggered and Continuous need a change data feed on the source, so either enable write-time CDF on the Delta table or use automatic change data feed, which computes row-level changes at read time and lets sources such as Iceberg tables sync incrementally. Sync pipelines run as managed Lakeflow pipelines and each sync uses up to 16 connections; Lakebase supports up to 1,000 concurrent connections with transactional guarantees.
There is also an accelerator. LTAP Direct Writes is a beta capability that writes bulk loads straight into the storage layer behind your branch instead of routing them through the live compute endpoint. It is off by default, you opt in per synced table in the Create synced table dialog, it needs Postgres 16, 17 or 18, and you can't add it to an existing synced table: delete and recreate. It accelerates the initial load in every mode and every subsequent sync in Snapshot mode (Synced tables).
The other direction is Lakebase Change Data Feed, in Public Preview. Every insert, update and delete on a Lakebase Postgres table is captured from the write-ahead log and written as a new row in a Unity Catalog managed Delta table, batched and flushed about every 15 seconds (Lakebase Change Data Feed). No external CDC stack. If your operational data starts in an ERP or CRM instead of in the app, that is a different problem with different tooling, and we cover it under data integration.
Governance, identity and who does what
Registering a Lakebase database in Unity Catalog creates a read-only catalog representing the Postgres database, which brings permissions, lineage and audit logs to the operational data alongside the lakehouse data, and lets you query both from one SQL interface through a serverless SQL warehouse (Register a Lakebase database in Unity Catalog).
Authentication has two shapes. Native Postgres password roles still exist, but projects created since 21 May 2026 have password authentication off by default, so turn it on in the project settings if a client needs it (release notes). OAuth roles map Databricks identities to Postgres roles, so users, service principals and groups connect with OAuth tokens:
CREATE EXTENSION IF NOT EXISTS databricks_auth;
SELECT databricks_create_role('8c01cfb1-62c9-4a09-88a8-e195f4b01b08', 'SERVICE_PRINCIPAL');
The function creates the role with LOGIN only, so grant privileges afterwards. Group roles are powerful here: any direct or indirect member of the Databricks group can authenticate as the group role with their own token, which means you manage membership in the workspace instead of maintaining per-user grants in Postgres (Create Postgres roles). Group names are case-sensitive and must match the workspace exactly.
Four roles show up around a Lakebase project. Application developers connect Databricks Apps or an external service with a standard driver and own the schema their app writes to. Data engineers own the synced tables and their modes, and the CDF feeds going the other way. Platform engineers own projects, branch policy, expiration defaults, restore windows and the project permission model, with CAN MANAGE for creating and deleting branches and CAN USE for creating databases and roles inside one (Manage branches). AI engineers use it as agent state: Databricks offers managed agent sessions and managed agent memory, both backed by Lakebase and usable from agents built on any framework (Agent memory). That pattern is close to the work described on our AI agents page.
When to keep your existing database
Keep your existing operational database when your region isn't on the list. Lakebase projects are created in your workspace region, and the supported set is twelve AWS regions including us-east-1, eu-west-2 and ap-southeast-2 (Manage projects). Azure has its own region list, generally available since 2 March 2026 (Azure release notes), so check the one for your cloud.
Keep it when your workload sits outside the limits. Each branch has a storage quota, adjustable on request, and when a database hits it write performance drops until you reclaim space; only actual tables and indexes count toward it, and the retained history doesn't (Manage projects).
Keep it when the application needs something a managed service won't give you, such as operating system access on the host (PostgreSQL compatibility). And keep it when the database has no relationship to your lakehouse at all. The argument for Lakebase is proximity to governed analytical data, so a standalone Postgres behind a product that never joins to a gold table gains branching and scale to zero, and little else.
Finally, synced tables are read-only by design. Databricks recommends running only read queries against them to protect integrity with the source (Synced tables). If your application needs to write to the same rows the lakehouse owns, model that explicitly, with app-owned tables for writes and synced tables for reads, joined in one query.
What to do first
Start with one project, one application, and one gold table you already trust. Sync it in Snapshot mode, join it to an app table, and measure the query latency your users will actually see. Then decide whether Triggered at 15 minutes is good enough before you reach for Continuous, because the cost difference is real and most dashboards don't need seconds.
Set the restore window and the branch expiration defaults on day one. Both are project-wide settings that get awkward to change once people depend on them, and branch deletion is permanent with no recovery.
If you are weighing it against the Postgres offering from another warehouse vendor, the branching model is where they diverge, which we worked through in Lakebase vs Snowflake Postgres. Pricing is usage-based with a 14-day free trial (Lakebase pricing), and snapshot storage became billable on 1 June 2026 (release notes), so put snapshot retention in the same budget review as the restore window.
Frequently asked questions
Is Databricks Lakebase generally available?
Lakebase is generally available on AWS, and on Azure since 2 March 2026, with GA covering autoscaling, scale to zero, instant branching, automated backups and point-in-time recovery (Databricks blog, Azure release notes). Lakebase Change Data Feed, which stores Postgres changes as Delta tables, is in Public Preview (Lakebase Change Data Feed).
Is Lakebase real Postgres or a Databricks dialect?
It runs the open source Postgres engine, not a fork and not a new SQL dialect, with existing clients, libraries and extensions such as pgvector and PostGIS (Databricks product page). Versions 16, 17 and 18 are supported, with 17 the default (Postgres compatibility).
How do I get Delta tables into Lakebase Postgres?
Create a synced table from a Unity Catalog Delta or Iceberg table, view or materialized view, and it appears in Postgres as a read-only table in a schema named after the Unity Catalog schema (Synced tables). Choose Snapshot, Triggered or Continuous depending on how fresh the data needs to be and what you are willing to pay.
What does a point-in-time restore do to my running database?
A point-in-time restore leaves your running database untouched. The restore creates a new root branch containing the data from the chosen moment, the original branch keeps operating, and existing connections stay live until you repoint your application at the new branch (Point-in-time restore).