# Lakebase vs Snowflake Postgres: branching is the difference

> Why Databricks Lakebase beats Snowflake Postgres for teams already on Databricks. Git style branching, scale to zero, one governance model.

Published: 2026-10-08
Author: Preetham Reddy
Tags: Databricks, Knowledge Share, Thought Leadership
Canonical: https://www.techfabric.com/blog/lakebase-vs-snowflake-postgres-branching

---

A developer on one of our product teams asked for a copy of production to reproduce a bug. On most stacks that request starts a conversation about disk, about how stale the snapshot can be, about who pays for the extra instance and when we tear it down. On Lakebase it was a branch, and he had it before the conversation finished.

That is the whole argument in one sentence, and the rest of this post is why it holds up against the obvious alternative. We run Lakebase on several products now. It is the operational database under applications we build for clients, and the branching model has changed both how we develop and how we handle production incidents. Snowflake shipped its own managed Postgres in the same window. People keep asking me which one to pick. If your platform is Databricks, the answer isn't close, and the reason is architectural, which is a deeper thing than a feature checklist.

## Two different bets on what Postgres should be

Snowflake bought Crunchy Data in June 2025 and turned it into Snowflake Postgres. Crunchy was a serious Postgres shop, the acquisition made sense, and the product is a competent managed Postgres. Read Snowflake's own documentation for what it actually is: "Each instance runs a Postgres database server on a dedicated virtual machine managed by Snowflake" with "attached disks to deliver best-in-class transactional performance" ([Snowflake docs](https://docs.snowflake.com/en/user-guide/snowflake-postgres/about)). A dedicated VM with attached disks, PgBouncer in front, private networking, Postgres 16 through 18.

That is a good description of open source Postgres on a VM with the operational work taken off your plate. Patching, backups, failover and connection pooling are real value, and if you are lifting and shifting an existing application into Snowflake it will do the job without code changes. What it doesn't do is change the shape of the thing. Compute and storage are still bound together on one machine. You pick an instance size, you pay for it while it idles, and a heavy query still competes with your live traffic for the same CPU.

Databricks took the other bet. Lakebase separates the compute layer from the storage layer and puts the data in cloud object storage, which is the definition Databricks uses for the [lakebase category](https://www.databricks.com/blog/what-is-a-lakebase) itself. Everything interesting downstream follows from that one decision. You can't branch a database in seconds if the data lives on a disk bolted to a VM. You cannot scale to zero if the engine and the data share a machine. Snowflake didn't ship a worse version of branching; branching isn't available to the architecture they chose.

```mermaid
%% caption: Dedicated-VM Postgres binds compute to its disks. Lakebase separates compute from lake storage, which is what makes instant branches possible.
flowchart TD
  A["Managed Postgres on a dedicated VM"] --> B["Compute bound to attached disks"]
  B --> C["Fixed size and no instant branches"]
  D["Lakebase serverless Postgres"] --> E["Compute separated from lake storage"]
  E --> F["Copy on write branches and scale to zero"]
  class A legacy
  class F accent
```

## What branching actually does to a working week

The Databricks docs describe branching as working "similar to branching your code in Git" ([Branches](https://docs.databricks.com/aws/en/oltp/projects/branches)), and the analogy survives contact with real work, which is more than most analogies manage.

A project starts with a single `production` branch. You create children off it: a development branch, a staging branch, or one branch per developer if you want complete isolation. Each branch has its own compute. Changes on a child never touch the parent. The isolation goes further than data, which matters more than it sounds. Roles created, GRANTs and REVOKEs applied, and role attributes changed on one branch have no effect on any other branch. So a developer testing a permissions change is testing it properly, on clean ground, without a shared staging database where somebody else's grant is still lying around from last Tuesday.

Branches appear instantly because of copy-on-write. The branch inherits the parent's schema and data through pointers to the same storage, and only writes new data when you modify something. The docs state two consequences plainly. Branch creation time doesn't depend on the size of your database, and creating a branch has no performance impact on production.

The second one is the part teams underestimate. Cloning a production database on a conventional stack is an event. You schedule it, you watch the source, you do it at night. Here it is a command, and it costs production nothing, so people do it casually, which is exactly the behaviour you want.

Branch reset is the other half. When your development branch has drifted and you want today's data, you reset it from the parent and it updates in place, instantly, with your connection string unchanged. No seed scripts, no teardown, no overnight refresh job that somebody has to own.

```sql
-- illustrative: what a developer touches after a branch is created.
-- The connection string is the only thing that changes.
psql "postgres://...@ep-dev-branch.../appdb"

-- run the migration against production-shaped data
ALTER TABLE quotes ADD COLUMN decision_reason text;

-- wrong? drop the branch and make another one. Production never knew.
```

## Troubleshooting production without touching production

This is where branching earns its place beyond developer convenience, and it is the part I would sell hardest to anyone on the fence.

Something goes wrong in production at a known time. A job wrote bad values, or an operator deleted rows, or a release shipped a migration that corrupted a column. The usual response is forensics against the live database under pressure, with a lock on a large table one careless query away from making the incident worse.

Lakebase lets you create a branch from a point in time inside your restore window. The docs give the exact case. If a table was dropped yesterday at 10:23 AM, you create a branch set to 10:22 AM and pull the missing data out of it. That branch is a new root, fully functional, and your production branch keeps serving traffic the entire time, untouched. You can compare before and after side by side. You can run the expensive diagnostic query against the historical branch with no risk, because nothing you do there can reach production.

Point-in-time recovery at this granularity isn't unique; most managed Postgres offerings will restore you to a timestamp. What is different is that recovery here produces a branch you can keep and query alongside the live system, instead of a restore operation you perform on the database. One is an investigation tool. The other is a last resort.

The same mechanism does observability work. Create a branch at the state before a release, run the new query plan on both, compare. That is a cheap experiment on Lakebase and a project on a VM.

## Serverless is a cost argument and a behaviour argument

Lakebase computes autoscale and scale to zero when idle, and the GA notes confirm new instances get autoscaling, scale-to-zero, branching and instant restore by default ([release notes](https://docs.databricks.com/aws/en/release-notes/lakebase/)). You pay for active compute hours. Storage is billed separately, and for an expiring branch you pay only for what changes: modify 1GB on a branch of a 100GB database and you pay for roughly 1GB, nowhere near a second 100GB copy.

Do the arithmetic on a normal engineering org. Development, staging, QA, a branch per feature, a branch per CI run. On dedicated instances each of those is a machine sized for peak and running at 3am on a Sunday for nobody. The Databricks docs put it in one line: a development branch you use occasionally costs far less than running a dedicated development server around the clock.

The behaviour change is the bigger prize. When a non-production environment costs almost nothing while idle and nothing to create, engineers stop rationing them. They stop sharing one staging database and stepping on each other. They stop leaving a feature branch's schema change in staging for three weeks because tearing it down is work. Cheap environments get used properly; expensive ones get hoarded and abused.

One caveat worth knowing before you plan around it. A permanent branch, one with no expiration, is billed differently from an expiring one. Set expirations on throwaway branches. That is the whole discipline.

## The part that only matters if you are already on Databricks

If your analytics, your governance and your models already live in Databricks, Lakebase is no second platform. Access control and auditing run through Unity Catalog, the same model as the rest of your estate, so an application's operational data is governed exactly like its analytical data. Sync tables keep operational and lakehouse data in step without you maintaining a pipeline for it. GA brought Postgres 17 support alongside 16, pgvector, automated backups, point-in-time recovery, and up to 8TB per instance ([GA announcement](https://www.databricks.com/blog/databricks-lakebase-generally-available)).

That integration is the thing a feature comparison misses. On a Databricks stack, the alternative to Lakebase is a Postgres somewhere else plus the pipelines, the second set of credentials, the second audit trail and the drift between them. We have built that before. It works until the day somebody asks which copy of a row is authoritative.

Snowflake Postgres makes exactly the same argument on the Snowflake side of the fence, and for a Snowflake shop it is a reasonable one. The question is which platform you are consolidating onto, and then whether the database you add to it was designed for the cloud or ported into it. If you are weighing that decision more broadly, we wrote about the lock-in half of it in [the lock-in argument for leaving Snowflake](/blog/the-lock-in-argument-for-leaving-snowflake-is-over).

## What I would tell a team starting on Lakebase

Design the branch hierarchy on day one, before anyone connects an application. Production sits as the protected root, with a long-lived development branch under it and short-lived branches carrying expirations under that. Protected branches cannot be deleted or reset from their parent and they block project deletion, which is the guardrail you want on the one branch that matters.

Make branch creation part of the pipeline instead of something a person does. A branch per pull request, a branch per CI run, deleted when the work finishes. If creating an environment stays manual, people will keep sharing one.

Know what doesn't come with a branch. Logical replication slots and subscriptions aren't inherited, so if the branch needs to publish or subscribe, set that up again. Find that out in a design session, well before a failover.

And pick the restore window deliberately. Point-in-time branching only reaches as far back as your retention, and the branch you wish you had created is always one from before you noticed.

If you want to talk through what this looks like on your stack, that is the work we do on [Databricks](/databricks).
