# Lakebase database branching: a database per coding agent

> Databricks Lakebase branches a Postgres database in under a second. Here is how that gives every coding agent and every PR its own isolated database.

Published: 2026-10-09
Author: Ivan Vydrin
Tags: Databricks, AI/ML, Knowledge Share
Canonical: https://www.techfabric.com/blog/lakebase-database-branching-coding-agents

---

Databricks published a workflow on 8 October 2026 for giving every coding agent its own database: [Lakebase copy-on-write branching](https://www.databricks.com/blog/lakebase-and-agentic-sdlc-branching-databases-coding-agents), which creates a full Postgres branch in under a second regardless of database size, wired into Git worktrees and GitHub Actions. It matters because the shared dev database is the thing that breaks first when three agents work in parallel, and it's the one piece of the agentic toolchain that skills, hooks and MCP servers don't fix.

Most teams running coding agents have already solved code isolation. Git worktrees give each agent its own directory with its own branch checked out, so nothing collides at the file level. The database is still one box that everyone writes to. Databricks is blunt about what happens next. Agents conflict on schema changes, they interfere with each other's data, or they quietly fall back to mocks that don't reflect real-world data. Those were developer problems before. Agents just make them happen faster and at higher concurrency.

## What copy-on-write branching actually does

A Lakebase branch inherits the schema and data of its parent but shares the underlying storage through pointers. Only modified data gets written separately. Two consequences follow from the [Databricks branching documentation](https://docs.databricks.com/aws/en/oltp/projects/branches), and both are the point of the whole thing.

Branches appear instantly, and the size of your database has no impact on branch creation time. A 100GB database branches as fast as a 100MB one. Creating a branch has no performance impact on the production workload, because nothing is copied.

Storage billing follows the same logic, with one condition attached. An expiring branch is billed only for the data that changes on it. Modify 1GB on a branch of a 100GB database and you pay for roughly 1GB. A permanent branch, one with no expiration, is billed for its full data size. If your CI creates permanent branches by default, you have built a copy machine with extra steps.

Compute is separate and scales to zero when idle. A branch that an agent touched for twenty minutes this morning costs nothing this afternoon.

```mermaid
%% caption: Each agent gets its own worktree and its own Lakebase branch; CI creates an ephemeral branch per pull request from production.
flowchart TD
  A["Shared dev database"]
  B["Git worktree per agent"]
  C["Lakebase branch per agent"]
  D["Pull request opened"]
  E["Ephemeral branch from production"]
  F["Migration run and preview app"]
  G["Merge then delete branch"]
  A --> C
  B --> C
  C --> D
  D --> E
  E --> F
  F --> G
  class C accent
  class A legacy
```

## The loop, with the syntax

The Databricks workflow has two halves. Locally, a post-checkout hook in the repository fires whenever Git creates a worktree, and the hook creates a Lakebase branch. The agent runs `claude -worktree feature-123`, Git makes the worktree, the hook makes the database, and the agent now has its own code directory and its own Postgres. Agent behaviour around all this is guided through repository instruction files, `AGENTS.md` or `CLAUDE.md`.

In CI, GitHub Actions creates an ephemeral branch as a child of production for each pull request, runs the migration tool against it, deploys a preview app on Databricks Apps, and posts a schema diff on the PR. Because the branch starts from production, the migration gets tested against production schema before it reaches production.

The CLI call underneath is one command. From the [postgres command group reference](https://docs.databricks.com/aws/en/dev-tools/cli/reference/postgres-commands):

```bash
databricks postgres create-branch projects/my-project-id pr-123 \
  --json '{
    "spec": {
      "source_branch": "projects/my-project-id/branches/production",
      "no_expiry": true
    }
  }'
```

For a PR branch you want the opposite of `no_expiry`. Expiration is mandatory on creation. You set `ttl`, `expire_time`, or `no_expiry: true`, and there is no fourth option. The Python SDK form, from the [branch management docs](https://docs.databricks.com/aws/en/oltp/projects/manage-branches):

```python
# Illustration of the documented SDK call, not a tested snippet
from databricks.sdk import WorkspaceClient
from databricks.sdk.service.postgres import Branch, BranchSpec, Duration

w = WorkspaceClient()

branch_spec = BranchSpec(
    ttl=Duration(seconds=604800),  # 7 days
    source_branch="projects/my-project/branches/production",
)
result = w.postgres.create_branch(
    parent="projects/my-project",
    branch=Branch(spec=branch_spec),
    branch_id="development",
).wait()
```

Databricks suggests expiration durations by use. CI/CD pipelines get 2 to 4 hours, demos 24 to 48 hours, feature development 1 to 7 days, and long-term testing 30 days. Thirty days is also the hard ceiling; you extend by updating the timestamp before it fires. Reset a branch from its parent and the TTL countdown restarts from the original interval.

## Branches don't merge, and that changes your migration discipline

This is the part that catches teams who reason by analogy with Git. Lakebase branches aren't merged back into the parent. Parent and child both change independently, so reconciling their data gets impractical fast. Branch reset only runs one direction, parent to child.

So schema changes live in code, next to the application logic, and get promoted through migrations. Databricks names Drizzle, Flyway, Liquibase and Alembic; the example repository uses Drizzle. The deployment automation applies the migration when it deploys the preview app, and applies it again when the change merges to main.

If your team currently makes schema changes by connecting to the dev database and typing SQL, branching will not help you. You need the migration file to be the artifact. That's a prerequisite more than a side effect, and it's worth sorting out before you wire up any hooks.

## What it requires, and the limits that will bite

You need CAN MANAGE permission on the project to create, delete or update branches, and CAN USE to list them or create databases and roles inside one. That split is the right shape for CI, because the pipeline's service principal needs CAN MANAGE, and it will be creating and destroying real infrastructure on every PR.

The [project limits](https://docs.databricks.com/aws/en/oltp/create/) are generous but finite. 500 branches per project. 20 concurrently active computes, with the default branch exempt so production stays reachable. 6 read replicas per branch, 500 databases and 500 roles per branch, 30 days maximum history retention, and scale-to-zero configurable between 60 seconds and 7 days. The concurrent compute limit is the one that an enthusiastic CI pipeline will find. Computes with scale to zero enabled suspend themselves after inactivity, which is what keeps you under it.

Some expiration restrictions are structural, and advisory language doesn't cover them. You cannot expire a protected branch, cannot expire a default branch, and cannot expire a branch that has children or create children from an expiring branch. That last one shapes your hierarchy. If PR branches hang off production and production is protected, you are fine, but a two-level ephemeral scheme will not work.

Branch deletion is permanent. All associated data and compute go with it.

Pricing is usage-based and billed in [Lakebase capacity unit hours](https://www.databricks.com/product/pricing/lakebase), with a 14-day free trial and committed-use discounts available. Endpoints take an autoscaling range, so a CI branch can be pinned small:

```bash
databricks postgres create-endpoint \
  projects/my-project-id/branches/pr-123 primary \
  --json '{
    "spec": {
      "endpoint_type": "ENDPOINT_TYPE_READ_WRITE",
      "autoscaling_limit_min_cu": 0.5,
      "autoscaling_limit_max_cu": 4.0
    }
  }'
```

Lakebase is generally available on AWS and Azure and in Beta on GCP as of 15 June, across twelve AWS regions including us-east-1, eu-west-1 and ap-southeast-2. Your project is created in your workspace region.

## Where this fits beyond the agent loop

Three other branching workflows come out of the same mechanism, and for an operations or finance team they are the more immediately useful ones.

Point-in-time branching. You can create a branch from any point inside your restore window. Databricks gives the example of a critical table dropped yesterday at 10:23 AM: branch at 10:22 AM and extract the missing rows, with the original branch untouched and still serving traffic. The same move gets you a database as it stood on a specific date for a financial reconciliation, a regulatory audit or a forensic review. That's a different thing from a backup restore, because nothing goes offline.

Schema migration testing against real production shape, before the migration runs on production.

Production-derived data with Unity Catalog masking, which is how you let an agent work against realistic data without handing it PII. Databricks notes that most teams run one workspace per environment for security and compliance, and branch from a seeded database instead of production for exactly this reason. Their walkthrough uses a single workspace for simplicity; the concepts carry over.

## Where the fit gets thin

If you have one developer and no agents, branching solves a problem you don't have. The value scales with concurrency.

If your application's state lives mostly outside Postgres, in object storage, a message queue, a third-party SaaS, branching the database gives you a partial environment and a false sense of isolation. Your tests will pass against data that doesn't match what the other systems hold.

And this is a Lakebase feature, not a Postgres feature. Lakebase runs open source Postgres with no fork and no new dialect, so your existing drivers, pgAdmin, DBeaver and psql all work. But the branching comes from the [Lakebase architecture](https://www.databricks.com/product/lakebase) of separated compute and object storage. An RDS instance will not do this. Moving an operational database to get agent isolation is a large decision to make for a development workflow; make it because you want transactional and analytical data in one governed copy, and take the branching as the thing that improves your week.

## What I would tell a team starting here

Start with CI before you start with the agents. An ephemeral branch per pull request is a contained change with a clear payoff, because reviewers validate against a real database and migrations get tested against production schema before they touch production. You will learn the permission model and the expiration rules on something that fails safely.

Set TTLs from the first branch you create, and audit for permanent branches monthly. The billing difference between expiring and permanent is the difference between paying for your diffs and paying for full copies, and nobody notices the drift until the invoice arrives.

Get your migrations into code before you give an agent a database. Branching rewards teams whose schema changes are already files in a repository and punishes teams whose schema lives in someone's psql history.

For how Lakebase branching compares with the alternatives in this category, see [Lakebase vs Snowflake Postgres](/blog/lakebase-vs-snowflake-postgres-branching). For the agent side of this, how teams scope what an agent can touch and prove it afterwards, we write about that at [/ai/agents](/ai/agents).

## Frequently asked questions

### How fast is a Lakebase branch, and does database size matter?

Branch creation is sub-second and the size of your database has no impact on creation time, because copy-on-write shares the parent's storage through pointers instead of copying bytes. Creating a branch also has no performance impact on the production workload, which is what makes per-PR branching viable in CI.

### What does a Lakebase branch cost?

A Lakebase branch costs whatever its compute and storage come to, billed separately and usage-based. You pay only for active compute hours and computes scale to zero when idle, so an occasionally used branch costs far less than a dedicated dev server running 24/7. On storage, an expiring branch is billed only for data that changes on it, while a permanent branch with no expiration is billed for its full data size, so set a TTL on anything ephemeral.

### Can you merge a Lakebase branch back into production?

No. Lakebase branches are not merged back to the parent, because parent and child both change independently and reconciling their data becomes impractical; branch reset works in one direction only, parent to child. Schema changes are tracked in code and promoted to the parent through migration tools such as Drizzle, Flyway, Liquibase or Alembic.

### How many branches can one project have?

500 branches per project, with a limit of 20 concurrently active computes (the default branch is exempt so it stays always available), 6 read replicas per branch, and 500 databases and 500 roles per branch. Contact Databricks Support if you need the concurrent compute limit raised.
