Skip to content

Databricks · Lakebase · 2:20

Test database changes on a full copy of production

Lakebase is the managed Postgres that Databricks runs beside the lakehouse, and its branches give your team a full copy of production instantly, generally available on AWS and Azure. We design the schema, sync it with your lakehouse tables, build the application on top, and wire branching into its releases, so every migration is tested on production data and a bad write can be undone. One of our senior engineers will walk through your last hard release with you.

Illustrative example. Lakebase branching is GA on AWS and Azure. Captions are on by default.

Transcript

  1. 0:00

    What this is about

    More data teams now run the operational apps that use their data, on Postgres inside Databricks. The releases that go wrong are usually the ones that change that database, and this is the workflow we set up on Lakebase so they stop going wrong.

  2. 0:17

    Staging passed, production blocked orders

    Say the app tracks orders. Friday, a developer adds a column and an index to the orders table. It passed on staging, but staging was restored three weeks ago with a fraction of the rows, so nobody saw that the index build would block new orders for six minutes.

  3. 0:33

    Why teams test on stale copies

    Everyone knows the fix is testing on real production data. A full copy takes hours, though, doubles the storage and is behind by the time it's ready, so teams live with the old one.

  4. 0:45

    A full copy, made instantly

    This is where Lakebase, the Postgres inside Databricks, changes things. You branch the database the way you branch code. The branch has all of production, tables and data, and it's ready in an instant, even for a big database, because it shares production's storage until something changes.

  5. 1:04

    Run the migration on a branch

    So on Friday the migration runs on a branch first. A test copy of the app points at it, the column goes in, the index builds, and you see how long that lock holds. Production keeps taking orders throughout.

  6. 1:19

    Ship it through your normal deploy

    When it looks right, you compare the two schemas. Branches don't merge back, so the same migration goes through your normal deploy, and this time the index builds concurrently, so orders keep flowing.

  7. 1:32

    Expiring branches and recovery

    Test branches can expire after a day, and expiring ones are billed only for what changes on them. If a bad write reaches production, you branch from a minute before it and copy the rows back.

  8. 1:45

    What we set up

    That's how we work when we build an app on Lakebase. We design the schema, keep it in sync with your lakehouse tables and build the application on top, and every pull request gets a branch with production-shaped data, so each migration is timed before it ships.

  9. 2:03

    Test database changes on a full copy of production

    If a database change has held up a release on your team, send us the migration, and one of our senior engineers will walk through how a branch would have caught it. We're a Databricks partner, with 80 Databricks-certified engineers.

Talk to someone who has built this.

A technical conversation with a senior engineer about what you are trying to build.