Skip to content

Replit on Databricks: governed apps with Lakebase Postgres

Andrey Taran8 min read

Replit Agent can now build an application from a plain-language prompt, read live data from your Databricks SQL warehouse, and deploy to Databricks Apps with a managed Lakebase Postgres database provisioned behind it. Databricks announced general availability of the integration with native Lakebase support and published a walkthrough of the setup, which means the awkward part of internal app building, where the app's own data has to live somewhere outside the governed platform, has a supported answer.

That awkward part is the whole story. Reading from a lakehouse was never the hard bit. An operations manager who wants a franchise performance tracker, or a finance analyst who wants an invoice intake queue, needs somewhere to store approvals, comments, submissions and workflow state. Before this, that meant an external Postgres or MySQL instance sitting outside the Databricks perimeter, with its own permissions model, its own compliance work and its own on-call. Databricks' own comparison in the announcement is blunt about it. An external database means a separate model to maintain, manual provisioning and security, and separate compliance work, while Lakebase sits inside the perimeter, uses Unity Catalog at runtime, and is auto-provisioned by Replit Agent on deploy.

What the four steps actually look like

The setup takes three roles, and Databricks puts it at roughly fifteen minutes.

A Databricks admin creates a service principal in the account console, generates an OAuth client ID and secret, assigns the principal to the target workspace, and grants it the permissions the app needs. The Databricks documentation for connecting Replit is specific on two points people get wrong. The secret is shown only once, so save it when it appears, and the service principal needs Can Manage on the SQL warehouse the app will query. You also need the warehouse's server hostname and HTTP path.

Then a Replit organisation admin adds a Databricks (Service Principal) connector in the org's connector settings, pastes in the client ID and secret, and optionally sets RBAC so only certain Replit users can use that connector. A builder on Replit Enterprise activates the connector in their project, supplies the HTTP path and hostname, and starts prompting.

The grant side is ordinary Unity Catalog work. Illustrative only, since the exact objects depend on what the app reads:

GRANT USAGE ON CATALOG ops TO `<service-principal-application-id>`;
GRANT USAGE ON SCHEMA ops.franchise TO `<service-principal-application-id>`;
GRANT SELECT ON TABLE ops.franchise.daily_sales TO `<service-principal-application-id>`;

There is a second path. Databricks says Replit also supports a user-to-machine (U2M) connector, where each builder connects with their own Databricks identity and permissions instead of a shared service identity. For anything touching data with row-level or franchise-level sensitivity, take U2M. A shared service principal is a single identity whose grants are the ceiling for everyone building with it, and reasoning about who can see what through a shared identity is the kind of work this integration is meant to remove.

Plain language prompt in Replit

Replit Agent builds front end

Deployed as a Databricks App

Reads governed tables via SQL warehouse

Writes app state to Lakebase Postgres

Unity Catalog grants and audit

External Postgres outside the perimeter

Where each part of a Replit-built app lives once it is deployed to Databricks Apps.

Lakebase is the part that changes the decision

Lakebase Postgres is managed PostgreSQL running inside your Databricks workspace, co-located with workspace data and services. The developer documentation describes what that buys you. There's no VPC peering and no cross-cloud credential management, you get autoscaling within a configured min/max range, and the database scales to zero when idle, resuming on the next query. The inactivity timeout defaults to 24 hours and can be set anywhere from 60 seconds to 7 days.

Set that timeout deliberately. An internal approvals app used during business hours will happily idle overnight, and a short timeout turns idle compute into no compute. The default is generous because it protects latency, and your bill is your own problem.

That maps neatly onto Replit's Automated Preview Deploys, which give teams a separate test environment to review changes without touching production. Pair the two and a schema change from an agent gets tried somewhere survivable. Supervised AI Migration ensures that any database schema change proposed by Replit Agent requires team approval before it goes into production. Turn it on. An agent that can alter a production schema unsupervised is the failure mode everyone will point at afterwards.

Lakebase is for what the app writes: user state, sessions, conversation history, transactional records, logs. Databricks is explicit that read-only queries against large datasets belong in Unity Catalog instead of Lakebase Postgres. If someone's first prompt is "copy the warehouse into the app database so it's faster", that is the wrong shape.

Where it fits in operations work

The three patterns Databricks describes are all recognisable. A performance app for an operations leader, reading live governed data from a warehouse, with each user's token passed through so a franchise owner sees only their own numbers. A document-processing app where an analyst pushes invoices, forms and contracts through Databricks AI functions to turn PDFs and images into structured fields, then queries the results stored in Lakebase. A workflow app that replaces spreadsheets and email threads with an approval engine, reading analytics from Databricks and capturing new operational data in Lakebase as people use it.

What these have in common is that none of them is a dashboard. A dashboard shows you the number; these capture a decision. That is why the transactional database matters more than the front end, and why the integration is interesting at all.

If the data those apps need lives in an ERP or CRM and hasn't landed in the lakehouse yet, that pipeline is still your problem and comes first. Business Central, for one, keeps its data in its own database and reaches Databricks through a built pipeline, either the standard API pages read incrementally with a lastModifiedDateTime filter or tables exported by the open-source bc2adls extension. We cover those routes in data integration. No agentic builder shortcuts that step.

What it requires and what it costs

Three requirements gate this. Your Databricks workspace must be a standard workspace without a compliance security profile. Your organisation must be on the Replit Enterprise plan, which is custom-priced through sales and carries SSO/SAML, single-tenant environments and advanced privacy controls, alongside the SCIM, RBAC and audit logging described on Replit's Enterprise page. And you need a SQL warehouse the service principal can manage.

That compliance security profile exclusion is the sharpest edge. Databricks Apps itself is turned on by default for workspaces with the compliance security profile enabled, so the restriction lands on the Replit path while Apps stays available. Regulated teams running HIPAA or PCI workspaces should check this before scheduling anything.

On cost, you are paying three ways. There's the Replit Enterprise contract, Databricks Apps compute, and Lakebase. Databricks Apps bills per hour of compute while the app is running, based on provisioned capacity; a stopped app preserves its configuration and incurs no cost. Lakebase is serverless and scales to zero. Both have published rates on the Apps and Lakebase pricing pages, with a 14-day free trial.

There is a hard ceiling worth knowing before you invite a department to build: resource limits cap Databricks apps per workspace at 100. Agentic builders produce apps enthusiastically. Decide early who gets to deploy and what happens to an app nobody has opened in a quarter, or you will hit that number with abandoned prototypes.

The limits of what this builds

This is a builder for internal, governed, data-backed applications. It won't give you a customer-facing product, it won't replace a BI layer that hundreds of people read daily, and it won't excuse you from thinking about data models. An agent that explores your schema and writes queries is only as good as the schema it finds, so a catalog of inconsistently named tables with no constraints will produce an app that looks finished and is wrong.

It also doesn't remove the governance conversation; it moves it earlier. The grants on the service principal, the choice between U2M and M2M, and the approval rule on schema migrations are decisions somebody owns before the first prompt. The pitch is that you make them once and reuse them across every app.

Decide U2M or service principal before you enable the connector

Pick the connector mode first, because it is the one choice that is painful to reverse once people have apps in production. If builders should see only their own permitted data, use the user-to-machine connector and let each identity carry its own Unity Catalog grants. If a small platform team is building on behalf of others against a known set of tables, the service principal is simpler, and scope its grants to exactly those schemas.

Then set Supervised AI Migration on, set the Lakebase inactivity timeout to something shorter than a day for internal tools, and give one person the job of retiring apps. If you want help designing the data layer those apps read from, that is what we do.

Frequently asked questions

What is the Replit Databricks integration?

It is a generally available integration that lets Replit Agent build applications from plain-language prompts against live Databricks data and deploy them as Databricks Apps with a Lakebase Postgres database auto-provisioned for the app's own data, according to Databricks. Every deployed app inherits automatic user authentication, secure access controls and Unity Catalog integration because it runs as a Databricks App.

What does it cost?

It requires a Replit Enterprise plan, which is custom-priced through Replit sales, plus Databricks consumption for SQL warehouse compute, Databricks Apps compute and Lakebase. Apps are billed per hour of compute while running and cost nothing when stopped, and Lakebase scales to zero when idle.

Can a Replit-built app write data back?

Yes, to Lakebase. The app reads analytical data from Databricks through a SQL warehouse and writes operational data such as submissions, approvals and state to its Lakebase Postgres database, which sits inside the Databricks perimeter instead of in an external instance.

Does it work in a compliance-profile workspace?

No. Databricks states the workspace must be a standard workspace without a compliance security profile, even though Databricks Apps on its own is enabled by default in compliance-profile workspaces.