Expertise
A governed lakehouse your executives can question
A data platform is only useful if people can ask it a question and believe the answer. We stand Databricks up with Unity Catalog first, then ingestion, then the layers and the Genie space. The six-week Lakehouse Launchpad is the named build.
80
Databricks-certified engineers on the bench today, which proves the floor rather than the ceiling. Fifteen years of average production experience is the number that decides whether the thing still works in month six.
- 27
- source platforms with migration playbooks that execute, from Snowflake and Synapse to Teradata and SAP
- 4
- engagements with a fixed shape and a fixed fee, so you know the cost before the first meeting
- 10 to 3
- the delivery team size Fabric took us to, on the same Databricks stack we build yours on
Ingestion
Where most platforms acquire the debt they never pay off. A source connected badly in month one is still being worked around in year three, usually by a job nobody wants to touch.
- Lakeflow Connect for the sources it covers, and a written reason when we build instead
- Change data capture off operational databases without asking the DBA for a nightly window
- Files, APIs and queues landed once, with the schema drift handled rather than discovered
- Backfill and replay treated as normal operations rather than as incidents
Modelling
Bronze, silver and gold is a naming convention rather than a design. What decides whether the platform is usable is which grain the gold layer settles on and whether two teams can agree on it.
- Grain and conformed dimensions agreed with the people who will be questioned about the numbers
- Slowly changing history where the business genuinely reconstructs the past, and nowhere else
- Incremental processing that does not rebuild the estate to add a day
- Tests on the tables that finance reads, running on every load
Semantic layer and Genie
A Genie space is only as good as the definitions under it. Ours are written down, versioned and scored against ground truth, because the failure mode is not an error message. It is a plausible answer that disagrees with the close.
- Metric definitions held in the semantic layer rather than in each dashboard
- Synonyms and business vocabulary captured from how people actually ask
- An evaluation suite with expected answers, run on a schedule
- A named owner for every metric, so a disagreement has somewhere to go
Lakebase
Serverless Postgres inside the lakehouse, which is what lets an application share the governed data rather than copy it. This is the part where a decade of building applications matters more than a decade of building pipelines.
- Transactional state next to the analytical estate, with the sync direction stated
- Branching for environments, so a test does not run against production rows
- Unity Catalog permissions carried through instead of re-implemented in the app
- Applications and agents reading from it directly, under their own service principal
Cost and operations
The bill is an engineering output. Warehouse sizing, cluster policy, table layout and job scheduling account for most of it, and all four are decisions somebody made rather than facts about the platform.
- Serverless where the workload is spiky, classic where it is steady, with the arithmetic shown
- Liquid clustering and file sizing on the tables that are actually scanned
- Budget policies and tagging, so a spike has an owner within the hour
- Alerting on freshness and reconciliation, not only on job failure
What we build
- Unity Catalog structure before the first pipeline
- Ingestion through Lakeflow or Jobs, whichever suits the source
- Bronze, silver and gold layers for the domains you name at the start
- A Genie space with definitions you can score, not a demo that dies in week three
Accelerators doing the work.
We stopped rebuilding the same few parts of a Databricks programme. These are what we bring: adopted in your workspace, under your Unity Catalog.
How this is delivered
Migrations to Databricks
Off Snowflake, Synapse, Teradata and SQL Server, onto Lakehouse and Lakebase, with a cutover you can reverse.
FAQ
Questions we get asked
How do you stand a Databricks lakehouse up?
The Lakehouse Launchpad at /databricks/launchpad is the six-week build: identity and environment strategy, Unity Catalog, ingestion, medallion layers, dashboards, and a Genie space with real semantic definitions. Scope is fixed at the start, and in practice the thing that moves a date is how long it takes to get access.
We are still on a warehouse. Do we start with a build?
Not if the scope is still an argument. The Migration Readiness Sprint at /databricks/migration-readiness inventories the estate, agrees the exclusions, and hands you a wave plan. TechFabric Airlift at /accelerators/airlift is the migration factory that sprint runs on.
What stops Genie contradicting finance?
Definitions that live in the semantic layer, and a suite that scores answers against ground truth. TechFabric Experiments at /accelerators/experiments keeps that suite running. When a rollout has already lost the room, Genie Accuracy at /databricks/genie-accuracy is the two-to-four week engagement.
Do we have to rebuild a lakehouse that already works?
No. The Launchpad is for standing something up. If you already have a workspace, the Databricks Health Check at /databricks/health-check assesses what is there before anyone proposes rebuilding it.
How can we help?
A technical conversation with a senior engineer. If a two-week readiness sprint is the honest first step, we will say so.