Skip to content
TechFabric

Expertise

A governed lakehouse your executives can question

A data platform is only useful if people can ask it a question and believe the answer. We stand Databricks up with Unity Catalog first, then ingestion, then the layers and the Genie space. The six-week Lakehouse Launchpad is the named build.

80

Databricks-certified engineers on the bench today, which proves the floor rather than the ceiling. Fifteen years of average production experience is the number that decides whether the thing still works in month six.

27
source platforms with migration playbooks that execute, from Snowflake and Synapse to Teradata and SAP
4
engagements with a fixed shape and a fixed fee, so you know the cost before the first meeting
10 to 3
the delivery team size Fabric took us to, on the same Databricks stack we build yours on

Ingestion

Where most platforms acquire the debt they never pay off. A source connected badly in month one is still being worked around in year three, usually by a job nobody wants to touch.

  • Lakeflow Connect for the sources it covers, and a written reason when we build instead
  • Change data capture off operational databases without asking the DBA for a nightly window
  • Files, APIs and queues landed once, with the schema drift handled rather than discovered
  • Backfill and replay treated as normal operations rather than as incidents

Modelling

Bronze, silver and gold is a naming convention rather than a design. What decides whether the platform is usable is which grain the gold layer settles on and whether two teams can agree on it.

  • Grain and conformed dimensions agreed with the people who will be questioned about the numbers
  • Slowly changing history where the business genuinely reconstructs the past, and nowhere else
  • Incremental processing that does not rebuild the estate to add a day
  • Tests on the tables that finance reads, running on every load

Semantic layer and Genie

A Genie space is only as good as the definitions under it. Ours are written down, versioned and scored against ground truth, because the failure mode is not an error message. It is a plausible answer that disagrees with the close.

  • Metric definitions held in the semantic layer rather than in each dashboard
  • Synonyms and business vocabulary captured from how people actually ask
  • An evaluation suite with expected answers, run on a schedule
  • A named owner for every metric, so a disagreement has somewhere to go

Lakebase

Serverless Postgres inside the lakehouse, which is what lets an application share the governed data rather than copy it. This is the part where a decade of building applications matters more than a decade of building pipelines.

  • Transactional state next to the analytical estate, with the sync direction stated
  • Branching for environments, so a test does not run against production rows
  • Unity Catalog permissions carried through instead of re-implemented in the app
  • Applications and agents reading from it directly, under their own service principal

Cost and operations

The bill is an engineering output. Warehouse sizing, cluster policy, table layout and job scheduling account for most of it, and all four are decisions somebody made rather than facts about the platform.

  • Serverless where the workload is spiky, classic where it is steady, with the arithmetic shown
  • Liquid clustering and file sizing on the tables that are actually scanned
  • Budget policies and tagging, so a spike has an owner within the hour
  • Alerting on freshness and reconciliation, not only on job failure

What we build

  • Unity Catalog structure before the first pipeline
  • Ingestion through Lakeflow or Jobs, whichever suits the source
  • Bronze, silver and gold layers for the domains you name at the start
  • A Genie space with definitions you can score, not a demo that dies in week three
Built already

Accelerators doing the work.

We stopped rebuilding the same few parts of a Databricks programme. These are what we bring: adopted in your workspace, under your Unity Catalog.

How this is delivered

Migrations to Databricks

Off Snowflake, Synapse, Teradata and SQL Server, onto Lakehouse and Lakebase, with a cutover you can reverse.

FAQ

Questions we get asked

How do you stand a Databricks lakehouse up?

The Lakehouse Launchpad at /databricks/launchpad is the six-week build: identity and environment strategy, Unity Catalog, ingestion, medallion layers, dashboards, and a Genie space with real semantic definitions. Scope is fixed at the start, and in practice the thing that moves a date is how long it takes to get access.

We are still on a warehouse. Do we start with a build?

Not if the scope is still an argument. The Migration Readiness Sprint at /databricks/migration-readiness inventories the estate, agrees the exclusions, and hands you a wave plan. TechFabric Airlift at /accelerators/airlift is the migration factory that sprint runs on.

What stops Genie contradicting finance?

Definitions that live in the semantic layer, and a suite that scores answers against ground truth. TechFabric Experiments at /accelerators/experiments keeps that suite running. When a rollout has already lost the room, Genie Accuracy at /databricks/genie-accuracy is the two-to-four week engagement.

Do we have to rebuild a lakehouse that already works?

No. The Launchpad is for standing something up. If you already have a workspace, the Databricks Health Check at /databricks/health-check assesses what is there before anyone proposes rebuilding it.

How can we help?

A technical conversation with a senior engineer. If a two-week readiness sprint is the honest first step, we will say so.