Skip to content
TechFabric

Fix the foundation

Make the estate answer for itself

Grants that hold, lineage that survives a refactor, and an evaluation suite that says whether the thing got better. So the audit is a query.

You can see the data. You cannot explain it.

The platform works. The questions that stall a programme are the ones about it rather than in it.

Two dashboards disagree

Both are computed from real tables and neither is obviously wrong. Deciding which to believe takes a person a week and the answer does not stick.

The service principal sees everything

It was granted broadly during a deadline in 2023 and nobody has revisited it. What an application can reach is now wider than what any of its users can.

Nobody can score the agent

It answers. Whether it answers correctly, and whether last month's version was better, is a matter of opinion.

Governance gets retrofitted because it is invisible until the quarter it is not, and retrofitting it onto a live estate is most of what makes it expensive.

The shift

Three ways to end up defensible

Two of them produce a document rather than a property.

Write the policy down

True on paper

Policy written
Team briefed
Estate grows
Drift
Nobody knows

Audit it periodically

Accurate once a quarter

Audit runs
Findings raised
Some fixed
Drift resumes
Audit again

Make it a property of the platform

True continuously

Grants through groups
Lineage survives refactor
Evaluation on every change
Gate blocks a regression
Audit is a query

What we put in place

The governance and the evaluation together, because a governed estate producing unscored answers is only half defensible.

Catalogue and schema design

Structured around the organisation rather than the storage layout, so a grant model can follow how teams actually work.

Grants through groups

Joining a team is what changes access. Individual grants are how an estate becomes unexplainable one exception at a time.

Hive metastore migration

In stages, with external tables where a full move costs more than it returns. Sequencing rather than whether.

Lineage that survives

Traceable through the transformations you actually run, so the question of where a number came from has an answer.

An evaluation suite

Ground truth you own, running on every change, so a regression is caught by a gate rather than by a user.

The written record

What was decided and why, which is the thing the next platform lead needs and almost never gets.

A query

Rather than a reconstruction

The bar is whether it holds when asked

Every engagement here ends with a query somebody can run in front of an auditor rather than a document somebody wrote about the estate. That is a harder thing to deliver and it is the only version worth paying for.

FAQ

Explainable estate: common questions

We already have Unity Catalog. Is this still relevant?

Usually more so. Most rollouts stop halfway: the catalogues exist, half the estate is still external tables against the old metastore, and there is a service principal everybody has stopped asking about.

Can you do the governance without the evaluation?

Yes, and it is a common start. They are separated here because they answer different questions, and most estates need the first before the second is worth doing.

How do we know it stayed true?

Because the checks run rather than being scheduled. A grant model expressed through groups and an evaluation suite in the pipeline are both continuously true or continuously failing, which is the property a periodic audit cannot give you.