Skip to content
TechFabric

Finance said ninety days. Operations said open account.

Andrew Ripley8 min read

The board pack had a Genie screenshot on slide four. Revenue by active customer, asked in plain English, answered in a table that looked like it had always existed. The CFO asked the same question in the meeting. The number was different.

Finance had been using ninety days of transactions since the last audit. Operations had been using "open account" since the billing system was rewritten. Both definitions were in the semantic layer. Genie retrieved one of them. Nobody had decided which.

Most of the arguments I walk into come down to a definition. Two teams mean different things by the same word, both are right, and the system is stuck between them. A better prompt has never once fixed that.

The rollout fails before anyone admits it

Genie demos beautifully. Then it contradicts finance, and the people who were supposed to use it start opening the old dashboard again without filing a ticket. They just stop asking.

By the time someone calls it a failed rollout, the damage is social. An executive who has been embarrassed once in front of their peers will not ask a second time. You can retrain the space. Retraining will not unteach that feeling.

We used to treat this as a retrieval problem: go look at the tables, tune the instructions, add a few more example questions. Sitting with both teams, getting a decision made, and encoding the decision so a machine can score it is what actually moves the number.

Two to four weeks, then a number

That engagement is Genie Accuracy. We write metric and definition templates for the domains that matter first, harden the space and the semantic layer, sit with the people who already know the right answer, write the questions they actually ask, and record what the answer should be. That set becomes a regression suite.

You also get a before-and-after accuracy report on your own questions, and a named owner who can run the playbook after we leave.

Fabric Experiments is what keeps that suite alive. A change to a definition should fail a test before it reaches a board meeting. Native MLflow runs stay in Databricks. The gate is a policy checkpoint, and a failed evaluation blocks promotion.

Preetham wrote a longer version of the scoring argument in Nobody could tell whether the agent was any good. This is the Genie-shaped case of the same failure.

Sometimes the answer is a dashboard

Some questions should never have gone to natural language. A regulated number that has to match a filing, a definition that changes with the legal entity, a metric that three committees have already fought over: those want a curated tile and a caption.

We will say that after two weeks if it is true. Telling you Genie is the wrong tool is cheaper than a rollout nobody trusts. The suite still helps, because the dashboard has to be right too. The difference is who is allowed to type the question.

What we need from you

The people who already argue about the number. I sit in those sessions for exactly this reason, and I will not take the briefing from a data engineer whose job was to "set up Genie." An engineer who has never watched operations on a Tuesday does not know that Tuesday's number is always wrong because of when the batch runs, and will encode that error into the ground truth.

If executives have already gone quiet, start at Genie Accuracy. If you are earlier, and the job is standing the lakehouse up with definitions from day one, that is the Launchpad.