# Genie Accuracy

> Genie demos beautifully and then contradicts finance, because the same word means different things in different domains and nothing tests the semantic layer when it changes. We harden the definitions and put an evaluation harness behind them, so an answer can be checked rather than believed.

For: Teams whose natural-language analytics lost the room
Length: Two to four weeks
Canonical: https://www.techfabric.com/databricks/genie-accuracy

---

When executives stop trusting Genie, the rollout has already failed

## What we inspect

- The same metric resolves differently across two domains, and both look right
- Definitions live in people's heads rather than in the semantic layer
- Nothing regression-tests the semantic layer when someone changes it
- Teams quietly stop using it instead of reporting that it is wrong

## What you get

- Metric and definition templates for the domains that matter first
- A hardened Genie space and semantic layer
- An evaluation harness scoring answers against ground truth
- A before-and-after accuracy report on your own questions
- An operating playbook, and a named owner who can run it

## Questions

### How do you measure accuracy without a ground truth?

Building the ground truth is part of the work. We sit with the people who already know the right answer, write the questions they actually ask, and record what the answer should be. That set becomes the regression suite, and it is yours.

### Is this a Databricks problem or our problem?

Almost always the semantic layer rather than the model. Genie answers the question it was asked against the definitions it was given; when two teams define active customer differently, both answers are correct and one of them is wrong to the person reading it.

### What if the answer is that Genie is the wrong tool for us?

That is a legitimate outcome and we will say it. Some questions want a curated dashboard, not natural language. Telling you that after two weeks is cheaper than a rollout nobody trusts.

