Skip to content

Unity Catalog for AI agents in practice: four FHIR scenarios

Sushruth Aeluguri14 min read

Picture a care coordinator asking an internal agent for a patient summary. The agent answers, and the summary includes a substance use record covered by 42 CFR Part 2, which that coordinator is not cleared to see.

No bug was involved. The row filter on the Condition table worked. The column mask on SSN worked. Both evaluated correctly against the identity that ran the query, and that identity was the agent's service principal, which had been granted broad read access so it could serve every coordinator.

Agent identity decides what your row filters do. The four scenarios below show how that plays out in a lakehouse that lands CCD documents (a C-CDA clinical summary of a patient's care, exchanged between providers) and FHIR bundles, then normalizes them into FHIR R4 tables for Patient, Encounter, Condition, Observation, MedicationRequest and Coverage. Two users with identical job titles must get different answers from the same agent.

The names below are examples; adapt them to your catalog.

Care coordinator

Agent

Unity Gateway

Model services

MCP services and UC functions

FHIR tables in Unity Catalog

Traces and system tables

Unity Catalog governs the objects, Unity Gateway governs the runtime interactions, and telemetry records both.

Scenario one: a coordinator chats with the agent

This is the on-behalf-of-user case.

Start with tags, because attribute-based access control decides access by evaluating governed tags on securable objects, with policies attached at catalog, schema or table level and evaluated dynamically (ABAC overview, GA; metastore-level policies are Beta). The taxonomy is the real design work. Tag PHI columns, and carry a sensitivity column on the FHIR tables derived from each resource's meta.security labels, which is where a Part 2 record is marked when the bundle arrives. Tag that column too, so a policy's table condition and the filter logic both have something to match on.

ALTER TABLE health.fhir.patient
  ALTER COLUMN ssn SET TAGS ('phi' = 'true', 'pii_type' = 'ssn');

ALTER TABLE health.fhir.condition
  ALTER COLUMN sensitivity SET TAGS ('phi' = 'true', 'sensitivity_label' = 'true');

Row filter and column mask policies are GA. A column mask passes the value through a UDF that returns the original or a redacted version, and the return type must be castable to the column's type (policies). Creating one needs MANAGE on the securable or ownership, EXECUTE on the UDF, and Databricks Runtime 16.4 or above or serverless compute.

CREATE FUNCTION health.gov.mask_ssn(ssn STRING)
RETURNS STRING
RETURN CASE
  WHEN is_account_group_member('phi_full_access') THEN ssn
  ELSE 'XXX-XX-' || right(ssn, 4)
END;

Attach that as a column mask policy scoped to the schema, with a table condition matching the phi tag, so new FHIR tables inherit the rule as they land. Creating the first policy in Catalog Explorer shows every field (name, principals, exemptions, scope, table condition, policy type); after that, SQL or Terraform is faster.

The Part 2 row filter is the same shape with a boolean return. Rows where it returns FALSE disappear from results. Filter on the sensitivity column, since FHIR Condition.category carries values like problem-list-item or encounter-diagnosis and says nothing about Part 2 status.

CREATE FUNCTION health.gov.part2_visible(sensitivity STRING)
RETURNS BOOLEAN
RETURN is_account_group_member('part2_consented')
   OR (sensitivity IS NOT NULL AND sensitivity = 'unrestricted');

That filter is deliberately closed. A row whose sensitivity is NULL is hidden from anyone outside the consented group, because a missing label on a source that carries meta.security means the normalization job failed, and the safe reading of a failed label is that the row might be Part 2. The cost is that a mapping gap looks like missing data to a coordinator, which is the failure you want, since it gets reported.

That UDF is only as good as the column it reads. The normalization job that writes sensitivity from meta.security is part of the access control, and treating it as a convenience field is how the whole thing breaks quietly.

Then the tool. Make it a Unity Catalog function with a narrow signature instead of a general SQL escape hatch, and grant EXECUTE the way you grant data access.

CREATE FUNCTION health.agent_tools.get_active_conditions(patient_id STRING)
RETURNS TABLE (code STRING, display STRING, onset_date DATE)
RETURN SELECT code, display, onset_date
       FROM health.fhir.condition
       WHERE subject_id = patient_id
         AND clinical_status = 'active';

GRANT EXECUTE ON FUNCTION health.agent_tools.get_active_conditions
  TO `care_coordinators`;

Notice what the function body doesn't contain. No consent logic at all. The row filter handles that, which means the tool author cannot forget it and cannot route around it.

How the agent reaches that function matters as much as the grant. Unity Catalog functions used as agent tools are governed as custom tools under Unity Gateway, with the same privileges you use for data (Unity Gateway). For a Python agent, the documented path for calling your registered functions is UCFunctionToolkit, which Databricks lists as the recommended replacement for the old workspace functions MCP endpoint (Databricks-provided MCPs). If the agent needs to run SQL instead of a fixed function, system.ai.dbsql runs it on a warehouse using the caller's Unity Catalog and warehouse permissions, and it is reached at https://<workspace-hostname>/ai-gateway/mcp-services/system.ai.dbsql. That MCP is behind the Unity Gateway beta, which an account admin enables from the account console Previews page.

The two-identity test

Call the same tool twice against the same patient. Once through an agent running on behalf of a coordinator who is not in part2_consented. Once through the agent's service principal, which holds broad read access.

The service principal call succeeds. No error, and no empty result either. It returns the rows the service principal is allowed to see, which includes Part 2 conditions and unmasked SSNs, and the agent hands them to a user who should never have received them. There is nothing to alert on, no stack trace, no permission denial in a log.

So my recommendation is to run interactive agents on behalf of the user by default. The reasoning is mechanical. Row filters, column masks and ABAC policies all evaluate against the identity running the query, so on-behalf-of-user means the agent sees exactly what the user sees, with no second access model to keep in sync, and every read in the audit trail is attributed to a real person. Databricks says the same thing. If per-user access control or user-attributed auditing is required, use on-behalf-of-user (agent authentication).

The setup is small and easy to miss. The feature is disabled by default and a workspace admin turns it on, you need MLflow 2.22.1 or above, and you declare the REST API scopes the agent needs when you log it. Databricks recommends Databricks Apps when you need broader OAuth scopes.

Coordinator asks for a chart summary

Agent runs on behalf of user

Unity Gateway checks service policies

UC function tool executes

Row filter and column mask evaluate for the coordinator

Filtered rows return to the agent

Summary and trace attributed to the coordinator

One agent call on behalf of the user, with the policy check running against the coordinator's identity.

Scenario two: the nightly prior-auth batch

A scheduled agent that drafts prior authorization packets overnight has no end user to impersonate. This one runs as a service principal, and the question becomes how narrow the grant is.

The tempting shortcut is GRANT SELECT ON CATALOG health to the job's principal, because then nothing ever fails at 2am. What you get instead is a principal that sees every Part 2 record in the lakehouse, produces output nobody can attribute to a person, and sits outside every consent group you built in scenario one. If that principal's token leaks, or if someone later points an interactive agent at the same identity because it already works, the blast radius is the catalog.

Scope it to the schema or the specific tables the job reads, and prefer giving the principal EXECUTE on the same UC functions a human would call. If the batch genuinely must see Part 2 rows, put the principal in part2_consented deliberately and write down why, so the row filter stays the single place that decision lives. A service principal that is exempt because it was never subject to the policy is a different thing from one you consciously admitted.

For models and MCP services the batch calls, ABAC GRANT policies dynamically grant privileges to objects whose governed tags match a condition, so you stop writing one GRANT per model (GRANT policies). GRANT-policy support is GA for models, model services, model provider services and MCP services, and Beta for agent services and skills, which need an account admin to switch on the upgraded Unity AI Gateway tier from the account console Previews page. SQL management of GRANT policies requires classic compute on Databricks Runtime 18 LTS or above.

CREATE POLICY grant_production_model_access
ON SCHEMA health.models
COMMENT 'Grant EXECUTE on production MLflow models'
TO `batch_agents`
GRANT EXECUTE FOR MODELS
WHEN has_tag_value('lifecycle', 'production');

One behavior to internalize before you rely on this. Effective privilege is the union of direct grants and GRANT policies. A selective policy takes nothing away from a principal who already holds a direct grant on the model, its schema or its catalog. Inventory the direct grants first, or your policy is a second opinion instead of the control.

Scenario three: RAG over chart notes

You want the agent to retrieve from unstructured clinical notes, so you reach for a vector search index. Then you hit the limit: you cannot create a vector search index from a table that carries ABAC row filters or column masks (policies). The protection you spent scenario one building is exactly what blocks the index.

There are three designs, and each one costs something.

Index a de-identified source table. Run the notes through de-identification during normalization, write the safe version to a separate table with no filters, and index that. Retrieval becomes a single index for everybody, which is operationally the simplest thing you will ever run. The cost is recall and fidelity. De-identification removes the names, dates and identifiers that sometimes carry the clinical signal, and you now own a de-identification process whose failures are silent and whose output is permanently less useful than the source.

Separate indexes per access tier. One index over non-sensitive notes, one over the Part 2 subset, and the agent picks based on the caller's group membership. Permissions live on the index, so enforcement is coarse and easy to reason about. The cost is duplication: storage, sync pipelines, and drift between tiers, plus the fact that tiers are a fixed taxonomy. Two tiers is maintainable. The day someone asks for per-patient consent, this design stops scaling.

Filter after retrieval. Index everything, then join retrieved chunk IDs back to a governed table that carries the row filter and drop anything the caller cannot see. Enforcement stays in one place. The cost is that the embeddings themselves were computed over data the caller cannot read, your retrieval quality degrades unpredictably as results get dropped, and you need a hard guarantee that no chunk text leaves the retrieval step before the filter runs. Pick it only with the post-filter inside the tool function, never in the agent prompt.

Pick based on how your consent model actually works. Fixed tiers point at separate indexes; per-patient consent points at de-identification plus a governed structured lookup for anything sensitive.

Scenario four: sending PHI to an external model

The summarization model is hosted by an external provider. Unity Gateway is where that call becomes governable. It routes requests from a central control plane, applies rate limits, attaches service policies as guardrails on requests and responses, sets per-user budget thresholds and hard spend caps that stop requests when budgets are exceeded, and monitors usage through system tables (Unity Gateway). Contextual service policies are Beta. Older material calls this AI Gateway; the current name is Unity Gateway.

Two controls matter here. Service policies govern how each request and response proceeds based on its content and on who is making the call, which is where PII detection and prompt-injection guardrails attach. And because model provider services are Unity Catalog securables with governed tags, a GRANT policy can express the approval rule directly: EXECUTE only on models whose provider or country-of-origin tag is on your approved list, which means a newly added model is denied by default instead of silently available.

A content guardrail is a classifier, so it has a false negative rate, and PHI in a free-text clinical note does not look like a credit card number. Policies govern the call, not the provider's retention or training terms, which are a contract question. And the gateway sees what the agent sends, so if your tool already returned unmasked data because of a scenario-one identity mistake, the guardrail is deciding whether to forward data that should never have been retrieved.

If an external provider is in the path, set the hard spend cap and not only a budget alert. An agent in a retry loop against a frontier model is a cost incident that resolves itself at the cap and does not resolve itself at the alert.

For what happened after the fact, Unity Gateway observability tracks requests, token usage and latency through system tables, attributes cost to services, target models, principals and tags, and can log requests and responses to Unity Catalog Delta tables. Lakewatch, the agentic SIEM Databricks launched on March 24, 2026, is where you investigate an incident across that activity (press release). It is in Private Preview.

One trap: legacy workspace MCP endpoints, including the UC functions server at https://<workspace-hostname>/api/2.0/mcp/functions/{catalog}/{schema}/{function_name} with the unity-catalog OAuth scope, bypass Unity Gateway. Their calls do not appear in gateway tracing tables and they do not support MCP-level grants, policies or guardrails (Databricks-provided MCPs). Reusing an old integration quietly removes the entire runtime layer, which is why scenario one routes the function through the toolkit or a system.ai MCP instead.

Performance impact

Row filter and mask UDFs are part of the query plan and run per row. A simple SQL expression over a column plus is_account_group_member is cheap. A filter that joins a consent mapping table, or calls a Python UDF, is a per-row lookup on every scan of that table by every user, including the wide scans your BI tools run. If you need consent mapping, denormalize the decision into a column during the pipeline and let the filter read that column.

The bigger effect is on data skipping and predicate pushdown. A filter expression that the optimizer cannot push down forces more of the table to be read before rows are dropped, which turns a partition-pruned query into something closer to a full scan. Databricks publishes guidance on exactly this under performance considerations for row filter and column mask policies, covering UDF complexity, predicate pushdown and query optimization (ABAC overview). Read it before you write the filter, while the design is still cheap to change.

On-behalf-of-user changes the shape of caching, because the result of a query is now a function of who asked. Two coordinators in different consent groups running the identical tool call produce different result sets, so anything you cache has to be keyed by identity or not cached at all. Check the current docs for what Databricks caches across users before you assume either way, and measure the difference on your own workload.

The gateway adds a network hop to every model and tool call. For a chat agent making several sequential tool calls per turn, that hop is paid several times per user turn, and it is additive to model latency. Measure it against a direct provider call on your own traffic, in your own region, before deciding it is or isn't acceptable.

One more limit that shapes designs: only one distinct row filter per table and one column mask per column can resolve for a given user, and conflicts throw an error.

Tradeoffs, and when to skip this

This design concentrates governance in one vendor. The policies, the tags, the gateway and the traces are all Databricks, and the abstraction is not portable.

Several pieces are Beta or Preview: metastore-level policies, DENY policies, contextual service policies, GRANT policies for agent services and skills, parts of the managed MCP set, and Lakewatch. Beta features change. Do not build a compliance control whose only implementation is a Beta feature without a fallback.

Debugging interacting policies is the hard part. Tags inherit down the hierarchy, policies attach at several levels, and effective privilege is the union of direct grants and GRANT policies, so "why can this person see this row" becomes a multi-step trace. Keep the taxonomy small enough that a person can hold it in their head.

Unity Catalog doesn't cover everything either. It doesn't stop the model from hallucinating, it doesn't sanitize what a user pastes into a prompt, and it doesn't govern data that has left the platform. A column mask protects a query result; a screenshot is outside its reach.

Freshness sits outside all of it too. A chart summary built from stale vitals is wrong in a way no policy catches, so declare an expectation on Observation freshness in your Lakeflow Declarative Pipelines definition and choose deliberately whether violations warn, drop the row, or fail the update.

A single-user prototype on synthetic data doesn't need any of this. A read-only agent over already public data doesn't need it. The moment the data is regulated or the agent can act on something, the calculus flips.

What to check before anyone uses it

  • Every PHI column carries a governed tag, and the taxonomy is written down somewhere other than the policy definitions.
  • Each sensitive column resolves exactly one mask and each table exactly one row filter per user, because conflicts throw an error.
  • The sensitivity column on each FHIR table is populated from meta.security by the pipeline, and a test fails when it is null on a resource that carried a security label.
  • On-behalf-of-user is enabled, the agent is logged with MLflow 2.22.1 or above, and the declared OAuth scopes match the MCP services it actually calls.
  • The two-identity test runs in CI, asserting that the agent's answer for a given user matches what that user gets querying directly, and failing the build if the agent returns anything extra.
  • Direct grants on models and MCP services are inventoried, so GRANT policies carry the authority.
  • A hard spend cap is set on any external model path, not only a budget alert.
  • An audit query answers "what did this agent read for request X" and returns a person's name.

What I'd tell someone starting this

Decide the identity question first, before the tags, before the tools, before anyone writes an agent. Everything downstream inherits that choice, and changing it later means re-auditing every policy written in between.

Then run the two-identity test on one table and one tool. It takes an afternoon and it tells you whether the governance is real or decorative. For the catalog groundwork that comes before any of this, Unity Catalog before the first pipeline covers the earlier layer, and our Databricks practice page has how we approach the work.