# Voice agents on Databricks for auto refinance

> How a compliant outbound voice agent for auto refinance is built on Databricks. Consent gating, feature lookups, the agent runtime and what to measure.

Published: 2026-10-08
Author: Preetham Reddy
Tags: Databricks, Applied AI, Thought Leadership
Canonical: https://www.techfabric.com/blog/voice-agents-databricks-auto-refinance

---

An auto refinance lender buys leads from aggregators and sits on a book of existing loans, and the economics of both come down to how fast somebody calls. A borrower who filled in a form at 9pm on an aggregator site has filled in three others. The lender who reaches them first, with a real payment number instead of a promise to call back, writes the loan. The same lender is also sitting on equity mining data, borrowers whose vehicle is worth more than the payoff or whose credit has improved eighty points since origination, and almost none of those get a call because a human dialer costs too much per attempt to work a list that long.

That is the case for a voice agent. It is also the case for being extremely careful, because an outbound dialing program in the United States is one of the most heavily regulated things a lender can build, and a voice agent is an automated dialer no matter how conversational it sounds.

So the architecture question isn't "which speech model". It's where consent lives, how the agent reads it in under a second, and what evidence exists afterward that the call was allowed.

## The call loop and the data loop are different systems

A voice agent runs a hard real-time loop that takes audio in, transcribes it, runs a model turn and sends speech back out. Humans notice silence at around a second, and the loop has to hold that budget on every turn. Databricks measures LLM latency as time to first token plus time per output token times the number of tokens generated, and the [benchmarking guidance](https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/prov-throughput-run-benchmark) is blunt about the tradeoff, since concurrency raises throughput and raises latency with it, so a low latency use case means sending fewer concurrent requests to an endpoint instead of saturating it. For a dialer that means capacity planning per concurrent call, measured against concurrency instead of daily volume.

Around that loop sits everything that makes the call worth making. Who to call, whether you may, what their payoff is, what the vehicle is worth, what happened the last three times anyone reached them. That is the lakehouse. Keeping the two loops separate is the whole design.

```mermaid
%% caption: The real-time call loop stays outside the lakehouse; Databricks holds the agent, its tools and the record of every call.
flowchart TD
  A["Carrier and speech layer"] --> B["Consent and DNC gate"]
  B --> C["Agent on agent runtime"]
  C --> D["Online feature store on Lakebase"]
  C --> E["Unity Catalog tools"]
  C --> F["MLflow traces and scorers"]
  D --> G["Offline Delta feature tables"]
  F --> G
  class C accent
```

The telephony carrier and the speech-to-text and text-to-speech layer stay outside Databricks. Databricks is where the brain and the record live. The agent loop runs there, along with the tools it may call, the governed data behind those tools, and every trace of what it did.

## Consent is a gate before the dial, not a prompt in the script

Under the TCPA, an artificial or prerecorded voice call to a cell phone needs prior express consent, and for marketing content prior express written consent. A synthesized voice agent is squarely in that category. Two rule changes make the engineering harder than a boolean column.

First, revocation. The FCC's consent revocation rules took effect on 11 April 2025, and under them a consumer can revoke consent in any reasonable manner that clearly expresses a desire not to receive further calls or texts. A keyword counts, a form counts, and so does plain speech. The Commission [granted a limited waiver](https://docs.fcc.gov/public/attachments/DA-25-312A1.pdf) delaying one piece of this, the requirement to treat revocation in response to one type of message as applying to all future robocalls and robotexts from that caller on unrelated matters, which bought callers time to modify their systems. The direction of travel is clear. A "stop calling me" spoken mid-call is a revocation event, and it has to propagate.

Second, the Eleventh Circuit vacated the FCC's one-to-one consent rule in January 2025, which removed a requirement that would have forced per-seller consent on aggregator lead forms. You still have to prove, per lead, who the consumer consented to hear from and what the disclosure said.

So build consent as its own governed table in Unity Catalog, with one row per consent event instead of one row per borrower. You want the source, the timestamp, the exact disclosure text shown and the channel, plus the TrustedForm or Jornaya token if the lead came from an aggregator, with the revocation events in the same stream. Current state is always derived. When a regulator or a plaintiff's lawyer asks why you called someone on 14 March, the answer should be a query.

The gate runs before dial, and the LLM has no say in it. Model it as a deterministic function. Consent present and not revoked, number off the internal DNC and the National DNC, local time at the called party between 8am and 9pm, attempt count within your frequency policy, state rules applied for the states that are stricter than federal. If any check fails the call never happens. A language model is a poor place to decide whether a call is legal.

The agent also has to disclose, in the first seconds, that it's an AI assistant calling on behalf of the lender, and hand off to a human on request. Several states now require the AI disclosure outright. It is the right default everywhere.

## Features the agent needs in milliseconds

Mid-call, the agent needs the payoff balance, the current rate and term, the vehicle year and mileage band, the estimated value, the borrower's state, and the rate tier the pricing engine would quote. Reading those from Delta tables at call time is too slow.

This is what [Databricks Online Feature Stores](https://docs.databricks.com/aws/en/oltp/projects/feature-store) are for. They're powered by Lakebase, and when you create an online store with the Feature Engineering client, Databricks provisions a Lakebase project as the storage backend to give low latency access for real-time inference. Features sync from the offline Unity Catalog tables in one of three publish modes. TRIGGERED does an incremental sync on a schedule or by API, CONTINUOUS streams as new data lands, and SNAPSHOT takes a one-time full copy. Payoff balances and credit refreshes belong on CONTINUOUS or a tight trigger. Vehicle valuation tables can sync nightly.

The payoff is consistency. The offline table that trained the propensity model and the online store the agent reads at 2pm are the same feature definitions, and a model logged with `FeatureEngineeringClient.log_model` gets [automatic feature lookup](https://docs.databricks.com/aws/en/machine-learning/feature-store/automatic-feature-lookup) at serving time, so the endpoint pulls feature values from the online store without the caller passing them in. One fewer place for the agent's view of the borrower to drift from the lender's.

For equity mining, the candidate list itself is a batch job. Borrowers whose loan-to-value crossed a threshold, whose FICO band moved, whose term has enough months left for a refinance to beat the payoff. Score it nightly, push the eligible population through the consent gate, and hand the dialer a list that is already legal to call.

## Building the agent on Agent Bricks

[Agent Bricks](https://developers.databricks.com/docs/agents/overview) is where the agent itself lives. You write the agent loop in code with any framework, LangGraph or the OpenAI Agents SDK or your own, and Databricks provides the production infrastructure around it, including hosting, durable execution, memory, tools, model access, tracing and governance.

The deployment stack has three layers. Your framework runs the loop. `DurableAgentServer`, part of the `databricks_agentkit` library, wraps it in an HTTP server and exposes the [invocation API](https://developers.databricks.com/docs/agents/runtime). The agent runtime hosts that server on Databricks Apps. The minimal handler looks like this, from the Databricks documentation, as an illustration of the shape:

```python
from databricks_agentkit import DurableAgentServer, InvocationContext

app = DurableAgentServer()

@app.invoke
async def invoke(input, context: InvocationContext) -> dict:
    await context.emit({"type": "status", "message": "Looking that up"})
    answer = await run_my_agent(input, session_id=context.session_id)
    return {"answer": answer}
```

Two properties of that server matter more for voice than for chat.

Sessions. Invocations sharing a `session_id` run one at a time, in order. A phone call is a session, turns must not interleave, and that ordering guarantee is free. When you run more than one instance with `--instances`, send the session ID in an `X-Routing-Key` header so every turn of a call reaches the same instance.

Idempotency and recovery. Clients send a UUID with every invocation, and resending the same request returns the existing invocation instead of running the agent again; reusing an ID for a different request returns 409. Register a `@app.recover` handler and when a worker's heartbeats stop, the server starts a replacement attempt within seconds and calls recovery with the original input. The documentation is explicit that a replacement attempt can repeat side effects from the interrupted one, so make external calls idempotent. In a voice agent the side effects are things like logging a promise to pay or sending a disclosure SMS. Give each one an idempotency key derived from the invocation ID.

Give the agent tools instead of table access. A Unity Catalog function for payoff quote, one for rate eligibility, one for scheduling a callback, one for logging a revocation. Each is governed and each is auditable, and the model's job is to decide which to call and how to say the answer, while the loan book stays off limits to model-composed SQL.

## What to govern and what to measure

[Unity Gateway](https://docs.databricks.com/aws/en/ai-gateway/) is the control point for the model side. One place to grant access to models and MCP services, apply rate limits, attach service policies as guardrails on request and response content, and track spend by service and principal. For a regulated dialer, you want the guardrail on the output, so every quoted rate comes from the pricing engine, every APR statement carries its accompanying terms, and approval stays unpromised. Write that as a service policy attached to the model service so it holds regardless of which prompt version is deployed.

MLflow Tracing records each step the agent takes, and traces can be stored and governed in Unity Catalog. For a call this means the retrieved features, the tool calls, the model turns and the final text, tied to the call ID and retained alongside the recording. That trace is the compliance artifact. Build the retention schedule for it at the same time as the agent, well before someone asks about it a year later.

Then score it. MLflow [evaluation](https://docs.databricks.com/aws/en/generative-ai/agent-evaluation/) runs scorers, LLM judges and code-based checks, against traces, from the UI on a handful of traces and then programmatically through `mlflow.genai.evaluate()` on a dataset. For auto refinance the scorers worth writing are specific to the domain. Did the agent make the AI disclosure in the opening turn, did it honour a stop request on the turn it was spoken, did every quoted figure match the tool output, did it hand off when asked. Run the same scorers on a sample of production traffic.

The business metrics are the ones the lender already tracks, measured against a human-dialed control. Contact rate, consent-to-transfer rate, application completion, funded loans per thousand attempts, cost per funded loan. Keep a human control group running. It's the only way to know whether the agent is doing better or just dialing more.

## Where this goes wrong

The first failure mode isn't the model saying something odd. It's the consent table being a snapshot that updates nightly while the agent calls people who revoked at 10am. Treat revocation as a streaming event into the online store with the lowest sync latency you can run, and apply it at the gate on every attempt.

The second failure mode is cost blindness. A voice agent that holds a conversation burns tokens per turn, and at a thousand concurrent calls that compounds fast. Rate limits and budget caps in Unity Gateway, with spend attributed by service tag, are how you find out before the invoice does.

The third is scope. Start with one campaign, one state set, one script, outbound only, with a human transfer on anything the agent hasn't been explicitly built to handle. Equity mining on your own book beats aggregator leads as a first campaign, because you already have the loan data, the consent chain is yours instead of a third party's, and the population is large enough that even a small contact-rate lift pays for the build.

## If you're starting this

Build the consent and suppression layer first, before any speech work. It's the part that decides whether the program can run at all, it's the part that takes longest to get right across state rules, and it's useful on its own the day it ships, because your human dialer can use the same gate.

Then get one call working end to end with a scripted agent and no model in the loop, to prove the telephony, the latency budget, the session routing and the trace capture. Add the language model last, to a pipeline that already records everything it does.

And write the scorers before the prompts. If you can't state what a good call looks like as a check that runs against a trace, you won't be able to tell whether a prompt change helped. More on how we think about the decisioning side in [decision models in auto refinance](/blog/decision-models-auto-refinance-underwriting), and on the platform work in [/databricks](/databricks).
