Databricks released Funke, an open-source library that parses HL7v2 messages straight into native Spark types inside the lakehouse, keeping every segment, field, repetition, component and subcomponent addressable. For any organisation that receives admit, discharge and transfer feeds or lab results, that removes the two expensive detours teams normally take. One is converting everything to FHIR first. The other is paying a third-party engine to flatten messages into wide tables somewhere outside the platform.
HL7v2 is the standard that moves routine clinical traffic between systems. A message says a patient was admitted, an order was placed, a result came back. It is delimiter-encoded and nested, with the separators declared in the message header itself, and the specification is loose enough that two sending systems produce meaningfully different messages for the same event. That is why the data usually arrives in a data team's hands already mangled by somebody else's opinion about which fields matter.
What actually changed
Funke is the successor to Smolder, the Scala Spark data source Databricks open-sourced in 2021. Smolder predates Unity Catalog, Declarative Automation Bundles and Spark Declarative Pipelines, and wasn't built to fit cleanly with the modern Databricks platform. Funke is Python and PySpark, deploys as a bundle, ingests through a declarative pipeline and stores everything in Unity Catalog.
Look at the type. A parsed message becomes a map from segment name to the repetitions of that segment, and each field is a map addressable by field number, then repetition, then component, then subcomponent. From the repository:
# Spark type for an HL7v2 field
HL7v2FieldType = T.MapType(
T.IntegerType(),
T.ArrayType(
T.MapType(
T.IntegerType(),
T.MapType(T.IntegerType(), T.StringType()),
)
),
)
Because that is a normal Spark column, nothing is hidden behind a connector. Funke supports every HL7 message type and version, and the parser handles the separators and escape sequences the MSH header declares. The claim in the announcement is that parsing is lossless and the structure survives as far into the pipeline as possible. So the decision about which fields matter stays with the people who know the clinical semantics, taken at the point where they build a gold table instead of upstream in a vendor's mapping config.
That inversion is the whole business case. Most HL7 integration spend goes on negotiating a schema with a middleman before anyone has seen the data. Funke lets you land everything and argue later.
The pipeline, concretely
Then parsed_messages reads that stream and applies the parser, adding a single hl7 column of the type above. Gold tables are yours to build.
Volumes are the right landing zone here. Unity Catalog recommends them for registering landing areas for raw data produced by external systems, governed the same way tables are, reachable at /Volumes/<catalog>/<schema>/<volume>/<path>.
Under Spark Declarative Pipelines the bronze and silver tables are streaming tables, which carries one consequence worth planning for. A streaming table processes each row once, so if you change the parsing or extraction query later, old rows are not reprocessed unless you trigger a full refresh. For a clinical history table that matters. Decide early whether a given gold table is append-only history or something you expect to be able to restate.
Pulling a field out
The raw addressing is positional. Taking the message type out of MSH-9, from the repository example:
SELECT
md5Hash AS messageId,
hl7.MSH[0].fields[9][0][1][1] AS messageType,
hl7.PID[0].fields[5][0][1][1] AS patientLastName,
hl7.PV1[0].fields[2][0][1][1] AS patientClass
FROM silver_table
Precise, and unreadable six months later. Funke ships accessor helpers for that reason. get_value takes the segment, its repetition, the field, the field repetition, the component and the subcomponent and returns a column, with get_segment, get_field, get_component, get_subcomponent and a parse_hl7_version helper that reads the version out of the MSH header. My recommendation is to use the helpers everywhere except in throwaway exploration, and to put a comment next to every extraction naming the HL7 field in standard notation, because PID-5.1 is the thing an integration analyst can check and fields[5][0][1][1] is not.
The same extraction works in PySpark and in Spark SQL, which is the practical argument for this over a flattening vendor. An analyst who writes SQL can build a gold view without waiting on a data engineer to add a column to a mapping file.
You can also use the parser outside the pipeline. pip install . from the repo, then HL7v2Msg(raw) in plain Python, or parse_hl7v2_msg(HL7v2Schema()) as a Spark UDF on any DataFrame column.
What it costs and what it requires
Funke is open source under the Databricks License, so there is no licence fee. What you pay is Databricks compute, meaning pipeline DBUs for the streaming ingestion plus storage, billed per second with no up-front cost. The cost shape is the same as any Auto Loader pipeline, and HL7 feeds tend to be high message count and small message size, so file discovery is what to watch more than data volume.
Deployment provisions the library, the pipeline, the schema and the landing volume. databricks bundle deploy -t dev targets a named environment, and bundle deployment bind links bundle resources to pipelines that already exist, so adopting Funke into a workspace that already runs HL7 ingestion doesn't mean recreating data.
The demo's gold layer turns the parsed ADT stream into an adt_events history table and maintains a live current_census table using change data capture keyed on the visit number, which a bed_utilization table joins against facility capacity to feed a dashboard.
The edges of what Funke covers
It is a parser and an ingestion accelerator, and an interface engine is a separate job. It doesn't listen on MLLP, doesn't acknowledge messages, doesn't route, doesn't retry to a sending system. If your current spend is on Mirth, Rhapsody or Cloverleaf doing message routing between clinical systems, Funke replaces none of that; it sits downstream and takes files off a volume. Keep the interface engine, point a copy of the feed at a volume.
It also doesn't interpret. The mapping from a segment and field to a business concept is a decision made with clinical domain knowledge, and no library takes that off you. Funke's contribution is that the data is complete and reachable when you make it.
And it is HL7v2 only. FHIR-first estates get nothing from it, and organisations whose clinical data arrives as flat extracts from a vendor's reporting database have a different problem, closer to getting ERP and CRM data into Databricks.
One downstream use is worth naming. Once parsed messages sit in Unity Catalog with governance attached, they become a retrieval source for enterprise search and chatbots over company data, with the same row and column controls as everything else in the catalog.
Run the demo against your own messages first
Message variation between sending systems is where HL7 projects go wrong, and querying parsed_messages across your own sending facilities tells you more about your feed than schema discussions do. The test messages in the demo come from the HL7 v2-to-FHIR project, so they are clean; yours will not be, and that is the point of looking.
Frequently asked questions
What is Funke?
Funke is an open-source Databricks library that parses HL7v2 messages into native Spark types on the lakehouse, preserving the full segment, field, component and subcomponent hierarchy. It is the successor to Smolder, rebuilt in Python and PySpark around Unity Catalog, Declarative Automation Bundles and Spark Declarative Pipelines, and ships with a runnable demo.
Does Funke replace an interface engine like Mirth?
No. Funke is a parser and ingestion accelerator without interface engine duties, so it doesn't handle MLLP transport, acknowledgements or routing between clinical systems. It reads HL7 files that land in a Unity Catalog volume and makes them queryable.
How much does Funke cost?
Funke itself is free and open source under the Databricks License; you pay only for the Databricks compute and storage the pipeline consumes, billed per second with no up-front cost. The workload is a standard Auto Loader streaming pipeline, so cost scales with message volume and file count.
Do I still need to convert HL7v2 to FHIR?
For analytics on the lakehouse, no, you don't. Funke exists so you can skip the FHIR translation layer, which adds a conversion step and can drop detail that has no clean FHIR equivalent. You will still need FHIR for interoperability with systems that require it.