A request reaches the data integration center of a university hospital: a research project needs all creatinine values from the past five years – together with diagnoses, medication, and the associated case contexts. What sounds like a simple database query turns out to be detective work. The values originate from two LIS generations and three sites, an analyzer was replaced in the meantime, the units are coded inconsistently – and whether the local code “KREA-S” at site B measures the same thing as “Krea” at site A is ultimately known only to a colleague who has maintained the system for fifteen years.
In the laboratory, the same question arises in the opposite direction: an analyzer is due to be replaced. Which interfaces, rules, evaluations, and downstream systems depend on the values from this device? Which reports, which statistics, which notification pathways are affected? Here, too, an elaborate search often begins – through system documentation, interface specifications, and the memory of experienced staff.
Both situations share the same root cause. The problem is not a lack of data – it is the lack of a reliable answer to three simple questions: What data do we have? Where does it come from and where does it flow? And what does it mean?
The problem: data landscapes without a map
Hospitals and laboratories today manage highly complex data landscapes: laboratory information systems, HIS and PMS, analyzers, middleware, registries, research databases, and increasingly AI applications. The data is distributed, inconsistently coded, and rarely documented end to end. The knowledge of what a field means, where a value comes from, and who is responsible for a dataset does exist – but it sits in interface specifications, in historically grown configurations, and above all in the heads of individual people.
The consequences are part of everyday life: every integration project starts with a new inventory. Every research request requires rounds of manual clarification. Every audit means laboriously compiling evidence. And every staffing change carries the risk that knowledge about one’s own data landscape is irretrievably lost. Diagnostics – this becomes apparent here once again – suffers not primarily from a lack of data, but from its fragmentation.
Metadata: a catalog and map of your data
This is precisely the gap that intelligent metadata management closes. It is expressly not another data silo, but a descriptive layer above the existing data – a catalog and city map of one’s own data landscape. It provides structured answers to what otherwise has to be painstakingly reconstructed: What data exists – organized into clinical domains such as laboratory values, vital signs, microbiology, medication, or diagnoses, described down to the level of individual data elements with meaning, unit, and value format. Where it comes from and where it flows – from the source through interfaces and transformations to every destination in which a value is reused. And what it means – anchored in international standards such as LOINC, SNOMED CT, UCUM, and FHIR, which turn local labels into comparable, machine-readable information.
Concepts such as data catalogs, data lineage, and data products have long been established in the general data world. The decisive step is to transfer them consistently to the diagnostic and clinical world – with medical terminologies, regulatory requirements, and the real-world interface formats of laboratory and hospital, from HL7 and LDT to FHIR.
Origin and destination: lineage as the backbone of a resilient data landscape
The heart of this layer is continuous origin and destination tracking – lineage. Each data element is linked to its source systems, the transformations it passes through, and all of its destinations – ideally down to field level. A potassium value can thus be traced from the analyzer through the middleware and the LIS all the way to result reporting, analytics, and the electronic patient record.
From this linkage, almost as a by-product, emerges perhaps the most valuable capability: impact analysis. The question “What depends on this device, this interface, this field?” turns from weeks of research into a single query. The analyzer replacement from the opening example becomes plannable, because all affected systems and evaluations are visible at a glance.
Herein lies an often underestimated contribution to resilience. A data landscape is robust when it can absorb change: system replacements, migrations, outages, staff turnover, new regulatory requirements. Transparent metadata makes this knowledge available independently of individual people, documents dependencies traceably for audits, and creates the basis for acting quickly and confidently in the event of a disruption or migration – instead of searching in the dark.
Research data as first-class datasets
An essential design principle: an intelligent metadata layer describes not only live data streams from connected systems, but treats standalone datasets as equals – especially those created in research and innovation projects. A pseudonymized cohort, a curated evaluation dataset, or a training dataset for a model receives the same catalog entry as a live feed from the LIS: with schema, origin, time period, consent basis, curation status, version, and responsible persons.
For institutions that connect routine care and research – such as data integration centers at university hospitals – this is a significant lever: research datasets become findable and reusable instead of fading away tied to individual projects, their origin in routine care remains traceably documented, and the boundary between care context and research context is explicit rather than implicit. In this way, data is unlocked in the best sense: what already exists becomes visible, linkable, and usable for new questions – a foundation for research and innovation that carries far beyond the individual project.
Security and governance: protection begins with an overview
Metadata makes data not only more usable, but also more secure. Sensitive data can only be reliably protected if it is known where it resides and where it flows. Intelligent metadata management therefore classifies every data element and every dataset: sensitivity (identifiable, pseudonymized, anonymized), legal basis and consent, retention periods, notification obligations, and access rules.
This includes clear responsibilities – clinical data owners and operational data stewards per domain – as well as a defined lifecycle with versioning and change tracking. Governance thus shifts from a downstream control exercise to a built-in property of the data landscape: anyone planning an analysis can see immediately which data they may use, and under which conditions.
AI-supported auto-modeling – with a human in the loop
The most common objection to metadata initiatives is the maintenance effort – and it is justified if the catalog has to be filled manually. The key therefore lies in automation: the platform automatically reads structures from HL7 messages, FHIR bundles, GDT/LDT, CSV files, or even free-text documents and proposes the data model using AI – detected fields, their presumed meaning, matching standard codes, and the assignment to domains, each with a confidence score and source reference.
The decisive principle here is human-in-the-loop: the AI generates transparently substantiated proposals, the professional approval remains with humans, and every decision is documented in an audit trail. This corresponds to a whitebox approach, indispensable in a medical setting – and it makes it realistic to build a catalog that would otherwise fail due to sheer volume.
Conclusion and outlook
Intelligent metadata management does not create new data – it unlocks the data that already exists. It makes data landscapes findable, understandable, secure, and manageable: for laboratory operations that want to carry out changes in a plannable way; for hospital IT, which needs to keep dependencies and risks in view; and for research and innovation, which depend on traceable, reusable datasets.
We are building this metadata layer consistently on top of our established data management and harmonization layer – with semi-automatic mapping to LOINC, SNOMED CT, and FHIR, terminology services, and a certified environment in accordance with ISO 13485 and ISO/IEC 27001. We are currently actively developing the approach further and preparing the first fields of application.
For this next phase, we are looking for institutions that would like to put the approach into practice together with us – laboratories, hospitals, and data integration centers alike. If you recognize your own challenges in the questions described here, or are simply curious, we would be delighted to have a no-obligation conversation.

