A clinical data integration program copies values by default: pull a result out of its source system and write it into a shared record. USCDI is built the same way, defining a lab result as the finding itself rather than a pointer to where it lives. Imaging is where that breaks. USCDI v7 adds a separate element for reaching a study rather than carrying it, and a program scoped on the copy habit usually finds that out after imaging is already in scope.
That habit is not a mistake. Labs, pharmacy, scheduling and billing systems mostly hold discrete values: a result, a dose, a slot, a charge. Copying a value into a shared record is cheap, and the copy stays useful even after the source system changes or goes offline for maintenance.
Imaging holds something structurally different: a large binary object, not a value. A program that tries to integrate a study the same way it integrates a lab result runs into a seam nobody drew on the original architecture diagram.
What a Clinical Data Integration Program Copies by Default
The United States Core Data for Interoperability is the data set the Office of the National Coordinator uses to define a shareable clinical fact. It treats a lab result as a value in its own right.
Its Laboratory data class has carried a value element since USCDI v1, defined there as “documented findings of the analysis of a tested specimen. Includes both structured and unstructured (narrative) components.” Later versions shorten that wording and v7 renames the element Value/Result, without changing what it holds. The finding itself, not a pointer to where it lives, is what a receiving system is built to hold.
That is the shape a clinical data integration program defaults to across most of what it touches: an order, a result, an encounter, a charge. Each one is small enough to copy cheaply, and each copy stands on its own once it lands. A program that extracts, transforms and loads those values into a shared record is doing exactly what the underlying data model expects of it.
The same standard repeats that shape elsewhere. USCDI’s Medications data class defines its core data element as a coded drug value drawn from RxNorm, the normalized drug terminology the National Library of Medicine produces. That value is not a reference back to a pharmacy system’s own record. Whatever the data type, the pattern is consistent: when a fact is small enough to carry directly, the standard defines it as a value, and a data integration program copies it accordingly.
Where Imaging Breaks the Pattern
The Diagnostic Imaging data class sits in the same USCDI data set, defined as “tests that result in visual images requiring interpretation by a credentialed professional.” Its two original elements carry the test’s name and the interpreted findings. Their wording tightens across versions, and v7 renames the report element Diagnostic Imaging Result/Report.
USCDI v2 defined the Diagnostic Imaging Report element as “interpreted results of imaging test that includes the study performed, reason, findings, and impressions. Includes both structured and unstructured (narrative) components.” Neither element is the image. Both describe it, or describe what a radiologist found by reading it.
A data set that stopped there would leave a program with a description of a study and no standard way to reach the study itself. ONC addressed that in USCDI v7 with a separate element, Diagnostic Imaging Reference, defined only as “information that can be used to access a diagnostic imaging study.”
Its own examples read: “imaging study endpoint weblink, unique identifiers, and contextual information needed to retrieve a diagnostic imaging study.” A lab result has no equivalent element, because a lab result does not need one. The value already traveled.
The American College of Radiology had pressed for exactly that element. Its 2023 comment on the same data class page, filed while Imaging Reference was still a Level 2 candidate, argued that “the ability to reference the relevant DICOM image files themselves is critical to many patient care scenarios where the Diagnostic Imaging Report itself would be inadequate”, and named surgical planning, disease staging and emergency transfers as those scenarios. A report can state what a radiologist found. It cannot substitute for the study itself when the next clinician has to look at the images.
That is the structural difference a clinical data integration program has to design around. Everything else in the record moves as a value. A study moves as an address, resolved on demand against the system that still holds it.
The Two Patterns a Program Ends Up Running
In practice, a program ends up running two integration patterns side by side instead of one. Structured clinical facts move by copy: extract the value, transform it, load it into wherever the program keeps its unified record. That order and result feed usually runs on HL7 messaging, a standard What HL7 Is covers on its own terms. An integration engine is usually the thing doing the copying at scale, receiving, transforming and routing each message as it arrives.
Imaging moves by reference instead. The study stays on the system built to hold it, and every other system that needs it gets an identifier and an endpoint rather than a duplicate. FHIR in Plain Terms covers how a newer, REST-based standard formalizes that same reference pattern with its own resource model, separate from the messaging feed carrying everything else. The DICOM side of that handoff, connecting a modality, a RIS and an EHR to the same archive, is what PACS Integration: Connecting Modalities, RIS, and the EHR walks through surface by surface.
Neither pattern is the correct one in general. Each is correct for the data type it was built around, and a program that applies the wrong one to a given data type pays for the mismatch later rather than up front.
What Applying the Wrong Pattern Costs
A program that tries to copy imaging the way it copies a lab value runs into byte counts a relational warehouse was never sized for. A single study can outweigh every other record type in the same patient encounter combined. Sizing an imaging archive has to account for that skew on its own, separately from whatever the rest of the program copies.
The mismatch runs the other way too. A program that defaults to reference-only for every data type adds a network round trip and a dependency on a source system staying reachable. That is a real cost for data that was cheap to copy in the first place. Forcing a lab value through a retrieve-on-demand pattern built for a multi-gigabyte study solves a problem that value never had.
Ownership follows the same split. A copied value can outlive the system that produced it, because the copy is now independent. A referenced study cannot: whatever system holds the archive stays the system of record for as long as any pointer to it keeps resolving. That is usually longer than the contract that built the integration.
What This Means for Scoping the Integration
A product or engineering lead scoping a clinical data integration project can turn the pattern into a short set of questions. Ask them before the architecture is fixed, not after a study fails to load somewhere it was assumed to be:
- Which data types in this integration actually need to move by value, and which are large enough that a resolvable reference is the only workable answer?
- Who owns the system the reference resolves against, and what happens to every downstream pointer if that system is migrated or decommissioned?
- Does the receiving system, a warehouse, an EHR, a partner’s own product, need the study itself, or does a report plus a resolvable link satisfy the actual clinical use case?
- If pixel data does get copied somewhere, a research extract, a de-identified data set, who decided that copy was necessary, and what does it cost to keep two copies in sync?
The answers vary by program. What does not vary is that a program is better off answering them once, deliberately, than discovering the answer by watching an integration fail under real volume.
The Decision a Program Has to Make
A clinical data integration program has to decide which pattern each data type gets before the architecture is set, not discover it after imaging is already in scope. For imaging specifically, that decision comes down to one question. Does the rest of the program need a report and a resolvable pointer, or does it need the pixel data replicated into a general-purpose store built for rows?
EBM mAIn PACS® answers the addressing half of that question with a published, bounded surface: native DICOM C-STORE, Query/Retrieve, Modality Worklist, and Storage Commitment, documented on its integrations page. That is the layer a study is stored on and addressed from. Which systems consume that reference, and how far the copy pattern extends into HL7 or FHIR territory around it, stays an architecture decision for the partner building on top of it.
The standards already answer the question for anyone who checks them. Most programs back into replication anyway, because nobody asked before the contract was signed. The bill for that arrives later, with whoever inherits two copies of the same study and the job of keeping them in step.
