DICOM De-Identification: Removing PHI Without Breaking the Study

Glowing line-art illustration on a deep navy field of a medical study passing through a filter that strips identifying marks while the threads linking its parts stay unbroken

Run a strip-PHI script against a study and the fast path looks convincing: Patient Name goes blank, Patient ID gets replaced, and the study opens with no name banner in sight. DICOM de-identification looks finished at that point, and that is exactly where it usually goes wrong.

A DICOM object scatters identity across more places than a quick script touches: undocumented private tags, text burned into the pixel data, an observer’s name inside a structured report, and a web of UIDs. Those UIDs have to stay internally consistent, or the study stops behaving like one study. Remove too little from any of those places and PHI survives the pass. Remove too much, or remove it inconsistently, and the study can no longer open, relate to its own series, or reconcile against the archive it came from.

Where DICOM De-identification Actually Has to Look

Standard attributes are the part most tooling already handles. Patient’s Name, Patient ID, Patient’s Birth Date, Referring Physician’s Name, Accession Number: these sit at fixed tags the DICOM standard defines, and any competent parser can find them by tag alone. A de-identification pass that only touches this layer is not wrong. It is just incomplete.

Private tags are where a generic script runs out of certainty. Modality and workstation vendors are free to define their own tags outside the standard dictionary, and nothing requires a private tag to announce that it carries a name, an operator ID, or a referring clinic. A de-identifier that only recognizes standard attributes passes every private tag through untouched by default, with no way to know what any of them mean without the vendor’s private dictionary or a manual review.

DICOM’s own confidentiality profiles call this out directly. They warn that “identifying information may be contained in Private Attributes, new Standard Attributes, Retired Standard Attributes and additional Standard Attributes not present in Standard Composite IODs (as defined in PS3.3) but used in Standard Extended SOP Classes.” That is the annex admitting its own attribute list is not a complete map of where identity can hide.

UIDs, burned-in pixel text, and structured report content each behave differently enough that treating them as one problem is how PHI slips through.

De-identification Is Not the Same as Anonymization

The two words get used interchangeably outside a standards document, and they should not be. De-identification removes or replaces the elements a de-identifier is scoped to touch. Anonymization is the stronger claim most people mean by it: that nothing left in the object can lead back to the patient. A de-identified study can still fall short of that second bar.

DICOM’s own confidentiality profile says so plainly, in a note attached to the annex that defines it: “de-identification of the Attributes does not imply de-identification of the Information Object.” Touching every attribute the profile lists does not guarantee the resulting object carries no identifying information, because the profile only reaches what is encoded as an attribute, and pixel data is the clearest counter-example. DICOM’s guidance on UID replacement warns that “if any useable pixel data is retained during de-identification, then re-identification is nearly always possible if one has access to the original images.” That risk comes from the pixel content itself, compared against the original.

None of that argues against de-identifying a study. It argues for being precise about what a given pass actually achieves, and for treating “de-identified” and “safe to disclose to anyone” as two claims that overlap most of the time, not always.

Why UIDs Get Remapped, Not Deleted

Every DICOM object carries a stack of UIDs that exist purely to say what belongs with what. A Study Instance UID ties every series in an exam together, a Series Instance UID groups the images acquired in one pass, and a SOP Instance UID identifies one object. None of them are, by themselves, a name or an identifier of a person. The standard makes the distinction in the same breath: “Though individuals do not have unique identifiers themselves, studies, series, instances and other entities in the DICOM model are assigned globally unique UIDs.”

A UID still carries a trace of where it came from. UIDs are built on a hierarchical scheme of roots, and a knowledgeable person can often trace a root back to whoever was assigned it: typically the device manufacturer, sometimes the organization using the device. Anyone holding the original images, or a database that still lists their UIDs, can match the two and recover the identity. That is why the Basic Application Level Confidentiality Profile replaces instance UIDs instead of keeping them.

Replacement, though, is not deletion. A viewer, an archive, or a PACS reconstructs a study by grouping objects that share a Study Instance UID and a Series Instance UID. Strip those UIDs, or generate a fresh one for every object in what used to be a single series, and the receiving system sees a pile of unrelated single-image studies rather than one exam.

Nothing about the pixel data changed. The relationships the archive depends on did.

The confidentiality profile’s own action code for this situation calls for a replacement that is “a non-zero length UID that is internally consistent within a set of Instances”. Every object that originally shared a UID gets the same replacement, not a fresh one each. That consistency is what keeps a de-identified study behaving like a study rather than a folder of orphaned images. The Retain UIDs Option does the opposite: it keeps the original values, for the narrower case that needs an audit trail back to the original images.

Consistent replacement has limits of its own. The standard cautions implementers to “take care not to remove UIDs that are structural and defined by the Standard as opposed to those that are instance-related”. A naive pass that treats every UID the same way risks breaking something the format itself depends on, like the SOP Class UID that says what kind of object this even is.

Burned-In Text: What Standard Attributes Can’t Reach

Some modalities write patient information directly into the image itself, not just into an attribute alongside it. Ultrasound is the routine offender: a scanner overlays patient name, exam date, or facility name onto the same pixels being scanned, a habit left over from printed film. DICOM’s Clean Pixel Data Option exists because a de-identifier that only edits tagged attributes leaves that text untouched inside the pixel data.

The standard’s own comparison is useful here: “CT images do not normally contain such burned in annotation, whereas Ultrasound images routinely do.” A pipeline tuned against CT alone will happily pass ultrasound studies through with a name still visible on screen.

Finding that text is harder than it sounds. Optical character recognition can locate candidate regions, but the standard is blunt about the limit: “deciding whether or not that text is identifying information or some other type of information may be non-trivial.” A caliper measurement and a patient name can look identical to an OCR pass, and only one of them is PHI.

Whatever a pipeline decides to remove, it has to actually change the stored pixel values, not hide them. The standard is direct: “The Stored Pixel Values are to be changed (blacked out); it is not sufficient to superimpose an overlay or graphic annotation or shutter to obscure the Stored Pixel Data Values, since those may not be ignored by the receiving system.” A shutter is a suggestion to a DICOM viewer. The pixel data underneath it is what actually gets stored, copied, and re-shared, overlay or not.

Structured Reports and Other Non-Pixel Content

A study is not only images. Structured reports carry findings and observations as their own tagged content, and that content can name the person who made it. DICOM’s Clean Structured Content Option exists because, as the standard puts it, “Instances of Structured Report SOP Classes may contain identifiable information in a Content Sequence (0040,A730) encoded in Content Items.” That Content Sequence easily holds an observer’s name, a referring clinic, or a dictated note nobody thought to scrub separately from the image attributes.

A de-identification pass that clears Patient Name from the top-level data set but skips this layer has not finished. The standard does not hedge: “A de-identifier that does not implement this Option creates significant risk when attempting to de-identity a Structured Report unless it is only used to de-identify instances that are known to have no identifying information in the Content Sequence.” Acquisition context and specimen preparation content carry the same risk in other SOP classes: wherever the standard lets a system encode free-form content alongside an image, identity can ride along with it.

Safe Harbor Versus Expert Determination, at a High Level

DICOM’s confidentiality profiles are a technical toolkit, not a regulatory determination. The standard makes that boundary explicit, noting that using the profiles “does not replace a de-identification process, but should be part of it.” Whether a given result actually satisfies HIPAA’s de-identification standard is a separate question the Privacy Rule answers, not DICOM.

HHS’s guidance lays out two paths. Under Safe Harbor, information counts as de-identified once a defined list of identifiers is removed: names, most elements of dates, and categories such as “Medical record numbers” and “Full-face photographs and any comparable images”. Removal is only the first condition. The Privacy Rule adds a second at 45 CFR 164.514(b)(2)(ii): the covered entity must have no actual knowledge that what remains could still identify the individual, alone or combined with other information.

Expert Determination works differently. A qualified person applies accepted statistical and scientific methods. That person “determines that the risk is very small that the information could be used, alone or in combination with other reasonably available information, by an anticipated recipient to identify an individual who is a subject of the information”, then documents the analysis. There is no fixed checklist, only a documented judgment call about a specific recipient.

A DICOM object run through the Basic Profile with the Clean Pixel Data and Clean Structured Content Options covers much of what Safe Harbor’s list asks for, but does not automatically satisfy either method. Safe Harbor requires every listed category gone, and a pass tuned only to the most common DICOM attributes may not fully account for identifying dates or codes embedded inside private or structured content. Expert Determination requires a documented determination by a qualified person, something no automated tag-stripping pass performs by running. And once a de-identified study crosses an organizational boundary to reach that recipient, the identity and access questions that apply to any handoff still apply.

What Breaks When De-identification Goes Wrong

Too little removal is the obvious failure, and the one every rule above addresses. It shows up as a name readable in an ultrasound frame, an observer’s name inside a structured report, or a private tag carrying a referring clinic’s name nobody’s de-identifier knew to look at.

Too much removal, or removal applied inconsistently, breaks the study in a different, quieter way. Strip a Study Instance UID and a Series Instance UID without replacing them consistently across every object that shared the original value, and the receiving archive scatters one exam into single-image studies. Generate a fresh, independent UID per object instead of one consistent replacement per series, and the same failure shows up for a different reason: technically present UIDs that no longer agree with each other.

Strip the File Meta Information’s Transfer Syntax UID rather than replacing the header as the standard requires, and the receiving system loses the attribute that says how to decode the pixel bytes that follow. It is the same fragile dependency that shows up whenever a DICOM object’s header structure is treated as disposable. Neither failure mode is exotic. Both come from treating de-identification as a single find-and-replace instead of a set of layers, each with its own handling.

What This Means for Your Evaluation

For a product or engineering team building or buying a de-identification step, the practical question is not whether a tool touches Patient Name. Every tool does.

The real questions are narrower: does it handle private tags, and does it locate and black out burned-in pixel text rather than overlay it? Does it reach into structured report content items, and does it remap UIDs consistently within a study rather than issuing a fresh one per object? Walking a vendor through those four questions separates a real implementation from a script that only edits the obvious tags.

Whatever runs upstream, a de-identified study still has to move through the same DICOM services as any other object once it reaches an archive. Storing it, querying for it, retrieving it, and relating it back to its own series all run on the identifiers the pass just rewrote. Preserving those on the way in and out is the archive’s job.

EBM mAIn PACS® implements native DICOM C-STORE, Q/R, MWL and Storage Commitment against a documented services list, of which the storage, query/retrieve and commitment services are the ones an archived study touches. The platform’s own posture is designed to support HIPAA compliance, alongside the DICOM standards conformance that any PACS software evaluation should be checking regardless of which vendor built it.

Getting de-identification right is a detail question, answered one layer at a time. Getting it wrong tends to show up the first time someone tries to use the result.