A de-identification pipeline can clear every tagged attribute in a DICOM study (Patient Name, Patient ID, Referring Physician) and still hand back an identifiable file. A name, a medical record number, and a scan date can sit in plain text inside the image itself. Ultrasound frames, secondary capture screen grabs, and digitized film routinely carry that text burned directly into the pixel data, not into a field a parser can find by tag. DICOM does define an attribute meant to flag this risk, Burned In Annotation, but that flag is optional, self-reported, and absent from most studies that actually need it.
Real DICOM de-identification guidance treats burned-in annotations as a separate problem from attribute scrubbing, with its own detection step, its own standard option, and its own failure modes. Here is where that risk concentrates, what the standard actually requires to clear it, and what a detection pipeline gets wrong along the way.
Which Modalities Burn Text Into the Pixel Data, and Why
Ordinary CT and MR acquisitions rarely carry burned-in text. The standard draws the contrast with ultrasound directly: CT images do not normally carry this kind of annotation, while ultrasound images routinely do. Ultrasound scanners commonly burn an overlay of patient name, exam date, and facility directly onto the frame itself. Many vendors have kept that layout even though the same information also lives in proper DICOM attributes elsewhere in the object.
Secondary Capture is the other concentration point, and it exists precisely because not every image source speaks DICOM natively. The Secondary Capture Image IOD wraps non-DICOM sources into a standard object without constraining pixel content. Video interfaces digitizing an analog signal, film digitizers converting printed film, workstations producing a screen dump, and scanned paper documents all land here.
Whatever text, timestamp, or UI chrome was visible on the original source comes along in the pixels: an endoscopy tower’s overlay, a legacy modality’s printed report, a monitor screen grab. There was never a structured field for any of it to live in instead. A study that entered a PACS as a Secondary Capture object is the one most likely to carry identity where an attribute-only de-identification pass cannot reach it.
The Burned In Annotation Attribute Cannot Be Trusted Alone
DICOM gives the Burned In Annotation attribute (0028,0301) exactly one job: say whether an image “contains sufficient burned in annotation to identify the patient and date the image was acquired.” Its allowed values are YES and NO. The catch sits in how the standard defines its absence: “If this Attribute is absent, then the image may or may not contain burned in annotation.”
That is not a conservative default. It is an explicit acknowledgment that a missing value carries no information at all.
In the General Image Module, which a Secondary Capture object carries as a mandatory module, Burned In Annotation is Type 3, meaning optional. An image built on that module can leave the attribute out entirely and still conform, and most legacy Secondary Capture sources do exactly that. Even where a system does set the value, it is typically the acquiring device or software making that call about its own output, not an independent check. A de-identification pipeline that filters, routes, or skips pixel cleaning based on this attribute alone is trusting a self-report that was frequently never made, and, when it was made, was never verified by anything downstream.
DICOM’s Clean Pixel Data Option: What Actually Counts as Cleaned
The standard’s answer to this gap lives in PS3.15 Annex E, as the Clean Pixel Data Option attached to the Basic Application Level Confidentiality Profile. Implementing it means any identifying information burned into the Pixel Data attribute has to actually be removed. The standard is candid about what that takes: “This may require intervention of or approval by a human operator.” Full automation is not assumed.
What makes the Option worth naming separately from ordinary attribute scrubbing is the direction it runs. A pipeline is only entitled to assert the study is clean after doing the work, not before. The standard directs a cleaned image to carry Burned In Annotation (0028,0301) set to NO once the Option has actually been applied. A value of NO is not a default a receiving system should assume on its own.
The standard also gives implementers real latitude in how far to go, stating that “compliance with this Option requires that identifying information is removed, regardless of how that is achieved.” It adds that “the most conservative approach of removing any and all burned in text would be compliant.” Taking that approach can cost a useful localizer marking or a manual annotation along with the identity.
The scope is specific too: the Option targets the Pixel Data attribute in the top level Data Set of an image object, the same structure that carries a DICOM object’s identifiers and encoding rules. Pixel data tucked inside a private attribute gets removed outright rather than evaluated, because a private attribute is never known to be safe by default.
Detection Approaches: OCR and Region Masking, and Where They Fail
Two approaches dominate in practice, and each fails differently. Optical character recognition scans the full frame for text, flags candidate regions, and blacks them out or routes them to a reviewer. It generalizes across sources, which matters given how varied Secondary Capture content actually is, but it misses low-contrast text, small fonts near a screen edge, and rotated overlays. It also cannot reliably tell an identifying label from a clinical one: a caliper measurement and a patient’s date of birth can look identical to a pattern matcher trained only to find text.
Region masking takes the opposite approach: a fixed template blacks out a known coordinate range where a specific device or software version burns its overlay. It is fast and precise when the template is right, and it breaks the moment it is not.
Several things defeat a template silently: a firmware update that repositions the overlay, a new modality model added to a fleet without an updated template, an operator repositioning an annotation before capture. No error gets thrown, because the pipeline has no way to know the assumption it was built on no longer holds. Scanned film and heterogeneous screen captures compound the problem further, since there is often no fixed layout at all to template against.
For a team evaluating a vendor’s pixel-cleaning step, the useful question is not whether it uses OCR or masking. Most real implementations combine both. The question is what happens when detection is uncertain: does the pipeline default to a conservative full black-out of the frame, which the standard expressly permits, or does it pass an unclear region through unchanged.
Why Blacking Out an Overlay Is Not the Same as Cleaning the Pixel Data
DICOM treats these as two separate problems on purpose. Text burned into the Pixel Data attribute is one category, addressed by the Clean Pixel Data Option above. A different category covers identification “encoded as graphics, text annotations or overlays” attached to an image or a presentation state. The standard resolves that through a separate Clean Graphics Option, also noting that this “may require intervention of a human operator.”
The Annex is explicit that a de-identifier has to handle both categories, because neither Option covers what the other addresses. A de-identified study needs its burned-in pixel text removed under the Clean Pixel Data Option and, separately, any identifying graphics or overlays removed or cleaned under the Clean Graphics Option.
The practical trap follows directly from that split. In a current object, an overlay sits in its own Overlay Data attribute and a graphic annotation sits in a separate Presentation State instance, both outside the pixel array. Older objects are the exception: DICOM once described storing overlay bits in the unused bit planes of Pixel Data itself, and retiring that usage did not change the objects already written. Either way, a viewer configured not to render a layer changes nothing about what the file stores, and the study looks clean on screen while the Pixel Data element sits untouched underneath.
The standard closes that door directly: under the Clean Pixel Data Option the stored pixel values themselves have to be changed, blacked out. Superimposing an overlay, a graphic annotation, or a shutter over the text does not satisfy it, because a receiving system is under no obligation to ignore the values sitting underneath. A black box painted on at display time is a rendering choice, and the next viewer or a raw parser need not honor it. Neither Option is a rendering instruction: both require the underlying attribute or the underlying pixel values to actually change.
Research and Teaching Datasets: Where This Risk Concentrates
The Confidentiality Profile exists for a specific downstream use, and PS3.15 Annex E names it directly. De-identified instances are useful “in creating teaching or research files, performing clinical trials, or submission to registries where the identity of the patient and other individuals is required to be protected.” That is precisely where Secondary Capture and ultrasound content concentrates. A teaching file is disproportionately likely to be a screen capture, a cine loop, or a piece of scanned historical film, the same categories most likely to carry burned-in text in the first place.
A multi-site research collection or a registry submission also crosses an organizational boundary by design, so a pixel-level failure is no longer contained inside the institution that made it. Scale turns a rare miss into a systemic one: the standard’s note that cleaning may need a human operator’s review is manageable for a handful of cases and expensive across thousands of frames. Where that review does not run end to end, nothing downstream catches what the pipeline missed, so the frame-level miss rate is the rate the finished dataset inherits. A curation process that assumes automated cleaning caught everything, instead of budgeting for the pipeline’s detection failure rate, is how a burned-in name lands in a dataset that already passed an attribute-level compliance check.
What This Means for Your Evaluation
For a team building or buying a de-identification step, the Burned In Annotation attribute is a place to start asking questions, not a place to stop.
Worth confirming directly with any vendor: does pixel cleaning run as a real step at all, separate from attribute scrubbing? Is detection OCR, template masking, or both, and what happens on an uncertain match? Does Burned In Annotation get set to NO only after a pass that actually touched the pixels, rather than carried through from whatever the source device wrote? And are graphics and overlay attributes addressed as their own step, rather than assumed to be covered by pixel cleaning?
Whatever runs upstream to clean a study, the archive receiving it still has to store, index, and relate that study through the same DICOM services as anything else. The study has to stay correctly linked to its series and retrievable without re-deriving identity from scratch.
EBM mAIn PACS® implements native DICOM C-STORE, Q/R, MWL, and Storage Commitment against a documented services list. The platform is designed to support HIPAA compliance underneath whatever de-identification discipline runs before a study reaches the archive. That discipline is still the harder half of the problem, and it starts with treating an optional, self-reported attribute as a hint rather than a guarantee.
