Capturing Video and Findings in Endoscopy Reporting

Glowing line-art illustration on a deep navy field of a continuous ribbon of video frames unspooling from a scope tip, one still frame breaking free of the ribbon and resolving into a branching tree of coded findings, a single warm red point marking the captured frame.

An endoscopy procedure runs on video, not stills. A clinician watches a continuous feed for the length of the exam, captures a handful of frames and clips at the moments that matter, then produces a report once the scope comes out. Endoscopy reporting software has to carry the feed, the captured video and the finished report as one linked object. DICOM turns out to define that job in more specific detail than most product teams expect.

Three Endoscopy IODs and One Modality Code

A radiology study is a fixed set of slices, captured once, that a radiologist pages through later. An endoscopic procedure behaves more like a broadcast: the clinician is watching a live feed in real time and deciding, moment by moment, what is worth keeping. DICOM’s Visible Light Image IODs exist for sources built around exactly this kind of capture. The standard names the source equipment directly: “Rigid and flexible endoscopy equipment” sits first on its own list of the cameras and sensors this family of objects was written for.

Within that family, DICOM does not treat endoscopy as one object type. It defines three, because a still frame, a saved clip and a live transmission are three different engineering problems, not three names for the same file.

  • The VL Endoscopic Image IOD covers a single still frame. DICOM’s own description: “the Attributes of Single-frame VL Endoscopic Images.”
  • The Video Endoscopic Image IOD covers a stored clip. DICOM’s description: “the Attributes of Multi-frame Video Endoscopic Images.”
  • The Real-Time Video Endoscopic Image IOD covers a live transmission. DICOM’s description: “the Attributes of Multi-frame Video Endoscopic Images transmitted in real-time.”

Both the still and the stored-clip versions carry the same modality code. Content Constraints for each IOD are explicit: “The Value of Modality (0008,0060) shall be ES.” A viewer or archive that recognizes ES as endoscopy, rather than a generic image, can route, tag and query a captured frame the same way it already handles a CT or an MR series.

That single code has a documented limit worth building around rather than discovering later. DICOM’s own note on the Video Endoscopic Image IOD explains why: the same acquisition equipment does laparoscopy and colonoscopy alike, so “Modality is not useful to distinguish one type of endoscopy from another when browsing a collection of Studies.” The standard’s answer is to lean on two other fields, Procedure Code Sequence and Anatomic Region Sequence, populated on every instance rather than left to whatever the modality happened to default to.

Reporting software that only stores ES has a folder full of endoscopy. Reporting software that also captures the procedure and anatomy codes can tell a colonoscopy from a bronchoscopy without opening the file.

Endoscopic video carries something a radiology series never does, too. The standard’s note on the Video Endoscopic Image IOD says the video “may include audio channel(s) for acquiring Patient voice or physiological sounds, healthcare professionals’ commentary, or environmental sounds.” A CT or MR viewer has no audio track to play, sync or store. An endoscopy viewer built from the same assumptions drops the commentary a clinician recorded while the procedure was happening, which is exactly the kind of context a report is supposed to capture.

A Saved Clip Is Not a Longer Still Image

Treating a stored endoscopy clip as a bigger version of a still frame breaks the moment engineering has to move it. A video object does not carry pixel values the way an uncompressed CT slice does. It encapsulates an already-compressed bitstream, wrapping the output of a video encoder such as MPEG2 inside the object rather than storing pixel-by-pixel data in native format.

DICOM’s own Transfer Syntax definitions spell out what that means at the fragment level. An MPEG2 transfer syntax UID corresponds to “MPEG2 Main Profile / Main Level option of the ISO/IEC MPEG2 Video standard encoded in one or more Fragments.” A CT series is a stack of independent instances, each with its own header. A video clip is one instance holding a compressed stream broken into fragments end to end.

That distinction matters past the storage line item too. Sizing medical imaging storage for real study volume already establishes that a cine loop or procedure recording lands as a single multi-frame object whose size scales with frame count and duration. An endoscopy tower generating hours of procedure video a day is not a marginal add to a plan built around cross-sectional slice counts. It is its own encoding and throughput profile, built from fragments rather than frames counted one at a time.

The Secondary Capture Trap

Here is where a lot of integration work quietly goes wrong. DICOM defines a generic fallback for exactly the kind of hardware an endoscopy tower often is: a video interface with no native DICOM encoder behind it. The Secondary Capture Image IODs specify “images that are converted from a non-DICOM format to a modality independent DICOM format.” Among the equipment types DICOM names as a source for this path: “Video interfaces that convert an analog video signal into a digital image.”

Secondary Capture is tempting precisely because it asks nothing of the source device. Wrap whatever comes out of the video interface, tag it as an image, and move on. DICOM’s own description of the original single-frame version of this IOD is blunt about what that buys. It calls it “one relatively unconstrained, single-frame, Secondary Capture Image IOD”, retained because it is still in common use, not because it is a good fit for a specific modality.

Unconstrained is the problem. A frame written through the VL Endoscopic Image IOD or the Video Endoscopic Image IOD carries Modality fixed at ES, plus a mandatory VL Image Module and Acquisition Context Module. The video object adds mandatory Cine and Multi-frame Modules on top of those. An archive can index all of that the way it indexes every other named modality.

Secondary Capture requires none of those four modules. Its own SC Equipment Module makes Modality (0008,0060) a Type 3 Attribute, a definition the standard says overrides the one in the General Series Module. Two vendors’ towers can both output valid Secondary Capture, arrive with no modality value at all, and carry nothing that ties a frame to a procedure or an anatomic region.

Getting a Finding From the Procedure to the Report

Capturing the video is only half the job. The other half is connecting what the clinician saw to what ends up in the report, and DICOM alone does not own that connection. IHE runs a dedicated domain for it: “The IHE Endoscopy domain was formed in 2005 in Japan and addresses information sharing, workflow and patient care in digestive endoscopy.”

Two of its named profiles map directly onto the capture-to-report path. Endoscopy Image Archiving “defines a workflow focusing on the image information communication during the endoscopy procedure”, covering how captured frames and clips move from the acquisition modality into an archive.

Endoscopy Report and Pathology Order covers what happens next: it specifies “a series of workflows where gastroenterological endoscopy is conducted on the order from the hospital information system located outside of the endoscopy department and the endoscopy report returned to the system.”

That second profile names the actor doing the writing. Once the procedure finishes, “the Report Creator provides the observation report to the Order Placer”, the hospital-side system that generated the original order. The profile builds on both DICOM and HL7 together, because the image side and the order-and-report side are different standards doing different jobs on the same procedure.

Encoding what goes inside that report as coded, queryable data rather than a paragraph of prose is a separate problem again. DICOM Structured Reporting already solves it at the content-tree level for any specialty, endoscopy included. A finding tied to the still frame it was drawn from uses the same SCOORD-anchored mechanism a radiology measurement does. What endoscopy adds on top is the workflow layer above that content tree: getting the report back to the ordering system that has been waiting on it.

Judging Endoscopy Reporting Software: Capture and Report Separately

A partner evaluating a platform’s endoscopy support should be asking about the capture layer and the report layer separately. A platform can pass one and fail the other without anyone noticing until a real procedure runs through it.

  • Does the platform write the VL Endoscopic, Video Endoscopic or Real-Time Video Endoscopic Image IOD, or does an unconstrained video interface fall back to plain Secondary Capture.
  • Is captured video handled as an encapsulated, fragmented bitstream with its own throughput profile, not budgeted as if it were another CT series.
  • Does a finding captured during the procedure reach the ordering system through a documented workflow, or does someone retype it once the report is finished.
  • If an archive consolidates studies across departments, does it treat endoscopic video as its own object type, since a department that captures video and still frames every procedure is exactly the kind of source that archive was built to absorb correctly.

All of that still applies once a partner’s tower ships to a GI or pulmonary department. The dedicated IODs, the encapsulated transfer syntaxes and the IHE workflow profiles are already written, and none of them treats endoscopic video as a radiology image wearing a different label. A platform that writes Secondary Capture for its video and re-keys its findings by hand is not failing at endoscopy. It is simply not reading the sections of the standard that already define the job.

Those sections are the checklist a partner takes to any imaging platform it is considering building on, including ours: EBM Fabric™ is the partner platform behind EBM mAIn PACS®.