Optical Character Recognition (OCR) is automatic extraction of text from images within DLP. OCR is a feature within Advanced DLP.
Supported File Formats for OCR
"BMP_Fmt",
"GIF_87a_Fmt",
"GIF_89a_Fmt",
"ISO_JPEG2000_JP2_Fmt",
"ISO_JPEG2000_JPM_Fmt",
"ISO_JPEG2000_JPX_Fmt",
"JNG_Fmt",
"JPEG_2000_JP2_File_Fmt",
"JPEG_2000_PGX_Fmt",
"JPEG_File_Interchange_Fmt",
"JPEG_XR_Fmt",
"MS_DIB_Fmt",
"PNG_Fmt",
"TIFF_Fmt"
Considerations
At this time, English is the only supported language for extraction with OCR.
Some cases where the output may be impacted include the following:
- Small font size and low DPI
- Certain Background colors including highlighted text
- Handwriting in cursive or fonts that mimic cursive handwriting
- Unclear or blurry images and artifacts like noise in the image

