
by Marcia Spicer
One thing almost every digitization client has in common is a desire to share their collection. That’s great, and it’s a goal digitization works toward almost naturally. But someone has to think about the how. Yes, we want to share, facilitate research, create opportunities for spontaneous discovery of the collection… but how?
For collections that want to contain full-text document search, a key part of the how is OCR.
Learn more about Optical Character Recognition in our explainer.
Like most things, OCR (Optical Character Recognition) exists on a spectrum. Simply checking the OCR box doesn’t mean that the job is done. Without fully considering the ways OCR can work for your collection, it is easy to make these five crucial OCR mistakes.
-
“Automated OCR will do the job.”
-
“We’ll skip the review of proper names.”
-
“Numbers aren’t essential.”
-
“Since OCR is what’s really important, let’s use the scans we already have.”
-
“Just the text, thanks.”