Quick answer: Open the scanned PDF in PDFCore, choose OCR, select Make PDF searchable (recommended), choose the correct recognition language and save a new PDF. PDFCore keeps the scanned page image and adds an invisible selectable text layer locally.
Does the PDF need OCR?
Try selecting a word or searching for a phrase you can see. If nothing is found and the whole page selects like one photograph, the PDF is image-only and needs OCR. A born-digital PDF can also contain damaged or incomplete character mappings, but blindly OCRing good text can introduce new errors. Use native text when it is reliable.
Common OCR candidates include:
- Documents scanned from a flatbed or office copier
- Phone photographs saved as PDF
- Fax archives
- Historical material and printed forms
- Mixed documents with image-only appendices
Choose the right OCR output
| Output | What you get | Best use |
|---|---|---|
| Searchable PDF | The original page image plus an invisible recognized text layer | Archives, review, search and copy/paste while retaining the scan appearance |
| TXT | Recognized plain text without page design | Search indexing, quoting, analysis and lightweight notes |
| Editable Word | A DOCX reconstructed from recognized text and page geometry | Rewriting, correction and reuse |
| Recognize current page | A quick recognition result for the visible page | Testing language and image quality before a full job |
Extract text from an image without uploading it
If the source is a JPG or PNG rather than a PDF, first create a one-page PDF with Images to PDF. Choose Use original image size when preserving the photograph's proportions matters; choose A4 only when the output is meant to become a standard printed page. Open that PDF and use Recognize this page for a quick result, or To text when you need a TXT file. PDFCore sends the page to the OCR components already installed in Windows; it does not send the image to an online conversion service.
- Prepare one representative image. Rotate it upright and crop empty borders, but keep the untouched original.
- Create and open the PDF. For several photographs, put them in a deliberate order before conversion.
- Select the document language. The chooser lists OCR languages actually available to Windows. A language that is not installed cannot be selected.
- Recognize one difficult page first. Use a page containing small text, numbers, punctuation or multiple columns rather than an easy cover page.
- Choose the deliverable. Copy the quick result, export TXT for plain content, create a searchable PDF to preserve the image, or create Word when editing matters.
Windows language requirement: Microsoft documents that Windows OCR can use only recognition languages whose OCR language capability is installed. PDFCore reads that available-language list. Installing a display keyboard alone does not prove that the OCR component is present.
A five-minute reproducible OCR check
Create a small test image containing a heading, one paragraph, a date such as 2026-08-21, a negative amount such as -1,234.56, and a two-column list. Run it once with the correct installed language. Then compare the output character by character and record errors in four groups: letters, numbers, punctuation and reading order. Repeat with the same image after rotating it slightly or reducing its resolution. This shows whether an apparent improvement came from the OCR setting or simply from a better source image.
Do not report a single “accuracy percentage” unless the counting rule and source are fixed. For practical checking, count the total characters in the reference transcription, substitutions, insertions and deletions, and keep the test image with the result. Names, account numbers, dates, decimal separators and minus signs deserve manual review even when the surrounding paragraph looks correct.
Create a searchable PDF
- Open the source PDF. Preserve the original as an image record.
- Choose OCR. Select “Make PDF searchable (recommended).”
- Select the recognition language. The list is based on OCR languages installed in Windows. Choose the language used by most of the page text.
- Save to a new filename. A suffix such as
_searchablemakes the result easy to distinguish from the untouched scan. - Let OCR finish. PDFCore recognizes image-only pages and adds the text layer without sending the document to a web service.
- Search and select text. Test names, numbers, headings and a phrase from the middle of the document.
- Compare appearance and page count. The searchable PDF should still look like the original scan because the image remains the visible page.
Improve OCR accuracy before recognition
Resolution
A sharp scan around 300 DPI is a strong general starting point for ordinary printed text. Very low resolution loses character detail; extremely high resolution increases processing and storage without guaranteeing better recognition.
Orientation and skew
Rotate pages upright. Even a few degrees of skew can make line and table detection harder. If a phone photo uses perspective, correct the trapezoid before OCR.
Contrast and background
Text should be clearly separated from the page. Remove dark scanner borders, shadows, show-through and uneven lighting where possible. Avoid aggressive thresholding that closes small counters inside letters.
Language
Choose the correct Windows OCR language. A mismatched language model can confuse accented letters, punctuation and word boundaries. Mixed-language documents may need separate passes or extra review.
Source typography
OCR is less reliable on handwriting, dot-matrix output, decorative fonts, tiny footnotes, mathematical formulas, vertical writing and low-quality photocopies. These need focused human verification.
Windows OCR language availability
PDFCore uses languages exposed by the local Windows OCR platform. If a needed language does not appear, install the corresponding Windows language features/OCR capability, restart the app, and check again. Language availability depends on the Windows installation and edition; PDFCore does not download a private cloud OCR model.
This architecture keeps documents local, but OCR output still inherits the limitations of the installed Windows recognizer. Record the selected language and software version when reproducibility matters.
Convert a scan to editable Word
Choose “Convert scan to editable Word” or use PDF to Word and set the OCR language. Then select among three output goals:
- Editable layout: reconstructs paragraphs and tables while preserving graphics.
- Precise layout: keeps text placement closer to complex forms and photographed pages.
- Flowing text: simplifies layout for easier rewriting and reading order.
OCR does not know the original font, style names or table model. The Word output is an inference from pixels and coordinates. See the detailed format-preservation guide before processing a complex form or report.
Verify OCR where errors matter
Do not review OCR by merely glancing at the page image; a searchable PDF can look perfect while its invisible text layer contains mistakes. Test the recognized text itself:
- Search for a heading, a person’s name and a phrase containing punctuation.
- Copy a paragraph into a plain-text editor and compare it character by character.
- Check
0/O,1/l/I,5/S, decimal points, commas and minus signs. - Verify dates, amounts, account numbers, legal citations and medication or part codes.
- Check column order by selecting text across the page.
- Inspect tables row by row rather than trusting visual alignment.
OCR is not proof of correctness. For legal, financial, medical, historical or safety-critical use, preserve the scan and have a qualified person verify the transcription.
Troubleshooting
The page is searchable, but selection highlights the wrong area
The OCR coordinates may not align because of rotation, crop or unusual page scaling. Correct the page geometry before OCR and try the intended language again.
Words are recognized in the wrong order
Multi-column layouts, sidebars and tables can confuse reading order. For Word output, compare Editable and Precise layout. For plain-text extraction, plan to repair column boundaries.
Only some pages become searchable
Some pages may already contain native text while others are images. Test an image-only page directly. Check for very large, corrupt or protected pages and record the first page that fails.
The desired language is missing
Install the corresponding Windows language/OCR features, then reopen PDFCore. Availability is controlled by Windows.
Scan directly to a searchable PDF
With a compatible Windows Image Acquisition (WIA) scanner, PDFCore can choose 150, 200, 300 or 600 DPI, color/grayscale/black-and-white, flatbed or automatic document feeder, and duplex where the driver exposes it. Choose “Searchable PDF (OCR)” to create the scan and add text locally. Hardware and driver support determine whether feeder and duplex controls work.
Questions about OCR
Does OCR change how the scanned PDF looks?
A searchable-PDF workflow retains the scanned page image as the visible page and adds invisible text. The look should remain close, but page geometry and text alignment still need verification.
Can OCR read handwriting?
General printed-text OCR may recognize some clear handwriting but should not be relied on as an accurate handwriting transcription system.
Does local OCR work without internet access?
PDFCore uses installed Windows OCR components and does not upload the document. Required language capabilities must already be installed.
Make scans searchable without uploading them
Use PDFCore with the OCR language packs already installed on your Windows PC.
Download PDFCore