7 Ways to Get Better OCR Results from Photos and Scans
August 8, 2026 · updated August 15, 2026
When OCR output comes back full of garbage, people blame the engine. Fair enough — but most recognition errors are decided before the software runs, by the quality of the input. Here’s what actually matters, roughly in order of impact.
1. Resolution: give each letter enough pixels
The single biggest factor. As a rule of thumb, lowercase letters should be at least 20 pixels tall in the image. If you’re photographing a page, get closer or use optical zoom; if you’re screenshotting, zoom the page to 150% first. Upscaling a blurry image afterwards does not add information — capture it bigger.
2. Contrast: dark text, light background
Engines are trained mostly on black-on-white text. Faded thermal-paper receipts, gray-on-gray UI text, and watermarked documents all hurt. If you control the source, boost contrast; if you’re photographing, avoid dim light — noise from high ISO looks like texture to the model.
3. Keep it flat and straight
A small tilt is fine; a 30° perspective shot of a page is not. Photograph documents from directly above, and flatten curved book pages as much as you can. For whiteboards, stand square to the board rather than shooting from a seat at the side.
4. Kill the glare
Glossy paper, laminated cards and screens photographed with a flash all produce highlights that erase the text underneath. Turn the flash off and use ambient light, angling slightly until the glare moves off the letters.
5. Screenshots beat photos of screens
Photographing a monitor adds moiré, pixel grids and focus problems. If the content is on a screen you control, take a screenshot instead — it’s pixel-perfect, and accuracy jumps. Then paste it into the screenshot to text tool.
6. Crop to the text you want — or select it after
OCR reads everything it finds: navigation, buttons, captions, posters in the background. Cropping to the relevant region speeds things up and keeps the output clean. Don’t crop too tight: clipping the tops or bottoms of letters costs accuracy.
If you already dropped a full page or a messy screenshot, you don’t have to recapture it. On the preview, drag a box around the paragraph you want, or click a block to auto-detect it. Only that region is recognized.
7. Know the hard cases — and what the engine can read
The default engine is built for printed English, German, French, Japanese and Korean. Mixed lines are fine: a German abstract with English citations, or a Japanese dialog with English buttons, usually comes through in one pass. Korean loads a second recognizer only when the first pass looks weak.
These still give any OCR a hard time, local or cloud:
- Handwriting — printed-text models degrade badly on cursive.
- Stylized or decorative fonts — logos, certificates, fancy menus.
- Low-contrast overlays — subtitles over busy video frames.
- Very small type — footnotes and 8pt scans, even when the language is supported.
Tables are less of a dead end than they used to be: if the boxes look grid-like, the image to text tool also tries to reconstruct the table. Proofread the cells; the text view stays editable.
The part the tool does for you
Raw OCR still arrives with mid-sentence line breaks and hyphen-ated words from the original layout. Clean up text joins those lines and strips stray hyphens, while keeping the raw output one click away.
If a quote, parenthesis or accent still lands in the wrong place, edit the line in the result box. That is faster than recapturing the image.