Searchable Scan PDF Maker
Image-only PDFs have no searchable letters. This tool recognizes words on each page locally, then adds invisible text at their measured positions for search and copying. It keeps the original scan image streams rather than recompressing them. You still need to review recognition and positioning errors.
Key features
- Run page-by-page OCR in the browser using bundled Apache-2.0 English and Korean Tesseract models
- Keep original PDF scan image streams and append invisible text at OCR word positions
- Report words below engine confidence 60 and words unsupported by the embedded font
- Download a PDF and JSON review report without uploading the source PDF
How to use
- Choose an image-only scan PDF up to 15 MiB and 5 pages.
- Select Korean + English, Korean only, or English only.
- Build the searchable PDF and wait for page-by-page OCR.
- Test search and selection in a PDF viewer, then compare low-confidence results with the original scan.
Use cases
- Find a date or name in an archival scan
- Select text in an image PDF with mixed Korean and English
- Make a reviewable OCR copy while retaining the original scanned page appearance
Frequently asked questions
Does this change scan image quality?
The original image streams inside the PDF are retained while text drawing commands are appended. No displayed page image is recompressed. PDF metadata and internal structure do change.
Is Korean OCR always accurate?
No. Blur, skew, small type, handwriting and complex tables can cause errors. Confidence is an engine score, not a calibrated probability. Compare output with the source scan.
Will search and selection align perfectly?
Invisible words are placed at OCR bounding boxes. Extraction positions were checked on fixtures, but individual viewers can differ in character selection, especially around Hangul glyphs.
Why are some PDFs rejected?
This version only handles image-only scans. Existing text, encrypted or form PDFs, rotated or cropped pages, files over 15 MiB, and more than 5 pages are outside its scope.
Are files or OCR data sent to an external service?
Processing runs in the browser. OCR code, English/Korean models and font are served by this site; no external OCR API is called. Initial loading can require about 25 MiB of static assets and device memory.
Privacy
PDF bytes and OCR words remain in browser memory and are not uploaded. You manage the downloaded PDF and JSON report. OCR models, PDF code and font are static assets served by this site.
Comments & questions