DO TRUNG DUNG

OCR / Automation

OCR Searchable PDF Workflow

OCR workflow from my internship that turns scanned documents into searchable PDFs while keeping recognized text aligned with the original page.

Project access

View live project

Public demo unavailable

A public live/demo URL is not currently available for this project. You can still review the case study and interface evidence below.

Role / contribution
Development Intern
Area
OCR / Automation
Project type
Work project / internal product
OCR to searchable PDF workflow diagram

About the project

Turning scanned documents into searchable PDFs

During my internship, I worked on a document-processing workflow that detects text regions, recognizes them with Tesseract OCR and EasyOCR, and places the recognized text back at the corresponding PDF coordinates.

The result keeps the scanned page visually unchanged while adding a searchable text layer.

Workflow

01

Text-region localization

Locates text regions on the scanned page and keeps their coordinates for the later PDF text layer.

Region coordinates

02

OCR recognition

Runs Tesseract OCR and EasyOCR on the detected text regions.

Tesseract OCR + EasyOCR

03

Searchable PDF layer

Places recognized text back at the matching page coordinates to make the PDF searchable.

PDF text-layer placement

OCR
Tesseract / EasyOCR
Output
Searchable PDF
Frontend
React
Automation
Selenium / Python Requests