OCR / Automation
OCR Searchable PDF Workflow
OCR workflow from my internship that turns scanned documents into searchable PDFs while keeping recognized text aligned with the original page.
Project access
View live project
Public demo unavailable
A public live/demo URL is not currently available for this project. You can still review the case study and interface evidence below.

About the project
Turning scanned documents into searchable PDFs
During my internship, I worked on a document-processing workflow that detects text regions, recognizes them with Tesseract OCR and EasyOCR, and places the recognized text back at the corresponding PDF coordinates.
The result keeps the scanned page visually unchanged while adding a searchable text layer.
Workflow
01
Text-region localization
Locates text regions on the scanned page and keeps their coordinates for the later PDF text layer.
Region coordinates
02
OCR recognition
Runs Tesseract OCR and EasyOCR on the detected text regions.
Tesseract OCR + EasyOCR
03
Searchable PDF layer
Places recognized text back at the matching page coordinates to make the PDF searchable.
PDF text-layer placement
- OCR
- Tesseract / EasyOCR
- Output
- Searchable PDF
- Frontend
- React
- Automation
- Selenium / Python Requests