TableVision OCR — Convert Images & PDFs to Structured Data
An OCR and document intelligence system that extracts structured and unstructured data — including tables — from financial and medical PDFs/images.

Workflow & Architecture
How this system works
Financial or medical PDFs and images submitted for processing.
Text, bordered/borderless tables, and handwriting extracted.
Raw extraction normalized into consistent row/column structures.
Clean structured spreadsheet ready for downstream systems.
Overview
Combined modern OCR with Vision-Language Models to accurately extract text, bordered/borderless tables, and handwritten content from PDFs and images, converting the results into clean, structured Excel output.
The Problem
Financial and medical documents contain tables and handwriting that standard OCR handles poorly, blocking automated data entry.
The Solution
Built an OCR and VLM-powered extraction system supporting bordered/borderless tables and handwritten content, exporting clean structured Excel files.
Key Features
- Handles both bordered and borderless tables
- Extracts handwritten content alongside printed text
- Supports financial and medical document types
- Outputs clean, structured Excel files ready for downstream use
Outcomes & Business Value
- Automates conversion of unstructured documents into structured data
- Reduces manual data entry for financial and medical workflows
- Handles complex table layouts standard OCR misses


