This project processes a 62-page consolidated resume PDF (Consolidated Resumes.pdf, ~31MB) containing resumes from 38 UNC Charlotte students/professionals. The goal is OCR extraction, structured data building, and resume database creation.
Resume Sort/
├── CLAUDE.md # This file
├── Consolidated Resumes.pdf # Source: 62-page PDF (~31MB), pages are landscape-scanned
├── resume_database.csv # Final output: 38 people with name, summary, email, phone, accomplishment
├── build_resume_table.py # Generates resume_database.csv from hardcoded extracted data
├── ocr_resumes.py # OCR script (runs from Windows host, had networking issues)
├── ocr_inside_container.py # OCR script (runs inside Docker container - this one works)
├── arrow_pad.pyw # Unrelated: PyQt6 floating arrow pad utility
├── page_images/ # Raw PDF→PNG conversions (landscape, 200 DPI) - 62 files
├── page_images_rotated/ # Portrait-corrected PNGs for OCR - 62 files
└── ocr_output/ # DeepSeek-OCR results
├── page_001.txt ... page_062.txt # Individual page OCR text
├── all_resumes_ocr.txt # Consolidated text (215KB)
└── all_resumes_ocr.json # Structured JSON (210KB)
- 38 unique people across 62 pages (pages 42-62 are duplicates of pages 1-21)
- 36/38 have email, 33/38 have phone numbers
- Missing contact info: Kareem Saifeldawla (page 1) and Vincent Ma (page 10) - header cropped in scan
- All candidates are UNC Charlotte students/graduates (mix of BS, MS, PhD, and experienced professionals)
# 1. Start vLLM container (must use MSYS_NO_PATHCONV=1 for Git Bash on Windows)
MSYS_NO_PATHCONV=1 docker run -d --name deepseek-ocr \
--runtime nvidia --gpus all \
-v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
-v "/c/Users/Smirk/Downloads/Resume Sort:/workspace" \
-p 8000:8000 --ipc=host \
vllm/vllm-openai:latest \
--model deepseek-ai/DeepSeek-OCR \
--logits_processors vllm.model_executor.models.deepseek_ocr:NGramPerReqLogitsProcessor \
--no-enable-prefix-caching --mm-processor-cache-gb 0 \
--max-model-len 8192 --enforce-eager
# 2. Run OCR script INSIDE the container (not from Windows - networking doesn't work)
MSYS_NO_PATHCONV=1 docker exec deepseek-ocr python3 /workspace/ocr_inside_container.py
# 3. Cleanup
docker rm -f deepseek-ocr-
MSYS_NO_PATHCONV=1 is REQUIRED when running
dockercommands from Git Bash on Windows. Without it, Git Bash mangles Unix paths (e.g.,/workspace→C:\Program Files\Git\workspace), breaking volume mounts and exec paths. -
Run scripts INSIDE the container, not from the Windows host. Docker Desktop on Windows/WSL2 has unreliable localhost port forwarding - the vLLM API at
localhost:8000is accessible inside the container but Python on the Windows host hangs trying to connect. -
Images must be portrait orientation. The PDF pages are landscape-scanned (rotated 90° CCW). DeepSeek-OCR produces garbage/hallucinated repetitive text on sideways images. Rotating to portrait fixed everything - OCR went from ~31s/page of garbage to ~10s/page of clean text.
-
max_tokens must be < max_model_len. Setting
max_tokens=4096with--max-model-len 4096causes a 400 error because input tokens (prompt + image) need room too. Usemax_tokens=4096with--max-model-len 8192. -
DeepSeek-OCR model is ~6.2GB on GPU (3B params, BF16). Fits easily on RTX 4090 (24GB) but needs ~20GB+ free VRAM to avoid OOM during KV cache allocation. Close VRAM-heavy apps (After Effects, games, etc.) before starting.
-
--enforce-eageris needed on this setup to avoid CUDA graph memory overhead that causes OOM. -
Use
python3notpythoninside the vLLM container. -
The
"Free OCR."prompt works well for resume text extraction. The ngram logits processor withngram_size=30, window_size=90and whitelist tokens[128821, 128822]is required per the model docs.
- GPU: NVIDIA RTX 4090 (24GB VRAM)
- RAM: 64GB (Docker/WSL2 allocated ~31GB)
- OS: Windows 11 Pro
- Docker: Docker Desktop with NVIDIA runtime (WSL2 backend)
- Shell: Git Bash (hence the MSYS path issues)
- ~10 seconds per page average on RTX 4090 (eager mode, no CUDA graphs)
- Simpler pages (few lines): 2-5 seconds
- Dense pages (publications, detailed experience): 12-16 seconds
- Total for 62 pages: ~10 minutes