Skip to content

Latest commit

 

History

History
82 lines (64 loc) · 4.66 KB

File metadata and controls

82 lines (64 loc) · 4.66 KB

Resume Sort Project

Overview

This project processes a 62-page consolidated resume PDF (Consolidated Resumes.pdf, ~31MB) containing resumes from 38 UNC Charlotte students/professionals. The goal is OCR extraction, structured data building, and resume database creation.

Project Structure

Resume Sort/
├── CLAUDE.md                    # This file
├── Consolidated Resumes.pdf     # Source: 62-page PDF (~31MB), pages are landscape-scanned
├── resume_database.csv          # Final output: 38 people with name, summary, email, phone, accomplishment
├── build_resume_table.py        # Generates resume_database.csv from hardcoded extracted data
├── ocr_resumes.py               # OCR script (runs from Windows host, had networking issues)
├── ocr_inside_container.py      # OCR script (runs inside Docker container - this one works)
├── arrow_pad.pyw                # Unrelated: PyQt6 floating arrow pad utility
├── page_images/                 # Raw PDF→PNG conversions (landscape, 200 DPI) - 62 files
├── page_images_rotated/         # Portrait-corrected PNGs for OCR - 62 files
└── ocr_output/                  # DeepSeek-OCR results
    ├── page_001.txt ... page_062.txt  # Individual page OCR text
    ├── all_resumes_ocr.txt            # Consolidated text (215KB)
    └── all_resumes_ocr.json           # Structured JSON (210KB)

Key Data Facts

  • 38 unique people across 62 pages (pages 42-62 are duplicates of pages 1-21)
  • 36/38 have email, 33/38 have phone numbers
  • Missing contact info: Kareem Saifeldawla (page 1) and Vincent Ma (page 10) - header cropped in scan
  • All candidates are UNC Charlotte students/graduates (mix of BS, MS, PhD, and experienced professionals)

OCR Pipeline (DeepSeek-OCR via vLLM Docker)

What Works

# 1. Start vLLM container (must use MSYS_NO_PATHCONV=1 for Git Bash on Windows)
MSYS_NO_PATHCONV=1 docker run -d --name deepseek-ocr \
    --runtime nvidia --gpus all \
    -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
    -v "/c/Users/Smirk/Downloads/Resume Sort:/workspace" \
    -p 8000:8000 --ipc=host \
    vllm/vllm-openai:latest \
    --model deepseek-ai/DeepSeek-OCR \
    --logits_processors vllm.model_executor.models.deepseek_ocr:NGramPerReqLogitsProcessor \
    --no-enable-prefix-caching --mm-processor-cache-gb 0 \
    --max-model-len 8192 --enforce-eager

# 2. Run OCR script INSIDE the container (not from Windows - networking doesn't work)
MSYS_NO_PATHCONV=1 docker exec deepseek-ocr python3 /workspace/ocr_inside_container.py

# 3. Cleanup
docker rm -f deepseek-ocr

Critical Lessons Learned

  1. MSYS_NO_PATHCONV=1 is REQUIRED when running docker commands from Git Bash on Windows. Without it, Git Bash mangles Unix paths (e.g., /workspace → C:\Program Files\Git\workspace), breaking volume mounts and exec paths.

  2. Run scripts INSIDE the container, not from the Windows host. Docker Desktop on Windows/WSL2 has unreliable localhost port forwarding - the vLLM API at localhost:8000 is accessible inside the container but Python on the Windows host hangs trying to connect.

  3. Images must be portrait orientation. The PDF pages are landscape-scanned (rotated 90° CCW). DeepSeek-OCR produces garbage/hallucinated repetitive text on sideways images. Rotating to portrait fixed everything - OCR went from ~31s/page of garbage to ~10s/page of clean text.

  4. max_tokens must be < max_model_len. Setting max_tokens=4096 with --max-model-len 4096 causes a 400 error because input tokens (prompt + image) need room too. Use max_tokens=4096 with --max-model-len 8192.

  5. DeepSeek-OCR model is ~6.2GB on GPU (3B params, BF16). Fits easily on RTX 4090 (24GB) but needs ~20GB+ free VRAM to avoid OOM during KV cache allocation. Close VRAM-heavy apps (After Effects, games, etc.) before starting.

  6. --enforce-eager is needed on this setup to avoid CUDA graph memory overhead that causes OOM.

  7. Use python3 not python inside the vLLM container.

  8. The "Free OCR." prompt works well for resume text extraction. The ngram logits processor with ngram_size=30, window_size=90 and whitelist tokens [128821, 128822] is required per the model docs.

Hardware

  • GPU: NVIDIA RTX 4090 (24GB VRAM)
  • RAM: 64GB (Docker/WSL2 allocated ~31GB)
  • OS: Windows 11 Pro
  • Docker: Docker Desktop with NVIDIA runtime (WSL2 backend)
  • Shell: Git Bash (hence the MSYS path issues)

Performance

  • ~10 seconds per page average on RTX 4090 (eager mode, no CUDA graphs)
  • Simpler pages (few lines): 2-5 seconds
  • Dense pages (publications, detailed experience): 12-16 seconds
  • Total for 62 pages: ~10 minutes