GSoC 2026
17th Century Spanish text recognition with transformer models
This project will improve the RenAIssance OCR pipeline for 17th-century Spanish historical documents by strengthening preprocessing, removing marginalia, optimizing TrOCR inference, and adding a flexible correction and deployment stack. The plan is to convert raw scanned pages into cleaner, layout-aware inputs, improve printed-text OCR, extend the system to handwritten OCR with the available handwritten datasets, and integrate local LLM/VLM-based correction with external fallback support when needed. The final deliverables will be a modular OCR pipeline, a deployed Flask application, benchmarking against existing OCR tools, and a paper-style technical report with reproducible documentation.
Project details
Technologies
Not listed in the archive