ID
ReadyIndic document intelligence benchmark
A reproducible OCR, retrieval, and answer extraction benchmark for synthetic English, Hindi, and Kannada administrative documents.
- Python 3.11
- Tesseract 5
- scikit-learn
- Pillow
- Docker
- Project ID
- GP-DA-1GYWHUI
- Domain
- Data and AI
- Complexity
- Advanced
- Catalogued
- 21 Aug 2026
- Licence
- Single buyer
- Repository
- Private handover
- Buyer record
- Never public
Included in the project
- Python source and command-line tools
- Synthetic English, Hindi, and Kannada benchmark documents
- Tesseract OCR, TF-IDF retrieval, and extractive answer pipelines
- Retained CSV and JSON results, nine figures, and offline dashboard
- Complete source code in a private GitHub repository
- 83-page project documentation in PDF and editable Word formats
- 21-page setup and usage guide in PDF and editable Word formats
- 55 annotated references