The RVL-CDIP dataset consists of scanned document images belonging to 16 classes such as letter, form, email, resume, memo, etc. The dataset has 320,000 training, 40,000 validation and 40,000 test images. The images are characterized by low quality, noise, and low resolution, typically 100 dpi.
Source: Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation LearningPaper | Code | Results | Date | Stars |
---|