Docparser was primarily designed to process "transactional business documents", such as invoices, purchase orders, work orders, etc. We try to support all common file formats used to exchange business documents between entities, as well as file formats used by scanning software.
At the time of writing, Docparser can read the following file formats:
- Native PDF documents with text
- Scanned PDF documents with images only
- Microsoft DOCX and DOC files
- JPG image
- PNG image
- TIFF image
- CSV
- Microsoft XLS and XLSX files
- TXT
- XML
Note: Currently, for XLS and XLSX files, we are only able to parse the formula in a cell. Docparser does not currently display the results of a formula.
Also, it is only possible to extract the data from the first worksheet in a XLS or XLSX file.
The only solution to resolve both situations would be to save each worksheet as a PDF and upload them into Docparser.