Payhawk extracts invoice data by reading the document as an image. As a result, most files are processed automatically, including copy-protected PDF files. If data is missing or incorrect, the most common causes are password-protected PDF files, low-quality images, or incorrect supplier settings such as a default VAT rate.
Password-protected PDF files
PDF files that require a password to open cannot be read by the data extraction tools. To process such invoices, you can:
Enter the required data for the invoice manually.
Ask the supplier to send a copy of the invoice without password protection.
Copy-protected PDF files
Some PDF files open normally but show random symbols when you copy and paste their text. Payhawk reads such files as images, so the data is extracted correctly in most cases. If a value is missing or incorrect, upload the document again or enter the value manually.
Low-quality images
Blurry photos or low-quality images can also prevent accurate data extraction. Ensure that documents are scanned or photographed clearly, with all text legible and properly aligned.
Correcting extracted data
If you notice that the OCR tool has extracted incorrect information for any of the expense fields, enter the required data from the invoice manually. You can also try re-uploading the document to improve OCR accuracy. If a default VAT rate is saved for the supplier, verify and adjust the supplier settings as needed.
Using the Financial Controller AI Agent
If your company uses the Financial Controller AI Agent, the Agent reads the uploaded document and fills in the amounts, the supplier, and the dates on the expense. Copy protection does not affect the Agent, only password-protected PDF files cannot be read.