PDF Conversion

How to Convert a PDF Table to Excel

Last updated · By the PDFSolution team

A PDF does not contain a table. It contains text positioned at coordinates, and a table is something a human infers from that arrangement. Any PDF-to-Excel conversion is therefore a reconstruction: the tool looks at where the words sit, guesses where the columns and rows are, and writes cells accordingly.

That works well on clean, machine-generated statements and reports. It works less well on documents with merged cells, wrapped text, multi-line headers or ruled boxes drawn as graphics. Treat the output as a strong first draft, not a finished spreadsheet.

Ready to do it now? Open the tool this guide describes.

PDF to Excel

Step by step

  1. 1

    Open PDF to Excel and add the file

    Extraction runs in your browser; the document is not uploaded.

  2. 2

    Limit the page range if you can

    If only pages 4 to 9 hold the table, restricting the range keeps stray headers and footnotes out of the sheet.

  3. 3

    Run the extraction

    Text positions are analysed and grouped into rows and columns, then written to a spreadsheet file.

  4. 4

    Open it and check the edges

    Look at the first and last row, any column of numbers, and anywhere the original had wrapped text. Those are where mistakes appear.

Useful tips

  • If the PDF is a scan, there is no text to position — run OCR first, then extract.
  • Numbers formatted with thousands separators or trailing minus signs may arrive as text. A quick find-and-replace in the spreadsheet usually fixes a whole column.
  • Where the original document exists as a spreadsheet somewhere, asking for that file is always better than reconstructing it.

Common problems

Everything landed in one column
The column gaps in the original were too narrow to detect, often because the table is drawn with graphic lines rather than spacing. Try extracting the text instead and splitting on the delimiter in your spreadsheet program.
Rows are split in two
Cells with wrapped text produce extra rows. Sorting that out by hand is normally faster than re-running the conversion.
The sheet is empty
The PDF has no text layer. It is a scan, and needs OCR before anything can be extracted.

Frequently asked questions

Will the formatting come across?
No. The point of the conversion is the data. Colours, borders and fonts are not reproduced.
Are formulas recovered?
No. A PDF only ever contained the calculated result, so that is all there is to extract.
Is the file uploaded?
No, extraction happens in your browser.
Can it handle multiple tables on one page?
Sometimes. Distinct tables separated by clear vertical space usually come out as separate blocks; tables side by side rarely do.
What about scanned bank statements?
Run OCR first to create a text layer, then extract. Accuracy depends heavily on the quality of the scan.

Related guides