How to convert PDF to Word here
For a report, a letter, a contract or a CV, the default settings are the right ones, so you can go straight from adding the file to pressing Convert.
- Add the PDF, or several. Try with a sample report loads a two-page report with a logo, a table and lists; the link under the box loads a scanned letter to show text recognition.
- Under Output, choose Editable text (the default) or Exact layout. Leave Rebuild tables, Headers and footers and Pictures and drawings on unless you want plain text only.
- If the PDF is a scan, open Scanned pages and pick the language the document is written in. English, Spanish, French, German, Portuguese and Italian are available.
- Press Convert to Word. One file downloads by itself with the same name ending in .docx; several arrive together as pdf-to-word.zip.
- Scroll through the preview, which is drawn from the same structure the Word file is written from, and read the conversion report. Where a page came out badly, choose Use exact layout for this page; the file is rebuilt in a moment, and you download it again.
What happens to your PDF: it stays on your device
The PDFs people convert to Word are often the ones they cannot get the original of: a contract from the other side, a bank statement, a medical report, a form from an office. Online converters usually upload the file, convert it on their servers and keep it for a while under their privacy policy. Here the PDF is read into this tab’s memory and rebuilt by code running in the tab, and the Word file is created on your device. The file, its name, its text and any password you type are never sent anywhere.
Text recognition works the same way. The first time you convert a scan, the page downloads the recognition engine and the language file you chose from this site (for English, an engine of about 1.5 MB compressed and a 2.9 MB language file) and then reads the pages on your device. You can check it in the Network tab of your browser’s developer tools: you will see those files, the site’s consent and analytics requests, and no request that carries your document. Are online PDF converters safe? explains the test step by step.
Editable text or exact layout: which to choose
Editable text is what most people want. The converter works out the paragraphs, headings, lists and tables and writes them as real Word content: type into a paragraph and the lines reflow, add a row to a table and it grows, and the headings appear in Word’s navigation pane. The page will not look identical to the PDF, because Word lays the text out again with its own spacing, but it is a document you can work on.
Exact layout places every block of text in a frame at the position it had on the page, and every picture at its place, on a page of the same size. It looks like the PDF, which suits forms, certificates, brochures and pages with text scattered around a design. It is harder to edit, because each paragraph sits in its own box and does not flow into the next. Exact layout is designed for Microsoft Word: Google Docs and LibreOffice treat Word’s text frames differently, so for those editors choose Editable text.
You do not have to choose one for the whole file. After converting in Editable text, the preview lets you switch a single page to exact layout, such as a cover page, a form page or a page whose table was not recognised, and keep the rest editable.
Why PDF to Word is hard, and what to expect
A Word file stores paragraphs, styles and tables. A PDF mostly stores instructions to draw characters at positions on a page: this word here, in this font, at this size. It usually does not say where a paragraph ends, which line is a heading, that two columns of text are columns, or that a grid of lines and numbers is a table. A converter has to infer all of that from the positions, and every converter, paid or free, gets some of it wrong on some documents.
This one reads lines by their baseline, splits the page into columns where a gap runs down it, joins lines into paragraphs where the next word would not have fitted on the line before, spots headings from size and weight (and from the PDF’s own bookmarks when it has them), rebuilds numbered and bulleted lists, finds tables from their ruling lines or from columns that line up, and moves text repeated at the top and bottom of each page into Word’s header and footer. Hyphens at line ends are rejoined when a word was split across lines.
We measured it the hard way: 19 documents made in Word were exported to PDF by Word, converted back here, and opened again in Word. All 19 opened without Word’s repair prompt. On 17 of them at least 99.7% of the words came back in the right order; the other two were a page of equations and a page of Greek text, where the characters survive but are grouped into words differently. Expect a clean draft for text documents, not a perfect copy of a designed layout.
Tables, columns, headers and lists
Each of these is detected on its own, and the report says what was found. What to look at in Word afterwards:
- Tables with ruling lines are rebuilt from the lines themselves, so each piece of text lands in its own cell, merged cells stay merged and shaded cells keep their colour. A table that runs over several pages becomes one Word table, and a header row repeated on each page becomes a real header row.
- Tables without lines are found from columns that line up across several rows. When the converter is unsure, it keeps the text as lines with tab stops instead of forcing a table, and the report says so; check these first.
- Two- and three-column pages are read column by column, so the text comes out in reading order and flows as one column in Word. Add Word’s own columns afterwards if you need them.
- Running headers and footers move into Word’s header and footer, a different first page included, and a page number becomes a live field, so it stays right when you edit.
- Bulleted and numbered lists become real Word lists with levels, so new items number themselves. Lists whose numbering the converter cannot follow keep their numbers as typed.
Scanned PDFs and text recognition (OCR)
A scanned PDF is a picture of each page, with no text inside. You can tell by trying to select a word in your PDF viewer: if you can only draw a box over the page, it is a scan. The converter checks every page and only runs text recognition on pages that are pictures, so a digital PDF with one scanned page attached is handled page by page.
Recognition uses Tesseract, the open-source OCR engine, running in your browser. Each scanned page is drawn at 300 dots per inch in grey, straightened if it was scanned at a slight angle, and read with the position of every word, which then goes through the same analysis as a digital page. The result is paragraphs, lists and headings rather than one block of text, and the report gives the average recognition confidence.
On three test scans made at 300 DPI with a slight tilt and speckle, 99.0%, 98.3% and 91.4% of the words were recognised correctly; the lowest was a page with a coloured table and small print. A page took 4 to 8 seconds on a desktop computer; a phone is several times slower, so a 50-page scan can take many minutes there. Handwriting is not recognised, and right-to-left scripts, Chinese, Japanese and Korean are not included yet. Always proofread names, amounts and dates.
Also make a searchable PDF returns the original PDF with an invisible layer of recognised text laid over each scanned page. It looks exactly the same, but you can search it, select and copy text, and a screen reader can read it. Many people who look for a way to convert a scan to Word only need that.
Fonts in the Word file
A PDF usually carries its fonts inside it, often cut down to the characters it uses and renamed, and a Word file cannot use them. The converter therefore sets the text in a font the reader of the Word file is likely to have: Arial, Times New Roman, Calibri, Cambria, Courier New, Georgia, Verdana and the other common Windows and Office fonts are kept by name, and anything else becomes the closest of them by style (serif, sans-serif or fixed-width). The report lists each replacement, such as a LaTeX paper’s Computer Modern becoming Cambria.
Bold, italic, size, colour, superscript and subscript are kept. Letter spacing and exact line positions are not, because Word lays out the text again; that is what makes the text editable.
Password-protected and restricted PDFs
When a PDF is locked with a password to open it, the page asks for that password beside the file; it unlocks the file in this tab only and is never stored.
Some PDFs open without a password but carry restrictions set by their author, such as no copying of text. Converting to Word copies all the text, so a PDF that forbids copying is refused with an explanation, and the converter never removes or works around the restriction. Ask the sender for an unrestricted copy or for the original document.
Fixing the result in Word: the five checks
A few minutes of checking turns a converted draft into a clean document. Look at these five things, in this order:
- Headings: open View, Navigation Pane. Every section title should be listed; apply the Heading 1, 2 or 3 style to any that are missing, and a table of contents will then work.
- Tables: click inside each table and check that the text sits in the right cells. Unsure tables kept as tabbed text can be turned into tables with Insert, Table, Convert Text to Table.
- Page breaks: text flows again in Word, so a heading can end up alone at the bottom of a page. Remove stray empty paragraphs and use Keep with next on headings.
- Headers and footers: double-click the top of a page to open them, and check the page number field and a different first page.
- Hyphenation and line ends: search for words with a hyphen in the middle that should be one word, and for line breaks inside paragraphs, with Find and Replace.
How to edit a PDF goes through each fix with a worked example, and explains the cases where converting is not the best way to change a PDF.
When not to convert
Converting to Word is the right move when you need to rewrite or reuse text. For many jobs it is not needed, and a direct tool keeps the PDF exactly as it is. To change the order of pages or drop some, use Split PDF; to put files together, Merge PDF; to turn sideways pages upright, Rotate PDF; to make a file smaller, Compress PDF. To fill in a form or add a signature, the PDF viewer built into your browser or computer does it without changing the rest of the document. When a picture of a page is enough, PDF to JPG saves it as an image.
When you have finished editing in Word, Word to PDF turns the document back into a PDF in the browser, with real text, embedded fonts, links and bookmarks.
Your data
Only your choices are remembered, in this browser’s local storage under dtc.filetools.v1: the kind of Word file, the table, header and picture switches, text recognition and its language, and the searchable PDF option. The PDFs, their names, their text, the Word files and any password stay in the tab’s memory and are gone when you close it. Nothing is kept once the tab is closed, and there is no account or history of your conversions.
General information about converting PDF files to Word. The Word file is rebuilt from the PDF by inference, and text recognition makes mistakes, so check the result against the PDF before relying on it, especially names, figures, dates and anything legal or financial, and keep the original. Only convert documents you have the right to copy. Microsoft Word is a trademark of Microsoft; this tool is not affiliated with Microsoft.