Best starting settings: choose Editable layout, keep pictures and page graphics, enable editable table detection, and detect repeated content as Word headers and footers. Use the original page size and review any page with columns, forms or uncommon fonts.
Why PDF formatting changes in Word
A PDF page usually says, in effect, “draw this glyph at these coordinates.” It may not say that a large bold line is a heading, that six aligned text fragments form a table row, or that repeated text belongs to a footer. Word needs those higher-level relationships because it lays out paragraphs, tables and sections again whenever text, fonts, margins or printer metrics change.
This means “keep formatting” contains two goals that can conflict:
- Visual fidelity: keep objects close to their original positions.
- Editability: create real paragraphs, headings, tables, headers and footers that respond naturally in Word.
A page made from hundreds of absolutely positioned text boxes may look close but be unpleasant to edit. A clean flow of paragraphs may be easy to edit but no longer match a two-column page. Good reconstruction selects a useful balance for the document.
Choose layout mode by failure cost
| If this matters most | Start with | Why |
|---|---|---|
| Editing reports while retaining pages and tables | Editable layout | Uses page geometry while building Word-native paragraphs and tables |
| Keeping a form, flyer or diagram visually close | Precise layout | Positions text lines over preserved visual regions |
| Rewriting the content or improving accessibility | Flowing text | Prioritizes headings, paragraphs, tables and reading order |
Test the most difficult page, not merely page one. A title page can be simple while page 14 contains a three-column table, rotated label or nested chart legend that determines the correct mode.
A repeatable workflow for fewer formatting surprises
Start with a short page range that contains the document's hardest features: one ordinary text page, the densest table, a multi-column page, a page with a large image, and a page where the header or footer changes. Convert that range with Editable layout, then compare the result at the same zoom in PDFCore and Word. If only the visually complex page fails, repeat that page with Precise layout. Use Flowing text only when easier rewriting and reading order matter more than matching page breaks.
- Record the source. Note the PDF page size, page count and whether text can be selected. Keep the original unchanged.
- Choose one representative range. Include the hardest table, columns, graphics and repeated page regions rather than testing only the cover.
- Convert with pictures, tables and headers/footers enabled. For a scan, select an OCR language installed in Windows.
- Compare structure, not only appearance. Click inside tables, inspect Word's Navigation pane, and turn on formatting marks to reveal manual breaks and empty paragraphs.
- Measure repair work. Count incorrect line wraps, merged or split cells, misplaced images, header/footer errors and reading-order corrections.
- Freeze the winning settings. Convert the full document only after the representative range passes your acceptance checks.
Practical acceptance test: a DOCX is not “format preserved” merely because its first page looks similar. Confirm that text remains editable, tables are real Word tables, column order is correct, repeated content is placed intentionally, and Word opens the file without a repair prompt.
What PDFCore can and cannot reconstruct
PDFCore 8.0.1 can use PDF text geometry to create editable paragraphs, detect table regions, preserve pictures and non-text page graphics, classify repeated regions as headers or footers, and use locally installed Windows OCR languages for image-only pages. Those are reconstruction operations, not recovery of the original authoring file. A PDF normally does not retain Word style names, tracked changes, formulas, spreadsheet logic, original image crops or the exact font license. Complex forms, handwritten notes, overlapping artwork, unusual encodings and nested tables may still require manual repair.
For evidence you can inspect, the layout-mode comparison publishes one source PDF and three DOCX outputs. The quality methodology explains the checks used for editable text, tables, images, page ranges, headers and footers. These files demonstrate specific tested behaviors; they are not a promise that every PDF will reproduce perfectly.
Preserving tables as editable tables
PDF tables are often just lines plus text fragments. There may be no stored concept of a row or cell. A converter has to infer a grid from coordinates, spacing and ruling lines. Enable table detection when you want data to become Word cells, then verify:
- Column count and column order
- Merged header cells
- Rows that continue across a page break
- Decimal alignment and negative values
- Footnotes placed below rather than inside the table
Borderless tables are harder because ordinary paragraphs can also line up. Dense financial statements, cells containing multiple paragraphs and tables nested inside forms may need manual repair even when the page looks right.
Keeping columns and reading order
Columns are another inference problem. The converter must decide whether text to the right continues the current sentence, starts a second column, labels a chart, or belongs to a separate sidebar. Editable layout uses word and block geometry plus reading-order analysis to avoid joining widely separated text on the same row. Flowing text intentionally simplifies page positioning, so it is not the first choice when the exact column arrangement matters.
After conversion, read the first sentence at the bottom of each column and the first sentence at the top of the next. A page can look plausible while the underlying paragraph order is wrong.
Fonts, line wrapping and pagination
If Word cannot use the same font, it substitutes another with different character widths and line heights. A small difference accumulates: a sentence wraps early, a paragraph grows by one line, and later content moves to the next page.
- Install the matching font only if you have the right to use it.
- Keep the original page dimensions and margins.
- Avoid changing the Word printer or compatibility mode before reviewing.
- For a handoff, embed fonts in Word only when licensing permits.
- Expect OCR text from scans to use an available system font rather than reproduce a photographed font exactly.
Images, charts and background artwork
Keep “Preserve pictures and page graphics” enabled when the page contains logos, photos, charts, shaded panels or decorative artwork. PDFCore separates native text from non-text visual regions so it can avoid using a single full-page screenshot as the entire Word page. That preserves editable text while retaining graphics that Word cannot sensibly reconstruct.
Headers, footers and repeated text
A repeated company name or page number can be handled in three ways: detect it as a Word header/footer, keep it in the body, or remove it. Detection is usually best for reports because Word can repeat the content consistently and keep the main body cleaner. Keep it in the body when each page intentionally varies. Remove it only when the repeated material is not needed and you have verified that legal notices, exhibit identifiers or revision codes will not be lost.
Formatting from scanned pages
A scan contains appearance but little structure. OCR recovers character hypotheses and coordinates; it does not recover the original word processor’s styles. Precise layout can retain the photographed appearance around editable text. Editable layout is better when you want reconstructed paragraphs and tables. For poor scans, prioritize recognition accuracy before typography:
- Straighten rotated pages.
- Use a sharp source near 300 DPI for normal print.
- Remove dark borders and show-through if possible.
- Select the correct OCR language.
- Verify names, codes, totals and punctuation manually.
Diagnose a formatting problem quickly
| Symptom | Likely cause | First response |
|---|---|---|
| Lines wrap earlier | Font substitution or different page metrics | Check font, page size and margins |
| Two columns interleave | Ambiguous reading order | Use Editable or Precise layout; test the page alone |
| Table became loose text | No explicit table structure in PDF | Enable table detection and rebuild unusual merged cells |
| Text is not selectable | Source is scanned or graphics were preserved as an image | Choose the correct OCR language and reconvert |
| Page header appears in body | Repeated region was not classified | Choose header/footer detection or move it in Word |
| Background repeats behind text | Text was not fully separated from artwork | Try another mode and report a synthetic sample |
Preflight before sharing the Word file
- Turn on Word’s formatting marks and inspect unexpected empty paragraphs or manual breaks.
- Use the Navigation pane to confirm headings are real headings when structure matters.
- Click tables and images to confirm they are the intended object type.
- Search for a phrase near each column boundary to verify reading order.
- Compare page numbers, revision codes, totals and footnotes to the PDF.
- Run Word’s accessibility checker when the DOCX will be published accessibly.
- Save a PDF from the finished DOCX and compare it visually to the source if appearance is contractual.
Questions about formatting
Why does a converted PDF use many text boxes?
Positioned text boxes are one way to keep visual geometry. They are useful for precise layouts but less convenient for rewriting. Choose an editable or flowing mode when semantic editing matters more.
Can a converter restore the original Word styles?
Usually not exactly. A PDF often omits the source document’s style names and logical structure. Converters infer headings and paragraphs from appearance and geometry.
Should I convert the whole PDF at once?
After a small representative test. Use a page range containing the hardest table, scan or multi-column page, compare modes, then convert the full document.
Convert for your editing goal
PDFCore lets you choose editability, visual precision or flowing text instead of pretending one reconstruction fits every page.
Download PDFCore