INVOLVE RESUME

How ATS actually works

DOCX or PDF: What Actually Parses, by System

Text PDFs parse fine in modern systems and Taleo still breaks the tie toward DOCX. Plus the wrapped bullet that invents jobs you never had.

Published 2026-10-08Involve Resume

The file format question gets a confident answer from everybody and evidence from almost nobody. Here is what the vendors actually publish, what breaks in practice, and the one failure mode nobody warns you about because it is invisible from the outside.

What the vendors document

SystemDocumented resume formatsSource
Taleo (Oracle).doc, .docx, .txt, .rtf, .pdf, .htm, .html, .wpd, .odt, and for parsing also .xls and .xlsxOracle Help Center
Workable.doc, .docx, .pdf, .rtf, .html, .odtWorkable Help
SmartRecruitersRecommends .docx "for maximum parsing compatibility"; PDF acceptable if exported from a word processorSmartRecruiters glossary
AshbyFull-text search runs over the resume; format is not the constraint, extractable text isAshby docs

Two things follow immediately. PDF is not banned anywhere. And the only vendor that expresses a preference in writing expresses it for DOCX.

The tie-breaker is Taleo, and it is a size limit

Oracle publishes a detail that almost no resume advice mentions. A resume submitted to Taleo for parsing "cannot exceed 100 kilobytes or the size defined by the system administrator".

One hundred kilobytes. A plain single-column DOCX is comfortably inside it. A one-page PDF exported from a design tool, with an embedded photograph and two subset fonts, is not. That is not a formatting opinion, it is arithmetic, and it explains a category of application that vanishes without an error message.

Taleo is also the oldest parser you are likely to meet and the one whose behaviour was built around native Word files. So the rule we use is narrow and testable: if both formats recover your record equally well, send the DOCX, because the DOCX has the wider margin against the strictest reader in the set.

What actually breaks, in order of how often we see it

  1. A scanned or image-based PDF. There is no text layer, so there is no text to extract. SmartRecruiters' guidance says it outright: do not scan the PDF as an image.
  2. A PDF exported by a design tool. Text can be stored as vector outlines or with mangled character mapping. It looks perfect and copies out as gibberish.
  3. Two columns. A sidebar in Word is usually a table cell, and cells are read in sequence rather than side by side. Your skills column arrives in the middle of a job title, or does not arrive.
  4. Headings the parser does not recognise. A parser matches section headings against a vocabulary. "Career Highlights" is a heading a person understands; if the vocabulary does not have it, everything underneath it is lost.
  5. Text inside a header or footer. Frequently skipped. Your phone number lives there in about one CV in ten.
  6. Dates a parser cannot place. A role with no readable date range cannot go on a timeline, so it counts for zero towards a years-of-experience filter.

How to tell whether your PDF has a text layer

Thirty seconds, no tools:

  1. Open the PDF in any reader.
  2. Select all, copy.
  3. Paste into a plain text editor, not into Word.
  4. Read what arrived.

If your name, your employers and your dates are all there as ordinary words, in a sensible order, the text layer is fine and PDF is a legitimate choice. If you get boxes, ligature damage, letters in the wrong order, or no text at all, the file has no usable text and no parser will rescue it.

Do the same test on the DOCX by saving a copy as plain text. Comparing the two is a more useful five minutes than any amount of argument about formats, because it is your document rather than a generalisation.

One outcome worth knowing about: an unreadable resume is not always a rejection. Greenhouse documents that when a candidate cannot be assigned a match score, for example because the resume cannot be read by the system, they are flagged as needing manual review. That is a better outcome than a silent zero. It also means leaving the ordered queue and depending on a human having time.

The failure nobody warns you about: the wrapped bullet

This one we found in our own code, on our own export, and it is the reason this cluster exists.

Every PDF hard-wraps. The bullet glyph is drawn once, at the start of the first visual line. Every continuation line arrives at the parser as a bare line with no marker on it. A parser can usually handle a long continuation, because a long line with no bullet is obviously a bullet that lost its marker. What it handles badly is the last line of a wrapped bullet, which is short by definition.

A short line, no bullet, no dates: that looks exactly like the start of a new job header.

We ran a four-role resume through our own reader and got six roles back. Two of them were inventions, titled with fragments of sentences: "year revenue of >20M EUR" was one of them. Both had no employer and no dates. In an employer's system those become two jobs you cannot explain, they drag your average tenure down, and they change the role count a filter might read.

The fix in our parser is a rule with three conditions: a line directly under a bullet, starting with a lowercase letter, with no date range under it, continues that bullet rather than starting a job. The subtle part is that the continuation has to stay "under a bullet" through the join, because bullets wrap onto three lines all the time and without that the third line breaks silently.

We are telling you about our own bug because the general point survives it. A parser is a set of heuristics about line shapes, and your CV is the input that decides whether they fire correctly. You cannot audit the employer's heuristics. You can make your document a less ambiguous input.

Practical consequences you can act on today

  • Keep bullets short enough not to wrap onto a stub. A bullet that ends with two words on its own line is the exact shape that breaks. Trim it or extend it.
  • Never start a bullet's continuation with a capitalised phrase that could read as a job title.
  • Put a date range on every role, in one consistent format. Dates are what tell a parser "this is a header".
  • Use one column. Every time.
  • Use the section headings the vocabulary already knows: Experience, Education, Skills, Certifications, Languages.
  • Keep the file small. Under 100 KB is the safe target, not an aesthetic preference.

The decision, in a table

SituationSend
Taleo or any older enterprise portalDOCX
The portal names a preferred formatWhatever it names
Modern system, both formats tested and equalDOCX, on the tie-break above
Modern system, your PDF tested cleanerThe PDF, and keep the test
Your CV was made in a design toolRebuild it in a word processor first
You are pasting into a text box as wellA plain text version you have read through

The thing to stop doing

Stop deciding this by argument. The formats behave differently on your document, with your bullet lengths, your headings and your date style, and a generic rule cannot know any of that.

Involve Resume builds the real DOCX and the real PDF bytes, parses each one back, and compares the recovered record against what went in, field by field: name, email, phone, location, role count, dated roles, employers, skills, qualifications. The format that recovered more is the one it tells you to send, and DOCX breaks a tie for the reason above.

The limit is stated on the screen as plainly as we can put it: this is our reader, not the employer's. A different parser will differ in detail. What the check proves is narrower than "this will parse everywhere", and it still counts for something. A document our own reader cannot recover is a document no reader will recover, and every fault it has found so far has been a fault in the file rather than in how we read it.

Upload the CV you were about to send and look at the recovered fields before you look at anything else. If your four jobs come back as six, you have found something worth an evening.

Sources: Oracle Taleo, Candidate Management, Workable, Importing candidate data, SmartRecruiters, What is resume parsing, Ashby, Candidate Search, Workday, What is an applicant tracking system?

Questions people also ask

Is PDF safe to send to an applicant tracking system?

A text-based PDF is accepted by every major system. Oracle lists PDF among Taleo's supported parsing formats, and Workable lists .pdf alongside .doc and .docx. The failures come from PDFs exported by design tools, scanned images, and multi-column layouts, not from PDF as such.

Why does DOCX keep winning the tie-break?

Because the least forgiving parser you are likely to meet is the oldest one. SmartRecruiters' own guidance tells candidates to submit .docx for maximum parsing compatibility, and Taleo was designed around native Word documents. If both formats look equal on your own testing, the format with the wider margin is the one to send.

Does a two-column CV really break parsing?

It can, and the mechanism is specific. In a Word document a sidebar is usually a table cell, and cells are extracted one after another rather than side by side, so a skills column can land in the middle of a job title or be dropped entirely.

How big can a resume file be?

Smaller than you think in older systems. Oracle documents that a resume submitted to Taleo for parsing cannot exceed 100 kilobytes, or a limit the system administrator sets. A design-heavy PDF with embedded images passes that in one page.

Check both formats before you send one

Involve Resume builds the DOCX and the PDF, reads each back, and tells you which one recovered more of your record.

Open Involve Resume