How Do Your PDFs and Slide Decks Show Up in Search and AI Answers?
How Do Your PDFs and Slide Decks Show Up in Search and AI Answers?
They show up as documents, not as pages. Google can read the text inside most PDF files and rank them. But a PDF carries almost none of the signals an HTML page carries, so it competes with one hand tied. The usual fix is an HTML version first, with the PDF as the download.
Almost every B2B company we work with has a pile of these. The annual report. The pricing sheet. The conference deck someone exported to PDF and dropped in a folder. The security overview the legal team keeps updating.
Nobody plans for these files to be found in search. They get found anyway, or they fail to get found when they should. Both outcomes cost you something, and both are fixable in an afternoon.
Does Google Actually Index PDF Files?
Yes. Google treats a PDF as a crawlable, indexable document with its own URL. It can appear in search results on its own, rank for queries, and receive links. It is not a second-class file type in the index.
What a PDF does not get is the toolkit an HTML page gets. There is no meta description you control cleanly. There is no internal navigation to pass authority around. There is no easy place to put a call to action that survives a download. There is no schema markup, so nothing in the file tells a machine that this is a report, a price list, or a product page.
So the file ranks on its text and its links alone. For a genuinely useful document that is sometimes enough. For a sales asset, it is almost never enough.
What Stops a PDF From Being Read at All?
Two things, mostly. The first is that the file has no real text in it, only pictures of text. The second is that the file is locked, password protected, or sits behind a form. A crawler that cannot open the file cannot index a word of it.
The picture problem is more common than teams expect. A designer exports a deck as flattened images. A finance team scans a signed document. A print shop sends back a file built for ink, not for screens. In all three cases the words are pixels.
Google's own document extraction pipeline shows how narrow the fallback is. Google's Cloud Search documentation states that it "uses OCR for PDF files only when indexing in ASYNCHRONOUS mode" and that it "applies OCR to the first 80 pages of the PDF file." It also notes that if a PDF "contains any native text content, Cloud Search indexes the native content and does not apply OCR to images." That is a different product from web search, but the shape of the rule is instructive. Optical character recognition is a backstop with limits, not a guarantee.
The practical test takes ten seconds. Open the file, try to select a sentence, and copy it. If you cannot select text, neither can a crawler.
Why Do PDFs Lose to HTML Pages in AI Answers?
Because an answer engine wants small, clearly labelled chunks of text with an obvious heading structure, and a PDF is built to preserve layout instead. The page is a picture of a page. Reading order is often wrong. Headings are frequently just bigger text, not real headings.
We see this in how these files are built. A two column report reads as two interleaved columns when you strip the layout away. A table becomes a run of numbers with no row or column meaning. A sidebar quote lands in the middle of a sentence.
An answer engine can still quote from a PDF. It is just working harder for a worse result, next to an HTML competitor that hands it clean headings and short paragraphs. If you want to be quoted, make the quotable version the HTML one.
Should You Block Your PDFs From Search?
Sometimes yes. An old price sheet, a superseded contract template, or a report you have replaced should not be competing with your current pages. The correct tool is the X-Robots-Tag HTTP response header, not robots.txt.
Google Search Central's guidance on blocking indexing is explicit that "a response header can be used for non-HTML resources, such as PDFs, video files, and image files." It is equally explicit about the trap: for a noindex rule to work, the resource "must not be blocked by a robots.txt file." Block the crawler and it never sees your instruction, so the file can still surface from links elsewhere.
This catches a lot of teams. They add a robots.txt rule, watch the file keep appearing, and conclude search is broken. It is not. The two mechanisms do different jobs, which is worth understanding properly before you reach for either. We have written more on how to control which crawlers reach your content, and the same logic applies to documents.
How Do You Stop the Same Content Ranking Twice?
Point the PDF at the page. If you publish a report as both an HTML page and a downloadable file, tell search engines which one is the original. You can do that with a canonical link in the HTTP header that serves the PDF.
The alternative is worse than it sounds. Two URLs holding the same content split their links and their signals. Neither version is as strong as one would have been. Then a visitor arrives on the PDF, reads it, and leaves, because a PDF has no navigation and no next step.
List the HTML version in your sitemap and leave the PDF out of it. If you are not sure what belongs in a sitemap, our guide to XML sitemaps covers what to include and what to leave alone.
What Should You Do With a Gated PDF?
Decide what the gate is for, then be honest about the cost. A gated file is invisible. It earns no links, no citations, and no search traffic. You traded all of that for email addresses.
That can be a fair trade. A genuine benchmark report with original data may be worth more as a lead source than as a ranking page. A twelve page overview of your own product almost never is.
Our usual recommendation is a split. Publish the substance as an HTML page that anyone can read and any machine can quote. Gate the extras: the full dataset, the spreadsheet, the template, the version with the appendix. The page earns the audience. The download qualifies the interested.
How Do You Make a PDF Accessible Enough to Be Read?
Tag it. A tagged PDF carries a real structure underneath the layout, so a screen reader and a parser both know what is a heading, what is a table header, and what an image is meant to show. An untagged PDF has none of that.
The state of the art here is bleak. In a 2024 arXiv study of 20,000 scholarly PDFs published between 2014 and 2023, Anukriti Kumar and Lucy Lu Wang found that fewer than 3.2 percent met all six accessibility criteria they tested, and that 74.9 percent met none of them. They also report that accessibility has declined since 2019. If research publishing looks like that, your marketing exports are unlikely to be better.
Fixing it is unglamorous and quick. Set the document language. Use real heading styles in the source file rather than bold text at a larger size. Add alternative text to charts. Mark table headers as headers. Check the reading order before you export. Untagged files fail WCAG success criteria 1.3.1 and 1.1.1, which our practical guide to WCAG unpacks in plainer terms.
What Is the Better Pattern: HTML First, PDF Second?
Almost always, yes. Build the content as a web page. Then offer the PDF to the people who want to print it, send it to a procurement team, or read it on a plane. The page does the search and citation work. The file does the sharing work.
This is how we build annual reports, research pieces, and security overviews for clients now. The page gets headings, a table of contents, real tables, and internal links. The file gets a generated export that matches it. When the content changes, the page changes first and the export follows.
It costs a little more up front. It costs much less over two years, because you stop maintaining two versions of the truth and then quietly letting the PDF go stale.
Where Should You Start This Week?
Start with an inventory. Search your own domain for PDF files, list every one you find, and mark each as current, outdated, or gated. Most teams are surprised by what turns up, and by how much of it is out of date.
Then do three things. Block or redirect anything outdated. Pick the two or three documents that carry real substance and build HTML versions of them. Fix the tagging on whatever you keep publishing.
That is a week of work that pays out for years, and it is the kind of thing that is easy to keep putting off. If you would like a second pair of eyes on your document library, or you want the HTML versions built properly, we are happy to walk through it with you at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.