83% of SharePoint PDFs not indexed despite EAP enrollment — possible folder depth or crawler issue | The place for Zendesk users to come together and share
Skip to main content
June 26, 2026
Question

83% of SharePoint PDFs not indexed despite EAP enrollment — possible folder depth or crawler issue

  • June 26, 2026
  • 6 replies
  • 158 views

Despite being enrolled in the PDF Content Allowance EAP, a large number of PDFs from our SharePoint connector are not being indexed by Zendesk.

Note: We have already confirmed that our SharePoint connection is listed as a search source under Knowledge Admin > Search Settings > Search Sources (237 items for the main portal, 5 items for the inquiry directory).

Here is a summary of our investigation.

== PDF Sync Status ==

  • Total PDFs in SharePoint: 591
  • Synced to Zendesk: 99
  • Not synced: 492 (~83%)

== Additional Context ==

The majority of unsynced PDFs already existed in SharePoint before we enabled the EAP (approximately 3 weeks ago). This suggests the issue is not limited to newly added files, but also affects pre-existing files that should have been indexed during the initial sync.

== Comparison of Synced vs. Unsynced Folder Structures ==

Folders with higher sync rates (confirmed on Zendesk side)

  • DocLib1 (legal document library): 17 files
  • Shared Documents/006_業務ガイド/情報セキュリティ/...: 41 files
  • Shared Documents/004_社内システム/...: 13 files
  • Shared Documents/005_制度/...: 14 files
  • Other: 14 files

Folders with low or zero sync rates (confirmed on SharePoint side)

  • Shared Documents/人事発令/ (FY2022–2026): 188 files
  • Shared Documents/200_規程・内規: 48 files
  • Shared Documents/INFOMATION添付資料: 29 files
  • Other unsynced folders: 227 files

== Questions ==

(1) Folder depth limit — most likely cause Synced PDFs appear to be concentrated in folders with a maximum depth of 5 levels. We have 29 PDFs located 6–7 levels deep in the SharePoint folder hierarchy, and none of them appear to be synced. Is there a folder depth limit for PDF ingestion?

(2) Scope of synced libraries and folders Large folders such as 人事発令 (personnel announcements), 規程・内規 (company regulations), and INFOMATION添付資料 (information attachments) are entirely unsynced. Are there any restrictions on which libraries or folders are included in the sync scope? Is there a way to specify individual folders in the connector settings?

(3) Impact of unsupported file types We have noticed that some folders contain unsupported file types (e.g., .png, .rar, .zip, .pptx). Based on a similar report in this community, it seems that the presence of unsupported files may cause the crawler to stop mid-process, preventing other files (including PDFs) in the same folder from being indexed. Could this be a factor in our case? Is there any logging available to identify which files were skipped or caused the crawler to abort?

(4) How to verify EAP activation status We have confirmed that we are enrolled in the EAP, but we are not sure how to verify whether the PDF ingestion feature is actually active and functioning correctly. Is there a way to check this?

== Background / How We Got Here ==

We were initially directed to post in the "Zendesk EAP: Knowledge: PDF Ingestion" group by Zendesk Support, who referenced a similar thread there. However, when we attempted to post, we received the message: "This user does not have access to the category." So here we are.

As for the support process itself — we reached out to Zendesk Support first, and after some back-and-forth, the response we received was essentially: "Please post in the community." The irony of a company that sells customer support tooling responding to support requests by redirecting users to a community forum — with no troubleshooting steps offered — was not lost on us. We understand EAPs come with limited SLA coverage, but a templated deflection response is a rather on-the-nose demonstration of exactly the problem Zendesk's own AI is supposed to solve.

We're posting here in good faith and hope someone from the product team can shed some light.

    6 replies

    henry_collins
    Newcomer
    June 26, 2026

    This doesn’t look like a strict folder-depth limitation in Zendesk.

    More likely it’s one of these: library-level scope/permissions differences, SharePoint path/length or traversal issues (which can silently skip folders), or batch failures during ingestion. Unsupported file types usually don’t block everything, but they can break a batch depending on how the crawler handles errors.

    Since some libraries are fully missing, first check ingestion logs for those specific libraries, you’ll usually see a clear reason like access denied, path issues, or skipped batches.

    Full-Stack Developer | SEO Strategist | Helping businesses build better support and web experiences
    sfujiwaraAuthor
    June 30, 2026

    Hi Henry, thank you — this is genuinely helpful, and it matches what we'd started to suspect: the folder-depth theory didn't hold up on closer inspection.

    The direction you point to (library-level scope/permissions, path-length/traversal issues, or a failed ingestion batch) makes sense. The blocker for us is the last step you mention: we haven't been able to find where the ingestion logs actually live.

    Could you (or anyone from the product team) clarify:
    1. Where exactly can we view the SharePoint connector's ingestion logs? We've looked under Knowledge Admin and the connector settings but haven't found a per-library/per-file log showing access-denied, path issues, or skipped batches.
    2. Is there any in-product indicator confirming the PDF Ingestion EAP is actually active and processing for our account (vs. just being enrolled)?

    For context, the unsynced files are concentrated in a few large libraries (人事発令 ~188, 規程・内規 48, etc.) that are missing almost entirely, which seems consistent with a library-scope or batch-failure cause rather than per-file issues.

    We'd really appreciate a pointer to the logs — that seems to be the key to diagnosing this. Thanks again.

    Tobias11
    Contributor
    June 30, 2026

    I am in Support with Zendesk right now, seems all PDF bigger 10MB can´t be indexed in Sharepoint.

    Tobias
    sfujiwaraAuthor
    July 2, 2026

    Thanks Tobias, useful data point.

    Our numbers suggest 10MB may not be the whole story though. In our sync, Word is ~97% and Excel ~80%, but PDF is only ~17% — so the failure is concentrated on the PDF type itself, matching the GA docs listing only Word/Excel/Markdown (PDF looks EAP-stage).

    Within PDFs, failures cluster by library, not evenly. Several libraries sit at 0%, and those are mostly scanned/image-only PDFs (no text layer); others sync 35-40%. A pure 10MB cap would let smaller files through, so a whole library at 0% points more to image-only PDFs or permissions than size — though size may add to it.

    We've asked Zendesk support to confirm the actual limits (size cap, image-only handling, per-library caps) and GA timeline, and can share back what we hear.

    Did you confirm the 10MB with support, or infer it from your failed files? And do your failures line up with particular libraries or scanned docs?

    Maddie Hoffman
    Community Manager
    July 2, 2026

    Hi ​@sfujiwara, just FYI I moved this post into the Knowledge PDF Ingestion EAP group, so that you can get the right eyes on your question/feedback from our Product Managers. 

    September 2, 2026

    The problem is that the EAP only allows up to 100 PDF files to sync.  Ran into the same issue.