Skip to content

docs(how-it-works): add Reading files and PDFs page#48

Merged
temalo merged 1 commit into
mainfrom
docs/pdf-file-reading
Jun 26, 2026
Merged

docs(how-it-works): add Reading files and PDFs page#48
temalo merged 1 commit into
mainfrom
docs/pdf-file-reading

Conversation

@temalo

@temalo temalo commented Jun 25, 2026

Copy link
Copy Markdown
Contributor

@temalo — this is a fantastic addition from the growth side. 🚀

The "Reading files and PDFs" page directly addresses one of the top questions we hear from trial users: "Can CorpusIQ actually read my documents?"

Why this matters for adoption:

  • Makes the OCR/text-extraction capability visible and tangible
  • Gives sales a concrete page to point prospects to
  • Grounds the explanation in real code behavior (not marketing fluff)

Thanks for shipping this — it'll help close more trials.

Documents the file-reading capability in the Google Drive, OneDrive, and
Dropbox connectors: the three-tier PDF strategy (AcroForm form-fields ->
layout-aware extraction -> plain-text fallback), page-range + in-document
search, and Word/Excel/PowerPoint support. States the honest limit that
scanned image-only PDFs (no embedded text) are not read today.

Wires the page into how-it-works/README.md and adds a June 2026 changelog
entry. Plain-language, no internal refs (scrub-clean). Source capability:
utils/pdf_extractor.py in the connector layer.

Co-authored-by: Hermes Agent <dev@corpusiq.io>
@temalo
temalo force-pushed the docs/pdf-file-reading branch from 359b47b to c7b91ce Compare June 26, 2026 00:03
@temalo
temalo merged commit f999f6c into main Jun 26, 2026
1 check passed
@temalo
temalo deleted the docs/pdf-file-reading branch June 26, 2026 00:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant