PDF to Markdown with OCR and Tables
Identify text and scanned PDFs, preserve page evidence, review OCR, and check complex tables before export.
Read the PDF guideDetailed workflows for recovering useful Markdown, packaging source material for AI agents, and preparing traceable datasets. Conversion happens in browser memory; each guide also explains the points where review or a separate cloud service may be required.
Identify text and scanned PDFs, preserve page evidence, review OCR, and check complex tables before export.
Read the PDF guideOrganize EPUB or long-form source material into chapters, references, and conservative indexes an agent can navigate.
Read the Book Skill guideClean source text, choose chunk boundaries, preserve metadata, and evaluate retrieval before adding more data.
Read the RAG guideDoGetSkill does not upload source documents to a conversion server. Guides distinguish local conversion from optional later steps such as OCR component downloads, embedding APIs, hosted vector databases, or model analysis so you can decide what leaves the device.