DeepSeek Harness plugin: content-aware PDF reading tools for vision models (pdf_scan / pdf_read_page / pdf_render_region). Detects figures (vector+raster), tables, math, and two-column layout per page, then routes text pages to extraction and figure/table/math pages to high-DPI region crops rendered for read_image — avoiding DeepSeek's ~800x800 whole-page token ceiling. Backed by PyMuPDF; returns a clear warning (with the exact install command) when Python or a dependency is missing.
README / EN
Package documentation
Registry summary
DeepSeek Harness plugin: content-aware PDF reading tools for vision models (pdf_scan / pdf_read_page / pdf_render_region). Detects figures (vector+raster), tables, math, and two-column layout per page, then routes text pages to extraction and figure/table/math pages to high-DPI region crops rendered for read_image — avoiding DeepSeek's ~800x800 whole-page token ceiling. Backed by PyMuPDF; returns a clear warning (with the exact install command) when Python or a dependency is missing.
dsh.pub verifies the pinned bundle contract, runtime facts, and distribution semantics. The complete README remains in the source repository.
- **`page.find_tables()` false-positives** on plot grids and diagrams, so tables
are primarily read by rendering (the reliable path); Markdown tables from
`pymupdf4llm` are a best-effort extra.
- Formula detection is heuristic (fonts + LaTeX producer). PyMuPDF decodes inline
math well, but stacked fractions can still be imperfect — use
`pdf_render_region` on an equation when exact structure is needed.
- A whole page renders to only ~83 DPI-equivalent at the budget; the tools never
do that for a two-column page — they crop regions instead, and body text uses
extraction.
- Large PDFs are parsed into memory.
Same capability topic
Related by capability
Ranked by shared capability topic, runtime face, tools or UI contributions, and distribution mode. This is not a similarity, quality, or compatibility ranking.