全部插件

DSH / BUNDLE / BUNDLES

dsh-pdf-reader

v0.2.0AngelosZou / dsh-pdf-readerffecd361b7

可安装组合包组合包与其他模块社区 · Topic 自动分析

概览

dsh-pdf-reader

源码级技术说明DeepSeek Harness plugin: content-aware PDF reading tools for vision models (pdf_scan / pdf_read_page / pdf_render_region). Detects figures (vector+raster), tables, math, and two-column layout per page, then routes text pages to extraction and figure/table/math pages to high-DPI region crops rendered for read_image — avoiding DeepSeek's ~800x800 whole-page token ceiling. Backed by PyMuPDF; returns a clear warning (with the exact install command) when Python or a dependency is missing.展开完整技术说明收起技术说明
DeepSeek Harness plugin: content-aware PDF reading tools for vision models (pdf_scan / pdf_read_page / pdf_render_region). Detects figures (vector+raster), tables, math, and two-column layout per page, then routes text pages to extraction and figure/table/math pages to high-DPI region crops rendered for read_image — avoiding DeepSeek's ~800x800 whole-page token ceiling. Backed by PyMuPDF; returns a clear warning (with the exact install command) when Python or a dependency is missing.

README / ZH

插件文档

目录摘要

DeepSeek Harness plugin: content-aware PDF reading tools for vision models (pdf_scan / pdf_read_page / pdf_render_region). Detects figures (vector+raster), tables, math, and two-column layout per page, then routes text pages to extraction and figure/table/math pages to high-DPI region crops rendered for read_image — avoiding DeepSeek's ~800x800 whole-page token ceiling. Backed by PyMuPDF; returns a clear warning (with the exact install command) when Python or a dependency is missing.

dsh.pub 核对固定版本的组合包契约、运行时事实与分发语义;完整 README 请查看源仓库。

在 GitHub 阅读完整 README

LIMITATIONS

已知限制

- **`page.find_tables()` false-positives** on plot grids and diagrams, so tables are primarily read by rendering (the reliable path); Markdown tables from `pymupdf4llm` are a best-effort extra. - Formula detection is heuristic (fonts + LaTeX producer). PyMuPDF decodes inline math well, but stacked fractions can still be imperfect — use `pdf_render_region` on an equation when exact structure is needed. - A whole page renders to only ~83 DPI-equivalent at the budget; the tools never do that for a two-column page — they crop regions instead, and body text uses extraction. - Large PDFs are parsed into memory.