全部插件

DSH / BUNDLE / BUNDLES

dsh-qwen38-ninfer-compaction-fix

v1.2.0zhubaohi / dsh-qwen38-compaction-fix080fa00992

可安装组合包组合包与其他模块社区 · Topic 自动分析

概览

dsh-qwen38-ninfer-compaction-fix

DSH plugin (NInfer engine only): fixes compaction failure on local qwen3.8-27b gateways served by NInfer — xhigh thinking burns the entire output token budget, so thinking is off for compaction-only, with the model's non-thinking sampling parameters; the same idea applies to other launch methods

README / ZH

插件文档

目录摘要

DSH plugin (NInfer engine only): fixes compaction failure on local qwen3.8-27b gateways served by NInfer — xhigh thinking burns the entire output token budget, so thinking is off for compaction-only, with the model's non-thinking sampling parameters; the same idea applies to other launch methods

dsh.pub 核对固定版本的组合包契约、运行时事实与分发语义;完整 README 请查看源仓库。

在 GitHub 阅读完整 README

LIMITATIONS

已知限制

- **Engine scope: NInfer.** The wire fields this plugin rewrites are the ones the NInfer gateway interprets. Gateways served by other engines (llama.cpp, vLLM, FastMTP, ...) may ignore or spell these fields differently; the same thinking-off idea can be ported to them, but that port is a different package. - Model matching is an **exact id match** against the id declared under `llm-pi-ai.providers.<provider>.models[].id` in `settings.yaml`. See the model name section above. - The HTTP layer signatures track specific dsh releases; see "How it works" for what happens when a signature stops matching. - This plugin shapes requests for the *local gateway* you run. It does not change the harness's own routing or the server's real capacity limits (the server still enforces them).