All plugins

DSH / BUNDLE / BUNDLES

dsh-qwen38-ninfer-compaction-fix

v1.2.0zhubaohi / dsh-qwen38-compaction-fix080fa00992

InstallableBundlesBundles & other modulesCommunity · Topic auto-analysis

Overview

dsh-qwen38-ninfer-compaction-fix

DSH plugin (NInfer engine only): fixes compaction failure on local qwen3.8-27b gateways served by NInfer — xhigh thinking burns the entire output token budget, so thinking is off for compaction-only, with the model's non-thinking sampling parameters; the same idea applies to other launch methods

README / EN

Package documentation

Registry summary

DSH plugin (NInfer engine only): fixes compaction failure on local qwen3.8-27b gateways served by NInfer — xhigh thinking burns the entire output token budget, so thinking is off for compaction-only, with the model's non-thinking sampling parameters; the same idea applies to other launch methods

dsh.pub verifies the pinned bundle contract, runtime facts, and distribution semantics. The complete README remains in the source repository.

Read the full README on GitHub

LIMITATIONS

Known limitations

- **Engine scope: NInfer.** The wire fields this plugin rewrites are the ones the NInfer gateway interprets. Gateways served by other engines (llama.cpp, vLLM, FastMTP, ...) may ignore or spell these fields differently; the same thinking-off idea can be ported to them, but that port is a different package. - Model matching is an **exact id match** against the id declared under `llm-pi-ai.providers.<provider>.models[].id` in `settings.yaml`. See the model name section above. - The HTTP layer signatures track specific dsh releases; see "How it works" for what happens when a signature stops matching. - This plugin shapes requests for the *local gateway* you run. It does not change the harness's own routing or the server's real capacity limits (the server still enforces them).