Overview
dsh-agent-harness
README / EN
Package documentation
dsh-agent-harness
Free, MIT-licensed bridge between Agent Harness and DeepSeek Harness (DSH).
Bring model-aware routing, agent lifecycle controls, execution telemetry, and qualification evidence to DSH without changing where your work runs.
The short version
Agent Harness is the orchestration and intelligence layer. It can plan work, coordinate agents, test outcomes, benchmark models, optimize settings, run a model laboratory, qualify task performance, and recommend routes.
DSH is the implementation layer. It owns providers, models, local inference, tools, sessions, and workspace execution.
dsh-agent-harness is the bridge between them. It translates public contracts
for routing, jobs, lifecycle, telemetry, capabilities, and optional evidence.
It does not own a model catalog. It does not hardwire local model choices. It uses the models configured in DSH or the explicit route returned by Agent Harness.
Experiment ownership is intentionally outside this repository. Model,
prompt, and settings campaigns, fixtures, randomized arms, independent
scoring, and promotion decisions live in the Agent Harness repository under
experiments/qwen3/. This plugin supplies only the execution profiles,
preset/tool wiring, telemetry, and public handoff fields that those campaigns
exercise. See the boundary contract.
Why this exists
Fast inference is useful, but raw tokens per second do not tell the whole story. A model can be fast and still fail to use tools, follow a task, or produce a verifiable artifact.
Agentic work needs a wider feedback loop. The loop must measure correctness, tool execution, latency, token use, failure modes, and task class—not just decode speed.
This connector gives Agent Harness a reliable path into real DSH execution. It also gives DSH users a safe path to routing and qualification improvements without replacing their local runtime.
The relationship
Agent Harness
planning · orchestration · testing · benchmarking · model lab
optimization · qualification · routing recommendations
|
v
dsh-agent-harness
public bridge and translator
|
v
DSH
implementation · tools · sessions · provider/model execution
workspace changes · local inference · configured fallbacks
The plugin is useful to both sides. It makes DSH more discoverable as a robust implementation harness, and it makes Agent Harness more useful as the system that plans, evaluates, and improves real agentic work.
What users get
- Task-aware routing through a versioned HTTP contract.
- Local fallback when Agent Harness is offline or unavailable.
- DSH-owned provider and model selection.
- Serialized model lifecycle for one-model-at-a-time machines.
- Protection against serialized tool-call text being mistaken for execution.
- Job spawn, status, wait, result, and cancellation controls.
- Execution telemetry for steps, tools, tokens, outcomes, and stop reasons.
- MCP stdio integration for external agent clients.
- An open connector contract for future hosts and community adapters.
How the bridge works
User task
|
v
Agent Harness plans and classifies the work
|
v
dsh-agent-harness requests an optional route
|
v
DSH selects its configured implementation model and executes the task
|
v
Independent verification checks tools, artifacts, and outcomes
|
v
Agent Harness receives optional aggregate evidence and improves future routes
The bridge stays useful at every point in that flow. If routing is unavailable, DSH continues with its own configured provider and model settings.
If evidence submission is unavailable, local execution still completes. A remote service is never required merely to install or start the connector.
Model ownership
The connector is intentionally model-agnostic.
Users select models in DSH. The connector reads DSH-owned settings such as:
flash_providerandflash_model;pro_providerandpro_model;- DSH presets and tool budgets.
Callers may also provide an explicit provider/model pair for an individual job. That explicit selection takes precedence over routing and tier configuration.
Agent Harness may maintain curated Quinn, Qwen, or other model profiles in its optimizer and benchmark systems. Those profiles remain outside this plugin.
This separation means users can change their DSH models without waiting for a connector release, while Agent Harness can improve recommendations over time.
The first release favors explicit configuration alignment because it is easy to inspect and verify. Over time, capability discovery and generic configuration handshakes can reduce duplication as the model laboratory learns how providers, models, and settings behave across task classes.
Installation and activation
The connector is designed to ship with the Agent Harness OS installation. It may also be installed from the DSH marketplace as a free MIT-licensed plugin.
Installation mounts the connector for discovery but does not change existing DSH behavior. The connector is disabled by default.
Enable connector execution in the DSH host environment:
DSH_AGENT_HARNESS_ENABLED=1
Enable task-aware Agent Harness routing separately:
DSH_AGENT_HARNESS_ROUTING_ENABLED=1
One-command local launch
From a checkout of this repository, the included launcher selects Node 22 when
nvm is installed, enables the connector and routing, and starts the existing
DSH web profile:
./plugins/dsh-agent-harness/scripts/start-dsh-agent-harness.sh
It uses ~/deepseek-harness-stable by default. Set DSH_ROOT to use another
installation, or set either connector flag to 0 for a disabled or
routing-off smoke run. Extra arguments are forwarded to DSH. The launcher
never selects or embeds a model; provider and model configuration remain
DSH-owned.
With both flags disabled, existing dsh-crew behavior is preserved.
Experimental Qwen2.5 tool bridge
Qwen2.5-Coder Q4_K_M is registered separately as lmstudio-qwen-bridge. The
bridge injects a Qwen-style tool-call contract and converts only
schema-validated serialized calls into native OpenAI tool calls:
DSH_QWEN_TOOL_BRIDGE_ENABLED=1 \
./plugins/dsh-agent-harness/scripts/start-dsh-agent-harness.sh
This route is experimental. The current matrix shows that the Q4 model still
often emits incoherent or malformed text, so the bridge refuses ambiguous
output and does not promote the model for executable work. Keep the normal
lmstudio route and Qwen3-4B as the validated tool-capable path.
The sidecar exposes GET /health. If port 3124 is already occupied by a
healthy bridge, a second launcher reuses it; a foreign occupant fails clearly
instead of silently starting DSH against the wrong adapter.
Bridge responses include shape-only x_qwen_bridge diagnostics. These identify
native calls, malformed serialized text, parser acceptance, and suspicious
placeholder argument keys without returning prompt or command contents.
The structured bridge mode is experimental. Direct LM Studio schema probes pass, but the current DSH preset/tool surface still exhausts Qwen2.5's context before a valid callable tool action; the minimal preset now reaches tools but tool-specific schemas still do not produce convergence. It is not a promoted route. A single targeted correction turn is now supported, but the current Qwen2.5 arm still needs precise edit-result feedback before it can converge. The registered edit tool now provides that result shape in a DSH-compatible form, but the latest live arm still did not verify and remains experimental.
Required configuration alignment
The bridge works best when DSH and Agent Harness are configured with the same provider and model identities.
Configure each local provider/model in DSH first. Then configure the matching provider/model in the Agent Harness settings panel. Use the same provider IDs, model IDs, endpoint mappings, and relevant runtime settings in both places.
For example, if DSH uses:
provider: ollama
model: qwen3.5:9b
the corresponding Agent Harness model entry should resolve to that same DSH provider and model identity. Do the same for every flash/pro tier or task-class route that Agent Harness may recommend.
Before enabling routing, verify:
- DSH can run the provider/model directly.
- Agent Harness can see and benchmark the matching provider/model.
- The route returned by Agent Harness uses identifiers DSH recognizes.
- Presets, context limits, tool settings, and optimization variants are aligned where they affect execution.
- A small tool-use task succeeds through the complete bridge.
For bounded execution, DSH may set max_tool_calls as the hard volume limit
and max_repeated_tool_calls as the consecutive identical-call limit (default
3). The connector reports tool_budget_exceeded and
stagnation_detected separately, and cancels the live session once per job.
These are safety controls and do not turn a worker completion message into a
verified task success; independent verification remains the authority.
Telemetry also exposes a stable failureClass for routing/evidence analysis,
including context_limit, model_runtime, timeout, workspace_access,
tool_budget, stagnation, and verification_missing.
For benchmarked jobs with an independent verifier, stop_after_verified_tool
can opt into cancellation immediately after a successful test/check command
tool result. This is deliberately opt-in because the DSH runtime cannot know
which command is authoritative for an arbitrary interactive task.
Qualification callers should additionally pass verifierCommand with the
exact command declared by the task contract, for example
node --test calc.test.mjs. Telemetry then accepts only that command (after
whitespace normalization) as verification and will not stop or mark the job
verified because a different worker-authored test happens to pass. Omit this
field only for exploratory jobs that intentionally use the broader
test-like-command detector.
If the configurations diverge, Agent Harness may recommend a model that DSH cannot resolve, or the two systems may apply different settings. The connector fails safely where possible, but it cannot reconcile incompatible provider or model identities automatically.
Offline-first behavior
The free connector does not require a paid account, remote credentials, or a reachable Agent Harness endpoint for local DSH execution.
When routing is enabled but unavailable, the bridge fails open to DSH's own configured tier settings. If DSH has no model configured for the requested tier, the connector reports that clearly rather than inventing a local model.
Runtime status includes a read-only modelPreflight report when a public
Harness model catalog is supplied as harness_models plugin configuration. It
compares DSH's flash/pro identities with that catalog without rewriting DSH
settings or disabling local fallback. Completed jobs expose routing.source,
tier, provider, model, and public route confidence/reason fields.
When supplied, the same status includes a sanitized, versioned
modelCapabilities snapshot with ranked task-class recommendations and an
explicit routePromotion gate. Unqualified classes are not promoted, and
catalog secrets or arbitrary fields are excluded.
Capability status distinguishes disabled, offline, connected, stale, and unsupported protocol states.
Paid routing is marked available only when a compatible handshake confirms an active entitlement, fresh evidence, and valid credentials.
The feedback loop
The intended optimization loop is:
- Agent Harness plans or classifies a real task.
- DSH executes it with the user-selected implementation model.
- The connector records observable execution telemetry.
- Independent verification checks the artifact and tool behavior.
- Aggregate results are optionally submitted as qualification evidence.
- Agent Harness updates task-class recommendations and optimizer profiles.
- Future tasks receive better route recommendations.
This prevents a fast but unreliable model from winning solely on raw throughput. It also makes recommendations sensitive to task type: bounded implementation, debugging, test authoring, multi-file work, and reasoning can favor different settings.
The experimentation phase is the product feedback loop. Verified outcomes feed the model laboratory, which improves optimizer profiles, configuration discovery, and future routing recommendations.
Community model laboratory
Agent Harness can become a community model laboratory, not just a private optimizer. Users can publish configurations, benchmark recipes, task suites, and independently verified results for others to download and try.
Shared work can be compared through task-class leaderboards and reproducible qualification records. A configuration that wins on raw speed but fails tool execution or independent verification should not outrank a slower, reliable configuration.
Community submissions should include model/provider identity, runtime settings, hardware and software context, task class, benchmark version, verification method, and provenance. Users should be able to inspect, download, fork, and retest a submission before applying it to DSH.
Downloads are recommendations, not automatic trust decisions. Imported configurations should be reviewed, scoped to the user's own DSH providers, and kept separate from credentials, source content, and executable code.
Evidence and privacy
Evidence submission is separate from execution and requires explicit consent. It is designed for aggregate qualification metrics, not source collection.
The allowlisted evidence shape can include task class, model/settings identifiers, independently verified success, wall-clock latency, tool-call count, token totals, and failure mode.
The connector excludes task instructions, source contents, workspace paths, and credentials from evidence payloads. Submission failure never blocks local work.
Public connector surface
The connector exposes versioned, host-neutral concepts for capability discovery, task classification, routing, job lifecycle, execution telemetry, offline fallback, optional aggregate evidence, and MCP stdio access.
Future connectors can adapt other agent hosts or orchestration systems to the same public ideas without importing private Agent Harness implementation code.
Community adapters should provide a stable manifest, explicit capabilities, bounded timeouts, cleanup handlers, data-flow documentation, and clear offline behavior.
Responsibilities
Agent Harness: orchestration and intelligence
- Plans and coordinates agentic work.
- Runs testing, benchmarking, and model-lab workflows.
- Owns optimization, qualification, and evidence history.
- May provide curated model/settings recommendations.
- May provide optional paid routing intelligence.
DSH: implementation and execution
- Owns local providers and model configuration.
- Owns inference, tools, sessions, and workspace execution.
- Provides the host lifecycle and configured fallback tiers.
- Remains useful when Agent Harness is unavailable.
dsh-agent-harness: the bridge
- Adapts DSH to public Agent Harness contracts.
- Translates task metadata, routes, jobs, and telemetry.
- Enforces safe fallback and false-success protection.
- Keeps installation separate from execution.
- Does not embed model defaults or paid-only code.
For marketplace users
Install the free connector when you want a clean bridge from DSH into Agent Harness planning, routing, benchmarking, and qualification workflows.
Keep it disabled while evaluating the package. Enable execution first, then enable routing when you are ready to compare recommendations with your DSH configured defaults.
The package is MIT-licensed and offline-capable. Agent Harness services are optional and may include separate paid features.
The marketplace listing should treat this README as the technical source of truth. Product copy must preserve the free-plugin and paid-service boundary.
For developers
The package is an open connector surface, not a private product shim.
Start with connector.json for the manifest contract, src/ for the public
adapter modules, and test/ for behavior-oriented examples.
Run the local gates:
npm test
npm run validate
The test suite covers routing, DSH configuration intake, fallback, host integration, lifecycle cleanup, tool execution validation, capabilities, service access, evidence privacy, job control, web routes, and MCP handlers.
Current status
This package is alpha software undergoing parity testing against dsh-crew.
The connector currently includes tested routing, task classification, DSH configuration intake, lifecycle controls, fallback, telemetry, job routes, MCP integration, capability discovery, service gating, and evidence contracts.
The deterministic-edit feedback loop now parses compact success/failure result envelopes and can issue one targeted reread correction. The Qwen2.5-Coder 7B result-gated live screen still produced only one edit call and no verifier, so this path is diagnostic and bounded rather than a promotion of that model.
Standalone cutover and marketplace publication remain release gates. The temporary dsh-crew compatibility path remains available during migration.
Roadmap
- Complete dsh-crew parity and stable-host boot validation.
- Validate offline, reload, shutdown, and evidence privacy gates.
- Publish a public connector contract for community adapters.
- Add additional host connectors over time.
- Expand task-class qualification and optimizer feedback loops.
- Publish the free plugin to the DSH marketplace after release details are confirmed.
License and provenance
This project uses the MIT License. See LICENSE and NOTICE for the upstream
dsh-crew attribution and the new integration work attribution.
The package is free, offline-capable, and does not embed Agent Harness credentials.
The plugin follows Cordis lifecycle rules: its entry point returns a disposer, and injected web routes return their own disposer. This keeps hot reload and profile shutdown safe.
LIMITATIONS
Known limitations
Submitted through a public pull request. Automated checks verified the public bundle contract and committed files, but did not inspect runtime capabilities. This is not a human review, security audit, publisher identity check, or official endorsement.
