- TypeScript 85%
- Shell 15%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| agent | ||
| plugin | ||
| completion.zsh | ||
| opencode.json | ||
| README.md | ||
Local vLLM Orchestrator + Worker Setup
An orchestrator/worker agent setup for opencode that runs entirely on a local OpenAI-compatible vLLM server — Qwen3.8-27B on 2× Intel Arc Pro B70 (TP=2). No cloud APIs anywhere in the config.
Components
opencode.json— provider/model config only (no agent bodies):- Provider
openai-compatible→ modelqwen38; the endpoint and API key come from environment variables (see below), not from this file. Every completion request carrieschat_template_kwargs: {"enable_thinking": true, "thinking_budget": 2048}(verified to pass through@ai-sdk/openai-compatibleinto the vLLM body). Request timeout: 600 s (long thinking blocks at ~50 t/s take minutes). default_agent: orchestrator(the agent bodies live inagent/*.md).
- Provider
agent/orchestrator.md(primary, default agent): decomposes the user goal into 1–2 (max 2) self-contained worker tasks, dispatches them concurrently toqwen-worker, reads the reports, verifies, and synthesizes — or dispatches one follow-up round for concrete gaps.agent/qwen-worker.md(subagent,task: deny— no sub-spawning): executes one well-scoped task end to end (investigate, implement, fix, verify) in long focused passes, and reports back in a fixed format:Summary / Files changed / Verification / Open issues.
Both agents run on the same model: openai-compatible/qwen38. Agent files are
the canonical, editable form of the definitions; after editing agent/*.md,
restart opencode (new sessions pick up changes).
Server facts (design around these)
| Fact | Value |
|---|---|
| Base URL | $OPENCODE_VLLM_BASE_URL (full address: scheme + host + port + /v1) |
| API key | $OPENCODE_VLLM_API_KEY (server accepts any non-empty string, e.g. local) |
| Max concurrent requests | 4 (max-num-seqs); practical sweet spot: 1–2 |
| Throughput | ~53 t/s per stream alone; ~45 t/s each at 2 concurrent streams |
| TTFT | ~4.5 s at 8K prompt, ~7 s at concurrency 2 |
| Efficient workload shape | long, continuous generation (workers work in long passes, not chatty exchanges) |
Environment variables
The endpoint and credentials are not hardcoded in opencode.json — they
come from two env vars, substituted by opencode via {env:...}:
| Variable | Meaning | Example value |
|---|---|---|
OPENCODE_VLLM_BASE_URL |
Full address of the vLLM OpenAI-compatible endpoint (scheme + host + port + /v1) |
http://127.0.0.1:8000/v1 |
OPENCODE_VLLM_API_KEY |
API key sent as Authorization: Bearer … (the local server doesn't check it; any non-empty string works) |
local |
They are exported in ~/.bashrc (block marked # opencode local vLLM endpoint). If a variable is unset, opencode substitutes an empty string and
the provider fails fast with "/chat/completions" cannot be parsed as a URL
— a useful canary that your environment is missing the export. To point the
agents at a different machine/port, change only the exports in ~/.bashrc
(no config edit needed), then start a new opencode session.
How to dispatch
Just talk to the orchestrator in natural language — in the TUI, or one-shot:
opencode run "Refactor src/parse.ts to stream instead of buffering, add tests for it,
and make sure the existing test suite still passes"
The orchestrator will decompose that into at most two self-contained worker tasks (e.g. implement + tests and run suite + verify), dispatch them in parallel, and hand you a synthesized report. Trivial asks (one command, one file read) are answered directly without dispatching workers.
The concurrency rule (why)
- Never more than 2 worker tasks in flight. Two concurrent streams is the measured sweet spot (~45 t/s each); a 3rd stream degrades all of them.
- Max 2 rounds of dispatch per goal: round 2 only for concrete gaps identified in the round-1 reports, then the orchestrator reports back.
- Failure handling: a failed/timed-out worker task is retried once with a narrower scope; if it fails again it is reported to you — no extra workers are spawned to cover for it.
- Worker prompts are always self-contained (paths, constraints, done criteria): workers cannot ask the user questions, and each round-trip costs server time.
Check server health before a multi-worker run
curl -s http://127.0.0.1:8000/health
# expect: HTTP 200 (empty body)
curl -s http://127.0.0.1:8000/v1/models | grep -o '"id":"qwen38"'
# expect: "id":"qwen38"
If /health is not 200, do not start a multi-worker run — restart the vLLM
server first. The server is single-user and local by design: no load
balancing, no server-level retries, no queueing.
Notes
- opencode loads config at startup: after editing
opencode.json, quit and restart opencode. - Backups of previous configs live next to it (
opencode.json.bak-*,opencode.json.pre-orchestrate.bak).