Agentic configuration for opencode putting my dual Intel Arc B70 setup to use.
  • TypeScript 85%
  • Shell 15%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-15 22:07:16 +02:00
agent Changed model definision intoQwen3.8-27B-Int4 2026-09-15 22:07:16 +02:00
plugin Add self-sync plugin to keep SOUL.md and MEMORY.md in sync across machines 2026-09-01 23:35:17 +02:00
completion.zsh Added completion for opencode 2026-09-01 22:47:15 +02:00
opencode.json Changed model definision intoQwen3.8-27B-Int4 2026-09-15 22:07:16 +02:00
README.md Read vLLM endpoint and API key from OPENCODE_VLLM_BASE_URL / OPENCODE_VLLM_API_KEY env vars 2026-08-30 23:11:16 +02:00

Local vLLM Orchestrator + Worker Setup

An orchestrator/worker agent setup for opencode that runs entirely on a local OpenAI-compatible vLLM server — Qwen3.8-27B on 2× Intel Arc Pro B70 (TP=2). No cloud APIs anywhere in the config.

Components

  • opencode.json — provider/model config only (no agent bodies):
    • Provider openai-compatible → model qwen38; the endpoint and API key come from environment variables (see below), not from this file. Every completion request carries chat_template_kwargs: {"enable_thinking": true, "thinking_budget": 2048} (verified to pass through @ai-sdk/openai-compatible into the vLLM body). Request timeout: 600 s (long thinking blocks at ~50 t/s take minutes).
    • default_agent: orchestrator (the agent bodies live in agent/*.md).
  • agent/orchestrator.md (primary, default agent): decomposes the user goal into 1–2 (max 2) self-contained worker tasks, dispatches them concurrently to qwen-worker, reads the reports, verifies, and synthesizes — or dispatches one follow-up round for concrete gaps.
  • agent/qwen-worker.md (subagent, task: deny — no sub-spawning): executes one well-scoped task end to end (investigate, implement, fix, verify) in long focused passes, and reports back in a fixed format: Summary / Files changed / Verification / Open issues.

Both agents run on the same model: openai-compatible/qwen38. Agent files are the canonical, editable form of the definitions; after editing agent/*.md, restart opencode (new sessions pick up changes).

Server facts (design around these)

Fact Value
Base URL $OPENCODE_VLLM_BASE_URL (full address: scheme + host + port + /v1)
API key $OPENCODE_VLLM_API_KEY (server accepts any non-empty string, e.g. local)
Max concurrent requests 4 (max-num-seqs); practical sweet spot: 1–2
Throughput ~53 t/s per stream alone; ~45 t/s each at 2 concurrent streams
TTFT ~4.5 s at 8K prompt, ~7 s at concurrency 2
Efficient workload shape long, continuous generation (workers work in long passes, not chatty exchanges)

Environment variables

The endpoint and credentials are not hardcoded in opencode.json — they come from two env vars, substituted by opencode via {env:...}:

Variable Meaning Example value
OPENCODE_VLLM_BASE_URL Full address of the vLLM OpenAI-compatible endpoint (scheme + host + port + /v1) http://127.0.0.1:8000/v1
OPENCODE_VLLM_API_KEY API key sent as Authorization: Bearer … (the local server doesn't check it; any non-empty string works) local

They are exported in ~/.bashrc (block marked # opencode local vLLM endpoint). If a variable is unset, opencode substitutes an empty string and the provider fails fast with "/chat/completions" cannot be parsed as a URL — a useful canary that your environment is missing the export. To point the agents at a different machine/port, change only the exports in ~/.bashrc (no config edit needed), then start a new opencode session.

How to dispatch

Just talk to the orchestrator in natural language — in the TUI, or one-shot:

opencode run "Refactor src/parse.ts to stream instead of buffering, add tests for it,
and make sure the existing test suite still passes"

The orchestrator will decompose that into at most two self-contained worker tasks (e.g. implement + tests and run suite + verify), dispatch them in parallel, and hand you a synthesized report. Trivial asks (one command, one file read) are answered directly without dispatching workers.

The concurrency rule (why)

  • Never more than 2 worker tasks in flight. Two concurrent streams is the measured sweet spot (~45 t/s each); a 3rd stream degrades all of them.
  • Max 2 rounds of dispatch per goal: round 2 only for concrete gaps identified in the round-1 reports, then the orchestrator reports back.
  • Failure handling: a failed/timed-out worker task is retried once with a narrower scope; if it fails again it is reported to you — no extra workers are spawned to cover for it.
  • Worker prompts are always self-contained (paths, constraints, done criteria): workers cannot ask the user questions, and each round-trip costs server time.

Check server health before a multi-worker run

curl -s http://127.0.0.1:8000/health
# expect: HTTP 200 (empty body)

curl -s http://127.0.0.1:8000/v1/models | grep -o '"id":"qwen38"'
# expect: "id":"qwen38"

If /health is not 200, do not start a multi-worker run — restart the vLLM server first. The server is single-user and local by design: no load balancing, no server-level retries, no queueing.

Notes

  • opencode loads config at startup: after editing opencode.json, quit and restart opencode.
  • Backups of previous configs live next to it (opencode.json.bak-*, opencode.json.pre-orchestrate.bak).