jcode: Fitting 20 Parallel Coding Agents into 91 MB with a Shared Daemon
jcode is a terminal coding-agent harness written in Rust under the MIT license for Linux, macOS and Windows. Its sessions share one daemon instead of one process each. In the author's own headless benchmark, 20 concurrent sessions use 90.6 MB against 3376.8 MB for Claude Code, and time to first frame is 14.0 ms. It also ships a vector memory graph, a swarm mode where the server notifies agents when files shift under them, and support for subscription OAuth and OpenAI-compatible endpoints. All numbers are self-reported and not yet independently reproduced.
Overview: a coding-agent harness that treats memory footprint as a first-class metric
jcode (1jehuang/jcode) is a terminal coding-agent harness written in Rust, released under the MIT license for Linux, macOS and Windows. The project describes itself in two lines: the most RAM-efficient harness, and the most intelligent harness. It competes in the same space as Claude Code, Codex CLI, OpenCode, GitHub Copilot CLI, Cursor Agent, pi and Antigravity CLI, but it makes a different bet. Most tools invest in prompts and tool chains. jcode first makes the runtime itself extremely light, then builds a memory graph, multi-agent collaboration (swarm) and a fast terminal renderer on top.
Installation is one line. On macOS and Linux run `curl -fsSL https://jcode.sh/install | bash`. On Windows 11 run `irm https://jcode.sh/install.ps1 | iex` in PowerShell. Inside the TUI, `/update` downloads the latest stable release in the background and reloads with the session preserved. From a shell, `jcode update` does the same. The README also documents a no-downgrade policy: older or equal versions are skipped, a development build compares its Git commit with the release tag, and if ancestry cannot be verified locally or through GitHub, the update stops instead of risking a downgrade. The default is `features.update_channel = "stable"`. Only an explicit `"main"` channel follows the source branch.
Core architecture: one shared daemon, many thin clients
The most important design choice is that sessions share a single daemon. The README's headless measurement says it plainly: jcode sessions share one daemon, while Claude Code runs one `claude -p --input-format stream-json` process per session. The process model decides the slope of the memory curve.
The author's headless benchmark works like this. Each session completes five real model turns (a file listing, a file read, a repository search, a summary and a reply) with tool calls. Then the total PSS of every process is measured. Both tools used claude-sonnet-4-6. The results: with 1 session, jcode uses 32.6 MB and Claude Code uses 261.0 MB (8.0x). With 5 sessions, 51.0 MB against 908.6 MB (17.8x). With 10 sessions, 66.7 MB against 1749.7 MB (26.2x). With 20 sessions, 90.6 MB against 3376.8 MB (37.3x). Each additional session costs jcode about 3.1 MB and Claude Code about 164 MB, roughly 54 times more. The run is dated 2026-09-29 and used jcode v0.89.19-dev (default build, local embeddings not compiled in) and Claude Code 2.1.267. It can be reproduced with `python3 scripts/bench_headless_memory.py`.
The slope matters more than any single number. If you treat agents as workers that you launch in bulk, for example 20 parallel tasks, the Claude Code model needs about 3.4 GB, while the jcode model needs about 91 MB. Memory stops being the limit on parallelism. Model quota and API cost take its place.
Interactive performance: 14 ms to first frame
Besides memory, the README measures startup latency over 10 interactive PTY launches. Time to first frame: jcode 14.0 ms (range 10.1 to 19.3 ms), Antigravity CLI 383.5 ms, pi 590.7 ms, Codex CLI 882.8 ms, OpenCode 1035.9 ms, GitHub Copilot CLI 1518.6 ms, Cursor Agent 1949.7 ms and Claude Code 3436.9 ms (range 2032.7 to 8927.2 ms), about 245.5 times slower than jcode. Time to first input: jcode 48.7 ms against Claude Code 3512.8 ms, about 72.2 times.
The interactive memory table is more nuanced, and more honest. With one session, jcode with local embedding off uses 27.8 MB. With local embedding on it uses 167.1 MB, compared with 144.4 MB for pi, 140.0 MB for Codex CLI and 386.6 MB for Claude Code. So once the local semantic embedding is enabled, jcode is not smaller than pi or Codex for a single session. The advantage appears as you scale. With 10 sessions jcode uses 260.8 MB (117.0 MB with embedding off), Codex CLI 334.8 MB, pi 833.0 MB, Claude Code 2300.6 MB and OpenCode 3237.2 MB. The extra PSS per added session is about 10.4 MB for jcode, 21.6 MB for Codex CLI, 76.5 MB for pi, 212.7 MB for Claude Code and 318.4 MB for OpenCode.
Memory system: embeddings plus a memory graph
Most of jcode's claimed intelligence comes from agent memory. According to the README, every turn and response is embedded as a semantic vector. Each turn queries a graph of memories with a cosine-similarity check to find related entries. Hits are fed into the conversation, or optionally a memory sideagent first verifies that they are relevant.
Writing is asynchronous too. When semantic drift appears, when K turns have passed since the last extraction, or when the session ends, a sideagent extracts memories and stores them in the graph. The harness also exposes explicit memory tools, so the agent can search or store actively and does not depend only on the passive background process. It adds session search, which is classic RAG over earlier sessions. Finally, ambient mode consolidates memories every so often: it reorganizes them and checks for staleness and conflicts. The design separates passive recall, active read and write, and background maintenance.
Swarm: server-aware multi-agent collaboration
Start two or more agents in the same repository and the server manages them for native collaboration. The key mechanism is conflict awareness. When agent A edits a file that agent B has read, so the code shifts under B's feet, the server notifies B. B can ignore the notice if it is irrelevant, or re-read the file. Agents can also call a swarm tool to spawn their own teammates. The main agent then becomes a coordinator and the spawned agents become workers.
Configuration separates root reasoning from worker effort. In `~/.jcode/config.toml`, under `[agents]`, `swarm_root_effort` (used by `/effort swarm`) and `swarm_deep_root_effort` (used by `/effort swarm-deep`) both default to `max`. Accepted levels are none, minimal, low, medium, high, xhigh and max, mapped to each provider's supported range. Environment overrides are `JCODE_SWARM_ROOT_EFFORT` and `JCODE_SWARM_DEEP_ROOT_EFFORT`. These settings do not change the worker `swarm_effort`. You can let the root think little and the workers work hard, or the reverse, depending on task cost.
Providers and ecosystem: subscription login plus OpenAI-compatible endpoints
jcode supports subscription-backed OAuth flows, so you reuse the model quota you already pay for and fall back to direct API providers when needed. Built-in logins cover Claude, OpenAI/ChatGPT/Codex, Google Gemini, GitHub Copilot, Azure OpenAI, Alibaba Cloud Coding Plan, Fireworks, Novita AI, MiniMax, Meta Model API, LM Studio, Ollama and any custom OpenAI-compatible endpoint. Built-in OpenAI-compatible profiles include openrouter, deepseek, zai, kimi, moonshotai and opencode. The native OpenAI providers use Responses WebSocket v2 with opportunistic background prewarming and an HTTPS fallback.
For scripts and agents, the `jcode provider add` command writes a profile in one step. It supports `--api-key-stdin` (the key stays out of shell history), `--api-key-env`, `--context-window` and `--no-api-key` for local servers such as vLLM or Ollama. The `extra_body` setting and the `JCODE_OPENAI_EXTRA_BODY` variable inject non-standard top-level fields into requests. One example in the README is NVIDIA NIM with DeepSeek-V4, which only enables thinking when `chat_template_kwargs` is present. Anthropic Messages-compatible gateways are supported through `type = "anthropic-compatible"` profiles with bearer, custom-header or no authentication. The default stream idle timeout is 180 seconds and scales with reasoning effort (high 2x, xhigh 3x, max 4x).
For MCP, jcode reads `~/.jcode/mcp.json` and a project-level `.jcode/mcp.json`. It also reads Claude Code's `~/.claude.json` and the repository `.mcp.json` live on every load, with no stale snapshot. Only stdio servers are supported today. HTTP and SSE entries are recognized and skipped. Each request times out after 30 seconds by default, adjustable with `timeout_secs`. Migrating from Codex CLI imports `~/.codex/config.toml` once, and the imported environment values may contain secrets.
UI engineering: custom components built for speed
The author wrote a mermaid rendering library with no browser or TypeScript dependency (mermaid-rs-renderer) and claims it renders diagrams 1800 times faster, so flow charts can appear inline in the side panel and the chat. The `panel` tool opens desktop panels from Markdown or PDF.
Info widgets use only the empty space on screen and step aside when there is none. jcode claims it can render at over a thousand frames per second, which avoids flicker. A custom scrollback adds capability, but a terminal-level limit prevents smooth partial-line scrolling, so the author also built a terminal of their own, Handterm.
Impact, limits and observations
For developers, jcode pushes the marginal cost of a parallel agent close to zero. Running dozens of sessions on a laptop is no longer a memory problem. For enterprises, OAuth subscription reuse, Anthropic-compatible and OpenAI-compatible gateways, custom headers and self-hosted vLLM make it possible to plug into internal gateways.
Several cautions apply. First, every benchmark was produced by the project author on their own Linux machine, and no independent reproduction is cited. The tables also use different versions: the headless test used jcode v0.89.19-dev and Claude Code 2.1.267, while the interactive memory rerun lists jcode v0.9.1888-dev and Claude Code 2.1.86. Read each table in its own context. Second, PSS measures memory, not task quality. The claim of the most intelligent harness is a self-description, and the README data does not show better problem solving than competitors. Third, single-session memory is not the lowest once local embeddings are on, so the advantage depends on the shared daemon and on the embedding-off configuration. Fourth, MCP over HTTP/SSE is not supported yet, and the README gives no separate numbers for Windows or macOS. Fifth, a shared daemon concentrates the failure domain: a daemon crash could affect every session, and the README does not discuss this trade-off.
Overall, jcode points in a useful direction: make the agent shell real lightweight infrastructure first, then talk about memory and collaboration. If you run large-scale parallel coding agents, batch reviews in CI, or work on a low-spec machine, run `scripts/bench_headless_memory.py` yourself and check the numbers on your own hardware.