Apache Maka: An Auditable Agent Workspace Where the Append-Only Event Log Is the Runtime
Apache Maka (Incubating) is a high-performance agent workspace that writes every model message, tool call, permission decision and termination as an append-only RuntimeEvent. The interface, the next prompt and crash recovery are all projections of that log. Desktop, TUI, CLI and Eval share one Runtime Host, data stays local, and you bring your own model. The project promises per-task benchmark results on the same model and official verifier. The current release is 0.2.0, and Windows and Linux are still previews. The README also lists clear build requirements: Node.js 22.19 or newer, npm, Git and ripgrep. Direct Peer features add a Rust toolchain. Model access is configured in Settings, with API, local and gateway options, so teams with strict data rules can keep code and records inside their own network.
Apache Maka (Incubating) is a high-performance agent workspace built around one promise: it keeps a complete record of everything the agent did. It is not another chat shell.
It is an engineering design for an agent runtime plus thin clients. This report draws on the project README and public descriptions, and it covers the architecture, the working mechanism, the fit for teams, and the limits.
The problem it targets
An agent harness exists to finish tasks. The README holds it to two measures: how many tasks it completes, and at what cost. Maka makes both measures public.
The project says it benchmarks against other harnesses on the same model with the same official verifier, and that the per-task results ship with every report in the docs/eval folder of the repository. That moves the claim from "we are better" to "read the result of every single task". Most agent projects publish only an aggregate score, so this habit is uncommon.
Core idea: the log is the runtime
The central design choice is that the log is the runtime. Every model message, tool call, permission decision and termination is written as an append-only RuntimeEvent. The user interface, the next prompt and crash recovery are projections of that log. None of them is the only copy of the truth.
This follows the event sourcing pattern, and it brings three benefits. First, crash recovery is simple, because state can be replayed from the log. Second, the interface and the prompt builder do not each keep a private state, so they cannot drift apart. Third, and most useful for agents: old tool output can leave the next prompt without leaving the log. In long tasks, tool output tends to fill the context window. Maka can trim it from the prompt to save tokens while the full audit trail stays on disk. Context management and auditability are two needs that usually pull against each other. Maka puts them in two different layers, so each one can be satisfied.
One Runtime Host
Maka has a single Runtime Host, which is the one execution authority. The desktop app, the TUI and CLI, and the Eval harness are thin clients of that host.
Eval owns only the experiment and its scores. This layering means the behavior you see on the desktop comes from the same execution path that produced the benchmark results. There is no second code path for evaluation, so the usual gap between "what we measured" and "what we ship" has less room to grow.
Local first, bring your own model
Sessions, settings and run records stay on your machine. Maka does not bundle a shared model account. On first launch you open Settings, then Models, add an API, a local model or a supported account connection, test it, and pick a default. The app separates three connection states: configured, send-ready and experimental.
An account flow that is not wired into the runtime is not shown as a usable model. That is a modest and honest choice, because it avoids overstating compatibility. The model can be a cloud API, a local model or a compatible gateway. A team with sensitive code can therefore point Maka at a local model and keep both the code and the record inside its own network.
Building and running it
Building from source needs Node.js 22.19 or newer (CI uses Node.js 24), npm, Git, and ripgrep, which the runtime Grep tool depends on. After npm ci, the command npm run dev starts the desktop development environment with hot reload. For the terminal, npm run cli:dev opens the TUI, and npm run cli:dev -- run runs a single turn.
A --graph option exists for graph-style tasks. The Direct Peer and Peer Mesh features need Rust stable 1.98 or newer and a platform linker, with separate dev:peer entry points. On platforms, macOS is the main target, while Windows and Linux are marked as previews.
Project status and ecosystem meaning
Maka is an Apache Software Foundation incubating project. The latest Apache release is 0.2.0 (incubating).
The signed source archive is the official release. Packages distributed elsewhere are convenience artifacts, and development builds are not approved Apache releases. The Apache 2.0 license and the foundation governance process help enterprise adoption: the license is clear, and the community process is documented.
Limits and risks
First, the README excerpt gives no concrete benchmark scores or cost figures. Readers must open docs/eval to inspect per-task results, and this report repeats no number that it could not verify. Second, the project is still incubating.
The version is 0.2.0 and two of three platforms are previews, so a production rollout needs time for validation. Third, a complete append-only log raises disk and privacy questions. The log stores model messages and tool output, which can hold secrets or private code, so teams need their own retention and redaction rules. Fourth, any claim of high performance should be reproduced on your own tasks, because a public benchmark is not your workload.
Where it may go
Auditable agents will matter more over time. When an agent can edit files and run commands, each permission decision and each step in the chain of cause must be traceable.
Maka treats that chain as a first-class part of the runtime, not as a logging feature added later. That is the most interesting part of the project. If the open evaluation practice holds up over releases, it could also push agent harnesses toward more transparent comparisons with one another.