Graphify: Turn Any Codebase into a Queryable Knowledge Graph Without a Vector Store
Graphify, a Claude Code skill, turns code, PDFs, screenshots and whiteboard photos into one persistent knowledge graph via /graphify. It claims 71.5x fewer tokens per query than raw files (self-reported) and separates found facts from guesses. Outputs: graph.html, Obsidian vault, GRAPH_REPORT.md, graph.json; a SHA256 cache skips unchanged files.
Codebases keep growing, and so do the piles of notes, papers and screenshots around them. Yet a coding agent usually starts every new session by re-reading the same raw files from scratch. That burns tokens, and it leaves no durable memory between sessions. Graphify is an open-source project aimed at exactly this problem. It is a Claude Code skill: you type /graphify inside Claude Code, it reads the files in a folder, builds a knowledge graph, and hands back structure you did not know was there. The project's own README claims 71.5 times fewer tokens per query than reading the raw files, a graph that persists across sessions, and an explicit distinction between what it found and what it guessed. The README also draws a comparison with Andrej Karpathy, who keeps a /raw folder where he drops papers, tweets, screenshots and notes. Graphify presents itself as an answer to the problem such a folder creates: material accumulates, but nothing can use it efficiently. A caution is needed here. The 71.5x figure comes from the project's own description. We have not reproduced it, so treat it as an order of magnitude and not as a guarantee.
The first defining trait is that the tool is fully multimodal. Its input is not limited to source code. PDFs, Markdown files, screenshots, diagrams, whiteboard photos and even images containing text in other languages can all be dropped in. Graphify uses Claude's vision capability to extract concepts and relationships from this mixed material and connects everything into a single graph. In practice, a hand-drawn system sketch, a design-review PDF and the source of a core module can become linked nodes in one structure, instead of living in three folders that never meet. For teams that constantly cross-check what the documents say against what the code does, a unified representation has real value. The design also explains the phrase in the project's title, no vector store. Queries rely on explicit nodes and edges, not on similarity scores in an embedding space. That choice trades the fuzziness of nearest-neighbour search for relationships that a human can inspect and challenge.
The second trait is the careful design of the outputs. Running /graphify on a folder produces a graphify-out directory with clearly separated roles. The file graph.html is an interactive view where you can click nodes, search, and filter by community. The obsidian folder can be opened directly as an Obsidian vault. With the --wiki option, the tool also writes Wikipedia-style articles meant for agent navigation. GRAPH_REPORT.md is a written report that lists the god nodes, which are the most densely connected core concepts, along with surprising connections and suggested questions to ask next. The file graph.json holds the persistent graph itself, so you can query it weeks later without re-reading the sources. Finally, a cache folder stores SHA256 hashes, so a re-run processes only the files that changed. Together these outputs serve both audiences: the human who wants to browse and the agent that wants to navigate. The incremental cache keeps the cost of daily iteration under control.
Installation is lightweight. The tool requires Claude Code and Python 3.10 or newer. You run pip install graphifyy, then graphify install. One detail can trip people up: the PyPI package is temporarily named graphifyy with two y letters, because the plain graphify name is still being reclaimed. The command line tool and the skill command keep the name graphify. On Windows, if the command is not recognised after installation, add the Python Scripts folder to PATH, or use pipx. On macOS, an externally managed environment error is also solved by pipx. A manual route exists as well. You download the skill file SKILL.md with curl into the ~/.claude/skills/graphify folder and register it in ~/.claude/CLAUDE.md. After that, you open Claude Code in any directory and type /graphify followed by a dot.
From an industry view, Graphify points to a route worth watching. Instead of handing long-term memory entirely to a vector database, it asks the model to compile the material once into an explicit graph, and later queries read only that graph. The benefits are that results can be audited, updated incrementally, and kept stable across sessions. The cost is that graph quality depends on how accurately the model extracts entities and relations, and a wrong edge, once written, will be reused. The project's stress on separating found facts from guesses is a response to that risk, but how well it works must be checked on your own corpus. The first build of a large folder also calls the model, especially for vision extraction, so the cost is not negligible, and the SHA256 cache only reduces later runs. Our recommendation is practical. Run it first on a mid-sized repository. Compare the god nodes in GRAPH_REPORT.md with your own understanding of the system. Sample a few edges and trace them back to their sources. Only then decide whether to adopt it in a daily agent workflow. If the results hold up, a graph like this could become a shared layer of context for code, documents and visual material.
Sources
FAQ
Why does Graphify say it needs no vector store?
It extracts the material once into an explicit graph of nodes and edges, saved as graph.json. Queries read those stated relations, not similarity scores in an embedding space, so results can be inspected and traced. The cost is that graph quality depends on how accurately the model extracts.
Can the 71.5x token saving be trusted?
It is the project's own claim, measured per query against reading raw files. We have not reproduced it. Real savings depend on corpus size and question type, so sample it on a mid-sized repository first.
How do I install it, and what are the pitfalls?
You need Claude Code and Python 3.10 or newer. Run pip install graphifyy, then graphify install. The PyPI name has a temporary extra y, but the command is still graphify. On Windows add the Scripts folder to PATH or use pipx. On macOS, pipx also fixes the externally managed environment error.