text-to-cad: Giving AI Agents a Parametric CAD Toolkit

Published · AI Daily — AI-assisted deep research, methodology & disclosure

text-to-cad is a library of agent skills that turns plain-language or image requests into CAD models. Each part is Python source built on build123d and the OpenCascade kernel. Running the file produces a STEP file as the main output, with STL, 3MF and GLB exports. The source records the design intent, so models can be reviewed, diffed and regenerated. The library ships twelve skills and plugins for four agent hosts. This analysis reads the repository's design and names its risks, including silent loss on unmigrated models and a Windows blocking issue. The tool itself was not run here.

Background and Problem Definition

Mechanical design still depends on graphical interfaces. An engineer opens a CAD program, sketches a profile, adds constraints, extrudes the shape, and exports a STEP file for the machine shop. A language model can write Python and read a specification, yet it cannot easily reproduce that sequence of clicks. The practical question is therefore clear: can an agent turn a plain-language request, or a photograph of a part, into a three-dimensional model that another person can check, change and manufacture?

text-to-cad answers with a deliberate design choice. The model is source code. Each part is a Python file written with build123d, a Python CAD library that runs on the OpenCascade kernel. Running the file produces the artifact. The STEP file is the main output, and the source records the design intent. Geometry becomes something that can be compared, reviewed and regenerated, rather than a file that can only be replaced.

The hard problem is verification, not generation. A model can look correct in a viewer and still place a hole in the wrong spot. The repository therefore treats measurement, snapshot review and validation as first-class skills, not as afterthoughts. That choice matters more than any single generation feature.

Architectural Core and Technical Principles

text-to-cad is a library of agent skills, not a single program. It ships twelve skills: CAD, step.parts, Engineering Drawing, DXF, URDF, SRDF, SDF, SendCutSend, DfAM Check, DFM, G-code and Bambu Labs. The CAD skill sits at the centre. Most of the others consume its output or describe the same parts for a different tool. The version constraints are explicit. The skill's requirements file pins cadgen[snapshot] to version 0.7.11. The README badges name build123d 0.11, Open CASCADE 7.9, Python 3.11 or newer, and Node.js 20 or newer. An agent that runs the skill therefore depends on a particular kernel and runtime. The pin lets a team reproduce a published release. The cost is that upgrades remain manual work.

The stack has three layers. At the top, the agent reads SKILL.md, which routes each task to a short reference file. In the middle, the cadgen command-line tool and its MCP server perform execution. The MCP server starts through uvx, so the first run downloads the runtime. At the bottom, the OCP binding exposes OpenCascade to Python, handling geometric operations and STEP input and output. The governing principle is that the source file is the truth. SKILL.md tells the agent to edit the model source and rerun it with python, not to patch an exported file. A change is therefore a diff in a Python file, while meshes and STEP files are derived results. This is the same reasoning that makes code reviewable, and it fits version control well. Distribution follows the same logic. Plugins exist for Codex, Claude Code, Cursor and Grok Build, and skills install with npx skills add. Each host displays models in its own way, but all of them call the same cadgen package.

Practical Evaluation and Applications

I did not install the tool or generate a model with it. Everything in this section comes from the repository's documentation, configuration and layout. It is therefore a reading of the design, not a measurement of its output. The CAD skill names four capabilities. It creates and edits parts and assemblies from a decorated Python model. It exports STL, 3MF and GLB, either through a mesh decorator or a one-off build command. It resolves named references in a saved STEP document through read_scene and scene.resolve. It renders snapshots for visual and motion review, which requires Chromium. A sensible workflow uses all four: generate the model, export it, check dimensions with a native build123d script, and inspect the snapshot before release.

Two warnings in the documentation matter for production use. First, a model that needs a version migration and is left unmigrated silently loses its kinematics, materials and animation. A silent loss is worse than a crash, so a team should make migration an explicit pipeline step. Second, on Windows 11 with Smart App Control enabled, the unsigned OCP wheel is blocked and imports fail with a DLL error. The repository's answer is to turn that feature off or run the skills under WSL. Teams that only have Windows workstations should test this before rollout. Two further practical points apply. Usage analytics stay off until a user opts in, and the DO_NOT_TRACK variable disables them. Plugin installation downloads a local runtime through uv on first start. Neither is a defect, but both belong on an organisation's install checklist. A workable adoption plan has five parts. Keep the generated Python under git. Pin cadgen and record the version. Run cadgen doctor as an environment check. Write one assertion script per critical dimension. Treat the snapshot as a review aid, not as proof of fit.

Industry Impact and Outlook

The project points in a direction more than it delivers a finished product. CAD becomes a set of engineering assets that code, tests and version control can constrain. For small hardware teams, robotics groups and makers who already write Python, parametric design becomes much more accessible. It also shifts responsibility. When a model generator produces a plausible dimension, the risk moves from the person who drew the part to the process that checks it. A team that adopts this approach needs measurement gates before manufacturing. The DFM, DfAM Check and SendCutSend skills point the same way, because they test a part against manufacturing rules instead of trusting the geometry.

The robot-description skills raise a second question, which is consistency across formats. URDF, SRDF and SDF describe the same robot for different consumers. Keeping CAD geometry, joint limits and simulator models in agreement over time remains an open engineering problem. Gathering them in one skill library is a sound structure, but the consistency checks remain the user's job. At the time of this snapshot the project has 16,697 stars, which signals strong interest. Interest is not validation. The useful evidence would be measured rework rates, check pass rates and delivery times on real projects. This analysis contains none of that data. Readers should treat the project as a promising interface to test, not as a proven replacement for a CAD workstation.

Sources

FAQ

What form does a text-to-cad model take?

Each part is a Python file built on build123d. Running it produces a STEP file as the main output, and the source is the design record.

How are the dependency versions controlled?

The CAD skill's requirements.txt pins cadgen[snapshot] to 0.7.11. The README badges name build123d 0.11, Open CASCADE 7.9 and Python 3.11 or newer.

What known risks does the documentation warn about?

An unmigrated model silently loses kinematics, materials and animation. On Windows 11 with Smart App Control on, the unsigned OCP module is blocked, so the feature must be disabled or the skills run under WSL.