CADFather: A Vision-Language Agent That Coordinates Tools to Rebuild Parametric CAD from Meshes
CADFather is an agent that rebuilds editable parametric CAD programs from 3D meshes without any extra training. A Qwen3.8-27B vision-language assistant inspects renders and routes work among three tools: learned operation proposals from CADENA-RL, algorithmic geometric proposals, and numerical parameter optimization, while it keeps the best valid result. On the full DeepCAD, Fusion360 and MCB test sets it reports zero invalid outputs and mean IoU of 0.987, 0.976 and 0.913, above CADENA-RL with sampling. On BenchCAD it reaches voxel IoU 0.968 at an estimated 0.046 US dollars per part.
Background: from a mesh to an editable CAD program
A mesh describes a shape. It does not record how the shape was built. When a part exists only as a scan or a legacy mesh, an engineer cannot change a bore diameter or a wall thickness, because no construction history exists. CADFather attacks this reverse engineering problem. Given a target mesh, it recovers an executable, structurally coherent, editable parametric CAD program. The paper comes from Lomonosov Moscow State University, Innopolis University and FusionBrain Lab. It is arXiv 2610.09127, submitted on 6 October 2026, and the code is public at kulibinai/CADFather on GitHub.
Existing methods each have a strength and a weakness. CAD-Recode, cadrille and CADEvolve generate a full program in one pass. CADReasoner revises a full program using discrepancy feedback. CADENA adds one operation at a time. CADFit fits geometry inside an optimization search. The paper's central observation is simple: no single source of CAD operations works equally well across all part geometries and all stages of reconstruction. A learned model proposes plausible operations but may miss a feature or use wrong dimensions. An algorithmic construction derives operations straight from the target geometry. Numerical optimization tunes the dimensions of a program that is already close. So the real question becomes: at each step, which candidate should be developed, and with which tool?
Core architecture: a vision-language assistant plus three tools
CADFather is an agent that needs no extra training. The assistant is Qwen3.8-27B with FP8 weights and explicit thinking turned off. The learned operation generator is the CADENA-RL checkpoint, served in bfloat16. Both models stay frozen. Every tool writes the same representation, the CadQuery-based DSL of CADENA. Because of this, a candidate made by one tool can be extended or refined by another.
The three tools have clear jobs. First, learned proposal generation: for a parent program, one request returns n alternative next operations. The assistant picks n, up to 32 per call. Second, algorithmic proposal generation: this tool uses geometric fitting on the target mesh and no learned model. On the empty start, it checks whether the part is a body of revolution. If not, it cuts the target with planes, treats each cross-section as a sketch, and finds the depth over which that sketch follows the target surface. It drops extrusions that add little overlap. On an existing candidate, it compares candidate and target, finds missing material and extra material, and builds extrusions for the first and cuts for the second. It returns a ranked list, four proposals per call. Third, parameter optimization: it changes numbers only and keeps the operations. The free parameters are sketch coordinates, radii and extrusion depths.
The decision loop and the candidate pool
For each part, the system keeps a pool of candidates. Each entry stores its identifier, parent, source tool, program, mesh, operation count, validity and measurements. Extending a candidate creates a child and never rewrites history. Code-level deduplication stops identical programs from entering twice. At each step the assistant sees a bounded table: up to twelve top-scoring candidates, plus the protected best and the empty root. Each row lists the producing tool, the operation count, IoU and GMS, the score of the first operation in its chain, and the tools still available. The assistant also sees renders of the target and of the latest candidates. It writes one sentence of intent, then issues up to four act calls, or calls finish with the reason shape (all features present) or stalled (no more progress).
A tool response is only a proposal. The execution component builds it, validates it and measures it. Then the assistant inspects the render and decides whether to admit it to the pool. The validity rule is strict: the program must build, and the whole mesh must be a watertight solid with positive volume. Target and candidate share one frame, and a candidate is not rescaled to its own bounding box, so a wrong size costs score. The score is the mean of IoU and normalized GMS, and an invalid candidate scores zero. The final output is the protected best result, not the last candidate the assistant touched. Fixed budgets guarantee termination: 24 assistant steps, at most 12 operations, 1000 seconds, 80 generator samples, 6 algorithmic calls, 10 optimizer calls and 120 program executions per part.
The mathematics of the optimizer
The optimizer works on signed distances: the distance from a point to the solid surface, negative inside. The program's distance is computed directly from its operations. The target's distance is computed from the mesh at sampled points. The mismatch is the mean squared difference of the two distances.
The gradient is estimated numerically: each parameter moves by a small step, and the change in mismatch gives its derivative. The optimizer first fits points spread over the whole part, then points near the surface. It keeps whichever of the starting, coarse and fine parameter sets matches best near the surface, so the program can come back unchanged. It supports extrusions, holes and revolves only. The authors state that they do not introduce a new numerical optimizer.
Benchmark results
On the full DeepCAD, Fusion360 and MCB test sets (8046, 1725 and 5000 parts), CADFather has an invalidity ratio of zero. Mean IoU is 0.987, 0.976 and 0.913, against 0.966, 0.952 and 0.895 for CADENA-RL with sampling. Mean GMS is 0.983, 0.960 and 0.766. Median Chamfer distance at 30000 points (scaled by 1000) is 0.039, 0.034 and 0.083, against 0.042, 0.036 and 0.089. The learned tool is the same CADENA-RL checkpoint, so the gain does not come from a stronger generator. One caution: the published CADENA rows use a more lenient validity rule than CADFather. The invalidity ratios are not directly comparable.
On CADENA-Bench (3396 parts), CADFather reaches IoU 0.909 and GMS 0.721, against 0.876 and 0.670 for CADENA-RL greedy. On CADBench, scored with the benchmark's own code over 18000 parts, it reaches IoU 0.930, surface IoU 0.787 and Chamfer distance 0.029, ahead of CADFit at 0.859, 0.679 and 0.038. Its valid shape rate, 0.968, is below the 0.981 of CADFit, because 576 parts exceed the evaluator's 30 second limit. On BenchCAD, voxel IoU is 0.968 with no invalid programs.
Ablations: algorithmic proposals matter most
On 1000-part subsets of each dataset, removing algorithmic proposals lowers the mean score by 0.012 to 0.110. On MCB the invalidity ratio rises from 0 to 0.124. The reason is concrete. On some parts, every first operation sampled from the generator gives an open surface, and the search cannot continue from a non-watertight base.
An extrusion of a target cross-section gives a closed solid to build on. Removing the optimizer costs only 0.004 to 0.006 and leaves invalidity at zero. So in this system, tool diversity mostly buys robustness, while the optimizer adds the last small gain in fit. The authors admit that the ablations do not isolate the value of the assistant's selection policy.
Cost and latency
Mean assistant steps per part are 6.37 on DeepCAD, 8.41 on Fusion360, 12.25 on MCB and 13.04 on CADENA-Bench. Wall time is 82.6, 119.9, 245.6 and 306.7 seconds. Generator samples range from 19.0 to 67.1. The assistant ends the part itself in 86.1% of DeepCAD cases but only 25.0% on CADENA-Bench, so complex mechanical parts are often cut off by the budget.
On BenchCAD, priced at Qwen3.8-27B OpenRouter rates (0.425 and 2.55 US dollars per million input and output tokens), CADFather reaches voxel IoU 0.943 after four steps at about 0.016 dollars per part, and 0.968 at full budget for about 0.046 dollars. The paper cites Claude Opus 5.5 and GPT-6 Astra with tools at 0.962 and 0.959, costing 5.94 and 1.75 dollars per part. Those frontier numbers are vendor self-reports with different inputs and subsets, and the authors call the comparison contextual. The CADFather estimate counts model tokens only, not geometry processing or execution.
What it means for developers and enterprises
For manufacturing and industrial software teams, this approach turns legacy scans and history-free meshes into parametric models that people can edit. It needs no custom-trained model.
It combines an existing generator, geometric algorithms and a general vision-language model, so it can run in a private environment on open weights. The modular design also means any tool can be swapped: a stronger generator or a faster fitting routine plugs in directly. The wider lesson is that in verifiable domains such as geometry, an agent earns its value through routing and budget allocation, not through the raw power of one model.
Limitations and future work
The authors list several limits. The system depends on the operations its CAD representation supports and on the proposals its tools can make. Parameter optimization cannot fix a wrong operation sequence, and keeping old candidates does not help if a good alternative was never generated. Search limits can stop a run that would have succeeded. The assistant decides from incomplete visual and numerical evidence, so results may depend on the prompt and the model. Typical failures include a wrong initial topology (a solid block rebuilt as a hollow frame), matched volume with missing surface detail (knurls, shallow pockets), and a helical spring rebuilt as a tube. The evaluation measures geometry and program validity only. It does not measure design history, manufacturing intent or engineering constraints, and it covers single parts, not assemblies, tolerances or simulation. A direct comparison with other CAD agents such as CAD-Assistant and IterCAD, and a controlled study of the assistant's tool-selection policy, remain future work.
One more point for careful readers: the final pick uses the same IoU and GMS metrics that the paper reports, so it is a best-of-N selection by the reported metrics. Keep this in mind when you quote the numbers.