GeoReform: Evolving the Formalization Itself for Multimodal Geometry Solving

Published · AI Daily — AI-assisted deep research, methodology & disclosure

Multimodal LLMs misread geometry diagrams. GeoReform shows text conversion can backfire: on 200 Geometry3K examples, structure injection fixes 28 errors but adds 13. It optimizes formalization from failures. Qwen3VL-2B rises from 42.0% to 56.0%.

Multimodal large language models have long struggled with geometry, and the reason is not mysterious. The decisive information sits inside the diagram: which lines are perpendicular, which angles are equal, which arc belongs to which circle. Models often misread these relations, and they use them wrongly even when they read them correctly. The common remedy of recent years is to translate the entities, relations and constraints of a diagram into explicit text, then let the model reason over that text. The idea is intuitive and it works. Yet a paper published on 8 October 2026, GeoReform (arXiv 2610.12391, by Jialu Wang, Ruichen Zhang, Xiaoou Liu, Hua Wei and Tianlong Chen), asks a sharper question: how good is the translation itself, and who checks it?

The paper offers a small but telling set of numbers. On 200 examples from Geometry3K, structure injection repairs 28 problems that the base model got wrong, but it also breaks 13 problems that the model previously solved. The net effect is positive, yet the cost is real. The authors explain the mechanism in plain terms. Redundant relations distract the model. Ambiguous references to diagram elements lead it to apply a constraint to the wrong object. The bottleneck is therefore not the extraction of more geometric facts. It is the organization of those facts into a representation that supports the next step of reasoning. Giving a model more information is not the same as giving it the right information. From this diagnosis, GeoReform changes the status of formalization. It is no longer the fixed output of a parser. It becomes a policy that the system can optimize. The loop is simple to state. The system executes the full reasoning pipeline with the current formalization policy. It collects the failed rollouts. It diagnoses defects in the current representation. Then it mutates the policy so that the representation selects, grounds, groups and presents geometric entities, relations, constraints and targets in a better way. Each of the four verbs maps to a distinct failure. Selection decides which facts survive. Grounding settles which diagram element a statement refers to. Grouping decides how facts relate to each other. Presentation decides how the final text is laid out for the model to read.

The key design choice is where the learning signal comes from. A classic parser aims to describe the image faithfully, but a faithful description is not always a useful one. GeoReform lets failed reasoning tell the system which representations mislead the model. In this sense it belongs to the wider family of reflective and evolutionary approaches to prompt and policy optimization, where reviewing past errors becomes a built-in part of the system. One caution is in order. The abstract does not describe the mutation operators, the search budget or any ablation of individual components. Those details sit in the full text, and we do not speculate about them here. The headline result favors small models. On Geometry3K, Qwen3VL-2B improves from 42.0% to 56.0% accuracy, a gain of 14 points. A model at the two-billion-parameter scale gains that much from a better representation alone. This suggests that a sizeable share of its errors do not come from a lack of reasoning ability. They come from a problem statement that was represented badly before reasoning began. The authors also report extensive experiments and analyses across geometry reasoning benchmarks, and they conclude that effective formalization is crucial for multimodal geometry reasoning. Still, one model on one benchmark is not a general law. Whether the same size of gain holds for larger models and other benchmarks needs independent replication.

For the wider industry, the lesson goes beyond geometry. Any task that turns visual or semi-structured input into text before reasoning faces the same tension: extracting information and organizing information are two separate jobs. Chart question answering, circuit diagram understanding and engineering drawing review all fit this pattern. GeoReform suggests that treating the intermediate representation as a first-class object to optimize may pay off more than simply scaling the model. A practical step for engineering teams is to keep failed cases in a structured log and use them to revise the input representation, not only the prompt. That path is cheap, auditable and easy to iterate, and it deserves a place on today's technology radar.

Sources

FAQ

What core problem does GeoReform address?

It argues that the bottleneck in geometry reasoning is not extracting more facts but organizing them into a useful representation. On 200 Geometry3K examples, structure injection fixes 28 errors yet adds 13 new ones. So the authors treat formalization as a policy that can be optimized from failures.

How does the optimization loop work?

The system runs the full reasoning pipeline, collects failed rollouts, diagnoses defects in the current representation, and mutates the policy. The mutated policy selects, grounds, groups and presents entities, relations, constraints and targets better. The signal comes from final reasoning results.

What do the results show, and what remains unknown?

Qwen3VL-2B improves from 42.0% to 56.0% on Geometry3K, so representation quality matters a lot for small models. The abstract gives no data on larger models, other benchmarks or component ablations, so independent replication and a full-text read are needed.