The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in LLMs

Published · AI Daily — AI-assisted deep research, methodology & disclosure

This paper introduces "Mathematical Primitives" to diagnose how deeply LLMs actually understand math, rather than judging them by final-answer accuracy alone. The authors build a four-dimensional benchmark spanning Discovery, Generation, Digestion, and Execution, finding that Discovery, identifying which primitive a problem needs, is the dominant bottleneck, while explicitly supplying the right primitive sharply unlocks latent execution ability. Building on this, they propose a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning to a student model, consistently beating baselines across model scales and difficult math benchmarks.

Background and Problem Definition

Large language models have posted striking scores on frontier mathematical benchmarks, from olympiad-style problems to multi-step proof assistance, yet accuracy alone tells us little about whether a model actually possesses structural mathematical understanding or is instead stitching together memorized patterns and heuristics that happen to produce the right final number. This paper introduces the notion of a "Mathematical Primitive" — the minimal structural unit of knowledge a solution depends on, such as a specific lemma, a transformation trick, or a pivotal intermediate construction — as a probe for that deeper understanding.

Building on this notion, the authors propose a diagnostic benchmark that decomposes mathematical reasoning into four distinct capability dimensions: Discovery (can the model identify which primitive a problem requires), Generation (can it construct that primitive from scratch), Digestion (can it correctly internalize a primitive once it is handed to it), and Execution (can it carry out the remaining computation once the right primitive is in hand). This framing directly targets a long-standing blind spot in the field: outcome-only accuracy collapses four very different failure modes into a single pass/fail signal, so post-training teams end up throwing undifferentiated data at a model without knowing which cognitive stage is actually broken.

Architectural Core and Technical Principles

The diagnostic methodology proceeds in two stages. First, capability profiling tests the same problem set under two conditions — unaided, and with the correct primitive explicitly supplied — and compares the resulting accuracy gap. This isolates two qualitatively different failure modes: a model that genuinely lacks the needed primitive versus one that has access to it but cannot execute on it correctly. Second, bottleneck localization measures the marginal contribution of each of the four dimensions to overall solution accuracy by holding the others constant.

Three findings stand out. First, solution accuracy masks a highly differentiated capability profile: two models with nearly identical final accuracy can have very different Discovery, Generation, Digestion, and Execution scores. Second, once the correct primitive is explicitly supplied, models show a sharp jump in demonstrated execution capacity, implying that a large amount of latent computational ability is being withheld not by an inability to compute but by a failure to identify which tool to invoke in the first place. Third, across all four dimensions, Discovery is the dominant bottleneck — models are frequently not failing to calculate, but failing to recognize which mathematical primitive the problem calls for.

Practical Evaluation and Applications

Building on these diagnostics, the authors conduct a post-training attribution analysis showing that Discovery-limited failures — cases where a model has the execution capacity but fails to self-identify the required primitive — are disproportionately easy to repair through targeted training, with a far better return on training investment than other failure categories.

This insight motivates a primitive-privileged self-distillation framework: rather than transferring an entire reasoning trace indiscriminately, as conventional self-distillation does, the teacher model explicitly annotates which mathematical primitive it is invoking at each step, and only this primitive-guided reasoning signal is selectively distilled into the student. Extensive experiments across multiple model scales and challenging mathematical benchmarks show the framework consistently outperforming baseline post-training methods, validating the broader diagnose-then-repair paradigm over undifferentiated fine-tuning.

Industry Impact and Outlook

The implications extend well beyond mathematics. The work offers a transferable methodology for converting black-box post-training tuning into targeted intervention grounded in an interpretable capability profile.

For teams building reasoning-focused models, the immediate takeaway is to diagnose whether failures stem from Discovery, Generation, Digestion, or Execution before blindly scaling training data or lengthening chain-of-thought traces. Looking forward, this four-dimensional diagnostic lens could extend naturally to other domains that depend on structured knowledge retrieval under reasoning, such as code generation and scientific hypothesis verification, positioning it as a candidate standard tool for post-training quality assessment.

Sources

FAQ

What are the four capability dimensions in the diagnostic benchmark?

Discovery (identifying which primitive a problem needs), Generation (constructing that primitive from scratch), Digestion (internalizing a given primitive), and Execution (completing the remaining computation once the primitive is known).

What is the paper's central diagnostic finding?

Solution accuracy masks a highly differentiated capability profile, and Discovery is the dominant bottleneck in mathematical reasoning; explicitly supplying the correct primitive sharply unlocks latent execution capacity that models already possess.

How does the proposed self-distillation framework differ from conventional self-distillation?

Instead of transferring an entire reasoning trace indiscriminately, the teacher model explicitly annotates which mathematical primitive it uses at each step, and only this primitive-guided reasoning signal is selectively distilled into the student model.