MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination
MARGIN (Multi-Agent Runtime Grading via Incremental Normalisation) is a single-author paper by Joss Armstrong. It learns a model-specific confidence correction at runtime from observed answer outcomes. It needs no retraining and no held-out calibration set. The author reports that on BigCodeBench, a model's mean confidence is negatively related to its accuracy, and that picking the more confident responder performs below chance. Corrected scores then weight candidate answers in a collective decision. The evaluation uses 18 models across code generation, question answering and mathematics.
A common shortcut in multi-agent systems goes like this: ask several models the same question, then keep the answer from the one that sounds most sure. The shortcut rests on one assumption. It assumes a model's stated confidence moves in the same direction as its real accuracy.
The MARGIN paper challenges that assumption head on. The name stands for Multi-Agent Runtime Grading via Incremental Normalisation. The author is Joss Armstrong, the arXiv identifier is 2605.22949, version 1 appeared on 21 May 2026, and the latest version, v4, was revised on 8 October 2026.
The problem: sure does not mean right
The author makes a plain observation on BigCodeBench: a model's mean confidence is negatively related to its accuracy. Put simply, the models that speak with more certainty are, on average, wrong more often. Now look at pairs. In each pair one answer is right and one is wrong. If you always pick the more confident answer, you do worse than chance. For any orchestration layer that arbitrates by who sounds surer, this is an uncomfortable result.
The cause is not mysterious. Different models do not share one scale of certainty. Some habitually report high numbers. Others lean cautious. Comparing those raw numbers side by side is like subtracting a reading in Celsius from a reading in Fahrenheit.
The core mechanism: incremental normalisation at runtime
According to the abstract, MARGIN learns a correction for each model, at runtime, from answer outcomes it has already observed. It does not retrain any model. It does not need a held-out calibration set. In practice, the system keeps two quantities per confidence band: the recent accuracy, and the confidence the model stated. The ratio of the two becomes a correction applied to the reported score in that band. The words in the name map roughly onto this design. The correction is incremental, because it updates with every new outcome. It is a normalisation, because it pulls every model's numbers onto one comparable scale.
The corrected confidence then weights the candidate answers in a collective decision, which is to say a vote. MARGIN changes neither the models nor the voting frame. It inserts a thin conversion layer between them. That makes it easy to slot into existing orchestration code. One caution: the specific hyperparameters, such as window length and band boundaries, are not stated in the page we could read. Check the full paper for them.
Setup and reported results
The paper uses a pool of 18 models and takes a nine-model subset for the distribution-shift experiments. Tasks cover code generation, question answering and mathematics. The author also sets up five online calibration baselines and gives them exactly the same feedback as MARGIN, so the comparison stays fair.
On expected calibration error after the shift, MARGIN is lower than all five baselines in two code-generation transitions. In one question-answering transition it is lower than four of them, and the remaining comparison is described as inconclusive. That is an honest statement: the method does not win everywhere. For answer-selection accuracy, the gains are 4.3 and 14.0 percentage points on two of the three benchmarks, relative to uncalibrated confidence weighting. The abstract gives no figure for the third benchmark, and it gives no latency or cost data.
What it means for developers and enterprises
First, it targets cascading hallucination in tool chains. In a long agent pipeline, a confident error upstream becomes a fact downstream. If the arbitration signal itself is distorted, the error grows as it travels. Calibrating confidence before you use it is a cheap line of defence.
Second, it does not depend on offline data. Many teams have no ready calibration set, or their model versions change fast and old calibrations go stale. MARGIN corrects itself from online feedback, which suits pools that change often.
Third, it gives abstention a firmer basis. A calibrated score can decide when to refuse an answer or hand it to a person. Setting a threshold on a distorted number would not do that job.
Limits and cautions
First, it depends on outcome feedback. The system must learn afterwards whether an answer was correct, for instance whether a test passed or a reference answer exists. For open-ended writing, or any task without a way to verify, this condition fails.
Second, the paper page discloses no latency or compute overhead, so engineering teams must measure it. Third, this is a single-author preprint with four revisions, and we saw no evidence of peer review. Fourth, the author admits that one question-answering comparison is inconclusive, so the method may not win under every kind of shift. Finally, 18 models and three task families are a decent spread, but they do not cover every production workload.
Summary
The value of MARGIN does not lie in a complex new model. It lies in taking an assumption everyone relies on, testing it in the open, and offering a light repair. For teams building multi-agent orchestration, the practical step is simple.
First measure on your own tasks whether confidence and accuracy move together. Then decide whether to add this calibration layer. If they do not move together, this paper deserves a careful read.