From Long Text to Predictive Features: LLM-BlockFE Turns Offline Program Search into Risk-Control Gains

Published · AI Daily — AI-assisted deep research, methodology & disclosure

LLM-BlockFE uses the LLM offline only. Immutable code blocks form feature programs, scored by a downstream model. Block rollback and interleaved search avoid greedy traps. AUC rises 0.0069 to 0.0358; KS rises 0.02 to 1.56 points in five live apps.

Industrial risk-control systems run on structured-data models for a good reason. They are fast, cheap to serve, easy to monitor, and they fit the approval and audit chains that regulated finance demands. Yet a great deal of valuable signal sits outside those tables, buried in unstructured long text such as application narratives, communication records and free-form descriptions. Getting that signal into a tabular model has historically meant manual feature engineering: domain experts write rules and extraction scripts, test them, and repeat. The work is slow, and its coverage is capped by what the experts happen to think of. The obvious alternative is to hand every incoming long text to a large language model (LLM) at inference time. In a real-time risk pipeline that option often fails on latency, throughput, cost and stability before accuracy even enters the discussion. A paper posted to arXiv on 8 October 2026, "Long Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program Search" by Ziming Dai and colleagues, proposes a third path called LLM-BlockFE.

The central idea is simple to state. Use the LLM offline, and make its output a program rather than a prediction. LLM-BlockFE builds a feature program by appending code blocks one at a time, and each block is immutable once written. After each append, the candidate features that the program now produces are scored by a downstream model, so the search is driven by measured predictive value and not by the LLM's own opinion of its work. When the search ends, the programs are frozen and deployed. At serving time they extract structured features from long text, and the existing downstream model consumes those features as usual. No LLM call happens online. The expensive intelligence is paid for once, during search, and amortised across every later request. This turns the LLM from an inference engine into something closer to a tireless feature engineer, and it leaves the production path as cheap and predictable as it was before.

The hard part is the search itself. Programs grow block by block, and conventional greedy search keeps whatever looks best at the current step. That habit is fragile here. An early block can look useful and still steer every later block into a poor region, and a greedy procedure has no way to back out. The authors address this with a block-level rollback mechanism driven by depth-calibrated credit allocation. The intuition is that the final gain of a program is not spread evenly over its blocks. Early blocks shape everything that follows, and naive credit assignment tends to misjudge them. Calibrating credit by depth is meant to correct that bias, so that the system can decide which block is the real culprit and rewind the program to just before it. The abstract names the mechanism and its purpose but not the calibration formula or thresholds, so those details should be read in the full paper and not guessed.

The second design choice concerns scale. LLM-BlockFE advances several independent search trajectories in an interleaved fashion. Multiple trajectories reduce the odds that a single unlucky path dominates the outcome, but a naive version wastes budget, because each trajectory may rediscover the same ideas. The paper reduces that redundancy by sharing a fixed description of each trajectory's exploration direction. Every trajectory can see where the others are heading, so effort spreads out instead of piling up. In effect, a short and stable piece of text works as a coordination signal between searches. This matters because every LLM call costs money, and a repeated exploration is a call that bought no new information. Rollback fixes the vertical problem of being stuck on a bad path. Interleaving with shared descriptions fixes the horizontal problem of paths colliding. Together they form a fairly complete search discipline for a space where each step is expensive to evaluate.

The reported results cover two public and two private datasets. In the full-dataset comparison, LLM-BlockFE improves AUC by an absolute 0.0069 to 0.0358 over the strongest baseline on each dataset. The more telling evidence comes from deployment. Across five financial risk-control applications that have gone live, post-launch monitoring shows absolute KS improvements of 0.02 to 1.56 percentage points over the existing manually designed strategy. In credit and fraud work, small gains in ranking power translate into real changes in loss rates or approval rates, so numbers in this range are not trivial. They are also not uniform. The lower end of each range is small, which suggests that the benefit depends heavily on how much usable signal a given text field carries. Readers should also note two limits. The private datasets cannot be reproduced outside the authors' organisations, and post-launch monitoring against an existing strategy is not the same as a controlled experiment.

Several design properties make the approach attractive for regulated settings. The output is executable code, which can be reviewed, versioned, tested and replayed. A feature in production is no longer the opaque result of a model call. It is a program that an auditor can read and an engineer can rerun on historical data. Freezing the programs after search also gives stable behaviour: the same text yields the same features tomorrow, and drift can be monitored with ordinary tools. Because the LLM never touches live traffic, sensitive text does not need to leave the serving environment for a third-party endpoint at request time, and an outage at the model provider cannot take down the decision path. These are practical advantages that accuracy figures alone do not capture. Open questions remain, and they are the interesting ones. How should frozen programs be maintained when the input text distribution shifts, and how often should the search be rerun? Could LLM-written extraction logic leak label information or encode unwanted bias, and how would a team detect that before launch? Do features discovered for one downstream model transfer to another, or is each search tied to the evaluator that guided it? How much of the gain comes from the search machinery, and how much from simply giving an LLM many tries with a reliable scoring signal? The abstract does not answer these, and the paper's ablations will matter. Even so, the broader lesson is clear. For latency-bound and compliance-bound domains, the most deployable role for a large model may be as an offline builder of auditable components, not as an online judge. LLM-BlockFE is a concrete, evidence-backed example of that pattern, and it deserves attention from anyone who owns a feature pipeline.

Sources

FAQ

Why does LLM-BlockFE avoid calling an LLM online?

The LLM works only offline. It writes executable feature programs block by block, and a downstream model scores them. After the search the programs are frozen. In production they extract structured features from long text on their own, so online inference needs no LLM call and keeps the latency and cost of a traditional feature pipeline.

What do block-level rollback and depth-calibrated credit allocation solve?

Plain greedy search keeps only the current best step, so an early block that looks useful can lock the program into a poor solution. Block-level rollback lets the system rewind to before the faulty block. Depth-calibrated credit allocation helps decide where to rewind, since early blocks shape everything after them. The exact formula is in the full paper.

What do the results show, and what are the limits?

On two public and two private datasets, AUC rises by an absolute 0.0069 to 0.0358 over the strongest baseline. Across five deployed financial risk-control applications, KS rises by 0.02 to 1.56 percentage points over the manual strategy. The private data cannot be reproduced, and post-launch monitoring is not a controlled experiment.