UniSkill: Actor-Aligned Skill Proposals Push ALFWorld Success to 98.4%
UniSkill trains one shared policy to act in the environment and to propose skillbank edits: Add, Update or No Edit. The actor learns from environment reward. The proposer learns from contrastive action feedback: how swapping the retrieved skill for the proposed one changes the actor's log-likelihood gap between successful and failed trajectories of the same task. No extra rollout is needed per proposal. The paper reports 98.4% success on ALFWorld and 84.7% on WebShop, and says the method still works with a smaller backbone. Authors tested it on ALFWorld and WebShop with Qwen2.5 backbones, and ablations show that joint training and aligned feedback are both needed.
UniSkill is a paper submitted to arXiv (cs.AI) on 7 October 2026 by Yifei Lu and co-authors. One clarification first: it studies how an LLM agent maintains and improves a skillbank.
It is not about robot arm manipulation. An LLM agent can store reusable skills, distilled from past interactions, in a skillbank and retrieve them on new tasks. The hard part is training 'doing the task' and 'writing the skill' together so that they help each other and do not interfere.
The problem: skill benefit and actor progress get mixed
Recent methods optimize task execution and skill extraction jointly. A natural way is to reward a skill proposal by the benefit it brings when it is reused later. The paper points out two troubles.
First, the benefit of the skill is confounded with the improvement of the actor. A rising success rate may mean the skill got better, or it may mean only that the policy got stronger. Credit cannot be separated. Second, testing each proposal directly needs an extra batch of rollouts per proposal, so the cost grows with the number of proposals.
Core design: one shared policy, two roles
UniSkill uses one shared policy for two jobs: act in the environment, and propose edits to the skillbank from the resulting trajectories. There are only three edits: Add, Update and No Edit. In the paper's setup the skillbank starts empty in every run and grows only through proposals made during training. The retriever is Qwen3-Embedding-0.6B and returns the top-1 skill. The shared policy backbone is Qwen2.5-7B-Instruct, with a smaller Qwen2.5-3B-Instruct variant.
The two roles learn from different signals. The actor learns from environment reward with a GRPO-style clipped loss, and the advantage is normalized within the group of rollouts for the same task. The skill proposer is guided by contrastive action feedback and trained with REINFORCE++.
Mechanism: an alignment reward that needs no new rollout
This is the key step. For a skill proposal, UniSkill does not run the environment again. It makes a counterfactual comparison on trajectories that already exist. It takes successful and failed reference trajectories of the same task. For each, it computes how the token-normalized log-likelihood of the actor's actions changes when the retrieved skill is replaced by the proposed one.
Call these changes Delta+ (successful trajectories) and Delta- (failed trajectories). The alignment reward is R_align = Delta+ minus Delta-. The intuition is simple: a good skill should make the actor favor actions from successful trajectories, and should not make it favor actions from failed ones. Because this needs only forward passes over stored trajectories, no new environment interaction is required per proposal. The paper also uses a frozen skill critic model (DeepSeek-V4-Pro). Sensitivity to the critic was compared across three critics, but only on WebShop.
Preventing exploration collapse: skill-edit support regularization
Proposal-level feedback has a side effect. It can push the probability of some edit operations down, for example until No Edit or Update is almost never sampled, and exploration collapses.
The paper adds L_sup, a squared hinge penalty on any edit operation whose probability falls below a floor p_min of 0.1. The proposer loss is L_proposer = L_R++ + lambda_sup * L_sup. The joint objective is L_UniSkill = L_actor + lambda_prop * L_proposer, with lambda_prop of 0.5 and lambda_sup of 0.01.
Training setup
The optimizer is AdamW with a constant learning rate of 1e-6 and a clip range epsilon of 0.2. Each step samples 16 tasks with G = 8 rollouts per task, for 250 training steps, after 3 warm-up steps that learn only the output format.
Training temperature is 1.0 and evaluation temperature is 0.4. Episodes have at most 50 steps on ALFWorld and 15 on WebShop.
Results
On ALFWorld, UniSkill reaches an overall success rate of 98.4% (plus or minus 0.8, three runs). On WebShop it reaches a score of 90.5 (plus or minus 1.4) and a success rate of 84.7% (plus or minus 0.5). For comparison, Table 1 of the paper lists GRPO at 77.6% on ALFWorld, GiGPO at 90.8%, Skill1 at 97.5%, RetroAgent at 94.9% and Evolving-RL at 93.0%.
On WebShop success, SkillRL is at 72.7%, Skill1 at 82.9% and RetroAgent at 82.3%. The text states gains of 20.8 points over GRPO and 7.6 points over GiGPO on ALFWorld, and 12.0 points over SkillRL on WebShop. With Qwen2.5-3B-Instruct, ALFWorld validation success still exceeds 90% and ends above GRPO with the 7B model, though the paper gives no exact final figure for the 3B run.
Ablation: how much comes from feedback and joint training
Table 2 (ALFWorld success) gives a clear breakdown. With only the actor trained and the proposer frozen, success is 84.4% without retrieval and 87.5% with it. With only the proposer trained and the actor frozen, it is just 34.4% and 32.8%.
UniSkill without the R_align reward gets 84.4% and 89.1%. Full UniSkill gets 96.1% and 98.4%. The lesson is that joint training and aligned feedback are both needed, and training the proposer alone is nearly useless.
Impact for developers and the ecosystem
For teams building LLM agents, the paper offers a reusable idea: treat 'update memory or skills' as an action of the policy, not as an external post-processing step, and score those actions with a cheap counterfactual signal.
This avoids the cost of an extra rollout per proposal, which matters most in expensive environments. A skillbank is also readable text, so it is easier to audit and to edit by hand than experience compressed into weights.
Limits and cautions
The paper has no dedicated limitations section, but several points stand out. First, the alignment signal is noisy: among proposals with positive R_align, 24.5% (13/53) at step 25 and 15.6% (10/64) at step 75 had negative observed rollout gains. Second, the Evolving-RL reproduction collapsed in extended training, which shows that joint training carries a stability risk. Third, most baseline numbers are quoted from earlier papers and were not rerun under identical conditions.
Fourth, some gaps are small. The margin over Skill1 on ALFWorld is only 0.9 points, while standard deviations are between 0.8 and 1.4. Fifth, the analysis of skill-edit distributions comes from one run per setting, and critic sensitivity was tested only on WebShop. Read these results as a strong signal, not as a final verdict.