arXiv:2610.06347 · Autonomous Post-Training

ImproveAnyTask

An Autonomous Post-Training Harness for Iterative Model Self-Improvement

Xingbo Yao*, Xiaoman Wang*, Zhengwu Lei, Tinghui Luo, YiLin Zhang, Yuefeng Wu, Yijie Xu, Tianfu Wang, Qingyuan Zhan, Ye Guo, Daoxin Zhang, Zhe Xu, Jian Liu†, Hui Xiong†
* Equal contribution · † Corresponding authors

Paper GitHub Project Page PDF
ImproveAnyTask results across eleven benchmarks and five capability domains

ImproveAnyTask improves Base and Instruct models across knowledge, reasoning, instruction following, coding, and tool-use tasks.

+18.29 ppMean gain on Qwen3.5-4B-Base
+11.97 ppMean gain on Qwen3.5-4B-Instruct
+41.96 ppMaximum gain across 11 tasks

Task-specific improvement with less manual iteration

Adapting a general-purpose LLM to a target task usually requires repeated human decisions about error analysis, data construction, and training. ImproveAnyTask autonomously organizes this process under a limited compute budget equivalent to eight NVIDIA H20 GPUs for 24 hours.

A connected optimization loop

The harness turns evaluation evidence into focused, research-backed, and executable model updates, then repeats the process as the model's error distribution changes.

01 · DIAGNOSE

Error Attribution

Combines metric-level and case-level evidence to identify the highest-priority weakness for the current iteration.

02 · RESEARCH

Update Direction

Compares candidate strategies around the same weakness using reported gains, task fit, and reproduction difficulty.

03 · EXECUTE

Model Updates

Builds training data and configurations, checks execution at small scale, trains, evaluates, and retains reusable assets.