Error Attribution
Combines metric-level and case-level evidence to identify the highest-priority weakness for the current iteration.
An Autonomous Post-Training Harness for Iterative Model Self-Improvement
ImproveAnyTask improves Base and Instruct models across knowledge, reasoning, instruction following, coding, and tool-use tasks.
Adapting a general-purpose LLM to a target task usually requires repeated human decisions about error analysis, data construction, and training. ImproveAnyTask autonomously organizes this process under a limited compute budget equivalent to eight NVIDIA H20 GPUs for 24 hours.
The harness turns evaluation evidence into focused, research-backed, and executable model updates, then repeats the process as the model's error distribution changes.
Combines metric-level and case-level evidence to identify the highest-priority weakness for the current iteration.
Compares candidate strategies around the same weakness using reported gains, task fit, and reproduction difficulty.
Builds training data and configurations, checks execution at small scale, trains, evaluates, and retains reusable assets.
Read the full paper on arXiv, explore the public project page, or follow development on GitHub.
@article{yao2026improveanytask,
title={ImproveAnyTask: An Autonomous Post-Training Harness for Iterative Model Self-Improvement},
author={Yao, Xingbo and Wang, Xiaoman and Lei, Zhengwu and others},
journal={arXiv preprint arXiv:2610.06347},
year={2026}
}