Ruilin Tong proposes MILES, a framework that dynamically expands step-wise memory for self-improving LLM reasoning at test time. The system stores modular memory units of sub-goal embeddings and sub-instructions with learnable selection heads, using a coarse-to-fine retrieval mechanism to expand memory and guide reasoning. Experiments show MILES consistently matches or outperforms prior methods while achieving superior accuracy-efficiency tradeoffs, demonstrating effectiveness, robustness, and transferability under realistic test-time constraints.
No score is assigned. Sources and their independence are shown in the citation chain below.