A study across eight models found that agentic memory benefits depend on model capability, with strong models gaining from full guideline sets and weaker models performing best with selective retrieval. DeepSeek-V3.2 improved by 9.5 percentage points in task completion with full guidelines, while gpt-oss-120b gained 16.1 points using curated retrieval at only 5% additional token cost. The ALTK-Evolve framework distills reusable guidelines from prior trajectories without weight updates, and already-saturated models showed no measurable improvement.
No score is assigned. Sources and their independence are shown in the citation chain below.