← Back to the wire

Your Agent Aced the Task. Will It Do It Again?

AchievementResearchSep 15, 2026

A ReAct agent using GPT-4.1 achieved 77.4% average success on AppWorld across five runs but completed all five runs for only 53.0% of tasks, a 24.4-point consistency gap. The authors introduce the Consistency Analyzer, which resamples recorded trajectories to identify flip-prone decisions, and consistency guidelines that halve the gap to 12.0 points without reducing average accuracy. They argue standard Mean@k benchmarks hide this unreliability, unlike Pass^k, which requires success on every run.

Receipt № 19241 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

GPT-4.1Model
Canonical: https://huggingface.co/blog/ibm-research/altk-evolve-consistency