Hex reduced its Quick Edits feature's wrong-edit rate from 21% to 3% by hill-climbing against 1,800 eval cases. The small model (GPT-6 Luna or Claude Haiku 4.5) proposes chart edits and hands complex requests to the Hex agent. Hex says over 50% of improvements came from expanding the validator, not prompt changes, and cites Anthropic's eval design guide as a starting point.
No score is assigned. Sources and their independence are shown in the citation chain below.