← Back to the wire

Hill climbing to glory: using evals to improve AI error rate by 7x

AchievementProductOct 8, 2026

Hex reduced its Quick Edits feature's wrong-edit rate from 21% to 3% by hill-climbing against 1,800 eval cases. The small model (GPT-6 Luna or Claude Haiku 4.5) proposes chart edits and hands complex requests to the Hex agent. Hex says over 50% of improvements came from expanding the validator, not prompt changes, and cites Anthropic's eval design guide as a starting point.

Receipt № 22941 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

AnthropicCompanyGPT-6 LunaModelHexCompanyClaude Haiku 4.5Model
Canonical: https://hex.tech/blog/we-used-evals-to-improve-ai-feature/