Quesma's benchmark of RTK on Terminal-Bench 2.1, covering 1,740 attempts, found token filtering does not reliably reduce coding costs: Claude Code costs fell 5% while OpenCode costs rose 5%, and pass rates dropped slightly. JetBrains's SkillsBench run found no savings. Quesma notes RTK's "rtk gain" metric counts removed output, not billed tokens, and terminal output is a small share of total cost.
No score is assigned. Sources and their independence are shown in the citation chain below.