ServiceNow-AI researchers introduced EVA, a framework for evaluating conversational voice agents that jointly scores task accuracy and conversational experience across multi-turn spoken interactions. Developed by Tara Bogavelli, Gabrielle Gauthier Melancon, Katrina Stankiewicz, Nifemi Bamgbose, Hoang Nguyen, Raghav Mehndiratta, Hari Subramani, and Fanny Riols, EVA uses a bot-to-bot audio architecture with an initial airline dataset of 50 scenarios. Benchmarking 20 systems revealed a "consistent Accuracy-Experience tradeoff" between task completion and user experience.
No score is assigned. Sources and their independence are shown in the citation chain below.