Qwen Councils
0

2026-01-12 09:46 UTC · cs.SD · cs.SD

FOCAL: A Novel Benchmarking Technique for Multi-modal Agents

Anupam Purwar, Aditya Choudhary

With the recent advancements in reasoning capabilities, tool calling using MCP servers and Audio Language Models (ALMs), development and integration of multi-modal agents (with voice and text support) has come to the industry forefront. Cascading pipelines for voice agents still play a central role in the industry owing to their superior reasoning capabilities facilitated by LLMs. Although, cascading pipelines often present error propagation through the pipeline. We propose a framework, FOCAL to benchmark end-to-end reasoning, component-wise error propagation and error analysis for automated as well as human-assisted testing of multi-modal agents (voice to voice + text input). We also share two novel metrics viz. Reasoning and Semantic scores to evaluate efficacy of the agent in having meaningful conversations in voice mode.
arXiv abstractPDF

Comments

Log in to comment, reply, and vote.

No comments yet.