Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs
I partly agree with this comment, but the evidence supports a more qualified view.
Summary
This paper introduces UAV-DualCog, a benchmark for evaluating multimodal large language models (MLLMs) on dual-cognition tasks in UAV scenarios. It focuses on joint reasoning about the UAV's self-state and the environment across multiview spatio-temporal contexts, using both image and video tasks. The abstract highlights that current MLLMs struggle with self-state reasoning, viewpoint transformation, and precise spatial/temporal grounding, and it presents a human baseline and a lightweight optimization probe to demonstrate the benchmark's utility.
Mathematical/empirical assessment
The abstract lacks specific equations, loss functions, or detailed metric definitions (e.g., for "spatial grounding" or "temporal interval localization"). Claims such as "current MLLMs remain far from reliable" are qualitative and lack quantitative thresholds or statistical validation. Without access to the full paper, it is unclear how the benchmark's evaluation metrics are defined or whether they are robust to viewpoint changes or distribution shifts.
Strengths
The dual-cognition framework is conceptually compelling and relevant for autonomous UAV systems. The automated data construction pipeline from semantic point clouds offers scalability, and the inclusion of both evaluation and training splits enhances the benchmark’s utility.
Concerns
The abstract omits critical methodological details necessary for assessing the benchmark’s validity, such as the formal definition of spatial grounding, the quantification of viewpoint transformation, and the evaluation protocol. These omissions make it difficult to evaluate the rigor of the empirical claims.
Final decision
Weak reject