Pith. sign in

REVIEW 5 major objections 8 minor 52 references

Current multimodal models cannot reliably keep degraded evidence, security rules, and safe UAV actions coupled at the decision point.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 13:45 UTC pith:RBI5XKEM

load-bearing objection Solid offline UAV decision benchmark with a real multi-model gap; the headline numbers are only as strong as unvalidated action/policy labels. the 5 major comments →

arxiv 2607.23870 v1 pith:RBI5XKEM submitted 2026-07-26 cs.MA cs.AI

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

classification cs.MA cs.AI
keywords cyber-physical securitysafe action decision makingsecurity-policy compliancesmart-city UAV agentsUAV benchmarkvision-language-actionmultimodal evidence arbitrationdegradation-aware reasoning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Smart-city drones are no longer just cameras; they must choose a next action under bad visibility, messy operator language, and explicit security rules. This paper argues that existing UAV and vision-language benchmarks miss that coupling and offers MulRobBench, an offline test that forces models to recover mission context, arbitrate multimodal evidence, handle degradation, and pick a controlled safe action. On 3,024 strict samples across 17 models, the best semantic protocol-decision score is only 0.5141 and the best strict dimension accuracy is 0.1599. Failures concentrate in modality trust, constraint extraction, collaboration or abstention thresholds, and action consistency—not coarse scene recognition. A sympathetic reader should care because a fluent scene description can still violate a restricted zone, ignore missing data, or skip a required reobservation, which is exactly the cyber-physical risk the benchmark isolates.

Core claim

Across 17 uniformly audited multimodal models on MulRobBench’s 3,024 strict samples, current systems remain far from reliable protocol-conditioned UAV decision making: the best semantic protocol-decision score is 0.5141 and the best strict mean scoring-dimension accuracy is 0.1599. Models handle coarse context and targets relatively well but break on modality-trust selection, constraint extraction, collaboration and abstention triggers, and risk-aware action planning. Modality-removal on a matched 20-anchor subset changes 4–15 action selections per model, showing both vision and text shape decisions while exposing unstable combination of those inputs.

What carries the argument

MulRobBench: an offline, protocol-conditioned Vision-Language-Action decision contract that binds real UAV multimodal observations, injected security-policy mission context, and a closed ten-action safety vocabulary, scored along four links (context, evidence arbitration, degradation-aware reasoning, risk-aware action) with 12 dimensions and both semantic scores and strict structural diagnostics.

Load-bearing premise

Treating airport boundaries, sensitive-place rules, privacy limits, and other policies as injected mission context—rather than labels verified in the source imagery—still fairly measures whether a model preserves security-policy compliance when it chooses an action.

What would settle it

Find a model that, on the same 3,024-sample strict split and closed action contract, simultaneously posts high semantic protocol-decision score, high strict mean dimension accuracy, low unsafe-action rate, and near-zero normalization failures, with modality removal no longer flipping many of the 20-anchor actions.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • UAV VLA evaluation must report semantic validity, structural parseability, and action safety side by side; a single accuracy score hides decision-chain breaks.
  • Improving coarse scene or target recognition will not close the gap; gains must target modality-trust, constraint extraction, and abstention or collaboration triggers.
  • Glare, missing data, and operator shorthand are priority stress cases because they systematically decouple evidence quality from allowed actions.
  • Offline protocol-conditioned next-action audits become a necessary gate before claims of smart-city UAV cyber-physical safety.
  • Human multi-expert pilots on stratified subsets remain far above models, setting a concrete gap for future systems work.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Training objectives that reward fluent captions or open-ended plans may actively work against the structured action contract this benchmark requires.
  • If injected policies are later grounded in verified maps and entity labels, the same four-link chain could become a live mission-audit layer rather than only an offline test.
  • The non-monotonic modality-ablation results suggest some models are already over-relying on text priors; denser paired vision-text counterfactuals would expose that shortcut more sharply.
  • Similar decision-contract thinking likely transfers to other rule-bound embodied settings (ground robots in restricted sites, inspection agents) where evidence degradation and forbidden actions coexist.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 8 minor

Summary. The paper introduces MulRobBench, an offline benchmark for protocol-conditioned safe next-action selection by multimodal UAV agents in smart-city settings. Built on UAVScenes-derived observations with injected protocol semantics (restricted zones, privacy/media rules, temporary mission rules), the benchmark comprises 3,024 strict samples organized into 17 primary-attribution taxonomy nodes and 12 scoring dimensions (D1–D12) across four decision links. Seventeen multimodal models are evaluated under a shared decision contract with separated semantic scoring (Eq. 10–11) and strict structural diagnostics (Eq. 17, 20), plus action-level metrics (SafeAcc, UnsafeRate, MAD, PACS). Headline results: best semantic protocol-decision score 0.5141, best strict mean dimension accuracy 0.1599; a 20-anchor modality-removal study changes 4–15/20 actions per model. The authors attribute failures primarily to modality-trust selection, constraint extraction, collaboration/abstention thresholds, and action-rationale consistency rather than coarse context recognition.

Significance. If the results hold, this is a useful and timely contribution. The decision-level endpoint (protocol-conditioned next action under a closed action vocabulary) is genuinely under-served by existing UAV benchmarks, and the paper's central methodological choice — reporting semantic scores, action-set agreement, and structural compliance side by side rather than collapsing them into one number — is well motivated and executed with unusual discipline. Specific strengths worth naming: a fixed, audited 3,024-sample strict set with a documented split contract (Tables VI–VII); 17 models under a uniform audit with normalization failures retained rather than silently dropped (Phi-4-Multimodal's 3,024 failures are reported, not removed); a 5% normalization-failure credibility rule that prevents action-only fallback numbers from being over-read; per-dimension, per-condition, and error-chain analyses (Tables XIII–XVII) that localize failures to the middle of the decision chain; and explicit, honest limitation statements about the injected-semantics premise. The multi-expert reference pilot, the conditional robustness matrix (Table XVI), and the modality-ablation study are all falsifiable, checka

major comments (5)
  1. [§III.B (Benchmark Generation, step four) and Table XI] Every metric in the paper (Eqs. 5, 10, 13–17) is scored against Γ_i = (A*_i, A+_i, A−_i) and the dimension projections g_id = π_d(x_i), yet the manuscript nowhere reports who produced these labels, under what protocol, with what adjudication, or with what inter-rater agreement. Many labels are judgment calls rather than facts: whether a dust-occluded frame warrants hover vs. reobserve vs. request-another-UAV; which constraints are 'active' under a given protocol; what the 'primary' degradation is when conditions co-occur (Table XIV shows samples carrying multiple condition labels). The Multi-Expert Reference row cannot substitute for label validation: it is a 600-sample pilot with round values (0.83/0.84/0.94/0.78), no rater count, no agreement statistic, and no statement of whether the reference annotators are disjoint from the original label authors — if they overlap, the row measures
  2. [§III.D / Fig. 5 and Eq. (10)] The controlled semantic score s_d(ŷ_im, g_id) is computed via a scoring prompt (Fig. 5), which implies an LLM-based judge, but the manuscript never states which model executes the judge, at what settings, or how the judge itself was validated against human scoring. Since all semantic metrics (S_md, P_m, the group scores in Table XII, and the conditional tables) derive from s_d, judge identity and judge–human agreement are load-bearing. Please report the judge model/version, decoding settings, and a judge-validation study (e.g., judge vs. human labels on a subsample, with agreement statistics), or clarify if scoring is rule-based.
  3. [Abstract / §IV.E.1 and Eq. (17)] The headline 'best strict mean scoring-dimension accuracy is only 0.1599' conflates two distinct failure modes the paper itself separates elsewhere: normalization/format compliance and decision competence. The manuscript documents an action-only fallback parser and a 5% credibility rule precisely because strict validity q_imd is partly a format-compliance measure. As reported in the abstract, 0.1599 reads as a decision-competence number and likely overstates the capability gap in that interpretation. Please decompose strict failures into (i) structural/normalization failures and (ii) semantically wrong but well-formed responses, at least for the top models, and temper the abstract framing accordingly (e.g., report strict accuracy alongside the share of failures attributable to format).
  4. [§IV.D and Table X] The modality-removal study supports an abstract-level claim ('confirming that both visual and textual inputs influence decisions'), but it rests on a 20-anchor subset with no stated selection procedure, no uncertainty quantification, and full results shown for only 2 of 17 models (Table X). 'Changes 4–15 of 20 action selections' on n=20 is compatible with a wide range of effect sizes. Either expand the anchor set (with a stated sampling scheme and confidence intervals or an exact test) and report all models, or downgrade the claim in the abstract to a pilot-scale sensitivity observation consistent with §IV.H's framing of the human-reference pilot.
  5. [Reproducibility (Abstract; §III–IV)] The abstract claims a 'reproducible benchmark,' but the manuscript contains no data/code release statement, no pointer to the full scoring prompt and normalization parser, and no per-model inference settings (prompts, temperature, max tokens) backing the 'uniformly audited' claim. For a benchmark paper the artifact is the contribution; please add an explicit release plan (samples, Γ_i labels, scoring code, parser, prompts) and an appendix with the evaluation prompt and per-model decoding configuration. The semantic-scoring prompt in Fig. 5 appears abbreviated; the full template should be available.
minor comments (8)
  1. [§III.A] Terminology overload: 'task families,' 'taxonomy nodes,' 'scoring dimensions,' 'dimension groups,' and 'evaluation links' are used near-interchangeably in places (e.g., 'four evaluation links' vs. the groups in Eq. 4). A single glossary or consistent naming would help readers track the taxonomy-vs-dimension distinction the paper (rightly) insists on.
  2. [§III.C, Eq. (9)] Normalized entropy H_d and the hard-case support n_hard_d are defined but never used in any subsequent table or figure. Either report them (they would be informative for class-balance auditing of Y_d) or remove the definitions.
  3. [Table XVII] Modality-trust mismatch (Gemma-4-E4B) and collaboration-decision mismatch (Qwen3-VL-4B) are both reported as 3,024/3,024. A 100% trigger rate is either a pipeline artifact (e.g., the diagnostic fires whenever the response does not exactly match the required field) or a substantive claim that needs discussion. As tabulated it is uninformative; please clarify.
  4. [Table XIV] The 'Clean baseline' row reports degradation recognition 0.0000. Presumably there is no degradation to recognize on clean samples, so the metric is undefined rather than zero; scoring it as 0 risks misleading readers. Consider a dash or an explicit note.
  5. [§IV.C, Eqs. (15)–(19)] The MAD 0.25 alternative penalty, PACS equal weights, and the 5% normalization cutoff are disclosed as conventions, which is good practice; a brief sensitivity check (e.g., do model orderings under PACS survive weight perturbation, or do rankings change if the alternative penalty ranges 0.1–0.5) would strengthen the claim that these choices are not driving conclusions.
  6. [Table I and §II.C] The coverage legend renders as garbled glyphs ('#', 'G #', blank) and 'α3-Bench' appears with a corrupted character; check symbol fonts. Also 'UA V' spacing artifacts appear throughout the extracted text — verify these are not in the source.
  7. [§IV.A] State whether model inference used zero-shot prompts identical across models, whether any model-specific chat templates were needed, and how ties/parse ambiguities in the action-only fallback parser were resolved.
  8. [§IV.E.2 / Table XII] SmolVLM2-2.2B scores 0.6255 on the Action group but 0.0992 on Context — a striking inversion worth one sentence of interpretation, since it bears on the paper's claim that action metrics alone do not establish protocol-grounded capability.

Circularity Check

0 steps flagged

Empirical benchmark paper with no derivation chain that reduces predictions or first-principles claims to their inputs by construction.

full rationale

MulRobBench is an offline evaluation benchmark: it constructs labeled samples (Oi, Ci, Ri, Γi), defines scoring dimensions D1–D12 and action metrics (SafeAcc, UnsafeRate, MAD, PACS, strict Q), and measures 17 models against those contracts. Reporting that the best semantic protocol-decision score is 0.5141 and the best strict mean dimension accuracy is 0.1599 is an empirical measurement against author-assigned ground truth, not a claimed derivation or prediction forced by the inputs. Metric definitions (Eqs. 10–19) are ordinary benchmark contracts—averages and indicators over labeled sets—not self-definitional reductions of a scientific claim. There is no fitted parameter re-presented as an out-of-sample prediction, no uniqueness theorem imported from overlapping authors to forbid alternatives, and no ansatz smuggled in via self-citation. Weaknesses in label provenance, inter-rater agreement, or semantic-scorer design (if any) are validity/correctness concerns outside this circularity pass. The paper is self-contained as an empirical leaderboard; steps is empty.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 3 invented entities

Load-bearing commitments are benchmark-design choices and evaluation conventions rather than physical laws: offline single-step decision points stand in for cyber-physical safety; injected protocol overlays define ground-truth constraints; a closed ten-action vocabulary defines safe/forbidden sets; semantic paraphrase scoring plus strict parse/action diagnostics jointly define success; and composite indices (PACS, MAD) use hand-chosen weights/penalties. Free parameters are few and explicit reporting conventions. Invented entities are benchmark constructs (taxonomy, dimensions, composites), not new physical mediators.

free parameters (4)
  • MAD acceptable-alternative penalty (0.25) = 0.25
    Eq. 15 assigns deviation 0.25 to safe alternatives vs 0/1 for exact/other; authors call it a reporting convention, not calibrated physical risk, yet it enters MAD and multiplicatively scales PACS.
  • PACS equal weights over four components = 1/4 each
    Eq. 18 averages Q_m4, Q_m5, SafeAcc_m, and Q_m11 with equal weight by design choice; different weights would reorder composite diagnostics.
  • Normalization-failure credibility cutoff (5%) = 5%
    Models above 5% normalization failure are barred from best-action markings; cutoff is stated as interpretability policy, not empirically justified safety threshold.
  • 20-anchor modality-ablation subset size = 20 anchors
    Modality-sensitivity claim rests on a hand-chosen matched 20-sample anchor set rather than the full 3,024.
axioms (5)
  • domain assumption Offline single-step next-action selection under a fixed observation and injected rule state is a valid diagnostic proxy for decision-level cyber-physical safety and security-policy compliance.
    Stated throughout §I, §III, and Limitations; closed-loop control, latency, and long-horizon missions are explicitly out of scope.
  • ad hoc to paper Protocol semantics (restricted zones, privacy/media rules, sensitive-place norms, temporary mission rules) may be injected as benchmark mission context without pixel-level entity verification in source imagery.
    §III.B and Limitations treat these as protocolised decision conditions, not visually verified labels; central compliance claims depend on this separation.
  • domain assumption A closed vocabulary of ten action semantics plus sample-level standard/alternative/forbidden sets is sufficient to audit safe vs unsafe UAV decisions.
    Table III and Eqs. 12–14; only seven standard actions appear in the strict set, with ten retained for alternatives/extensions.
  • domain assumption Controlled semantic scoring under paraphrase equivalence plus strict structural diagnostics together measure decision quality without requiring real flight certification.
    §III.D and §IV.C; authors explicitly say semantic scores are not direct flight-control certification.
  • domain assumption UAVScenes-derived multimodal observations plus readability/alignment filters yield a representative strict evaluation distribution for smart-city decision pressure.
    §III.B–C construction pipeline; long-tailed preset schedule rather than real incident frequencies (authors caution on this).
invented entities (3)
  • MulRobBench decision contract (Oi, Ci, Ri, Γi) with 17 taxonomy nodes and D1–D12 scoring dimensions no independent evidence
    purpose: Organize protocol-conditioned VLA evaluation into auditable context, evidence, degradation, and action links.
    Core benchmark ontology introduced in §III.A; not an external standard prior to this paper.
  • Protocol-decision semantic score P_m over dimensions {8,9,10,11,12} no independent evidence
    purpose: Aggregate constraint, target, safe action, abstention/reobservation, and collaboration semantics into one headline semantic metric.
    Eq. 11 defines a benchmark-local aggregate used for ranking claims.
  • Protocol-Action Composite Score (PACS) no independent evidence
    purpose: Summarize degradation judgment, modality trust, safe-set agreement, reobservation, and action deviation in one composite.
    Eqs. 18–19; authors state it is not an inherited standard metric and not a sole ranking.

pith-pipeline@v1.2.0-grok45-kimik3 · 32242 in / 4165 out tokens · 85031 ms · 2026-07-30T13:45:24.021243+00:00 · methodology

0 comments
read the original abstract

Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action risk remain coupled during critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents in smart-city environments. MulRobBench integrates real UAV multimodal observations, protocol-level security policies, and action-level cyber-physical safety into a unified evaluation framework. The benchmark contains 3,024 samples spanning 17 task taxonomy nodes and 12 scoring dimensions across four stages: operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. Evaluation combines semantic scoring with structural diagnostics, including policy compliance, format compliance, unsafe actions, parsing failures, and dimension-level validity. Across 17 multimodal models, the best semantic protocol-decision score reaches only 0.5141, while the best strict mean scoring-dimension accuracy is 0.1599. A controlled 20-anchor modality-ablation study changes 4-15 action selections per model, confirming that both visual and textual inputs influence decisions. Analysis identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability. MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints.

Figures

Figures reproduced from arXiv: 2607.23870 by Belal S. Alsinglawi, Izzat Alsmadi, Junyi Wu, Lianhai Lin, Merouane Debbah, Weizheng Wang, Yi Jiang.

Figure 1
Figure 1. Figure 1: Two-level UAV-VLA task taxonomy. The inner ring contains four task families, and the outer ring contains 17 primary-attribution task-taxonomy [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Primary-attribution taxonomy of the strict evaluation set. The figure [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Protocol-context distribution by source-domain bucket. The figure [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Observation-condition family distribution. Labels include the clean [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Semantic scoring prompt template. The template constrains physical observation, mission context, security-policy rules, degradation conditions, action [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Semantic decision contrast under ordinary patrol and emergency hazard. A readable boulevard observation permits slow conservative progress, whereas [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Boundary-restricted decision check near an airport perimeter. The case tests whether limited visual evidence and restricted-zone protocol semantics [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Degradation-aware decision check under dust-obscured boulevard patrol. The case tests whether degraded RGB evidence shifts trust toward nonvisual [PITH_FULL_IMAGE:figures/full_fig_p011_8.png] view at source ↗
Figure 10
Figure 10. Figure 10: Per-dimension result structure over D1–D12. The figure localizes model differences to context, evidence, degradation, or action links; these scoring dimensions are distinct from the 17 primary-attribution task-taxonomy nodes. can identify protocol context reasonably well, and target dis￾ambiguation is also relatively easy. In contrast, modality-trust selection and constraint extraction remain global bottl… view at source ↗
Figure 11
Figure 11. Figure 11: Result structure after scoring-dimension group compression. Context [PITH_FULL_IMAGE:figures/full_fig_p015_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Security-policy semantics and unsafe-action tradeoff. The plot [PITH_FULL_IMAGE:figures/full_fig_p015_12.png] view at source ↗
Figure 15
Figure 15. Figure 15: Degradation-specific robustness for Qwen3-VL 8B. The plot relates [PITH_FULL_IMAGE:figures/full_fig_p016_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Protocol-context action pressure for Qwen3-VL 8B. The plot [PITH_FULL_IMAGE:figures/full_fig_p016_16.png] view at source ↗
Figure 18
Figure 18. Figure 18: Unsafe action replacement failure under a fire response scenario. The [PITH_FULL_IMAGE:figures/full_fig_p019_18.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 13 linked inside Pith

  1. [1]

    UA VBench: An open benchmark dataset for autonomous and agentic AI UA V systems via LLM-generated flight scenarios,

    M. A. Ferrag, A. Lakas, and M. Debbah, “UA VBench: An open benchmark dataset for autonomous and agentic AI UA V systems via LLM-generated flight scenarios,” 2025, arXiv:2511.11252. [Online]. Available: https://arxiv.org/abs/2511.11252

  2. [2]

    Drones as a service (DaaS) for 5G networks and blockchain-assisted IoT-based smart city infrastructure,

    T. Garg, S. Gupta, M. S. Obaidat, and M. Raj, “Drones as a service (DaaS) for 5G networks and blockchain-assisted IoT-based smart city infrastructure,”Cluster Computing, vol. 27, pp. 8725–8788, 2024

  3. [3]

    Advancing UA V security with artificial intelligence: A comprehensive survey of techniques and future directions,

    F. Tlili, S. Ayed, and L. C. Fourati, “Advancing UA V security with artificial intelligence: A comprehensive survey of techniques and future directions,”Internet of Things, vol. 27, p. 101281, 2024

  4. [4]

    Toward secure complex UA V cyber-physical systems: A unified threat taxonomy and cross-layer survey of cybersecurity challenges,

    M. I. Umrani, B. Butler, A. O’ Driscoll, and S. Davy, “Toward secure complex UA V cyber-physical systems: A unified threat taxonomy and cross-layer survey of cybersecurity challenges,”Internet of Things, vol. 37, p. 101902, 2026

  5. [5]

    Cyber physical systems: Design challenges,

    E. A. Lee, “Cyber physical systems: Design challenges,” in2008 11th IEEE International Symposium on Object and Component-Oriented Real-Time Distributed Computing, 2008, pp. 363–369

  6. [6]

    Cyber-physical systems security: A survey,

    A. Humayed, J. Lin, F. Li, and B. Luo, “Cyber-physical systems security: A survey,”IEEE Internet of Things Journal, vol. 4, no. 6, pp. 1802–1831, 2017

  7. [7]

    AirCopBench: A benchmark for multi-drone collaborative embodied perception and reasoning,

    J. Zhaet al., “AirCopBench: A benchmark for multi-drone collaborative embodied perception and reasoning,”Proceedings of the AAAI Confer- ence on Artificial Intelligence, vol. 40, no. 2, pp. 1507–1515, 2026

  8. [8]

    Benchmarking neural network robustness to common corruptions and perturbations,

    D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,” inInternational Conference on Learning Representations, 2019. [Online]. Available: https://arxiv.org/abs/1903.12261

  9. [9]

    EmbodiedCity: A benchmark platform for embodied agent in real-world city environment,

    C. Gaoet al., “EmbodiedCity: A benchmark platform for embodied agent in real-world city environment,” 2024, arXiv:2410.09604. [Online]. Available: https://arxiv.org/abs/2410.09604

  10. [10]

    UrbanVideo-bench: Benchmarking vision-language models on embodied intelligence with video data in urban spaces,

    B. Zhaoet al., “UrbanVideo-bench: Benchmarking vision-language models on embodied intelligence with video data in urban spaces,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 32 400–32 423. [Online]. Available: ht...

  11. [11]

    CityNav: A large-scale dataset for real-world aerial navigation,

    J. Leeet al., “CityNav: A large-scale dataset for real-world aerial navigation,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2025, pp. 5912–5922. [Online]. Available: https://openaccess.thecvf.com/content/ICCV2025/html/Lee CityNav A Large-Scale Dataset for Real-World Aerial Navigation ICCV 2025 paper.html

  12. [12]

    CityNavAgent: Aerial vision-and-language navigation with hierarchical semantic planning and global memory,

    W. Zhanget al., “CityNavAgent: Aerial vision-and-language navigation with hierarchical semantic planning and global memory,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 31 292–31 309. [Online]. Available: https:...

  13. [13]

    MM-UA VBench: How well do multimodal large language models see, think, and plan in low-altitude UA V scenarios?

    S. Daiet al., “MM-UA VBench: How well do multimodal large language models see, think, and plan in low-altitude UA V scenarios?” 2025, arXiv:2512.23219. [Online]. Available: https://arxiv.org/abs/2512.23219

  14. [14]

    ESARBench: A benchmark for agentic UA V embodied search and rescue,

    D. Zhang, P. Chen, J. Zhou, and S. Yang, “ESARBench: A benchmark for agentic UA V embodied search and rescue,” 2026, arXiv:2605.01371. [Online]. Available: https://arxiv.org/abs/2605.01371

  15. [15]

    UA V-ON: A benchmark for open-world object goal navigation with aerial agents,

    J. Xiaoet al., “UA V-ON: A benchmark for open-world object goal navigation with aerial agents,” inProceedings of the 33rd ACM International Conference on Multimedia. Association for Computing Machinery, Oct. 2025, pp. 13 023–13 029. [Online]. Available: https://dl.acm.org/doi/10.1145/3746027.3758251

  16. [16]

    RT-2: Vision-language-action models transfer web knowledge to robotic control,

    B. Zitkovichet al., “RT-2: Vision-language-action models transfer web knowledge to robotic control,” inProceedings of The 7th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol. 229. PMLR, 2023, pp. 2165–2183. [Online]. Available: https://proceedings.mlr.press/v229/zitkovich23a.html

  17. [17]

    OpenVLA: An open-source vision-language- action model,

    M. J. Kimet al., “OpenVLA: An open-source vision-language- action model,” inProceedings of The 8th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol

  18. [18]

    ARViP: Adversarial regularization in visuomotor policies for robotic VLA purpose,

    H. Wang, B. Wu, and S. Zheng, “ARViP: Adversarial regularization in visuomotor policies for robotic VLA purpose,”IEEE Transactions on Industrial Informatics, vol. 22, no. 4, pp. 3275–3285, Apr. 2026

  19. [19]

    A survey on vision– language–action models for embodied AI,

    Y . Ma, Z. Song, Y . Zhuang, J. Hao, and I. King, “A survey on vision– language–action models for embodied AI,”IEEE Transactions on Neural Networks and Learning Systems, vol. 37, no. 7, pp. 3031–3051, Jul. 2026

  20. [20]

    UA VScenes: A multi-modal dataset for UA Vs,

    S. Wanget al., “UA VScenes: A multi-modal dataset for UA Vs,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2025, pp. 28 946–28 958. [Online]. Available: https: //openaccess.thecvf.com/content/ICCV2025/html/Wang UA VScenes A Multi-Modal Dataset for UA VsICCV 2025 paper.html

  21. [21]

    Replanning-oriented framework for efficient real-time decision-making in multi-UA V sys- tems,

    X. Hai, L. Tan, Q. Feng, H. Duan, and C. Wen, “Replanning-oriented framework for efficient real-time decision-making in multi-UA V sys- tems,”IEEE Transactions on Industrial Informatics, vol. 21, no. 7, pp. 5127–5137, Jul. 2025

  22. [22]

    Task offloading for multi-UA V asset edge computing with deep reinforcement learning,

    S. A. Zakaryia, M. A. Mead, T. Nabil, and M. K. Hussein, “Task offloading for multi-UA V asset edge computing with deep reinforcement learning,”Cluster Computing, vol. 28, 2025, article 462

  23. [23]

    Resilient event-triggered formation control and secure estimation of multi-UA V systems,

    Z. Gu, T. Yin, Q. Lu, and J. H. Park, “Resilient event-triggered formation control and secure estimation of multi-UA V systems,”IEEE Transactions on Industrial Informatics, vol. 21, no. 6, pp. 4915–4923, Jun. 2025

  24. [24]

    Authentica- tion framework for secure smart farming system deployed for sustainable development of smart cities: A review,

    A. Patwal, M. Wazid, D. P. Singh, A. K. Das, and V . B. K, “Authentica- tion framework for secure smart farming system deployed for sustainable development of smart cities: A review,”Cluster Computing, vol. 29, 2026, article 433

  25. [25]

    AerialVLN: Vision-and-language navigation for UA Vs,

    S. Liu, H. Zhang, Y . Qi, P. Wang, Y . Zhang, and Q. Wu, “AerialVLN: Vision-and-language navigation for UA Vs,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 15 384–15 394. [Online]. Available: https: //openaccess.thecvf.com/content/ICCV2023/html/Liu AerialVLN Vision-and-Language Navigation for UA VsICCV ...

  26. [26]

    MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,

    X. Yueet al., “MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2024, pp. 9556–9567. [Online]. Available: https://openaccess.thecvf.com/content/CVPR2024/html/Yue MMMU A Massive Multi-discipline Multimodal Unde...

  27. [27]

    MMBench: Is your multi-modal model an all-around player?

    Y . Liuet al., “MMBench: Is your multi-modal model an all-around player?” inComputer Vision – ECCV 2024, ser. Lecture Notes in Computer Science, vol. 15064. Cham: Springer, 2025, pp. 216–233. 22

  28. [28]

    Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,

    C. Fuet al., “Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2025, pp. 24 108–24 118. [Online]. Available: https://openaccess.thecvf.com/content/CVPR2025/html/Fu Video-MME The First-Ever Comprehensive Evalu...

  29. [29]

    Holistic evaluation of language models,

    P. Lianget al., “Holistic evaluation of language models,”Transactions on Machine Learning Research, 2023. [Online]. Available: https: //arxiv.org/abs/2211.09110

  30. [30]

    DecodingTrust: A comprehensive assessment of trustworthiness in GPT models,

    B. Wanget al., “DecodingTrust: A comprehensive assessment of trustworthiness in GPT models,” inAdvances in Neural Information Processing Systems, 2023. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2023/ hash/63cb9921eecf51bfad27a99b2c53dd6d-Abstract-Datasets and Benchmarks.html

  31. [31]

    Vision-based learning for drones: A survey,

    J. Xiao, R. Zhang, Y . Zhang, and M. Feroskhan, “Vision-based learning for drones: A survey,”IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 9, pp. 15 601–15 621, Sep. 2025

  32. [32]

    α 3-Bench: A unified benchmark of safety, robustness, and efficiency for LLM-based UA V agents over 6G networks,

    M. A. Ferrag, A. Lakas, and M. Debbah, “α 3-Bench: A unified benchmark of safety, robustness, and efficiency for LLM-based UA V agents over 6G networks,” 2026, arXiv:2601.03281. [Online]. Available: https://arxiv.org/abs/2601.03281

  33. [33]

    HUGE-Bench: A benchmark for high-level UA V vision- language-action tasks,

    J. Guoet al., “HUGE-Bench: A benchmark for high-level UA V vision- language-action tasks,” 2026, arXiv:2603.19822. [Online]. Available: https://arxiv.org/abs/2603.19822

  34. [34]

    A mathematical theory of communication,

    C. E. Shannon, “A mathematical theory of communication,”The Bell System Technical Journal, vol. 27, no. 3–4, pp. 379–423, 623–656,

  35. [35]

    SelectiveNet: A deep neural network with an integrated reject option,

    Y . Geifman and R. El-Yaniv, “SelectiveNet: A deep neural network with an integrated reject option,” inProceedings of the 36th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 2019, pp. 2151–2159. [Online]. Available: https://proceedings.mlr.press/v97/geifman19a.html

  36. [36]

    On calibration of modern neural networks,

    C. Guo, G. Pleiss, Y . Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” inProceedings of the 34th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 70. PMLR, 2017, pp. 1321–1330. [Online]. Available: https://proceedings.mlr.press/v70/guo17a.html

  37. [37]

    Qwen3-VL technical report,

    S. Baiet al., “Qwen3-VL technical report,” 2025, arXiv:2511.21631. [Online]. Available: https://arxiv.org/abs/2511.21631

  38. [38]

    Qwen3.5: Towards native multimodal agents,

    Qwen Team, “Qwen3.5: Towards native multimodal agents,” 2026, official model-family release page. [Online]. Available: https://qwen.ai/ blog?id=qwen3.5

  39. [39]

    SmolVLM2-2.2B-Instruct model card,

    Hugging Face, “SmolVLM2-2.2B-Instruct model card,” 2025, official model card. [Online]. Available: https://huggingface.co/HuggingFaceTB/ SmolVLM2-2.2B-Instruct

  40. [40]

    Phi-4-Mini technical report: Compact yet powerful multimodal language models via mixture-of-LoRAs,

    Microsoft, “Phi-4-Mini technical report: Compact yet powerful multimodal language models via mixture-of-LoRAs,” 2025, arXiv:2503.01743; includes Phi-4-Multimodal. [Online]. Available: https://arxiv.org/abs/2503.01743

  41. [41]

    Gemma 4 model overview,

    Google, “Gemma 4 model overview,” 2026, official model documenta- tion. [Online]. Available: https://ai.google.dev/gemma/docs/core

  42. [42]

    InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency,

    W. Wanget al., “InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency,” 2025, arXiv:2508.18265. [Online]. Available: https://arxiv.org/abs/2508.18265

  43. [43]

    GLM-4.6V overview,

    Z.AI, “GLM-4.6V overview,” 2026, official developer documentation for the GLM-4.6V series, including GLM-4.6V-Flash. [Online]. Available: https://docs.z.ai/guides/vlm/glm-4.6v

  44. [44]

    Qwen2.5-VL technical report,

    S. Baiet al., “Qwen2.5-VL technical report,” 2025, arXiv:2502.13923. [Online]. Available: https://arxiv.org/abs/2502.13923

  45. [45]

    MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe,

    T. Yuet al., “MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe,” 2025, arXiv:2509.18154. [Online]. Available: https://arxiv.org/abs/2509.18154

  46. [46]

    Aya Vision: Advancing the frontier of multilingual multimodality,

    S. Dashet al., “Aya Vision: Advancing the frontier of multilingual multimodality,” 2025, arXiv:2505.08751. [Online]. Available: https: //arxiv.org/abs/2505.08751

  47. [47]

    Building and better understanding vision-language models,

    H. Laurenc ¸on, L. Tronchon, M. Cord, and V . Sanh, “Building and better understanding vision-language models,” 2024, arXiv:2408.12637; includes Idefics3-8B. [Online]. Available: https://arxiv.org/abs/2408. 12637

  48. [48]

    LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,

    LLaV A Team, “LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,” 2024, official project release page. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/

  49. [49]

    Concrete problems in AI safety,

    D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mane, “Concrete problems in AI safety,” 2016, arXiv:1606.06565. [Online]. Available: https://arxiv.org/abs/1606.06565

  50. [50]

    Advanced security frameworks for UA V and IoT: A deep learning approach,

    N. Quadar, A. Chehri, and B. Debaque, “Advanced security frameworks for UA V and IoT: A deep learning approach,”Internet of Things, vol. 32, p. 101594, 2025

  51. [270]

    2679–2713

    PMLR, 2025, pp. 2679–2713. [Online]. Available: https: //proceedings.mlr.press/v270/kim25c.html

  52. [1948]

    Available: https://people.math.harvard.edu/ ∼ctm/home/ text/others/shannon/entropy/entropy.pdf

    [Online]. Available: https://people.math.harvard.edu/ ∼ctm/home/ text/others/shannon/entropy/entropy.pdf