{"id":"68e348cd-e625-46f8-9a96-8bc2d2f79eca","arxiv_id":"2506.12831","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A vision-RF fusion framework with squint-aware beam tracking improves the time-averaged communication-sensing tradeoff in sub-THz air-ground ISAC.","lead":"This paper presents a base-station design that uses cameras and a small amount of radio probing to align beams for both communication and sensing in the sub-terahertz band. The goal is to reduce the time spent on beam training, leaving more time for data transmission and drone tracking in air-ground networks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 1's proof assumes without justification that KL-based Cor(H,G) and beamspace peak overlap S(H,G) are monotonically linked; this unproven chain is the load-bearing step for the correlation-based loss and the claimed ISAC efficiency gains.","rationale":"The reader's weakest assumption is exactly the unproven chain Cor -> S -> SE/CRB, and I agree with that identification. This concern is load-bearing because the ViR-Net loss in Eq. (24) and the TTD correlation-regulation design are both justified by Proposition 1; without a valid proof, the superior efficiency in Figs. 7-13 is an empirical observation rather than a consequence of the proposed mechanism. The proposed Monte-Carlo test directly settles whether the monotonicity claim holds for the system configuration used in the paper, and it also tests the directional claim that TTD-induced increases in Cor improve the Pareto boundary. Because the reader already assigned a CONDITIONAL verdict and this stress-test does not identify a reason to move away from that verdict, the appropriate recommendation is UNCHANGED: the paper should either supply a rigorous proof or clearly demote Proposition 1 to a heuristic, and it should release the simulation artifacts so the quantitative claims can be independently reproduced.","tokens_in":20935,"tokens_out":6276,"duration_ms":72396,"concrete_test":"Using the Tables I-II system parameters, generate 10^4 random user/target angle configurations. For each (H,G), compute Cor(H,G) from Eqs. (12)-(13), S(H,G) from Appendix A, and the Pareto-optimal (R*(gamma),CRB*(gamma)) from Eqs. (28)-(30) over a fine gamma grid. Check (a) whether S is nondecreasing in Cor across all samples, and (b) whether, at matched CRB, R* is nondecreasing in Cor (equivalently, at matched R*, CRB* is nonincreasing in Cor). Any monotonicity violation is a counterexample to Proposition 1. To isolate the causal design claim, repeat with random TTD-delay perturbations on a fixed physical channel and test whether increasing Cor through TTD values moves the Pareto boundary in the claimed direction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, that SoM-ISAC achieves significantly higher frame-level ISAC efficiency than RF-only schemes, rests on Proposition 1, which states that as Cor(H,G) increases, SE improves and CRB decreases monotonically. The proof in Appendix A contains two unshown implications. First, 'a higher Cor(H,G) gives rise to a higher S(H,G)' is cited to [43], but [43] is about bounding KL divergence between Gaussian mixtures and does not establish a monotone relation between KL divergence and exact argmax coincidences of beamspace power distributions. This implication is not generally true: Cor(H,G)=1/KL(bh_c,bh_s) can increase by concentrating probability mass within the same beam index while S(H,G) stays fixed, or S can change discontinuously with an arbitrarily small KL change. Second, Eq. (31) only shows that the xi terms depend on the exact peak-overlap indicator S; it does not show that the Pareto-optimal R* and CRB* in Eqs. (29)-(30) are monotone functions of Cor, because S need not track Cor. Since the ViR-Net loss in Eq. (24) directly rewards Cor(H,G)/Cor*(H,G), the training objective may optimize a quantity that does not track the actual SE-CRB boundary. If Proposition 1 fails or remains unproved, the theoretical motivation for the TTD correlation regulation and the claimed efficiency advantage are unsupported, even though the empirical comparisons might still be reproducible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a vision-RF fused ISAC transmission framework for sub-THz air-ground networks. It uses RGB-D cameras for coarse user/target localization, a squint-aware cross-pattern beam tracking (SA-CP-BT) scheme with true-time-delay (TTD) arrays for fine angle refinement, and a transformer-based ViR-Net that maps positioning spectra and images to TTD/PS/digital precoder parameters. The training loss (Eq. (24)) rewards high communication-sensing channel correlation Cor(H,G)=1/KL(beamspace power distributions), and the evaluation metric \"ISAC efficiency\" combines time-averaged SE and CRB with frame-structure overhead. Simulations in Figs. 7-13 compare SoM-ISAC against RF-only, vision-only, BCD, squint-elimination, and squint-sensing baselines, reporting better ISAC efficiency boundaries, lower runtime, and robustness to spatial distributions, TTD count, pilot count, and subframe duration.","tokens_in":21207,"tokens_out":5122,"duration_ms":57723,"significance":"If the main proposition holds, the framework is a credible way to exploit sub-THz hardware DoF and out-of-band visual data, with a useful efficiency metric and extensive simulation support. Strengths: the paper provides a detailed system model, an explicit frame-structure overhead model, ablation studies (Table V), and a wide set of comparisons; the central simulation results are internally consistent. However, the theoretical justification for the correlation-based loss and TTD regulation rests on Proposition 1, whose proof in Appendix A has a missing monotonicity argument. Since the loss function in Eq. (24) directly optimizes Cor(H,G), the reported gains could be an artifact of optimizing a proxy that is not established to track the SE-CRB boundary. The contribution is therefore conditional on closing that proof gap or replacing it with numerical validation.","major_comments":[{"comment":"The proof of Proposition 1 does not establish the claimed monotonic chain. Eq. (31) defines the xi terms in terms of exact peak-coincidence indicators I(psi_{c,m}=psi_{s,k,m}); it shows that both xi terms depend on S(H,G), but not that R* and CRB* are monotone in Cor(H,G). The statement \"It can be proved that a higher Cor(H,G) gives rise to a higher S(H,G)\" is cited to [43], which addresses KL divergence approximation for Gaussian mixture models and does not imply a monotone relation between inverse KL and the number of coincident beamspace peaks. Moreover, Cor(H,G) can increase by concentrating probability mass within a fixed beam index while S(H,G) remains unchanged, and S(H,G) can jump discontinuously under an arbitrarily small change in Cor(H,G). Because the loss in Eq. (24) directly rewards Cor(H,G)/Cor^*(H,G), the training objective may optimize a quantity that does not provably track the actual SE-CRB boundary. Please either supply a rigorous derivation of both implications or state Proposition 1 as a conjecture and support it with numerical evidence.","section":"Section II-D and Appendix A (Proposition 1, Eqs. (13), (31))"},{"comment":"The loss function normalizes by Cor^*(H,G), described as the maximum achievable correlation for each sample, but Cor^* is never defined or computed. Without a precise definition, the normalization is ambiguous: Cor^* could depend on the TTD configuration, making the loss's optimization landscape unclear, or it could be a theoretical maximum that is not available at training time. Please define Cor^*, explain how it is obtained, and state whether it is updated during training or fixed beforehand.","section":"Section IV-C, Eq. (24)"}],"minor_comments":[{"comment":"The KL divergence in the definition of Cor(H,G) is asymmetric and can be infinite when the sensing beamspace distribution has zero mass on a bin where the communication distribution is positive. Please clarify how zero entries are handled in the numerical implementation and whether this affects the smoothness of the training loss.","section":"Section II-D, Eq. (13)"},{"comment":"In the proof of Proposition 2, the monotonicity of phi_m and theta_m with the subcarrier index is asserted with \"It is easy to prove\" and the extension from Q_t=N_t to general Q_t is stated without derivation. Please provide the explicit monotonicity argument and justify the averaging step for subarrays of size L_h x L_v.","section":"Section III-C, Appendix B"},{"comment":"The curve labeled \"Perfect Prior\" is not described in the comparison-methods list in Section V-C. Please state how this bound is computed and whether it uses the same precoder optimization as SA-Opt-ISAC but with perfect channel knowledge.","section":"Section V-A, Fig. 7(a)"},{"comment":"The horizontal-axis label in Fig. 12 is inconsistent with the caption and the clause in the text: the axis reads \"number of TTDs Qt\" with ticks 0 to 300, while Table I sets Q_t=256 and the text compares Q_t=N_t with N_t=256. Please align the axis, the caption, and the parameter range.","section":"Section V-E, Figs. 10-13"},{"comment":"In the sentence defining the TTD operation in Eq. (2), the expression uses vec(T ⊗ 1_{L_h x L_v}), but T is a Q_th x Q_tv matrix and the Kronecker product with the all-one matrix gives dimensions that are not explicitly mapped to the N_t x N_RF phase-shifter array. Please clarify the exact indexing used to place the TTD delays in the diagonal matrix.","section":"Notation, Section II-A"}],"recommendation":"major_revision","confidential_remarks":"The proof gap in Proposition 1 is the central blocking issue: it is load-bearing for the loss design and the claimed efficiency gains. The gap is plausibly fixable by adding a formal argument or by reframing Proposition 1 as an empirically validated heuristic; I would not recommend rejection if the authors choose the latter and provide strong numerical evidence. The manuscript also reuses substantial components from the authors' prior work [23,27] (SoM concept, TTD-based precoding, frame structure), and the editor may wish to ask for a clearer statement of incremental contribution relative to those papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid engineering paper with a real integrated framework, but the theoretical cornerstone (Proposition 1) has a genuine proof gap that should be fixed or explicitly demoted to a heuristic before publication.\n\nWhat's new: The combination of squint-aware cross-pattern beam tracking for a UPA, a vision-RF fusion network (ViR-Net) that outputs TTD/PS/digital precoders, and a frame-level ISAC efficiency metric that balances performance against overhead is not something I've seen in one place. The simulation pipeline is serious: SUMO for trajectories, AirSim for visuals, Wireless InSite for ray-traced channels, and a consistent set of baselines. The ablation studies in Table V and the sensitivity analyses in Figs. 10-13 are a plus. The authors are honest about the tradeoffs (e.g., SA-Opt-ISAC beats them in absolute SE/CRB, but at much higher runtime).\n\nSoft spots: The proof of Proposition 1 in Appendix A is the load-bearing piece. It claims that higher KL-based correlation Cor(H,G) yields higher beamspace peak overlap S(H,G), citing [43] — but [43] is about KL divergence between Gaussian mixtures, not about argmax coincidences of beamspace power distributions. That monotonicity is not generally true, and the rest of the proof just asserts that Eq. (31) shows performance improves with S. The step from S to Pareto-optimal SE/CRB is sketched, not derived. Since the loss function (24) directly rewards Cor, the training objective may be chasing a quantity that doesn't actually track the SE-CRB boundary. This doesn't invalidate the empirical comparisons, but it does mean the claimed theoretical motivation is unsupported as written. Appendix B's generalization to Q_t < N_t is also asserted, not proved. I'd also like to see code, data, or at least error bars, since the simulations are deterministic and we have no sense of variance.\n\nThe citation pattern is fine; the authors reuse their own prior results ([23], [27]) heavily, which is normal in a research program, but it does mean the incremental novelty over those papers should be made explicit.\n\nWho it's for: researchers working on sub-THz ISAC, hybrid precoding, and multi-modal sensing for 6G. The framework is plausible and the simulations are detailed enough to be worth engaging with. A serious referee should see it, but with a request for either a rigorous proof of Proposition 1 or a clear statement that it is a heuristic, plus artifacts/error bars.\n\nRecommendation: I'd accept for review with major revision. The core empirical claims may survive, but the theoretical foundation needs shoring up or honest reframing.","headline":"The integrated framework is genuinely useful, but Proposition 1's unproven monotonicity chain is a real soft spot that should be fixed or reframed as a heuristic.","tokens_in":21801,"tokens_out":2872,"would_cite":false,"duration_ms":27249,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By coupling RGB-D cameras with a small number of squint-aware radio probes, this paper claims, a sub-THz base station can align its communication and sensing channels in beamspace and achieve higher frame-level ISAC efficiency than…","keywords":["integrated sensing and communication","sub-terahertz","air-ground network","hybrid precoding","beam squint","true time delay","vision-RF fusion","ISAC efficiency"],"falsifier":"Run a Monte Carlo sweep over user-target angular separations in the paper's channel model, computing $\\mathrm{Cor}(H,G)$, the beamspace peak overlap $S(H,G)$, and the Pareto-optimal SE-CRB pair from Eqs. (27)-(31); if any channel pair has higher $\\mathrm{Cor}(H,G)$ but lower $S(H,G)$, or higher $S(H,G)$ but a worse SE-CRB trade-off, then Proposition 1 is refuted.","tokens_in":20662,"feed_emoji":"📡","tokens_out":10599,"duration_ms":110192,"temperature":0.7,"pith_summary":"This paper aims to establish that sub-THz integrated sensing and communication in air-ground networks can be made both faster to set up and more effective by fusing camera images with a small number of radio measurements. Its central proposition is that raising the correlation between the communication and sensing channels, measured in beamspace, monotonically improves both the achievable spectral efficiency and the angle-estimation Cramér-Rao bound; if true, this makes \"make the two channels look alike\" a design objective. The paper then shows a concrete way to pursue that objective: use the true-time-delay layers of a hybrid precoder to rotate the equivalent channels, guide the process with a neural network that maps RGB-D images and a positioning spectrum to precoder parameters, and replace exhaustive beam training with a squint-aware cross-pattern beam tracking that exploits the frequency-dependent squint of sub-THz beams. The payoff the authors report is that their SoM-ISAC scheme nearly halves the per-frame running time of a radio-only near-optimal benchmark while keeping a competitive SE-CRB trade-off, and it outperforms radio-only schemes under the new ISAC efficiency metric that accounts for time overhead. A sympathetic reader would care because the claim, if correct, turns sub-THz hardware limitations into an actionable degree of freedom and points toward zero-pilot, vision-aided ISAC operation.","feed_headline":"Sparse RF probes plus cameras beat radio-only sub-THz ISAC efficiency","feed_subtitle":"A vision-guided precoder skips beam training and channel estimation, turning saved time into data and sensing.","key_machinery":"The load-bearing object is the C-S channel correlation $\\mathrm{Cor}(H,G)$, defined in Eq. (13) as the reciprocal of the Kullback-Leibler divergence between normalized beamspace power distributions of the communication and sensing channels aggregated over subcarriers. Proposition 1 asserts that this scalar controls the Pareto-optimal SE-CRB pair; the proof links it to the overlap $S(H,G)$ of beamspace peaks. The hardware that converts correlation into a tunable quantity is the true-time-delay (TTD) layer of the hybrid precoder: TTDs introduce frequency-dependent phase shifts, so at sub-THz bandwidths the otherwise harmful beam squint can be directed along angular trajectories or used to align equivalent channels. The squint-aware cross-pattern beam tracking (SA-CP-BT) mechanism exploits this by partitioning a visually obtained angular range and sweeping squint beams in horizontal and vertical passes, so a user or target angle is read off from the subcarrier with maximal array gain. Finally, the ViR-Net architecture (spectrum encoder, vision encoder, feature-fusion transformer, and precoding head) maps the positioning spectrum and RGB-D images to TTD, phase-shifter, and digital precoder values, trained by the loss in Eq. (24) that rewards both high correlation and good SE and CRB, and the ISAC efficiency metric in Eqs. (25)-(26) weights these performance gains by the time spent to obtain them.","core_discovery":"The paper's central claim is Proposition 1: as the communication-sensing channel correlation $\\mathrm{Cor}(H,G)$ grows, the achievable spectral efficiency improves and the Cramér-Rao bound on target angle estimation decreases in a monotonic way. Here $\\mathrm{Cor}(H,G)$ is the inverse KL divergence between the normalized beamspace power distributions of the aggregated communication and sensing channels, so it is a measure of how much the strongest communication and sensing directions overlap. Relying on this monotonicity, the authors treat the TTD-delay network of a standard hybrid precoder as a channel modulator: by choosing delays that raise the correlation, the equivalent communication and sensing channels are rotated toward each other without adding hardware. The rest of the framework — visual detection of users and low-altitude targets from RGB-D images, squint-aware cross-pattern beam tracking for angle refinement, and the ViR-Net that outputs TTD, phase-shifter, and digital precoder values under a correlation-weighted loss — is a way to realize this correlation gain quickly. The reported simulations conclude that the proposed scheme achieves significantly higher ISAC efficiency than RF-only schemes at the frame level, while remaining close to a radio-only near-optimal benchmark in the absolute SE-CRB trade-off.","pith_inferences":["If Proposition 1's monotonicity holds in general, then other means of increasing beamspace overlap, such as subcarrier assignment or reconfigurable-surface phase profiles, could substitute for TTD-based rotation; the paper does not test this, but it follows from treating correlation rather than hardware as the fundamental resource.","The SA-CP-BT principle effectively uses OFDM subcarriers as spatial indices within a single beam, which suggests a testable extension: a two-dimensional squint pattern that resolves both azimuth and elevation from one slot, at the price of handling ambiguity near grid boundaries.","The reported efficiency numbers come from ray-traced simulations of one intersection; field experiments would need to show that camera depth errors and object-detection misses, which are absent in simulation, do not erase the time savings.","The ISAC efficiency metric treats SE and CRB as a pair but does not say how a network should weight communication versus sensing value; choosing such weights could change the optimal number of slots and pilots."],"forward_implications":["Any design knob that raises $\\mathrm{Cor}(H,G)$ — not just TTD delays — should improve the SE-CRB trade-off, making beamspace channel correlation a primary target for ISAC precoder design.","Skipping the synchronization-signal block, initial beam training, and instantaneous channel estimation lets the saved time be spent in data transmission and sensing, which is why the proposed frame structure yields higher time-averaged SE and lower time-averaged CRB at the frame level.","The ablation results indicate that vision alone is not enough in sub-THz: the vision-only variant performs poorly, so sparse RF refinement through SA-CP-BT is a necessary complement to visual priors.","The number of beam-tracking slots, the number of pilots, and the number of TTDs each exhibit an optimum under the efficiency metric, so the framework provides a concrete criterion for trading estimation accuracy against operational latency.","In quasi-static settings with long subframe periods the radio-only near-optimal scheme becomes competitive or better, so the proposed scheme's advantage is specific to dynamic air-ground scenarios where positions change quickly."],"supporting_citations":[{"why":"Supplies the subspace-correlation form of the optimal transmit covariance used in the proof of Proposition 1.","marker":"[42]"},{"why":"Provides the beamspace dictionary and sparsity argument on which the definitions of Cor(H,G) and peak overlap S(H,G) rest.","marker":"[31]"},{"why":"Is cited for the step that higher Cor(H,G) gives higher beamspace peak overlap, the bridge Proposition 1 depends on.","marker":"[43]"},{"why":"Provides the subspace-rotation interpretation used in Remark 1 to justify TTD-based equivalent-channel modulation.","marker":"[32]"},{"why":"Gives the SoM-driven squint-aware optimization used by the SA-Opt-ISAC benchmark and extended by the proposed framework.","marker":"[27]"},{"why":"Defines the sensing-centric squint-control benchmark that the paper compares against and builds on for proactive squint patterns.","marker":"[11]"},{"why":"Supplies the non-TTD BCD hybrid-precoding benchmark whose performance the paper uses to show the benefit of TTD-based correlation regulation.","marker":"[16]"},{"why":"Supplies the hierarchical beam training procedure used by SA-Opt-ISAC to acquire coarse angular information before squint-aware refinement.","marker":"[18]"}],"fun_headline_variants":["Vision-guided precoder aligns beams to improve sub-THz ISAC efficiency","Cameras and RF join forces to cut sub-THz ISAC latency and raise throughput","Squint-aware beams plus visual cues trim sub-THz ISAC training overhead","Correlation-aware delays sharpen sub-THz ISAC sensing without extra hardware"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole design rests on the unproven step that when the communication and sensing beamspace distributions look more similar, their strongest beams actually overlap more, and that this overlap alone improves both data rate and sensing accuracy; the paper asserts this link rather than deriving it.","fun_headline_variants_meta":{"raw":{"variants":["Vision-guided precoder aligns beams to improve sub-THz ISAC efficiency","Cameras and RF join forces to cut sub-THz ISAC latency and raise throughput","Squint-aware beams plus visual cues trim sub-THz ISAC training overhead","Correlation-aware delays sharpen sub-THz ISAC sensing without extra hardware"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001126,"raw_usage":{"total_tokens":4687,"prompt_tokens":953,"completion_tokens":3734,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":3662}},"tokens_in":569,"tokens_out":3734,"duration_ms":27752,"temperature":1.0,"reasoning_tokens":3662,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:36:50.600206+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a Monte Carlo sweep over user-target angular separations in the paper's channel model, computing $\\mathrm{Cor}(H,G)$, the beamspace peak overlap $S(H,G)$, and the Pareto-optimal SE-CRB pair from Eqs. (27)-(31); if any channel pair has higher $\\mathrm{Cor}(H,G)$ but lower $S(H,G)$, or higher $S(H,G)$ but a worse SE-CRB trade-off, then Proposition 1 is refuted.","supporting_citations":[{"cited_title":"On the performance gain of integrated sensing and communications: A subspace correlation perspective,","cited_arxiv_id":null,"evidence_quote":"Supplies the subspace-correlation form of the optimal transmit covariance used in the proof of Proposition 1."},{"cited_title":"Mutual information maximizing wideband multi-user (wMU) mmwave massive MIMO,","cited_arxiv_id":null,"evidence_quote":"Provides the beamspace dictionary and sparsity argument on which the definitions of Cor(H,G) and peak overlap S(H,G) rest."},{"cited_title":"Approximating the Kullback Leibler divergence between gaussian mixture models,","cited_arxiv_id":null,"evidence_quote":"Is cited for the step that higher Cor(H,G) gives higher beamspace peak overlap, the bridge Proposition 1 depends on."},{"cited_title":"RIS-assisted integrated sensing and communications: A subspace rotation approach,","cited_arxiv_id":null,"evidence_quote":"Provides the subspace-rotation interpretation used in Remark 1 to justify TTD-based equivalent-channel modulation."},{"cited_title":"Synesthesia of Machine (SoM)-Driven Analog Precoder Optimization for Enhanced ISAC Performance in Sub-THz Systems","cited_arxiv_id":"2412.13532","evidence_quote":"Gives the SoM-driven squint-aware optimization used by the SA-Opt-ISAC benchmark and extended by the proposed framework."},{"cited_title":"YOLO: An efficient terahertz band integrated sensing and communications scheme with beam squint,","cited_arxiv_id":null,"evidence_quote":"Defines the sensing-centric squint-control benchmark that the paper compares against and builds on for proactive squint patterns."},{"cited_title":"Low-complexity joint transceiver optimization for MmWave/THz MU-MIMO ISAC systems,","cited_arxiv_id":null,"evidence_quote":"Supplies the non-TTD BCD hybrid-precoding benchmark whose performance the paper uses to show the benefit of TTD-based correlation regulation."},{"cited_title":"A unified 3D beam training and tracking procedure for terahertz communication,","cited_arxiv_id":null,"evidence_quote":"Supplies the hierarchical beam training procedure used by SA-Opt-ISAC to acquire coarse angular information before squint-aware refinement."}],"review_version":1}