{"id":"c67d6663-3f92-4c35-8c09-c45b32fabcbf","arxiv_id":"2608.10898","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A sensing-aware adaptive source-channel coding and beamforming scheme is proposed for bi-static integrated sensing and semantic communications, and simulations show it outperforms DJSCC-WF-ZF and BPG-WF-ZF.","lead":"This paper proposes a joint coding-rate and beamforming design that lets one wireless transmitter serve semantic image communication and radar-style target sensing at the same time. Simulations on bird images show the design beats two baselines in image quality and position-estimation accuracy, especially at low transmit power.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The E2E distortion surrogate in Eq. (12) is an unvalidated regression fit; if it is biased in the BER/SINR regime where Algorithm I makes rate and beam decisions, the claimed gains over DJSCC-WF-ZF and BPG-WF-ZF may not survive.","rationale":"The central claim is an empirical performance comparison, so the objective function used to produce the optimized operating points must predict the actual metric. The paper's objective is Eq. (12), a fitted curve rather than a verified bound. The reader's weakest assumption identifies exactly this point, and I agree with that diagnosis. The concern is load-bearing because all algorithmic decisions, including model selection in P4.1, rate adaptation in P4.3 and P4.4, and the beamforming surrogate in P4.5 and P4.6, are driven by this fitted curve, and no check is reported connecting the curve to the actual MS-SSIM after optimization. That said, I do not see an internal inconsistency that would force rejection: the HCRB derivation and AO decomposition are structurally coherent, and the surrogate could in principle be accurate. The right disposition is therefore CONDITIONAL rather than REJECT. The condition should be a direct validation of Eq. (12) against the DNN codec at the optimized operating points, together with reporting of the fitting parameters and residuals. This leaves the reader's verdict unchanged.","tokens_in":25974,"tokens_out":7758,"duration_ms":77397,"concrete_test":"Re-run the Fig. 3 experiment at PT=-4 dBm and ABR=0.063. Record the Algorithm I outputs (Rs*, Rc*, wc*, w0*) and the resulting SINR and realized BER. Then evaluate the actual MS-SSIM on the CUB-200-2011 test set by running the trained semantic codec with random bit flips at that BER, and compare it with the value predicted by Eq. (12) at the same (Rs, rho_b), including whether the reported 2.68% gain over DJSCC-WF-ZF is preserved. Independently, recompute the model-selection step by exhaustively evaluating actual DNN distortion for all G=20 rates at that operating point; if the selected Rs* differs from the one chosen by the fitted surrogate, or if the realized versus predicted distortion differs by more than the reported gains, the central claim is not established.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The optimization in Algorithm I minimizes the fitted logistic surrogate in Eq. (12), not the true E2E distortion. Eq. (12) is obtained by \"regression-based fitting\" on CUB-200-2011 in Section III-A; it is not an analytically derived upper bound, and the fitted parameters, fitting errors, and validation details are not reported. The surrogate is used in every decision: P4.1 selects the source model by comparing fitted values, P4.3 and P4.4 convert rate optimization into minimization of the fitted curve, and P4.5 and P4.6 replace BER minimization by SINR maximization using the monotonicity of that same fitted curve in rho_b. If the fit is biased in the low-SINR/high-BER or high-Rs regime where Algorithm I makes its choices, the selected rates and beamforming may not minimize actual distortion. The paper reports final MS-SSIM at selected operating points in Figs. 3 and 4, but it never checks whether the surrogate-predicted distortion at those points matches the simulated MS-SSIM, so a mismatch could go unnoticed. With G=20 and at least four fitted parameters per Rs value, the surrogate has enough flexibility to fit the training points while misrepresenting off-grid BER values; without parameter reporting, error bars, or code, the central performance claim rests on an unvalidated proxy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a sensing-aware adaptive source-channel coding (SA-ASCC) framework for a bi-static integrated sensing and semantic communication (ISSC) system. The transmitter superimposes SemCom and sensing symbols via beamforming; the receiver tasks are image reconstruction at the communication RE and target position estimation at the sensing RE under imperfect time synchronization. The paper approximates the E2E semantic distortion by a logistic regression fit (Eq. 12), derives a hybrid Cramér-Rao bound for target position (Propositions 3.1 and 3.2), and formulates a mixed-integer, non-convex optimization problem (P3.1) that minimizes E2E distortion subject to HCRB, bandwidth, and power constraints. An alternating optimization algorithm (Algorithm I) is proposed, combining exhaustive model selection, successive convex approximation, and fractional programming. Numerical results compare the proposed scheme with DJSCC-WF-ZF and BPG-WF-ZF benchmarks in terms of MS-SSIM and HCRB.","tokens_in":26426,"tokens_out":9731,"duration_ms":93860,"significance":"If the results hold, the paper provides a useful resource-allocation formulation for bi-static ISSC that couples semantic source-channel rate selection with sensing-aware beamforming. The HCRB derivation in Appendix A is careful and self-contained, and the SCA/FP reformulations in Section IV are nontrivial and largely plausible. The main value is in extending adaptive source-channel coding from communication-only settings to an integrated sensing and semantic communication setting. However, the central performance claim rests on the unvalidated fitted surrogate in Eq. (12), and the simulation evidence is presented without error bars or code. The contribution is therefore significant but conditional on the surrogate being a faithful model over the operating region used by the optimizer.","major_comments":[{"comment":"The E2E distortion surrogate is introduced as a logistic regression fit on the CUB-200-2011 dataset, not as a derived upper bound as stated in the abstract and Section I. The fitting parameters \\hat d_s^o, \\hat d_c^o, a_1^{\\mathrm{mse}}, a_2^{\\mathrm{mse}}, the number of fitted points per (R_s, \\rho_b), and the goodness-of-fit are not reported. Because this surrogate is used in every decision of Algorithm I (P4.1 model selection, P4.3/P4.4 rate optimization, P4.5/P4.6 beamforming), a bias in the low-SINR or high-BER regime would directly invalidate the claimed gains over the benchmarks in Figs. 3 and 4. Please report the fitting quality and provide a validation comparing surrogate-predicted distortion with actual decoded MSE and MS-SSIM at the operating points selected by the algorithm, including off-grid BER values.","section":"Section III-A, Eq. (12); Algorithm I"},{"comment":"The simulation results are reported as single curves without confidence intervals or multiple random seeds, and the text makes quantitative claims such as a 2.68% MS-SSIM improvement and a 24.1% bandwidth reduction. In addition, the BPG-WF-ZF baseline uses fixed channel coding rates R_c=2.3 and R_c=1.9 without explaining how these values are chosen or whether they are optimized for fairness. Without error bars and a clear baseline rate-selection procedure, the statistical significance of the performance comparison is not established. Please add confidence intervals and describe how the baseline rates and power allocations are set.","section":"Section V-A, Figs. 3 and 4"},{"comment":"The transformation from the BER minimization problem to SINR maximization relies on the monotonicity of the fitted logistic curve in \\rho_b (Eq. 12) and on Lemma 4.2. The paper argues that all constraints in P4.3 must be active at the optimum, but only constraints (32) and (33) are forced to equality by monotonicity of the objective with respect to \\hat\\rho_b and e^\\nu; constraint (24), R_c \\le C(\\gamma), need not be active. This imprecision does not necessarily break the algorithm, but the equivalence argument in Section IV-B1 should be restated more carefully, and the convexity claims for the resulting feasible set should be verified explicitly.","section":"Section IV-B, P4.3-P4.6"}],"minor_comments":[{"comment":"The parameter list states that the HCRB threshold is set to \\Pi=0.01, but the captions of Figs. 3 and 4 report \\Pi=0.06 and \\Pi=0.065, respectively. Please reconcile these values.","section":"Section V, parameter settings"},{"comment":"The displayed expression for d\\hat\\gamma/d\\gamma in Eq. (84) does not appear to be the correct derivative of \\hat\\gamma=\\ln2\\sqrt{L(C(\\gamma)-R_c)}/\\sqrt{1-1/(1+\\gamma)^2}; the denominator and numerator differ from a direct differentiation. The monotonicity conclusion may still be true, but the derivation should be corrected.","section":"Appendix C, Eq. (84)"},{"comment":"The output variable w_o^\\star should be w_0^\\star to match the notation used elsewhere; this appears to be a typographical error.","section":"Algorithm I, line 19"},{"comment":"Figure 2 is described as showing log_{10}D_o versus log_{10}\\rho_b for different R_s, but the axis ticks and legend are not reproduced in the text, making it difficult to assess the claimed sigmoidal behavior or the number of fitted models.","section":"Section III-A, Fig. 2"},{"comment":"No code or data repository is referenced; given that the distortions in Eq. (12) are obtained from a fitted model, releasing the fitting code and the simulation code would substantially improve reproducibility.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the unvalidated E2E distortion surrogate in Eq. (12). The paper's algorithmic machinery and the HCRB derivation are substantial, but the performance claims are only as strong as this fitted curve. If the authors can provide the fitting parameters, goodness-of-fit metrics, a cross-validation against actual decoded MS-SSIM at the optimizer's operating points, and error bars for the simulations, the paper would be suitable for publication. The inconsistency in the HCRB threshold values and the baseline rate selection should also be addressed. I do not see a fundamental flaw requiring rejection, but the central empirical claim is currently under-supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper deserves a serious referee, but the abstract oversells the distortion model. The HCRB derivation for bi-static positioning under time synchronization error is genuine, self-contained work (Appendix A), and coupling it as a constraint in a joint rate-and-beamforming design for integrated semantic communication and sensing is a legitimate extension of the authors' ASCC line. The AO decomposition—exhaustive model selection plus SCA/FP for the continuous subproblems—is competent, and the convex reformulations look plausible.\n\nWhat is actually new: no prior work jointly optimizes semantic source-channel coding rates and transmit beamforming under an HCRB threshold with the timing offset modeled as a random nuisance parameter. That is the paper's contribution, and it is useful for 6G system-level design.\n\nWhere it gets soft: the E2E distortion in Eq. (12) is a logistic curve fit on CUB-200-2011, not an upper bound. The abstract says 'deriving an upper bound,' but the body describes regression-based fitting. The fitted parameters, residuals, and validation against held-out BER/SINR points are not reported. This surrogate drives every rate and beamforming decision in Algorithm I, and the paper never checks surrogate-predicted distortion against the simulated MS-SSIM at the reported operating points. If the fit is locally biased where the algorithm decides—low SINR, high source rate—the claimed gains over DJSCC-WF-ZF and BPG-WF-ZF could shrink. This is not fatal: final MS-SSIM is evaluated on actual decoded images, so there is an independent check, and the gains in Figs. 3–4 are consistent across settings. But the missing fit diagnostics are a real gap in an otherwise careful paper.\n\nMinor issues: no error bars on any curve, no code or data release, and the benchmark gains are modest (single-digit percent MS-SSIM at selected operating points). The lower HCRB in Fig. 5 is a direct consequence of adaptive power allocation versus fixed ZF/WF baselines; the paper could state that more plainly.\n\nWho benefits: researchers working on joint semantic-communication/ISAC resource allocation. My recommendation: send it to review, but require the authors to report fit parameters and fit error, validate the surrogate at held-out BER operating points, and provide error bars or code.","headline":"The HCRB under imperfect timing sync is the real contribution; the E2E distortion surrogate is an unvalidated fit and the abstract overclaims it as an upper bound.","tokens_in":26820,"tokens_out":2666,"would_cite":true,"duration_ms":26574,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that jointly choosing the semantic source-coding rate and transmit beamforming for both tasks enlarges the achievable region of an integrated sensing and semantic communication (ISSC) system.","keywords":["semantic communication","integrated sensing and communication","adaptive source-channel coding","Cramér-Rao bound","beamforming design","bi-static ISSC","time synchronization error"],"falsifier":"Run Algorithm I on a fixed channel realization, take its chosen $R_s$, $R_c$, $w_c$, and $w_0$, then measure the actual MS-SSIM or MSE of the DNN codec at that operating point by transmitting real images with the specified BER; if the measured distortion does not match Eq. (12)'s predicted ordering, for example if a neighboring source rate yields lower actual distortion than the algorithm's chosen rate, then the claimed gains rest on the surrogate rather than on the true codec.","tokens_in":25773,"feed_emoji":"📡","tokens_out":5883,"duration_ms":55988,"temperature":0.7,"pith_summary":"This paper argues that in a bi-static integrated sensing and semantic communication (ISSC) system, the source-coding rate for a deep semantic image codec and the transmit beamforming for both communication and sensing should be chosen together, not separately. The proposed sensing-aware adaptive source-channel coding (SA-ASCC) framework fits an end-to-end semantic distortion function to a logistic curve, derives a hybrid Cramér-Rao bound (HCRB) for target-position estimation under imperfect time synchronization, and minimizes distortion subject to an HCRB threshold, channel-use, and power constraints. Numerical results show SA-ASCC exceeds the DJSCC-WF-ZF and BPG-WF-ZF baselines in MS-SSIM across power budget and bandwidth settings, and consistently achieves lower HCRB. If correct, this demonstrates that jointly coupling rate adaptation with sensing accuracy requirements enlarges the achievable region between the two tasks.","feed_headline":"Adaptive coding and beamforming widen the ISSC performance region","feed_subtitle":"Jointly tuning source-channel rates and beams lifts both image quality and positioning accuracy.","key_machinery":"The load-bearing object is the regression-fitted distortion surrogate, Eq. (12), a generalized logistic function of $\\log_{10}\\rho_b$ at each discrete source rate $R_s$, which turns reconstruction quality into a smooth function of coding rates and SINR. The second object is the hybrid Fisher information matrix $J(\\chi)=J_D(\\chi)+J_B(\\chi)$, where $J_D$ is the observed FIM with entries proportional to $\\operatorname{Tr}(\\zeta_j\\Sigma)$ and $J_B$ contributes the prior $1/\\sigma_T^2$ for the time-synchronization error; the HCRB for target position is the upper-left block of $J(\\chi)^{-1}$, split by the Woodbury identity into the perfect-synchronization CRB plus an additional TS-induced error. These two objects make reconstruction distortion and sensing accuracy comparable in a single objective and a single constraint. The algorithm then alternates exhaustive search over a bank of $G=20$ source-rate models with successive convex approximation and fractional programming for the continuous variables.","core_discovery":"The central claim is that there exists a joint design variable, namely the pair of source and channel coding rates together with the two transmit beamforming vectors, whose optimization gives a larger achievable (distortion, HCRB) region than separate design. The paper approximates the SemCom distortion as the sum of a source-distortion part and a channel-distortion part, fits Eq. (12) to measurements on the CUB-200-2011 dataset, models the time-synchronization error as a Gaussian nuisance parameter, and derives the HCRB as the block inversion of the sum of the observed and prior Fisher information matrices. An alternating optimization algorithm solves the resulting mixed-integer non-convex problem by exhaustive search over discrete source-rate models and convexified joint rate and beamforming subproblems. The reported consequence is that SA-ASCC outperforms DJSCC-WF-ZF and BPG-WF-ZF in MS-SSIM at low signal-to-interference-plus-noise ratio and requires less bandwidth for the same reconstruction quality, while also achieving a lower HCRB.","pith_inferences":["The regression fit is the fulcrum: the method's portability to other image datasets, codecs, or tasks is untested, and a biased fit at low SINR would shift the operating point even if the true codec performs better there; the paper does not report the fitting error or validation curves.","A direct extension would be to replace the fitted surrogate with the actual measured distortion at each candidate operating point; if the two orderings diverge, the claimed region gains would need to be re-evaluated.","The HCRB assumes the time-synchronization error variance is known, so an adaptive estimator of $\\sigma_T$ or a robust worst-case bound would be a natural continuation of the framework.","Exhaustive search scales linearly with the number of pretrained models, so a continuous or learned rate-selection module would be needed when the DNN model library grows large."],"forward_implications":["The optimal source-coding rate is no longer a function of channel state alone: the HCRB threshold shifts the chosen $R_s$, $R_c$, and beam emphasis, so sensing requirements must enter rate adaptation.","At fixed power, resources moved to sensing do not cost one-for-one in reconstruction quality, because SemCom retains high MS-SSIM in the low-SINR regime; the paper's reported 24.1 percent bandwidth saving at equal MS-SSIM is one manifestation.","Imperfect time synchronization shrinks the achievable region, so bi-static designs that ignore TS error will overestimate positioning accuracy and should include the prior Fisher information term.","The decomposition into model selection and joint rate and beamforming subproblems makes the mixed-integer program tractable with exhaustive search over a discrete model bank plus standard convex tools."],"supporting_citations":[{"why":"Supplies the hyperprior-based DNN models used to obtain discrete source-coding rates and the pretrained feature extraction and recovery networks.","marker":"[34]"},{"why":"Provides the adaptive source-channel coding rate-adaptation formulation and the finite-blocklength BER lower bound used in Eq. (7).","marker":"[15]"},{"why":"Models the time-synchronization error as a Gaussian random variable and derives the CRB with TS error that the HCRB extends.","marker":"[26]"},{"why":"Defines the DJSCC baseline architecture that the proposed scheme is compared against in the simulations.","marker":"[39]"},{"why":"Establishes bi-static beamforming under a CRB constraint on target position, the comparison point for the sensing-oriented design.","marker":"[24]"},{"why":"Motivates digital deep joint source-channel coding and the regression-based E2E distortion approximation approach.","marker":"[14]"},{"why":"Provides the multi-user ASCC framework that jointly optimizes coding rates, power, and beamforming, from which the proposed sensing-aware extension departs.","marker":"[16]"}],"fun_headline_variants":["Joint coding and beamforming enlarge ISSC region","Adaptive rates and beams improve both sensing and semantics","Wider performance region via joint coding and beam design","SA-ASCC outperforms benchmarks in distortion and HCRB","Integrated design beats separate coding and beamforming"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole rate and beamforming optimization treats Eq. (12), a logistic curve fitted on bird images, as the true end-to-end distortion of the semantic codec at every BER and source rate, and the fitting error is not reported.","fun_headline_variants_meta":{"raw":{"variants":["Joint coding and beamforming enlarge ISSC region","Adaptive rates and beams improve both sensing and semantics","Wider performance region via joint coding and beam design","SA-ASCC outperforms benchmarks in distortion and HCRB","Integrated design beats separate coding and beamforming"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000425,"raw_usage":{"total_tokens":2239,"prompt_tokens":1064,"completion_tokens":1175,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":1099}},"tokens_in":680,"tokens_out":1175,"duration_ms":9368,"temperature":1.0,"reasoning_tokens":1099,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:42:21.492528+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Algorithm I on a fixed channel realization, take its chosen $R_s$, $R_c$, $w_c$, and $w_0$, then measure the actual MS-SSIM or MSE of the DNN codec at that operating point by transmitting real images with the specified BER; if the measured distortion does not match Eq. (12)'s predicted ordering, for example if a neighboring source rate yields lower actual distortion than the algorithm's chosen rate, then the claimed gains rest on the surrogate rather than on the true codec.","supporting_citations":[{"cited_title":"V ariational image compression with a scale hyperprior,","cited_arxiv_id":null,"evidence_quote":"Supplies the hyperprior-based DNN models used to obtain discrete source-coding rates and the pretrained feature extraction and recovery networks."},{"cited_title":"Adaptive Source-Channel Coding for Semantic Communications","cited_arxiv_id":"2508.07958","evidence_quote":"Provides the adaptive source-channel coding rate-adaptation formulation and the finite-blocklength BER lower bound used in Eq. (7)."},{"cited_title":"Coordi nated transmit beamforming for networked isac with imperfect csi and time synchronization,","cited_arxiv_id":null,"evidence_quote":"Models the time-synchronization error as a Gaussian random variable and derives the CRB with TS error that the HCRB extends."},{"cited_title":"Dee p joint source- channel coding for wireless image transmission,","cited_arxiv_id":null,"evidence_quote":"Defines the DJSCC baseline architecture that the proposed scheme is compared against in the simulations."},{"cited_title":"Sensing and communication optimal trade- off for mimo bistatic systems,","cited_arxiv_id":null,"evidence_quote":"Establishes bi-static beamforming under a CRB constraint on target position, the comparison point for the sensing-oriented design."},{"cited_title":"D 2-jscc: Digital deep joint source-channel coding for semantic communications,","cited_arxiv_id":null,"evidence_quote":"Motivates digital deep joint source-channel coding and the regression-based E2E distortion approximation approach."},{"cited_title":"Adaptiv e source- channel coding for multi-user semantic and data communicat ions,","cited_arxiv_id":null,"evidence_quote":"Provides the multi-user ASCC framework that jointly optimizes coding rates, power, and beamforming, from which the proposed sensing-aware extension departs."}],"review_version":1}