{"id":"0281fc48-6498-40fc-88c4-802675ae5eed","arxiv_id":"2507.12452","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A monolithically integrated 32-channel WDM PAM4 receiver reports 1.024 Tb/s aggregate rate and 0.38 pJ/bit, but the aggregate rate appears to be extrapolated rather than demonstrated concurrently.","lead":"A 32-channel silicon photonics receiver chip is reported to reach 1.024 Tb/s on a single optical fiber using PAM4 modulation and no digital signal processing, with energy efficiency below 0.38 pJ/bit. The significance hinges on whether all 32 channels really run at once, since the chip's own description shows some decoder circuits are shared.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Architecture text contradicts headline: shared 4-to-1 multiplexed PAM4 decoders allow only eight of 32 channels to be decoded simultaneously, so 1.024 Tb/s and 0.38 pJ/bit are extrapolated, not demonstrated.","rationale":"The reader's REJECT is well founded. The central claim is the 1.024 Tb/s aggregate PAM4 receiver at under 0.38 pJ/bit, and that claim requires all 32 channels to be decoded simultaneously at 32 Gb/s. The paper explicitly describes a shared decoder architecture: TIA outputs are 4-to-1 multiplexed into PAM4 decoders, 'enabling on-chip measurements of all channels, eight at a time.' A 4:1 MUX before the decoder means only one of four TIA outputs reaches a decoder at any instant; with eight decoder groups, only eight channels can be active concurrently. The BER plot in Fig. 5n is therefore a composite of sequential eight-channel measurements, not a simultaneous 32-lane result. The 0.38 pJ/bit efficiency is computed by summing 32 TIA powers (6.89 mW each) and 32 decoder powers (4.88 mW each) and dividing by 1024 Gb/s; if only eight decoders exist, the decoder power should be divided by eight, and the actual simultaneous throughput is 256 Gb/s, changing both the data rate and the efficiency metric. This is an internal inconsistency, not a disagreement with community consensus. The per-channel characterization appears thorough, with measured eye diagrams, sensitivity curves, and bathtub plots, and the O-DeMux tuning and crosstalk work is documented. None of that rescues the aggregate claim unless a decoder-per-channel path or a simultaneous 32-channel demonstration is provided. The single decisive check is counting decoder cores in the layout: eight cores invalidate the headline. Therefore the reader's REJECT verdict stands unchanged.","tokens_in":15626,"tokens_out":6960,"duration_ms":70186,"concrete_test":"Inspect the chip layout and micrograph (Fig. 5a, b) and count the number of PAM4 decoder cores and 4:1 MUX outputs. If there are only eight decoder cores fed by eight 4:1 MUXes, the demonstrated simultaneous decode rate is 8×32=256 Gb/s, invalidating the 1.024 Tb/s headline. If the design contains 32 decoders and the MUX is only a test tap, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires all 32 wavelength channels to be decoded concurrently at 32 Gb/s each. The paper's own architecture description (Results, 'PAM4 detection and decoding', Fig. 4a) says TIA outputs are '4-to-1 multiplexed and routed to PAM4 decoders', which 'enabl[es] on-chip measurements of all channels, eight at a time, while reducing the required chip area.' A 4:1 MUX placed before a PAM4 decoder means each decoder can process only one of its four TIA inputs at any instant; with eight such groups, at most eight channels are decoded simultaneously. The per-channel BER results in Fig. 5n were therefore collected sequentially in groups of eight, not as a 32-lane concurrent link. The abstract's claim of a 'concurrent electrical detection system' is contradicted by this shared-decoder arrangement. Since the headline 1.024 Tb/s aggregate rate, the 0.38 pJ/bit energy efficiency (computed as 32×(6.89+4.88) mW / 1024 Gb/s), and the bandwidth density all assume 32 simultaneously active channels, the central claim is not supported by the described demonstration. The receiver's per-channel performance may be solid, but the aggregate claim depends on an architecture that the text itself undermines.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a monolithically integrated 32-channel WDM silicon-photonics PAM4 receiver in the GlobalFoundries 45CLO process. The receiver combines a 1:32 optical demultiplexer (an MZI binary tree followed by cascaded ring resonators) with capacitive and thermal phase shifters for autonomous wavelength locking, and a CMOS electrical detection chain of photodiodes, TIAs, PAM4 decoders, deserializers, and on-chip BER testers. The authors claim an aggregate data rate of 1.024 Tb/s (32 channels x 32 Gb/s) on a single input fiber, BER below 10^-12, end-to-end latency under 100 ps, energy efficiency under 0.38 pJ/bit, and a bandwidth density of 3.55 Tb/s/mm^2, all without equalization or DSP. The central experimental evidence is per-channel BER and bathtub measurements taken from a single-channel test chip and from the 32-channel chip one selected channel at a time.","tokens_in":15902,"tokens_out":6944,"duration_ms":79045,"significance":"If fully supported, the demonstrated receiver would be a significant advance in monolithic WDM optical receivers, combining dense wavelength-division multiplexing, near-zero-power capacitive wavelength tuning, and DSP-free PAM4 detection at record aggregate bandwidth. The paper includes useful engineering strengths: a full 32-channel monolithic demonstration, measured ring-resonator process-variation statistics, autonomous locking with low tuning power, and on-chip BER measurement. However, the headline aggregate-rate claim is not supported by the architecture as described, because the PAM4 decoders are shared through 4-to-1 multiplexers and only eight channels can be decoded simultaneously. The significance of the work therefore hinges on a load-bearing claim that the manuscript itself contradicts.","major_comments":[{"comment":"The text states that 'the outputs of the TIAs are 4-to-1 multiplexed and routed to PAM4 decoders followed by de-serializer and BER measurement system blocks enabling on-chip measurements of all channels, eight at a time.' This means at most eight of the 32 channels can be decoded simultaneously: one channel per 4-to-1 multiplexed group. Consequently, the demonstrated simultaneous data rate is at most 8 x 32 Gb/s = 256 Gb/s, not 1.024 Tb/s. The per-channel BER results in Fig. 5n were collected sequentially in groups of eight, not as a 32-lane concurrent link. The abstract's phrase 'concurrent electrical detection system' does not resolve this, because detection without simultaneous PAM4 decoding cannot deliver the claimed aggregate output. The headline claim of operation at 1.024 Tb/s therefore needs either evidence that all 32 channels have dedicated concurrent decoders, or a revision of the aggregate-rate claim.","section":"Results, 'PAM4 detection and decoding' (Fig. 4a)"},{"comment":"The energy-efficiency calculation of about 0.38 pJ/bit uses 6.89 mW per TIA and 4.88 mW per PAM4 decoder summed over 32 channels and divided by 1.024 Tb/s. This implicitly assumes 32 decoders operating simultaneously at 32 Gb/s each. With the described 4-to-1 multiplexed decoder architecture, only eight decoders are active at once, so the efficiency at the claimed aggregate rate is not established. The bandwidth-density claim of 3.55 Tb/s/mm^2 also cannot be reproduced from the stated chip footprint (4.2 mm^2 in the introduction versus 4.72 mm^2 in Fig. 5b) and appears to depend on an unspecified area definition. Both metrics are load-bearing for the record comparisons in Fig. 1c-d and Extended Data Table 1.","section":"Discussion and Methods (energy efficiency; bandwidth density)"}],"minor_comments":[{"comment":"The paper reports BER below 10^-12 for all channels but does not state the number of bits examined or the confidence interval for each BER measurement; please add this information so the BER claim is verifiable.","section":"Results, Fig. 5n and on-chip BER"},{"comment":"The 'end-to-end latency of under 100 ps' claim is not supported by any described measurement or simulation in the manuscript; either provide the measurement setup and result or remove the claim.","section":"Abstract and Results"},{"comment":"The chip footprint is given as 4.2 mm^2 in the introduction and 4.72 mm^2 in the Fig. 5b caption; please reconcile these numbers and define the area used for bandwidth-density calculations.","section":"Introduction and Fig. 5b"},{"comment":"In the sentence about Fig. 5j, 'WDM OAM4 receiver' appears to be a typo for 'WDM PAM4 receiver'; please correct it.","section":"Results, 'System integration'"},{"comment":"There are minor typographical errors such as 'stat-of-the-art' and 'the stat-of-the-art' in the Discussion; please proofread the text.","section":"Discussion"}],"recommendation":"reject","confidential_remarks":"The per-channel photonic and electronic results appear credible and the monolithic integration is impressive, but the central claim of 1.024 Tb/s aggregate operation is directly contradicted by the described 4-to-1 multiplexed decoder architecture. This is not a presentation issue: without simultaneous decoding of all 32 channels, the headline data rate, energy efficiency, and bandwidth-density comparisons are not supported. If the authors can provide evidence that all 32 channels actually have concurrent decoders, or if they resubmit with the aggregate claim reduced to the demonstrated 256 Gb/s simultaneous operation, the work could be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nBottom line: the headline 1.024 Tb/s is not supported by the paper's own description of the chip. The Results section (Fig. 4a) says the 32 TIA outputs are 4-to-1 multiplexed into shared PAM4 decoders, which 'enables on-chip measurements of all channels, eight at a time.' So only eight channels can be decoded at any instant; the demonstrated simultaneous rate is 256 Gb/s, not 1.024 Tb/s. The abstract's 'concurrent electrical detection system' is doing a lot of work, and the energy-efficiency claim of 0.38 pJ/bit is computed assuming 32 decoders are active, so it inherits the same problem.\n\nThat said, the paper is not empty. The monolithic 1:32 O-DeMux built from an MZI binary tree plus ring array, with capacitive phase shifters and autonomous per-device locking at near-zero static power, is a genuinely useful integration. The measured per-channel BER (<1e-12 at 32 Gb/s), TIA power (6.89 mW), and decoder power (4.88 mW) are concrete, and the bathtub measurements look careful. The channel-wavelength mapping and the 150-ring process-variation study are also good supporting data. The as-stated per-channel receiver performance is believable.\n\nThe soft spots beyond the aggregate-rate issue: no public data or design files, so independent verification is limited to what the figures show. The comparison to prior PAM4 receivers uses the inflated aggregate, so the 'record' and 'bandwidth-density-energy product' claims should be re-benchmarked at 256 Gb/s or after a true 32-channel concurrent demo. I don't see a circularity problem—this is measurement, not derivation.\n\nWho is this for? People working on WDM silicon photonics receivers and autonomous wavelength control will get value from the O-DeMux and locking work. The paper deserves a serious referee rather than a desk reject, but the authors need to either show 32 channels operating simultaneously (e.g., with 32 decoders, or a mux that cycles fast enough to claim aggregate, which they don't) or scale the claims down to what was actually measured. I'd recommend major revision with the burden on the authors to reconcile the architecture text with the headline.\n\nRegards,","headline":"The 1.024 Tb/s aggregate claim collapses under the paper's own 4-to-1 multiplexed decoder architecture, but the integrated 32-channel O-DeMux and per-channel measurements are real and worth reviewing.","tokens_in":16484,"tokens_out":5630,"would_cite":false,"duration_ms":55018,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A monolithic 32-channel WDM PAM4 receiver chip achieves 1.024 Tb/s over one fiber at under 0.38 pJ/bit.","keywords":["silicon photonics","WDM receiver","PAM4","optical interconnects","monolithic integration","capacitive phase shifter","wavelength locking","energy efficiency"],"falsifier":"Drive all 32 PAM4-modulated carriers into the chip simultaneously and count errors on all decoder outputs together over a sustained run; if the aggregate error rate exceeds one in $10^{12}$ bits while all 32 channels are active, the reported 1.024 Tb/s error-free claim is not a simultaneous one. A simpler check is to capture four TIA outputs through the 4-to-1 multiplexer at once, rather than one channel at a time, to confirm concurrent decoding.","tokens_in":15431,"feed_emoji":"📡","tokens_out":10930,"duration_ms":114266,"temperature":0.7,"pith_summary":"The paper reports a single-chip optical receiver that demultiplexes 32 wavelength channels from one input fiber and decodes each as a four-level PAM4 signal at 32 Gb/s, for an aggregate rate of 1.024 Tb/s. It further reports a bit-error rate below $10^{-12}$ with no equalization, pre-distortion, or digital signal processing, an end-to-end latency under 100 ps, and a chip energy efficiency under 0.38 pJ/bit. The result matters for data-center and AI interconnects, where power per bit and packaging complexity are the main constraints; a receiver that reaches terabit-scale throughput on a single fiber and locks its wavelengths at near-zero static power targets both. The paper also reports a bandwidth density above 3.55 Tb/s/mm$^2$, which it identifies as the largest reported data rate for a WDM PAM4 receiver chip to date.","feed_headline":"Silicon photonics chip hits 1.024 Tb/s on one fiber","feed_subtitle":"A single receiver chip decodes 32 wavelengths at 32 Gb/s each with no DSP or equalization.","key_machinery":"The central mechanism is the 1:32 optical demultiplexer: a 1:8 MZI binary tree followed by an 8:32 drop-ring resonator array, with a capacitive phase shifter and sense-actuation-memory (SAM) control loop on each tunable stage. The MZI tree is what makes the dense 200 GHz grid manageable; it first groups carriers so each ring sees only carriers spaced 1600 GHz apart, relaxing the ring isolation requirement. The capacitive phase shifters, built from the CMOS gate stack, shift resonances without DC current, and the SAM loops lock each MZI and ring to its assigned carrier using only a small tapped monitor photocurrent. That tuning machinery is what allows the zero-equalization, zero-DSP receiver to hold its $10^{-12}$ error rate despite fabrication spread and temperature drift.","core_discovery":"The paper's central claim is that monolithic integration of the optics and electronics, rather than a faster serial electrical front end, is the route to terabit-scale single-fiber reception. The chip separates 32 carriers spaced 200 GHz apart using a 1:8 Mach-Zehnder interferometer binary tree followed by eight branches of three drop-ring resonators; the tree first widens the channel spacing to 1600 GHz to ease ring crosstalk. Each demultiplexed channel is detected by a differential silicon-germanium photodiode pair and transimpedance amplifier, then decoded by a 2-bit time-interleaved flash PAM4 decoder. Capacitive phase shifters in the MZI and ring stages give autonomous wavelength locking at zero static power, so the receiver needs no equalization or DSP. The measured outcome is a bit-error rate below $10^{-12}$ on all 32 channels at 32 Gb/s per channel, i.e. 1.024 Tb/s on one fiber, at 0.38 pJ/bit and above 3.55 Tb/s/mm$^2$.","pith_inferences":["The same MZI-tree-plus-drop-ring architecture should extend to 64 or 128 wavelengths by adding tree stages and ring branches, since each locking loop is power-neutral and the electrical channels are replicated.","Scaling per-channel line rates beyond 32 Gb/s would eventually force equalization or DSP as TIA bandwidth and PAM4 eye closure become limiting; the design's margin at 32 Gb/s does not reveal where that boundary sits.","A full end-to-end link test with a multi-wavelength comb or laser-array transmitter would exercise realistic inter-channel crosstalk and source line noise that single-carrier receiver tests do not capture.","The near-zero-power capacitive tuning scheme is portable to other wavelength-selective photonic circuits, such as optical switches or sensors, wherever drift correction must cost almost no static power."],"forward_implications":["A single fiber can deliver 1.024 Tb/s into a receiver with no DSP or equalizer, cutting both optical-module power and packaging cost relative to multi-fiber or DSP-heavy designs.","At under 0.38 pJ/bit the receiver is reported as more than five times more energy efficient than existing end-to-end CMOS PAM4 receivers above 100 Gb/s, a direct benefit for power-constrained AI and data-center racks.","Autonomous near-zero-power wavelength locking lets the chip track thermal and process drift without spending milliwatts on heater tuning, so the energy-efficiency advantage persists in deployed conditions.","The >3.55 Tb/s/mm$^2$ bandwidth density suggests the receiver could sit close to a compute die in co-packaged optics without dominating package area.","All 32 channels meet BER below $10^{-12}$ at 32 Gb/s with a >0.05 UI timing opening, giving integration margin for a real link rather than a single-channel demo."],"supporting_citations":[{"why":"Supplies the prior monolithically integrated transmitter with autonomous wavelength stabilization that this receiver design builds on.","marker":"[38]"},{"why":"Introduces the monolithically integrated autonomous demultiplexer with near-zero-power capacitive phase shifters used for the O-DeMux.","marker":"[39]"},{"why":"Defines the 45 nm CMOS-silicon photonics foundry platform and its device parameters, including transistor speed and photodiode bandwidth.","marker":"[23]"},{"why":"Serves as a state-of-the-art CMOS PAM4 optical receiver baseline for the bandwidth-density and energy-efficiency comparison.","marker":"[40]"},{"why":"Provides a recent DSP-based PAM4 transceiver datum that the paper compares against for energy efficiency and aggregate data rate.","marker":"[34]"},{"why":"Provides an eight-lane 800 Gb/s PAM4 transceiver baseline for aggregate-rate and packaging comparison.","marker":"[35]"},{"why":"Supplies the broadband directional coupler design used in the MZI stages of the optical demultiplexer.","marker":"[42]"},{"why":"Supplies the track-and-regenerate slicer topology used in the PAM4 decoder.","marker":"[43]"},{"why":"Supplies the parallel PRBS generator used by the on-chip bit-error-rate measurement system.","marker":"[41]"}],"fun_headline_variants":["1.024 Tb/s on one fiber from a single silicon chip","Silicon receiver demuxes 32 wavelengths at 32 Gb/s each","No DSP, no equalization: 1.024 Tb/s PAM4 receiver","Zero-static-power tuning in a 1.024 Tb/s photonic receiver","Single chip, 32 channels: 1.024 Tb/s optical receiver"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline data rate assumes the 32 channels can be decoded at the same time, but the on-chip error measurements were taken eight channels at a time through shared decoder hardware.","fun_headline_variants_meta":{"raw":{"variants":["1.024 Tb/s on one fiber from a single silicon chip","Silicon receiver demuxes 32 wavelengths at 32 Gb/s each","No DSP, no equalization: 1.024 Tb/s PAM4 receiver","Zero-static-power tuning in a 1.024 Tb/s photonic receiver","Single chip, 32 channels: 1.024 Tb/s optical receiver"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000691,"raw_usage":{"total_tokens":3199,"prompt_tokens":1087,"completion_tokens":2112,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":703,"completion_tokens_details":{"reasoning_tokens":2006}},"tokens_in":703,"tokens_out":2112,"duration_ms":14746,"temperature":1.0,"reasoning_tokens":2006,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:45:44.416485+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Drive all 32 PAM4-modulated carriers into the chip simultaneously and count errors on all decoder outputs together over a sustained run; if the aggregate error rate exceeds one in $10^{12}$ bits while all 32 channels are active, the reported 1.024 Tb/s error-free claim is not a simultaneous one. A simpler check is to capture four TIA outputs through the 4-to-1 multiplexer at once, rather than one channel at a time, to confirm concurrent decoding.","supporting_citations":[{"cited_title":"IEEE Journal of Solid -State Circuits, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the prior monolithically integrated transmitter with autonomous wavelength stabilization that this receiver design builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the monolithically integrated autonomous demultiplexer with near-zero-power capacitive phase shifters used for the O-DeMux."},{"cited_title":"45nm CMOS -silicon photonics monolithic technology (45CLO) for next - generation, low power and high speed optical interconnects","cited_arxiv_id":null,"evidence_quote":"Defines the 45 nm CMOS-silicon photonics foundry platform and its device parameters, including transistor speed and photodiode bandwidth."},{"cited_title":"IEEE Journal of Solid-State Circuits, 2021","cited_arxiv_id":null,"evidence_quote":"Serves as a state-of-the-art CMOS PAM4 optical receiver baseline for the bandwidth-density and energy-efficiency comparison."},{"cited_title":"7.1 A 2.69 pJ/b 212Gb/s DSP -based PAM-4 transceiver for optical direct -detect application in 5nm FinFET","cited_arxiv_id":null,"evidence_quote":"Provides a recent DSP-based PAM4 transceiver datum that the paper compares against for energy efficiency and aggregate data rate."},{"cited_title":"IEEE Journal of Solid-State Circuits, 2025","cited_arxiv_id":null,"evidence_quote":"Provides an eight-lane 800 Gb/s PAM4 transceiver baseline for aggregate-rate and packaging comparison."},{"cited_title":"Scientific reports, 2017","cited_arxiv_id":null,"evidence_quote":"Supplies the broadband directional coupler design used in the MZI stages of the optical demultiplexer."},{"cited_title":"-C., W.W.-T","cited_arxiv_id":null,"evidence_quote":"Supplies the track-and-regenerate slicer topology used in the PAM4 decoder."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the parallel PRBS generator used by the on-chip bit-error-rate measurement system."}],"review_version":1}