{"id":"fa99c1fb-b59f-4f0a-96c3-fd97aa6469bf","arxiv_id":"2509.06626","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"BGP hijacking and passive interception at a single AS can censor most IPFS content for most requesters, because content providers are concentrated in a few weakly protected cloud prefixes.","lead":"This paper shows that a single internet provider or nation-state could block a large share of IPFS content for most users by hijacking a handful of network routes. It maps where IPFS content actually lives and argues that the network's reliance on a few cloud providers makes it more centrally censorable than its design intends.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 75%/57% headline is per-CID max over 100 attacker ASes (Fig. 8); the paper never shows one fixed AS achieves both thresholds, so 'a single malicious AS' is not established.","rationale":"The reader's conditional verdict rests on the simplified BGP model and static topology, which is a legitimate external-validity concern. My stress-test identifies a more fundamental, internal issue: the headline statistic appears to take, for each CID, the maximum over the 100 attacker ASes, while the abstract's wording implies a single fixed AS accomplishes the 75%/57% result across all content. The paper never reports the fixed-AS version of the metric, so the central claim is not actually established by the evaluation as presented. This does not require rejecting the paper: the existential reading ('for each vulnerable CID there exists some AS') may be true, and the underlying measurements are valuable. But the abstract and the reader's strongest_claim should be corrected to match what was computed, and the per-AS recomputation should be reported. The verdict remains conditional because the required fix is a re-analysis of already-collected data, not new data collection; if the fixed-AS analysis fails to reproduce the headline, the abstract's central quantitative claim would need to be substantially weakened.","tokens_in":20929,"tokens_out":7779,"duration_ms":70389,"concrete_test":"Re-run the simulation output per AS rather than per CID: for each of the 100 attacker ASes, compute for every CID the requester-blockage rate using that AS alone. Then, for each AS, compute the fraction of the 3,000 CIDs with blockage >57%, and report the maximum across ASes and the value for AS174. If the maximum is below 75%, the abstract's 'single malicious AS' claim is false and should be reworded to per-CID existential form; if it is at least 75%, identify which AS(es) achieve it. As a secondary check, recompute the '62 prefixes / 70% effectiveness' cost result with a single fixed prefix set applied to all CIDs, since the greedy selection in Section 4.3 may otherwise reflect per-CID optimization rather than one attacker's simultaneous actions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on a metric computed per CID as the maximum success rate over the 100 candidate attacker ASes (Fig. 8 caption: 'maximal rate of successfully blocked requesters by a single AS'). The abstract converts this into 'a single malicious AS can censor 75% of the IPFS content for more than 57% of all requester nodes,' which implies one fixed adversary. Section 4.2 does not report the fixed-AS analogue: for a given AS, the fraction of CIDs for which that same AS blocks more than 57% of requesters. Figure 9 even shows that the strongest single AS (AS174) blocks only about 67% of all CID x Requester pairs on average, so it is doubtful that any one AS reaches the 75%-of-CIDs/57%-of-requesters bar. If the best AS differs by CID, the headline overstates a single adversary's reach. This is an internal aggregation issue, independent of BGP-model realism: even within the paper's own model, the claim as worded is not demonstrated.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies network-level censorship of IPFS content via passive BGP interception and active BGP hijacking. The authors collected 3,000 CIDs from the live IPFS network using Bitswap monitoring, crawled IPFS server nodes to build a requester set, and simulated attacks over CAIDA-based AS topology and RPKI classification. The main claims are that 0.5–8% of CIDs can be fully blocked from all requesters, that a single malicious AS can block more than 57% of requesters for about 75% of CIDs, and that hijacking only 62 IP prefixes achieves roughly 70% of the full attack effectiveness. The paper also proposes and simulates countermeasures based on IPFS clusters and global collaborative pinning, reporting that replicating content to about 0.5% of server nodes limits the strongest attacker to a roughly 20% interception rate.","tokens_in":21121,"tokens_out":7479,"duration_ms":62041,"significance":"If the headline results hold, this would be the first systematic analysis of BGP-based routing attacks against IPFS, extending prior work on Bitcoin, Ethereum, and other decentralized systems to an important Web3 storage layer. The paper's strengths include real-world CID collection, the use of externally sourced CAIDA topology and RPKI data, and the explicit acknowledgment that the model provides an upper bound on realistic attacker effectiveness. The countermeasure analysis is a useful first step, although its protocol-level incentive questions are only sketched. However, the central quantitative claim in the abstract is not currently demonstrated by the presented experiments because of an aggregation issue in how the per-CID attacker success rates are combined; the contribution is significant but needs a re-analysis of the headline metric before it can be accepted.","major_comments":[{"comment":"The headline claim that 'a single malicious AS can censor 75% of the IPFS content for more than 57% of all requester nodes' is not supported by the presented evaluation. Figure 8 plots, for each CID, the maximal success rate over the 100 candidate attacker ASes, so the maximizing AS may differ across CIDs; the paper does not report the fixed-AS analogue (for example, for each AS, the fraction of CIDs for which that AS blocks more than 57% of requesters). Figure 9 reports a different quantity, the average portion of CID×requester pairs per AS (67% for AS174), which does not by itself establish the headline threshold. In addition, Figure 8 is labeled 'dataset 1' only, while the abstract presents the result as if it covers all three datasets. Please provide the fixed-AS distribution for each dataset and revise the abstract and Section 4.2 to match the evidence.","section":"§4.2 / Abstract / Fig. 8"},{"comment":"The statement that 'a small set of only 62 hijacked prefixes' reaches '70% of the full attack effectiveness' is not substantiated in the text. Section 4.3 describes a greedy prefix-selection algorithm and Figure 11 shows average success-rate curves, but neither the number 62 nor a definition of 'full attack effectiveness' appears in the narrative. Please specify the target metric (e.g., average over CIDs of the fraction of blocked requesters), the greedy algorithm's stopping rule, and how the exact 62-prefix figure is derived from Figure 11 or the underlying data.","section":"§4.3 / Abstract"},{"comment":"The appendix's list of the 'top 100' attacker ASes contains duplicate entries (e.g., AS1221 appears at ranks 27 and 73, AS2635 at ranks 28 and 74, and AS16509 at ranks 39 and 85, among many others). If these duplicates are used in the simulation, the attacker set comprises fewer than 100 distinct ASes; if they are not used, the appendix should list unique ASes. This point must be clarified and corrected because the attacker set is a core input to the reported effectiveness numbers.","section":"Appendix 8.2"}],"minor_comments":[{"comment":"The phrase 'all requester nodes' should be qualified as 'all IPFS server nodes in the crawled topology' to match the actual requester set; the limitation is acknowledged in Section 6 but should be reflected in the abstract and introduction.","section":"Abstract / §1"},{"comment":"The sentence 'we considered thecid.contactdomain' should read 'the cid.contact domain'.","section":"§3.2"},{"comment":"The caption of Figure 12 states 'Mean of 5 measured samples'; please clarify in the text how the five samples were drawn and whether the plotted curve is the maximum over the 100 ASes or an average, as the y-axis label says 'Maximally intercepted Requester Fraction'.","section":"§5.2 / Fig. 12"},{"comment":"The paper says the full IPFS crawl takes about 5 minutes and that two crawls were executed on July 12 and 13, 2025. Please clarify whether the requester snapshot is contemporaneous with the CID datasets collected in July 2025, and if not, discuss the potential effect of topology drift on the results.","section":"§4"},{"comment":"References [8] and [28] contain 'Accessed: [Insert Date of Access]' placeholders that should be filled in.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The central issue is the aggregation mismatch between the per-CID maximum over attacker ASes (Fig. 8) and the abstract's 'single malicious AS' wording. This is fixable by re-computing the headline metric as a fixed-AS distribution. The duplicate ASNs in Appendix 8.2 are concerning because they suggest either a formatting error or a simulation-input error; please verify that the simulation uses 100 distinct ASes. If the authors can provide the corrected fixed-AS analysis, the paper's main claim would be substantially stronger."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWorth a look. This is the first analysis I've seen of BGP-level attacks against IPFS content, and the measurement effort is real. They collect 3,000 live CIDs, map providers/resolvers/bitswap peers to ASes and prefixes, classify RPKI status, and simulate passive interception and hijacking over a crawled server-node topology. The central qualitative finding—IPFS content is concentrated enough that an AS-level adversary can block a large fraction of requesters for many CIDs—looks plausible and is consistent with earlier centralization measurements. The countermeasure analysis (global random pinning plus RPKI-hardened provider nodes) is a useful first cut, even though the enforcement protocol is left open.\n\nThe main problem is the headline claim. Figure 8 plots, for each CID, the maximum success rate over the 100 candidate attacker ASes. The abstract converts that into 'a single malicious AS can censor 75% of the IPFS content for more than 57% of all requester nodes.' But the paper never shows a fixed AS that achieves that. Figure 9 says AS174 blocks 67% of all CID x requester pairs on average; that's a different metric, and it doesn't tell us what fraction of CIDs AS174 can block for >57% of requesters. So the abstract overstates what the results demonstrate. This is an internal aggregation issue, not a question of BGP realism. The fix is straightforward: report the distribution over ASes of the fraction of CIDs that each AS can block at the 57% requester threshold, and state the headline in terms of the actual metric.\n\nOther soft spots are acknowledged by the authors: simplified BGP, static topology, server nodes only as requesters, no confidence intervals. Those are worth noting but don't undercut the core contribution. Lack of released code/data is a bit annoying for a measurement paper.\n\nVerdict: send it to peer review, but with a request to fix the aggregation problem and temper the abstract. The work is novel, the measurements are useful, and the qualitative conclusion—IPFS is vulnerable to AS-level censorship—survives even if the exact percentages are upper bounds. Just don't let the current headline stand.\n\nBest.","headline":"First AS-level censorship study of IPFS, with solid measurement and simulation work; the 'single malicious AS' headline overstates what the per-CID max metric actually shows.","tokens_in":21669,"tokens_out":3223,"would_cite":true,"duration_ms":26583,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a single malicious autonomous system can censor most live IPFS content for most requesters, because IPFS providers and resolvers are concentrated in a few weakly protected cloud prefixes.","keywords":["IPFS","BGP hijacking","censorship","autonomous systems","content availability","RPKI","decentralized storage","routing attacks"],"falsifier":"Run a controlled live test: from a test AS, announce a hijack of one of the provider prefixes the simulation marks as vulnerable for a known set of CIDs, then measure retrieval success from distributed IPFS requesters over several days; if fewer than half the simulated requesters actually fail to retrieve the content, the model's upper-bound estimate is too high. Complement this with BGP route-collector data showing that operators' local preferences and filters divert the announced route back to the legitimate path.","tokens_in":20693,"feed_emoji":"🚫","tokens_out":6328,"duration_ms":54268,"temperature":0.7,"pith_summary":"IPFS was built as a decentralized, censorship-resistant file system, but its live network has drifted toward concentration: most content is stored and resolved by a small set of nodes sitting in cloud-provider prefixes. The paper argues that this concentration turns ordinary routing-plane attacks into a censorship lever. Simulating passive interception and BGP hijacking against 3,000 content identifiers (CIDs) on the measured IPFS topology, it finds that a single malicious autonomous system (AS) can block more than 57% of requesters from 75% of the collected CIDs, and that hijacking about 62 prefixes already reaches 70% of maximal effectiveness. The paper further shows that spreading content across about 80 well-distributed server nodes, or adding backup providers on RPKI-protected max-length prefixes, cuts the achievable interception rate to roughly 20%. If the simulation reflects reality, any well-connected AS—an ISP, transit provider, or nation-state operator—could selectively deny IPFS content at scale.","feed_headline":"One rogue network can censor most IPFS content","feed_subtitle":"Because providers cluster in a few cloud prefixes, one AS can block 75% of content for most requesters.","key_machinery":"The mechanism is the CID retrieval chain and its routing-plane weak points. Retrieving a CID involves three contact sets: Bitswap peers that cache the block, DHT resolver nodes holding provider records, and the content providers themselves. The adversary maps all three for each CID, then uses BGP hijacking or passive interception to cut the requester's routes to them. The simulation's engine is an AS-level routing model that computes shortest-path routing trees over business relationships and classifies every provider prefix by RPKI status, where unprotected and short-prefix entries are hijackable by any AS and RPKI max-length entries only by a closer AS. The greedy prefix-selection count (62 prefixes for 70% effectiveness) and the countermeasure simulations (replication fraction versus interception rate) follow from the same model.","core_discovery":"The central claim is that IPFS censorship does not require an attack on IPFS itself; it can be done by attacking the Internet routing that carries IPFS traffic. Because content providers and the DHT resolvers that map CIDs to providers are concentrated in a few ASes and IP prefixes, and because most of those prefixes are not protected at RPKI max-length, an adversary controlling one AS can drop or hijack the connections requesters need. On the paper's measurements, passive interception alone blocks fewer than 20% of requesters for most CIDs, whereas BGP hijacking achieves roughly 70% blockage for most CIDs, and the combined attack lets one AS block 75% of the collected content for more than 57% of requester nodes. Only 0.5–8% of CIDs can be blocked for every requester; the rest leak through some provider or resolver. The paper also claims that the defenses it simulates—global randomized collaborative pinning plus RPKI-hardened backup providers—can reduce the maximal interception rate to around 20%.","pith_inferences":["The requester set is limited to observable IPFS server nodes; real end users behind gateways and client nodes may face different, possibly higher, blockage, since gateways concentrate many consumers and are weighted equally here.","Because the paper's own model treats route selection as shortest-AS-path with no local preferences or operator filtering, the 75% and 57% figures are upper bounds; live deployments with RPKI filtering and hijack mitigation would likely show lower, but still non-trivial, rates.","A natural extension is to weight requesters by served users and to test the greedy prefix-selection against live BGP data, which would tell whether the top prefixes are also the ones an operator would notice hijacking.","The collaborative-pinning result suggests a protocol-level default: IPFS could replicate popular CIDs to a small random set of well-distributed server nodes automatically, without waiting for providers to opt in."],"forward_implications":["An ISP or transit provider that controls a single well-connected AS can deny a majority of requesters access to the majority of live IPFS content without touching DNS or any IPFS software.","Because 62 hijacked prefixes buy 70% of maximal effectiveness, the attack is cheap enough for a well-resourced actor to run continuously, not just as a one-off disruption.","Fully blocking a specific CID is hard (only 0.5–8% of CIDs were blockable from all requesters), so targeted censorship is easier than blanket censorship.","Replicating popular content to roughly 80 distributed server nodes, about 0.5% of the measured server population, caps the best single-AS interception rate at about 20%.","Adding at least one provider on an RPKI-protected max-length prefix makes a CID effectively unhijackable in the simulation, as seen with the three IPFS cluster projects tested."],"supporting_citations":[{"why":"Supplies the Bitswap-monitoring infrastructure and CID collection method used to gather the 3,000 CIDs and their providers.","marker":"[3]"},{"why":"Provides the IPFS crawler used to build the requester topology and documents cloud centralization that motivates the attack surface.","marker":"[4]"},{"why":"Supplies empirical evidence that more than 75% of IPFS nodes sit in a few cloud providers, the concentration the attack exploits.","marker":"[10]"},{"why":"Provides the AS business-relationship graph used to compute routing choices in the simulation.","marker":"[19]"},{"why":"Provides the routing-tree algorithm used to derive AS paths between requesters and targets.","marker":"[22]"},{"why":"Supplies the routing-attack simulation methodology and the assumption that shortest-path BGP approximates real routing.","marker":"[53]"},{"why":"Provides the implementation of passive interception path calculation that the paper adapts for its evaluation.","marker":"[55]"}],"fun_headline_variants":["BGP hijacking blocks 75% of IPFS content","A single AS can censor most IPFS data","IPFS censorship: one AS blocks 75%","62 hijacked prefixes hit 70% of IPFS censorship","BGP attacks threaten IPFS: 75% content at risk"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The attack numbers rest on the assumption that real Internet routing behaves like the simplified shortest-path model used in the simulation; if operators filter hijacked announcements or prefer other routes in practice, the real blockage rates will be lower than reported.","fun_headline_variants_meta":{"raw":{"variants":["BGP hijacking blocks 75% of IPFS content","A single AS can censor most IPFS data","IPFS censorship: one AS blocks 75%","62 hijacked prefixes hit 70% of IPFS censorship","BGP attacks threaten IPFS: 75% content at risk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000365,"raw_usage":{"total_tokens":2011,"prompt_tokens":1040,"completion_tokens":971,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":656,"completion_tokens_details":{"reasoning_tokens":888}},"tokens_in":656,"tokens_out":971,"duration_ms":6888,"temperature":1.0,"reasoning_tokens":888,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:14:29.472865+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled live test: from a test AS, announce a hijack of one of the provider prefixes the simulation marks as vulnerable for a known set of CIDs, then measure retrieval success from distributed IPFS requesters over several days; if fewer than half the simulated requesters actually fail to retrieve the content, the model's upper-bound estimate is too high. Complement this with BGP route-collector data showing that operators' local preferences and filters divert the announced route back to the legitimate path.","supporting_citations":[{"cited_title":"Monitoring data requests in decentralized data storage systems: A case study of ipfs","cited_arxiv_id":null,"evidence_quote":"Supplies the Bitswap-monitoring infrastructure and CID collection method used to gather the 3,000 CIDs and their providers."},{"cited_title":"The Cloud Strikes Back: Investigating the Decentralization of IPFS","cited_arxiv_id":"2309.16203","evidence_quote":"Provides the IPFS crawler used to build the requester topology and documents cloud centralization that motivates the attack surface."},{"cited_title":"Cen- tralization in the decentralized web: Challenges and opportunities in ipfs data management.Proceedings of the ACM Web Conference (WWW), April 2025","cited_arxiv_id":null,"evidence_quote":"Supplies empirical evidence that more than 75% of IPFS nodes sit in a few cloud providers, the concentration the attack exploits."},{"cited_title":"As relationships dataset","cited_arxiv_id":null,"evidence_quote":"Provides the AS business-relationship graph used to compute routing choices in the simulation."},{"cited_title":"How secure are secure interdomain routing protocols? Technical report, Microsoft Research and Yale University, 2010","cited_arxiv_id":null,"evidence_quote":"Provides the routing-tree algorithm used to derive AS paths between requesters and targets."},{"cited_title":"Rout- ing attacks on cryptocurrency mining pools","cited_arxiv_id":null,"evidence_quote":"Supplies the routing-attack simulation methodology and the assumption that shortest-path BGP approximates real routing."},{"cited_title":"Rev- elio: A network-level privacy attack in the lightning network","cited_arxiv_id":null,"evidence_quote":"Provides the implementation of passive interception path calculation that the paper adapts for its evaluation."}],"review_version":2}