{"id":"8f3eec35-2dac-4682-ba25-e5ea3066030b","arxiv_id":"2412.17944","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Visiting samsung.com triggered requests to Facebook, Twitter, TikTok, Pinterest, and ad networks, and Samsung ads later appeared on globo.com.","lead":"This paper set up a phone proxy to intercept its own web traffic, then lists the third-party companies that received requests when visiting samsung.com and doing Google searches. It is a small illustrative case study of web tracking, not a systematic measurement, and similar results are already documented in earlier research.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Section 4 causal claim that a Samsung ad on globo.com followed from visiting samsung.com is not supported: there is no control condition, no timeline, and no evidence that the ad would not have appeared anyway.","rationale":"The reader's weakest_assumption correctly identifies the causal attribution of the globo.com Samsung ads to the prior samsung.com visit as the central unsupported link, and I agree that a control condition is missing. I mark agreement as partial because the same lack of evidence also weakens the first half of the claim: listing third-party hostnames (facebook.com, twitter.com, tiktok.com, etc.) does not by itself demonstrate that personal data was harvested, unless the relevant payloads are shown to contain identifiers or user data. Still, the proposed controlled experiment would directly settle the causal question, and the paper's direction is plausible and internally coherent. Because the reader already conditioned acceptance on additional data and controls, my concern does not move the verdict; the manuscript should remain CONDITIONAL pending the release of raw captures and a controlled replication.","tokens_in":6029,"tokens_out":3706,"duration_ms":33637,"concrete_test":"Run a controlled experiment with a clean device and fresh browser profile per trial. Alternating randomly, perform (A) visit samsung.com then navigate to globo.com, and (B) navigate directly to globo.com without visiting samsung.com, with no other traffic. Using mitmproxy logs, record every request on globo.com whose Host header or payload contains 'samsung.com'. Repeat at least 20 trials per condition. If Samsung ad requests occur at comparable rates in condition B, the Section 4 causal attribution fails; if they occur only or nearly only in condition A, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central evidence for cross-site tracking and retargeting is the observation in Section 4 (after Figure 3) that, after visiting samsung.com, a later visit to globo.com 'eventually reveals ... a series of advertisements originating from samsung.com.' That conclusion requires the causal claim that the prior samsung.com visit caused those ads. The manuscript offers no control condition (e.g., visiting globo.com without the samsung.com visit on the same device/network state), no timeline linking the visit to the appearance of the ads, no statement that the browser profile was clean or that background traffic on the test device was excluded, and no check that globo.com does not show Samsung ads to all visitors. Figure 3, as described, only shows a packet payload containing the string 'samsung.com'; it is not shown to contain a unique identifier, cookie, or cross-site correlation. Without these, the conclusion that the ads 'show the efficacy of technologies in tracking user interests and behaviors across the web' overstates what the captures demonstrate. A similar gap affects the web-search case: accesses to beacons.gcp.gvt2.com and optimizationguide-pa.googleapis.com are labelled as profile-building without showing that any user-specific data was transmitted. The central claim, however, can be tested directly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a small-scale empirical study of web traffic captured with a man-in-the-middle proxy (Mitmproxy). The authors describe two case studies: visiting samsung.com and performing a web search, and they claim that the captured traces show third-party tracking services being contacted and, in one case, a later visit to globo.com displaying Samsung advertisements. The paper frames these observations as evidence for 'surveillance capitalism' and concludes with policy recommendations for data protection, particularly in Brazil. The central observation—that a page visit triggers requests to multiple external domains—is plausible, but the manuscript provides no quantitative data, no timestamps, no payload excerpts, and no control conditions to support the stronger causal claims about cross-site retargeting and profile building.","tokens_in":6190,"tokens_out":2482,"duration_ms":26692,"significance":"If properly supported, the paper would fill a useful niche by providing a concrete, reproducible demonstration of third-party data flows during ordinary browsing and searching, complementing the largely theoretical surveillance-capitalism literature. The methodological approach—intercepting HTTPS with a user-installed certificate and documenting the resulting connections—is appropriate for this purpose and is a strength of the design. However, as written, the empirical content is presented only through narrative and screenshots, with no machine-readable trace, no counts, and no causal controls, so the significance of the findings cannot yet be assessed. The paper's accessibility and relevance to privacy regulation debates are real, but the evidence currently falls short of the abstract's claims about 'specific data types' and 'concrete evidence.'","major_comments":[{"comment":"The claim that Samsung advertisements on globo.com appeared 'shortly after' visiting samsung.com and that this 'shows the efficacy of technologies in tracking user interests and behaviors across the web' requires a causal link that the manuscript does not establish. There is no control condition (e.g., visiting globo.com without having first visited samsung.com on the same device and network), no timeline of requests, no statement that the browser profile was clean or that background traffic on the test device was excluded, and no check that globo.com does not display Samsung ads to all visitors. Without these, the observation is consistent with the retargeting explanation but also with several alternative explanations. This paragraph is the paper's strongest evidence, so this gap is load-bearing.","section":"Section 4, paragraph after Figure 3"},{"comment":"Figure 3 is described only as showing a packet payload containing the string 'samsung.com.' This does not demonstrate cross-site correlation. A payload snippet containing the domain name could arise from many benign mechanisms (e.g., a same-site script, a referrer field, or an analytics beacon) and does not by itself show that a unique identifier, cookie, or user profile was transferred between samsung.com and the later globo.com advertising request. The manuscript should show the actual request chain with headers, cookies, and timestamps, and explain why the observed fields constitute personal data transfer.","section":"Figure 3 and Section 4"},{"comment":"The abstract states that the research 'reveals specific data types exchanged between users and web services,' but the manuscript never identifies any data type beyond domain names. In Section 4.1, accesses to beacons.gcp.gvt2.com and optimizationguide-pa.googleapis.com are labeled as contributing to 'profile building,' yet no payload content or user-specific fields are shown. These are generic Google endpoints that can be contacted for performance monitoring, A/B testing, or resource fetching without transmitting personal data. The interpretation that these accesses constitute surveillance-capitalism data harvesting is therefore not supported by the presented evidence.","section":"Abstract and Section 4.1"},{"comment":"The methodology omits essential reproducibility information: the mobile device model and OS version, browser type, whether the device was freshly reset or had an existing profile, the number of repeated runs, the duration of each capture, the specific filter criteria used in Mitmweb, and how certificate pinning or apps that bypass the proxy were handled. The paper states that 'all personal data captured during our case studies are available at https://github.com/antonyseabramedeiros/,' but that URL points to a user profile rather than a named repository, and no dataset identifier or archival record is given. Without a citable, inspectable trace, the empirical claims cannot be independently verified.","section":"Section 3 and data availability statement"}],"minor_comments":[{"comment":"There is a typo in the sentence beginning 'Following the capture of network traffic initiated by a visit tosamsung.com'—'tosamsung.com' should be 'to samsung.com.'","section":"Section 4, first paragraph"},{"comment":"Figure 4 ('Advertising') and Figure 5 ('Web searching for Paris 6 Hotels') have captions that are too vague; each should describe what is shown, what was captured, and what the reader should conclude from it.","section":"Figure captions"},{"comment":"The anecdote about advertisements appearing after verbal discussions near a smartphone is explicitly described as having no collected evidence. This is fine as a motivation for future work, but it should be clearly separated from the empirical findings of Sections 4 and 4.1 to avoid the impression that the study supports voice-capture claims.","section":"Section 5"},{"comment":"The claim that 89 percent of Alphabet revenues derived from Google's targeted advertising by 2016 is cited to Zuboff (2023) without a page or external source; providing a primary reference would strengthen the factual basis.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is very short for the claims it makes and currently reads more like an extended abstract than a full research article. The core idea is reasonable and the topic is timely, but the absence of any quantitative or trace-level evidence makes it difficult to place in a research journal. In revision, the authors should be asked to include an anonymized but genuine capture excerpt (at minimum request URLs, timestamps, and relevant headers), a control visit, and a clear statement of what was and was not observed. If those additions are made, the paper could become a modest but useful empirical contribution; as it stands, the evidence-to-claim ratio is too low."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a small, readable case study that shows Mitmproxy captures of third-party trackers after visiting samsung.com and a claimed later Samsung ad on globo.com. The raw observation is plausible and the method is standard, but the paper's central causal claim about retargeting is not supported by the evidence it shows, and the novelty claim about filling a gap in empirical surveillance-capitalism literature is wrong.\n\nWhat's good: the paper is honest about its limits in one place (the voice-activated anecdote in the conclusions), it provides a GitHub link for the captures, and the basic setup – phone through proxy, visit a site, list external domains – is transparent and reproducible. The list of trackers (Criteo, Facebook pixels, Bing, Taboola, etc.) is concrete.\n\nThe soft spots are significant. The 'targeted ad' evidence is a single anecdote: after visiting samsung.com, a later visit to globo.com showed Samsung ads. There is no control visit to globo.com without the prior Samsung visit, no timeline, no check that the ads wouldn't have appeared for any visitor, and no identifier connecting the two sessions. The paper claims these ads 'show the efficacy' of cross-site tracking; that's an overstatement of what a packet capture shows. Same for the search case: beacons.gcp.gvt2.com and optimizationguide are labelled profile-building without evidence that any user-specific data was transmitted.\n\nThe literature gap claim is the weakest part. There is a large body of empirical web-privacy measurement work – studies using webXray and other crawlers, plus countless analyses of third-party requests – that already documents exactly this pattern. The paper doesn't cite it. That undercuts the 'we didn't find any studies' statement.\n\nAlso, the abstract promises 'specific data types exchanged' but the manuscript shows only domain names, not payload excerpts of actual data. The GitHub repo may contain more, but the paper alone doesn't support that promise.\n\nNet: this is a fine teaching demo of a Mitmproxy trace, not a research contribution. The causal leap from a couple of captures to 'surveillance capitalism revealed' is too big. If the authors added a control condition, released raw traces with timestamps, and corrected the novelty claim, it could become a useful short paper for an education or privacy-policy venue. As is, I would not send it to a serious referee; I'd suggest rejection with encouragement to resubmit after addressing the evidence gaps.\n\nRecommendation: desk reject, but with clear feedback.","headline":"A well-intentioned but thin case study whose central retargeting claim lacks controls and whose novelty claim ignores the existing web-privacy measurement literature.","tokens_in":6769,"tokens_out":2723,"would_cite":false,"duration_ms":27052,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One visit to samsung.com triggered a burst of requests to Facebook, TikTok, Pinterest, Bing, and other ad services, and a later globo.com visit showed Samsung ads, presented as concrete evidence of surveillance capitalism's data flows.","keywords":["surveillance capitalism","web tracking","third-party cookies","network traffic analysis","man-in-the-middle proxy","targeted advertising","data privacy","online behavioral advertising"],"falsifier":"On a clean phone and clean network, visit a set of neutral news sites and record whether Samsung ads appear before ever visiting samsung.com; if they appear at the same rate, the advertisement observation does not support the causal link. A second check is to dump the actual third-party cookie identifiers and real-time bidding requests during the globo.com visit and see whether they contain identifiers previously observed during the samsung.com visit.","tokens_in":5769,"feed_emoji":"📡","tokens_out":11491,"duration_ms":96855,"temperature":0.7,"pith_summary":"This paper attempts to supply the concrete, empirical evidence that discussions of surveillance capitalism often lack: it captures the actual network traffic produced by ordinary web use. The authors set up a man-in-the-middle proxy on a test phone, visit samsung.com, and observe outgoing requests to many third-party services, including Facebook, Twitter, TikTok, Pinterest, Bing, Google ad services, Criteo, Taboola, Outbrain, AppNexus, and Intentiq. They then visit the news site globo.com and report Samsung advertisements appearing there, which they read as targeted advertising connected to the earlier visit. They also capture the traffic generated by a Google search and list the analytics and optimization endpoints contacted. The point of the study is to make the hidden data transfers of the advertising ecosystem visible and to argue that consent and transparency are missing.","feed_headline":"A single web visit feeds a dozen ad trackers","feed_subtitle":"Network-capture study finds one visit to samsung.com can lead to Samsung ads on a news site.","key_machinery":"The load-bearing mechanism is the man-in-the-middle proxy: the test device trusts a certificate installed by the researchers, so all HTTP and HTTPS traffic passes through the proxy and can be read in plaintext. This makes every third-party request triggered by a page visit observable, including connections to pixel servers, analytics beacons, ad exchanges, and identity-sync services. The other half of the machinery is the advertising and tracking stack embedded in commercial websites: third-party cookies, tracking pixels, and JavaScript that report user activity to external domains, exactly the requests the proxy records.","core_discovery":"The central observation is that a single, ordinary web action produces a burst of machine-to-machine data transfers to companies the user never contacted. Visiting samsung.com triggers connections to external domains that host tracking pixels, retargeting services, content recommendation engines, programmatic ad exchanges, and identity-resolution services; the packet payloads reference samsung.com, showing that the visit itself is being reported. In the second case, after that visit, a subsequent navigation to globo.com shows Samsung advertisements, which the authors take as evidence that the first visit fed the ad-targeting chain. A separate search for 'Paris 6 Hotels' produces traffic to Google analytics and optimization endpoints. The paper presents these traces as direct evidence of the data-harvesting mechanisms behind surveillance capitalism.","pith_inferences":["A direct test of the paper's causal reading would be to repeat the globo.com observation with a clean device that never visits samsung.com; if Samsung ads appear with similar frequency, the ads were not caused by the visit.","The same proxy method could be applied to other e-commerce sites; if the pattern generalizes, then the 'one visit, many trackers' structure is a property of the advertising ecosystem rather than of one retailer.","The microphone anecdote in the conclusion is not supported by data in this study, but it suggests a controlled experiment that compares ad delivery with the microphone blocked versus enabled to see whether ambient speech changes the ad stream."],"forward_implications":["If a single visit to samsung.com produces requests to a dozen external services, then ordinary browsing routinely distributes one user's activity across many companies the user never chose, making per-company consent in practice impossible to grant.","The globo.com Samsung-ad observation implies that recent browsing history can change the advertising a user sees on unrelated sites, so the ad environment is not neutral content but a personalized response to prior behavior.","Because HTTPS traffic is decrypted and logged at the proxy, the study demonstrates that encryption alone does not hide web activity from the parties that control the user's device or network path, a relevant fact for privacy engineering.","The search-trace results show that even a plain Google query contacts dedicated analytics and optimization endpoints, extending the tracking picture beyond third-party cookies to first-party telemetry."],"supporting_citations":[{"why":"Supplies the proxy tool whose certificate-based interception makes the HTTP and HTTPS packet captures possible.","marker":"[Mitmproxy 2024]"},{"why":"Defines the concept of surveillance capitalism that the study operationalizes through network captures.","marker":"[Zuboff 2019]"},{"why":"Provides the revenue statistic that frames the economic scale of the targeted-advertising model behind the observed data flows.","marker":"[Zuboff 2023]"},{"why":"Supplies the account of harms from online behavioral advertising that motivates why the observed tracking matters.","marker":"[Wu et al. 2023]"}],"fun_headline_variants":["One site visit triggers a dozen ad trackers","Visit a site, get tracked by firms you never met","Data from a single visit sells your attention","One web action feeds an invisible ad machine","Trackers follow you from Samsung to Globo"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's strongest evidence, Samsung ads appearing on globo.com after visiting samsung.com, assumes the ads were caused by that single visit, with no control visit, no request timeline, and no filtering of background traffic on the test device to rule out other explanations.","fun_headline_variants_meta":{"raw":{"variants":["One site visit triggers a dozen ad trackers","Visit a site, get tracked by firms you never met","Data from a single visit sells your attention","One web action feeds an invisible ad machine","Trackers follow you from Samsung to Globo"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2623,"prompt_tokens":760,"completion_tokens":1863,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":376,"completion_tokens_details":{"reasoning_tokens":1792}},"tokens_in":376,"tokens_out":1863,"duration_ms":13825,"temperature":1.0,"reasoning_tokens":1792,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:07:50.515024+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a clean phone and clean network, visit a set of neutral news sites and record whether Samsung ads appear before ever visiting samsung.com; if they appear at the same rate, the advertisement observation does not support the causal link. A second check is to dump the actual third-party cookie identifiers and real-time bidding requests during the globo.com visit and see whether they contain identifiers previously observed during the samsung.com visit.","supporting_citations":[{"cited_title":"How Mitmproxy works","cited_arxiv_id":null,"evidence_quote":"Supplies the proxy tool whose certificate-based interception makes the HTTP and HTTPS packet captures possible."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the concept of surveillance capitalism that the study operationalizes through network captures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the revenue statistic that frames the economic scale of the targeted-advertising model behind the observed data flows."}],"review_version":1}