{"id":"4d03a3d5-3795-464e-aaa2-02cd81a06775","arxiv_id":"2608.07069","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Across 2,208 AI queries, 85.6% of 4,776 food and drink venues were never recommended; entry into answers tracks documentation, while star rating only affects rank among recommended venues.","lead":"This study built a complete list of every restaurant, cafe, and bar in two Bali districts and measured how often four AI assistants recommended each one. It found that over 85% of venues were never recommended by any system, and that getting into an answer depends on documentation, not star rating.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"M1's unweighted case-control analysis samples venues by ever-mentioned status, not individual trials, so Prentice-Pyke does not justify the odds ratios; full-cohort re-estimation is required before the two-margin claim is secure.","rationale":"I read the paper as making two separable claims: a population-level invisibility rate and a factor-based account of entry versus rank. The invisibility rate is well supported: the census-frame floors, the adversarial probe, and the validation estimates give it real strength, and I would not object to it. The factor account is where the load-bearing weakness sits. The case-control design is described transparently, but the statistical justification is a misapplication of Prentice-Pyke to cluster-level sampling. Since the full cohort is available, this is directly checkable; the paper should be required to do it before the two-margin claim is published. This is an internal-validity concern, distinct from the reader's external-validity caveat about API proxies and query mix, so I disagree with the reader's weakest_assumption identification. If the full-cohort re-estimation reproduces the odds ratios, I would accept the factor claims conditional only on the minor clarifications the reader already noted; if it does not, the main contribution would need substantial revision.","tokens_in":20800,"tokens_out":9476,"duration_ms":96871,"concrete_test":"Re-estimate M1 on all 4,776 census venues (or, equivalently, apply inverse-probability weights for venue selection: weight 1 for case venues, and for controls the inverse of the stratum-specific sampling fraction), with the same covariates, fixed effects, and clustering. Compare the four headline odds ratios (website, review volume, rating, Foursquare presence). If any odds ratio moves by more than about 15% or crosses 1.0, the case-control sampling design is biasing the entry-margin conclusions and the two-margin claim requires revision accordingly.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 4.4 defines the analysis frame as all 749 venues with any matched wave mention plus 620 never-mentioned controls, and Section 4.5 (M1) fits a binomial GLM over venue x persona x engine trials. This is cluster-level outcome-dependent sampling: selection depends on the venue-level aggregate of the outcome, not on each trial's outcome. Prentice and Pyke (1979), cited to justify unweighted estimation, applies to individual-level case-control sampling; it does not cover sampling on a cluster sum. Because the predictors of interest (rating, website, review volume, Foursquare presence) are venue-level, the selection probability is a nonlinear function of those predictors through the cluster outcome, so unweighted logistic slopes can be biased in either direction. The paper's own controls and sensitivity analyses (M2, M3, M5) do not repair this: GEE and bootstrap address correlation, not selection. The data for a full-cohort fit already exist; the restriction to 1,275 venues is a design choice, not a data limitation. This matters because the paper's headline 'documentation admits, rating ranks' dissociation rests on M1 odds ratios (website 1.92, rating 0.89, Foursquare 0.84); if those shift under a correctly specified full-cohort or weighted analysis, the central factor story is not established. The invisibility floor (85.6%) is independent of this issue and survives.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a census-denominated audit of AI venue recommendation. The authors enumerate 4,776 Google-Places-listed cafés, restaurants, and bars in two Bali submarkets (Canggu and Ubud) and query four production AI systems (ChatGPT, Claude, Gemini, Perplexity) through their search-grounded APIs with 96 persona-conditioned queries over seven days, yielding 2,208 runs and 12,439 valid mentions, under a pre-registered protocol with a frozen analysis snapshot. Five findings are reported: (1) at least 85.6% of venues were never recommended by any system, and 72.6% among venues with fifty or more ratings; (2) entry into answers is associated with documentation signals (own website OR 1.92, review volume OR 1.64, price listed OR 1.54, web mentions OR 1.44) while star rating is null at entry (OR 0.89) yet predicts first position within answers (OR 1.17); (3) Foursquare presence is null at both margins; (4) fabrication is rare (0.08% of mentions) while 93 recommendations went to permanently closed venues; (5) cross-system top-20 overlap is low (Jaccard 0.33–0.54) and a two-week test–retest shows run-to-run churn without temporal drift. The paper also reports two methodological lessons: a matching remediation reversed a factor's apparent effect, and the hours-listed covariate encoded collection provenance. The protocol, instrument, code, and derived data are released, and all tables and figures are claimed to be byte-reproducible.","tokens_in":21083,"tokens_out":29990,"duration_ms":253998,"significance":"If the central claims survive, this is an important and unusually rigorous contribution to the algorithm-audit literature. The census denominator converts 'never recommended' from a relative-prominence statement into a population rate, a structural advance over catalogue-based audits. The two-margin dissociation (documentation admits, rating ranks) is specific, falsifiable, and offers a concrete reconciliation of the conjoint and observational audit traditions. The instrument-validation discipline is exemplary: pre-registration, a frozen analysis snapshot, double-annotated extraction with adjudication, an audited entity matcher whose revision is disclosed to have reversed a coefficient, and a covariate–provenance artifact reported in full. The byte-reproducible pipeline and the publication of null results (Foursquare; rating at entry) that run against the funder's commercial interest strengthen credibility. The principal caveats — English tourist/nomad queries, two Balinese submarkets, search-grounded APIs as proxies for consumer apps, and an operationally defined Google-Places frame — are acknowledged in Section 7.","major_comments":[{"comment":"The case-control frame in §4.4 samples at the venue level (all 749 ever-mentioned venues plus 620 never-mentioned controls), but M1 in §4.5 is a trial-level binomial GLM over venue × persona × engine cells. The cited Prentice–Pyke (1979) consistency result covers logistic regression under outcome-based sampling of individual units; it does not cover sampling on a cluster-level aggregate of the outcome. Because a venue's sampling probability (ever mentioned) is a nonlinear function of its trial-level success probabilities and hence of the venue-level predictors, the unweighted trial-level likelihood is misspecified with respect to the sampling design, and the slope estimates in Table 1 (website 1.92, rating 0.89, Foursquare 0.84) can be biased in either direction. The robustness analyses (M2/M3/M5) address correlation and subsets, not selection. Since trial outcomes for all census venues are recoverable from the collected corpus (never-mentioned venues have all-zero outcomes by construction), a full-cohort fit over all 4,776 venues is feasible and should be reported as the primary analysis; alternatively, weight venues by the inverse of their selection probability and show that the coefficients are insensitive. This is load-bearing because the paper's headline 'documentation admits, rating ranks' dissociation rests on the M1 odds ratios.","section":"§4.4–4.5, Table 1"},{"comment":"The pre-registered extraction gate was ≥95% accuracy, with the metric level left unspecified. After remediation, strict run-level agreement is 91.5%, below that nominal gate, whereas mention-level precision (97.6%) and recall (99.1%) clear it. The interpretation under which the gate is passed was not fixed in advance, so it needs an argument rather than an assertion: either show that the run-level disagreements involve extractions irrelevant to the outcome definitions, or treat the gate as not met and provide a sensitivity analysis restricted to runs with perfect extraction agreement. The disclosure of all three numbers is commendable, but the confirmatory status of the study depends on the gate being met under a pre-committed interpretation.","section":"§4.1"},{"comment":"The selection-artifact caveat is applied selectively in the interpretation of M7. Section 5.4 declines to interpret the negative Foursquare coefficient because 'within-set coefficients of entry-relevant variables can carry selection artifacts,' yet presents the positive rating coefficient (OR 1.17) as evidence that 'rating ranks.' The same conditioning logic applies to every M7 coefficient, since the choice sets are generated by an entry process that appears to depend on the documentation variables (and possibly on rating, if the entry-margin null is revised). Please apply the caveat consistently and demonstrate that the rating-rank finding is not an artifact of differential selection into choice sets, for example by checking the stability of M7 coefficients across choice-set sizes or by estimating a joint entry/rank specification.","section":"§5.4, Table 3"}],"minor_comments":[{"comment":"The text reports the star-rating p-value in M7 as p = .0002, while Table 3 reports p < .001; align the two statements.","section":"§5.4 vs Table 3"},{"comment":"Please clarify the exclusion of 94 venues from the case-control set (1,369 → 1,275): how many were excluded for operational status versus missing rating, and whether exclusion is associated with case/control status or with the outcome.","section":"§4.4"},{"comment":"The Foursquare ladder (M4) is fitted on only 480 venues with three non-significant terms and uncorrected p-values; report the minimum detectable effect or a power statement so the reader can gauge how informative the null actually is.","section":"§5.3, Table 2"},{"comment":"Make explicit that the invisibility rates are conditional on the 96-query instrument and on the search-grounded API access route; the current 'at least 85.6%' phrasing is easily misread as a claim about all user queries or about what consumer apps display.","section":"§5.2"},{"comment":"The Gemini citation domains are recovered from citation titles, with unrecoverable citations excluded; state whether the reported top-domain ordering is robust to alternative recovery assumptions.","section":"§5.7, Figure 9"},{"comment":"Consider qualifying 'census' as 'Google-Places-listed' in the title or abstract; the operational frame in Section 3.2 is clear, but the unqualified phrase overstates the frame at a glance.","section":"Title/Abstract"}],"recommendation":"major_revision","confidential_remarks":"The funding disclosure is exemplary: the study is funded by Norly, a vendor of AI-visibility and review-management tools, and the paper pre-registered its protocol, froze its analysis snapshot, and reports null results (Foursquare; rating at entry) that conflict with commercially convenient narratives. The editor may wish to verify the pre-registration timestamp and the byte-reproducibility claim, as the confirmatory status of the study depends on them. The related-work section relies heavily on 2025–2026 arXiv preprints (Baig et al.; Jack et al. 2026a,b; Zatuchin; Martinez; Kumar; Iannelli and Ai) for load-bearing comparative claims; I have not verified these citations and recommend a quick check of their availability and content. The manuscript fits the journal's IR evaluation scope well as a methodological contribution to algorithm auditing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. The census-denominated audit is the real thing: a complete enumeration of 4,776 venues against which the authors run 2,208 search-grounded queries across four systems. The headline invisibility floor — 85.6% never recommended, 72.6% among venues with ≥50 ratings — is well-supported and survives every completeness critique because missing venues only push the rate up. The two-margin dissociation (documentation admits, rating ranks) is an elegant and important claim, and the paper earns credit for the Foursquare null, the staleness-not-fabrication finding, and the careful disclosure of the matching-error reversal and the hours-listed artifact. Pre-registration and byte-reproducible scripts are in place. This is the most carefully executed audit I've seen in this space.\n\nBut there's a structural problem with the entry-margin model (M1) that the paper does not acknowledge. The analysis frame is all 749 venues with any matched mention plus 620 never-mentioned controls, and then the binomial GLM is fit over venue × persona × engine trials. That is outcome-dependent sampling at the cluster level: selection depends on the venue-level aggregate (ever mentioned), not on each trial's outcome. Prentice and Pyke (1979), cited to justify unweighted estimation, applies to individual-level case-control sampling. It does not cover this design, and the venue-level predictors (rating, website, review volume) are exactly the variables that determine the cluster sum, so the unweighted logistic slopes can be biased in either direction. Cluster-robust SEs and GEE only fix correlation, not selection. The full-cohort data exist — they ran 1.4 million exposure opportunities — so restricting to 1,275 venues is a choice, not a constraint. The 85.6% floor is independent of this and survives, and the rank-1 margin (M7, conditional logit within runs) is not affected. But the 'documentation admits, rating ranks' story rests on M1. That needs a full-cohort or properly weighted refit before we know whether it's real. I'd also like to see the extraction gate issue cleaned up: mention-level metrics clear the 95% bar, but strict run-level agreement is 91.5%, and the paper declares the gate cleared without quite saying why the run-level metric isn't the binding one.\n\nThat said, this is a serious paper by any standard. The design, transparency, and measured nulls are exactly what the audit literature needs. I'd send it to referees, but with a mandate to demand the full-cohort re-estimation. If the dissociation survives, it's a major result. If it doesn't, the invisibility floor and the stability findings are still worth publishing.","headline":"A genuinely new census-denominated audit with a solid invisibility floor, but the two-margin factor story rests on a case-control design that Prentice-Pyke does not justify; the entry-margin estimates need a full-cohort refit before the dissociation is secure.","tokens_in":21586,"tokens_out":2783,"would_cite":true,"duration_ms":24626,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 2,208 queries to four AI assistants, at least 85.6% of the 4,776 restaurants, cafes, and bars in two Balinese markets were never recommended; documentation decides entry, star rating only decides rank.","keywords":["AI recommendation audit","census-denominated measurement","venue invisibility rate","two-margin visibility","local food-and-drink discovery","generative engine optimization","entity matching validation","search-grounded assistants"],"falsifier":"Re-run the identical 96-query instrument through the four vendors' consumer-facing configurations, or through server settings matching production consumer apps, and compare the invisibility rate and factor estimates against the same 4,776-venue census; if the never-recommended share falls materially below 85.6%, or if star rating becomes a significant positive predictor of entry, the two-margin structure (documentation admits, rating ranks) is an artifact of the API access route rather than a property of the assistants users interact with.","tokens_in":20569,"feed_emoji":"🍽️","tokens_out":15146,"duration_ms":109200,"temperature":0.7,"pith_summary":"This paper sets out to measure what AI assistants actually surface when a traveler asks where to eat or drink, using a complete enumeration of an entire market as the denominator. It claims that invisibility is the norm — at least 85.6% of the 4,776 cafes, restaurants, and bars across two Balinese submarkets were never recommended by any of four production systems in 2,208 runs — and that visibility is governed by two margins with different drivers. Entry into an answer tracks documentation: review volume, an own website, listed price information, and third-party web mentions all raise the odds, while star rating has no detectable effect at this margin. Once a venue is in an answer, the pattern reverses: rating significantly predicts being named first. A sympathetic reader would care because the factor venues invest most in signaling — their star rating — appears to do nothing for AI discovery, while infrastructure like having a website governs who is surfaced at all.","feed_headline":"85.6% of venues never recommended by any AI assistant","feed_subtitle":"A census of 4,776 Bali venues shows documentation—not star rating—decides which venues AI answers mention.","key_machinery":"The load-bearing object is the census itself: a complete enumeration of 4,776 Google-listed food-and-drink venues across two bounded geographic polygons, built by adaptive grid subdivision of the Places Nearby Search API (recursively splitting saturated 20-result cells down to a 130-meter floor) plus a resolution pass that folded 70 mention-probed venues of atypical listing type into the frame. Because every mention is resolved against this full registry, 'never recommended' becomes a measurable population rate rather than a relative prominence judgment. The second piece of machinery is the two-margin outcome decomposition: the entry margin models whether a venue is recommended at all (a binomial regression model over venue x persona x engine exposure opportunities in a case-control frame, cluster-robust by venue, with Benjamini-Hochberg multiple-comparison correction), and the rank margin models which recommended venue is named first (a conditional logit over within-run choice sets). The dissociation between the two margins — documentation admits, rating ranks — is the identity that carries the argument, and the paper reads it through the retrieval-then-generation architecture of the audited systems.","core_discovery":"On the paper's own terms, the central discovery is a census-denominated measurement of AI venue visibility: across 2,208 search-grounded responses from ChatGPT, Claude, Gemini, and Perplexity to 96 persona-conditioned queries about two Bali submarkets, 4,087 of the 4,776 enumerated venues — 85.6%, a floor — were never recommended by any system; even among venues with fifty or more ratings the never-recommended share is 72.6%, and visibility among the surfaced is long-tailed rather than winner-take-all (the leading venue holds 1.9% of recommendations). The paper's explanatory claim is a two-margin dissociation: admission into an answer is associated with documentation — review volume (1.64x), an own website (1.92x), listed price (1.54x), and web mentions (1.44x) as odds multipliers — while star rating is null at entry (0.89x), but within-answer ordering reverses the pattern, with rating significantly predicting first position (1.17x). Presence in the Foursquare open POI dataset shows no positive effect at either margin. The systems almost never fabricate venue names (0.08% of mentions) yet recommended permanently closed venues 93 times, making staleness rather than hallucination the practical failure mode, and a two-week test-retest shows answer churn is sampling stochasticity, not temporal drift.","pith_inferences":["If the two-margin structure is a property of retrieval-then-generation architecture rather than of these four systems, the same dissociation — documentation admits, rating ranks — should appear in other verticals such as hotels and services, where conjoint experiments already show rating dominance within constructed choice sets; a census-denominated replication in a structurally different market w","The Foursquare null is the first direct test of a widely assumed visibility lever, but the observational design cannot settle whether the dataset simply goes unread by retrieval or is redundant with web presence; a clean follow-up would register randomly chosen thin-documentation venues in open POI datasets and measure whether visibility moves.","Because the invisibility rate only rises when the frame broadens (above 92% under the broadest defensible estimate), markets with thinner Google coverage than Bali plausibly show even higher invisibility, making the documented associations a lower bound on the discoverability penalty faced by small independent venues.","The paper's demonstration that a matching-error remediation reversed one factor's apparent effect from positive to null raises a testable suspicion that earlier brand-level audits without validated matching have unstable factor conclusions; re-running those designs with census denominators and audited matchers would be the direct check."],"forward_implications":["A 4.9-star venue with thin documentation is, to these systems, indistinguishable from an absent one: improving a rating without growing the documentation trail should not be expected to change whether a venue is surfaced at all.","Single-shot, single-engine visibility checks measure sampling noise; meaningful measurement requires repetition across engines and phrasings.","The clearest correctable harm is staleness: 93 recommendations of permanently closed venues show that closure signals propagate more slowly than reputation signals in the sources assistants retrieve.","Because top-20 agreement between engines is only 0.33-0.54 (just eight venues appear in all four engines' top-20 lists), optimizing visibility for one assistant is not optimizing for AI in general.","For independent venues the action hierarchy is to be documented before being excellent: an own website, complete platform profiles, review volume, and third-party mentions precede any payoff from rating."],"supporting_citations":[{"why":"The closest design precedent: a pre-registered conjoint audit of twelve LLMs choosing among hotels, whose within-choice-set rating effect the paper reconciles with its entry-margin null.","marker":"[Baig et al., 2026]"},{"why":"The 37,000-query, 533-brand catalog audit documenting long-tail 'catastrophic invisibility'; its lack of a market denominator is the gap the census fills.","marker":"[Jack et al., 2026b]"},{"why":"Causal evidence that a one-star Yelp rating increase raises independent-restaurant revenue by 5-9%, establishing why representation on discovery intermediaries carries direct economic stakes.","marker":"[Luca, 2016]"},{"why":"Shows an AI recommendation lifts same-brand searches by 4.3 percentage points among previously unengaged consumers, extending the revenue chain into the AI channel.","marker":"[Iannelli and Ai, 2026]"},{"why":"The GEO survey whose reproducibility bar — repeated measurement, paraphrase controls, human-validated extraction — the paper adopts as its own standard.","marker":"[Martinez, 2026]"},{"why":"The statistical result that logistic slope coefficients are consistent under outcome-based case-control sampling, justifying the paper's case-control estimation frame.","marker":"[Prentice and Pyke, 1979]"},{"why":"The critique that algorithm audits are weakened by undisclosed design choices; it motivates the paper's pre-registration and instrument-validation discipline.","marker":"[Bouchaud and Ramaciotti, 2024]"},{"why":"The multi-industry brand audit whose low cross-model agreement (41.6% top-brand) the paper's cross-engine top-20 Jaccard results extend to venues.","marker":"[˙Zatuchin, 2026]"},{"why":"The methodological foundation that names noninvasive sock-puppet audits as the legitimate way to study opaque ranking platforms.","marker":"[Sandvig et al., 2014]"}],"fun_headline_variants":["85.6% of venues never recommended by any AI system","AI food recs: documentation beats star rating for visibility","AI venue audit: staleness, not hallucination, is the failure","Only 14.4% of Bali venues ever appear in AI answers","Rating predicts AI rank, but not whether venue is listed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole measurement rests on two proxy choices: search-grounded production APIs standing in for the consumer assistants, and the Google Places listing set standing in for the market, so if provider-side serving configurations differ from the consumer apps, or the English tourist-and-nomad query mix under-represents real demand, the measured invisibility rates and factor associations may not match what users actually see.","fun_headline_variants_meta":{"raw":{"variants":["85.6% of venues never recommended by any AI system","AI food recs: documentation beats star rating for visibility","AI venue audit: staleness, not hallucination, is the failure","Only 14.4% of Bali venues ever appear in AI answers","Rating predicts AI rank, but not whether venue is listed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000707,"raw_usage":{"total_tokens":3320,"prompt_tokens":1211,"completion_tokens":2109,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":827,"completion_tokens_details":{"reasoning_tokens":2021}},"tokens_in":827,"tokens_out":2109,"duration_ms":14517,"temperature":1.0,"reasoning_tokens":2021,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:22:41.683782+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the identical 96-query instrument through the four vendors' consumer-facing configurations, or through server settings matching production consumer apps, and compare the invisibility rate and factor estimates against the same 4,776-venue census; if the never-recommended share falls materially below 85.6%, or if star rating becomes a significant positive predictor of entry, the two-margin structure (documentation admits, rating ranks) is an artifact of the API access route rather than a property of the assistants users interact with.","supporting_citations":[],"review_version":1}