NTH

Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census

AuthorsVladimir Pitenin

August 21, 2026 3 min read
Watch on YouTube
The one-line take

A large-scale audit finds that AI assistants systematically overlook most restaurants, cafes, and bars, with documentation and stale business data shaping which venues become visible.

Key results

4,776
Census venues

Food-and-drink venues enumerated across Canggu and Ubud.

2,208
Search-grounded runs

Responses collected from ChatGPT, Claude, Gemini, and Perplexity.

85.6%
Overall invisibility

Share of census venues never recommended by any audited system.

1.92
Website entry odds ratio

Association between having an own website and entering an AI answer.

What the paper found

This study introduces a census-denominated audit of local AI recommendation, comparing a complete registry of 4,776 cafes, restaurants, and bars in Canggu and Ubud, Bali, with 2,208 search-grounded responses to 96 persona-conditioned queries from OpenAI’s ChatGPT, Anthropic’s Claude, Google’s Gemini, and Perplexity. The systems produced 12,439 valid venue mentions, yet 85.6% of census venues were never recommended; even among established venues with at least 50 ratings, invisibility remained 72.6%. A preregistered binomial GLM reveals a two-margin process: entry into an answer is associated with documentation and discoverability—an own website has odds ratio 1.92, while review volume, listed prices, and third-party web mentions also increase visibility—but star rating is not significant at entry. Conditional ranking analysis reverses this pattern: rating predicts which already-retrieved venue appears first. Foursquare presence shows no positive visibility effect after controls, challenging a common optimization assumption. Recommendations are unstable across repeated queries and disagree across systems, but a two-week test–retest indicates sampling stochasticity rather than temporal drift. The practical failure mode is staleness, not hallucination: permanently closed venues were recommended 93 times, while likely fabrication accounted for only 0.08% of valid mentions. The study’s main methodological contribution is demonstrating that complete-market denominators, entity-matching audits, and repeated multi-system sampling are necessary to measure AI invisibility reliably.

Original abstract

AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an own website (OR 1.92), listed price information (OR 1.54), and third-party web mentions (OR 1.44) -- while star rating is null at this margin (OR 0.89). Rank within answers reverses the pattern: among recommended venues, rating significantly predicts first position (OR 1.17). Presence in an open POI dataset (Foursquare), a folk-theorized visibility factor, shows no positive effect at either margin. Outright fabrication is rare (0.08% of mentions), but systems recommended permanently closed venues 93 times -- staleness, not hallucination, is the practical failure mode. Cross-system agreement is low (top-20 Jaccard 0.33-0.54). A two-week test-retest shows cross-period answer similarity comparable to same-day rerun similarity: the churn is sampling stochasticity, not temporal drift. We release our protocol, registry construction method, and derived data.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis