The short answer

A hotel’s average guest rating is the single strongest reputation signal an AI travel planner responds to — and replying to reviews, the tactic most GEO vendors lead with, has no measurable effect on whether the machine picks you. That is the headline of the first pre-registered causal audit of LLM hotel selection, published in June 2026 by Baig, Gillani and Ali (arXiv, June 15, 2026). The researchers ran a randomized choice-based conjoint across twelve models — GPT-4o-mini, three Gemini versions, four Claude models, and four open-weight systems — presenting each assistant with sets of five hotel cards whose rating, review volume, review recency, management response, chain affiliation, price, eco-certification, and list position were independently randomized. Across more than 60,000 model calls, a top guest rating raised the probability of being the recommended hotel by 31.6 percentage points; a high price lowered it by 30.0 points; eco-certification added 11.6; review volume added 8.3; and a visible management response added 0.1 points — statistically indistinguishable from zero (arXiv, June 2026).

For hotels, tour operators and DMCs, this is the first time the folklore of “AI optimization” has been priced in causal units instead of anecdotes. Your reviews are no longer just persuasion for human readers — they are structured input to the machine that increasingly decides which property gets surfaced at all. Roughly 40% of US travelers used generative AI to plan a trip during 2025, up eleven points year over year, while the share starting trip planning at a conventional search engine fell from about half to roughly a third (Phocuswright, 2025). If the assistant is the new gatekeeper, reputation data is the currency it reads. A travel AI audit now needs a review-data component, not just a website one.

How you causally audit what an AI “values”

The method matters, because it is what separates this study from the “we asked ChatGPT 50 times” posts that circulate in hotel marketing groups.

The researchers used a choice-based conjoint design — the same technique long used to price brand attributes in consumer research. Each model received sets of five synthetic hotel profiles and one question: which hotel would you recommend? Because every attribute was independently randomized across thousands of trials (3,024 choice sets per model), the average marginal component effect of each signal — the causal change in recommendation probability when only that signal changes — falls out of the data. Nine prompt paraphrases, three traveler personas, and decoding temperatures from 0.7 to 0.0 were varied to stress-test robustness. The design, hypotheses, and analysis plan were specified and cryptographically hashed before any confirmatory data were collected (arXiv, June 2026).

That last detail is why the numbers deserve attention: this is a pre-registered audit, not a fishing expedition. It is also scoped deliberately to the selection stage — the assistant choosing among already-retrieved candidates — which is precisely where a property’s manageable reputation signals compete.

What the audit found, signal by signal

Reputation signalEffect on P(recommended)Practical read
Top guest rating (vs. low)+31.6 ppThe dominant lever. Defend your average first.
High price−30.0 ppAI assistants punish headline price at selection time
Eco-certification (Green Key)+11.6 ppOver-weighted vs. human behavior — exploit it
Review volume (high vs. low)+8.3 ppSecondary but real: more reviews, more machine confidence
Chain affiliation−1.8 ppNo brand bonus at the selection stage; slight penalty
Visible management response+0.1 ppStatistically zero. Reply for humans, not for the model
First list position≈ $12/night equivalentPure ordering artifact — position moves the machine

Sources: arXiv 2606.16344, pooled panel AMCE estimates, June 2026.

Three findings deserve unpacking.

1. The machine reproduces human “valence primacy” — at machine scale. Decades of electronic-word-of-mouth research on human bookers found review valence elasticity of 0.888 against a volume elasticity of just 0.055 — rating an order of magnitude more powerful than review count (Yang et al., 2018 meta-analysis, cited in arXiv, June 2026). The LLM panel landed in the same place: 31.6 points for rating versus 8.3 for volume. If you are choosing between pushing more review volume and defending your average, the audit says defend the average. A 4.6 from 240 recent reviews beats a 4.1 from 900 stale ones.

2. Eco-certification is over-weighted by machines. Human travelers treat green credentials as a modest tiebreaker. The pooled LLM panel elevated eco-certification to the third-ranked signal at +11.6 points (arXiv, June 2026). If your property holds a Green Key, GreenLeader, or regional eco-label and it lives only in a PDF on page four of your site, you are leaving a measurable, machine-legible advantage unused. Put it in your schema markup, your Google Business Profile, and the factual blocks AI systems actually parse.

3. Management replies do not move the needle — but the models don’t admit what does. The response signal came in at +0.1 points, statistically equivalent to zero, even though responding to reviews measurably improves human-facing reputation over time (Proserpio & Zervas, 2017). Meanwhile, list position — a content-free artifact of ordering — shifted recommendations causally, worth about $12 per night in price-equivalent terms, with one model showing an effect an order of magnitude larger. And the models’ stated reasons only loosely matched their actual behavior (Spearman +0.59 to +0.85): they acted on position and review volume without naming them, and over-cited brand, which barely matters (arXiv, June 2026). In plain terms: you cannot ask an AI why it didn’t recommend you and trust the answer. You have to measure what actually gets said — which is what a structured AI visibility program for hotels exists to do.

The other half of the story: AI summaries can bury your worst reviews

The audit covers what moves an AI’s choice. A July 2, 2026 investigation by UK consumer group Which? covers the other side: what AI summaries of your reviews choose to say.

Which? found that Tripadvisor’s AI-generated hotel summaries and its Ollie chatbot downplayed or omitted serious complaints — food poisoning, hygiene failures, even reports of sexual harassment (The Guardian, July 2, 2026). At one Cape Verde resort, the AI summary described cleanliness as “spotless” while underlying reviews contained 102 mentions of food poisoning, including 14 reviews between December 2025 and April 2026 from guests reporting serious illness. Another Dominican Republic property summarized as popular with “diverse restaurants” earning “rave reviews” had recent guests calling it “the worst place imaginable” (Which? via The Guardian and Euronews, July 2026).

Two implications for operators:

  • The flattening cuts both ways. Summarization compresses sentiment into themes — “friendly staff,” “dated bathrooms” — and the compression is lossy. Your review base is now the raw material for a summary you do not control and cannot edit. The only defense is a review profile whose dominant themes are unambiguous.
  • Monitor the summaries, not just the stars. Check what Tripadvisor’s AI summary, Google’s review highlights, and chatbot answers actually say about your property each month. Misrepresentation at the summary layer is invisible in your average score.

The review playbook the evidence supports

  1. Treat rating defense as your #1 GEO lever. Nothing else in the audit comes close to +31.6 points. Operational quality that holds your average — and recent reviews that reflect it — is the highest-yield AI visibility work you can do.
  2. Prioritize Google reviews first, then feed the long tail. 71% of travelers encounter Google review scores during booking research, versus 44% for TripAdvisor and 38% for Booking.com — a gap that has widened 12 points since 2023 (Sojern, 2025). Google’s review graph also feeds the entity data Gemini leans on. Then keep OTA and TripAdvisor volume growing: the audit prices volume at +8.3 points.
  3. Keep replying to reviews — for the humans. Replies shape future review composition and convert human readers even though the machine ignores them at selection time. Just stop paying anyone who claims reply programs are how you get recommended by ChatGPT.
  4. Make certifications machine-readable. Eco labels over-perform in LLM weighting. One afternoon of schema and listing updates can capture a signal worth more than chain affiliation.
  5. For tour operators and DMCs, the same logic applies with different platforms. Your Viator, GetYourGuide and TripAdvisor review profiles are the equivalent signal layer — and as we documented in the tour marketplace blind spot, most operators’ review data never reaches AI planners in a usable form at all.

And because persona weighting shifts — the budget family was most price-sensitive, the eco-conscious couple valued certification most, and the business traveler erased the chain penalty (arXiv, June 2026) — segment-level measurement beats one-size answers. Our earlier analysis of persona-driven AI recommendations covers how to build that view; a structured GEO agency engagement or a first AI visibility audit priced on our pricing page is where most properties start.

FAQ

Do management responses to reviews influence AI hotel recommendations? No. In the June 2026 conjoint audit across twelve LLMs, a visible management response changed recommendation probability by +0.1 percentage points — statistically zero (arXiv, June 2026). Replies still matter to human readers and to long-run review composition, so keep them; just don’t buy them as an AI-visibility tactic.

What matters more to AI travel planners: review rating or review volume? Rating, decisively: +31.6 percentage points versus +8.3 for volume in the pooled audit panel. This mirrors human behavior, where valence elasticity outweighs volume elasticity by more than 10:1 (Yang et al., 2018).

Which review platform should hotels focus on for AI visibility? Google first — 71% of travelers see Google review scores during research versus 44% for TripAdvisor (Sojern, 2025) — then the platforms AI assistants actually retrieve from for your segment: OTAs for hotels, Viator/GetYourGuide for tour operators.

Do AI summaries of reviews reflect what reviews actually say? Not reliably. Which?’s July 2026 investigation documented Tripadvisor AI summaries describing properties as “spotless” while underlying reviews contained over a hundred food-poisoning mentions (The Guardian, July 2, 2026). Audit the summaries, not just the scores.

How do I find out what AI actually says about my hotel? Structured measurement: run prompts across ChatGPT, Gemini and Perplexity for your property and compset, on a fixed cadence, and score mention share. Palmtree’s travel AI audit does exactly this — or see pricing to scope it for your property.

The evidence era of GEO has started. The hotels that win the recommendation layer will be the ones managing reputation as machine input — measured, causally understood, and audited monthly — rather than the ones buying folklore.