AI Showdown: Where You Sleep Changes Everything
WEEK 95 :: POST 4 :: THE JUDGE’S CHOICE
Directions Given To The A.I. This Week+
Instructions Given to each A.I. — Please provide 3 prompt variations that share this objective:
Each A.I. also received two static attachments: the blog post template (structure) and the authoring instructions (voice and standards). The text below is the week-specific assignment as sent — reflowed for the web; wording unchanged.
I'd like you to write this week's Ketelsen.ai post. Two files are attached: the blog post template (the structure to follow) and the authoring instructions (context, voice, and standards). Please read both before you begin, then produce the complete post in a single response.
This week's theme: "Where You Sleep Changes Everything" — Lodging and Neighborhood Intelligence.
This is Week 4 of an eight-week series on planning a vacation with AI. Week 1 established the reader's real constraints — a validated budget ceiling and a constraint profile. Week 2 chose the destination. Week 3 locked the flights, which fixed the dates and the arrival airport. This week they decide where they will actually wake up every morning.
Lodging is the decision travellers most often get wrong in a way they cannot undo. A flight that disappoints is four hours; a badly chosen neighbourhood is the whole trip. It is also the part of travel where the listing is written by someone whose interest is opposed to the reader's: the price shown is not the price paid, the photographs are chosen and cropped, the reviews are a mixture of the genuine, the incentivised, and the fabricated, and the words that sound most reassuring — cozy, vibrant, up-and-coming, steps from — are the ones most often doing concealment work.
The job this week is to give the reader a defensible lodging decision: this property, in this neighbourhood, at this true all-in cost, with these cancellation terms, chosen for reasons they can state.
The three prompts should help a reader work through:
- True total cost, not nightly rate. Resort fees, cleaning fees, service fees, occupancy taxes, parking, security deposits, mandatory "destination charges" — the drip-pricing pattern from Week 3's airfare work reappears here in a different costume, and it is usually worse. Two properties whose headline rates differ by 15% can invert once everything is added.
- Review forensics. How to read a review set rather than a review score: weighting recent reviews far more heavily than old ones, spotting clusters of suspiciously similar phrasing or timing, reading the three-star reviews where the honest detail lives, and noticing what a stream of positive reviews never mentions. A property whose complaints are all about the same thing is telling the reader something a 4.5 average conceals.
- Neighbourhood, decided by block and not by city. Walkability to what the reader actually plans to do, transit access at the hours they will use it, noise, and how the character of an area changes between a Tuesday afternoon and a Saturday night. The right question is never "is this a good area" but "is this a good area for this trip, these travellers, and these hours."
- Cancellation-policy risk as a priced feature. A non-refundable rate is a discount in exchange for accepting a risk. The reader should be able to say what that risk is worth to them rather than defaulting to the cheaper number.
- Decoding listing language. What cozy, charming, lively, convenient for transport, and partial view reliably mean, and which specific photograph or floor-plan absence should prompt a direct question to the host.
The output a reader should walk away with is a ranked shortlist they can act on: two or three properties scored on true cost, review credibility, location fit, and cancellation risk, with the reason each one ranks where it does.
A note on the strongest version of this week: at the advanced end, this is a lodging dossier matrix — candidate properties as rows; true all-in nightly cost, review-pattern credibility, location score against this trip's actual itinerary, and cancellation-policy risk as columns; with the reader's own weightings applied and the trade-offs made explicit. That structure is worth reaching for, and it is the natural successor to Week 3's fare decision framework.
A hard constraint, carried forward from Week 3 and just as binding here. AI models cannot see live listings, current rates, or today's availability, and their knowledge of specific properties is stale, thin, and — for anything below the level of a famous hotel — frequently invented. No prompt in this post may ask the AI to recommend a named property, quote a current rate, describe a specific listing it has not been shown, or assess a named hotel's present condition. A confidently hallucinated hotel recommendation is worse than a hallucinated airfare: the reader books it.
Neighbourhood character deserves particular care. A model will happily describe a district's safety or atmosphere from training data that is years stale and was never reliable — and this is the one place in the series where a wrong answer can put someone somewhere they should not be. Prompts should have the AI generate what to check and where to check it — the questions to ask, the hours to look at, the sources to consult — rather than deliver a verdict.
Design the prompts so the AI does what it is genuinely good at: structuring the comparison, exposing the fee layers, teaching the reader to read a review set, and naming what to verify. The reader supplies the listings; the AI supplies the judgement framework. Posts that blur that division should expect to be marked down on Practical Utility, exactly as in Week 3.
Series dependency chain, for the Metadata block: Week 4 consumes the confirmed routing and dates from Week 3, the destination from Week 2, and the budget ceiling from Week 1 — lodging is scored against what the flights left of the budget, and a nightly rate that breaks the ceiling is a signal to revisit the property tier or the neighbourhood, not to quietly raise the budget. Week 4 produces the booked lodging and its location, which Week 5's itinerary assumes as its starting point every morning: the itinerary is built outward from where the reader wakes up.
Because readers may arrive at this post without having read Weeks 1 to 3, the prompts should work for someone who knows their destination, dates, and rough remaining budget, while making clear they get far more from them with a real constraint profile, confirmed flights, and a live set of candidate listings in hand.
Three difficulty tiers as always — Beginner, Intermediate, Advanced — each a genuinely different approach to the same problem, not the same prompt at three lengths.
On examples: this is a consumer travel topic. The template lists tech startup / retail / freelance as suggested industry examples — those are marked MAY, and this week you should almost certainly adapt them. Families needing two bedrooms and a kitchen, couples choosing between a central hotel and a quieter rental, solo travellers weighing safety and walkability, older travellers for whom stairs and lift access decide everything, and anyone booking around a fixed-date event are the right contexts here. Choosing them over the suggested business examples is correct behaviour and will not be scored against you.
A note on supplied figures. Anything marked `[SUPPLIED — use as given]` above came from Ketelsen.ai's own research brief. Use it freely — you are not fabricating by repeating it, and you will not be marked down for leaving it uncited. Do not attach an invented source to it. (No supplied figures this week. Given the live-listing constraint above, this is a bad week to invent any — if you find yourself reaching for a typical resort fee or an average nightly rate, that is the signal to restructure the prompt so the reader supplies the real number instead.)
## BEFORE YOU SUBMIT — STRUCTURAL CHECK
(This block is identical every week. It exists because these specific items are the ones posts drop, and a dropped structural item costs compliance points for something that takes one minute to add.)
Your post is parsed by a script before any human reads it. Confirm all seven:
1. ☐ Response begins with `PLATFORM: <your name>` and `WEEK: 4` 2. ☐ `## Lead` present once, at the very top, before Variation 1 3. ☐ `## In one line` present in all three variations 4. ☐ `## What this prompt gives you` present in all three variations 5. ☐ `## The Prompt` present in all three variations, with the prompt in double quotes beneath it 6. ☐ `## Introductory Hook` and `## Current Use` present in all three variations (three of each — not one) 7. ☐ Every template heading written as `##`, none bolded instead; prompt breakdown is running text split on ` : `, with no `###` headings inside it
A complete post has 57 `##` headings. If your count is well short, a section is missing or was bolded instead of hashed.
One extra check this week: confirm no prompt asks the AI to name, rate, or price a specific property, or to pronounce on a neighbourhood's current safety. Those must be things the reader goes and verifies.
Scored across seven dimensions by this week's rotating judge. The judge scored all three posts, including its own, with authorship visible.
Claude takes Week 4 :: Vacations Series with 62 of 70, ahead of ChatGPT on 58.
Each dimension scored 1-10 by the judge. These are the judge's own scores, not measured data.
1. Prompt Quality & Creativity
Winner: ChatGPT
ChatGPT produced the most technically ambitious prompts in the set. Its Intermediate workflow does not merely score properties. It creates an evidence ledger, separates cash cost from refundable cash tied up, applies hard disqualifiers before ranking, adds confidence labels to scores, prices cancellation risk through scenarios, and runs a sensitivity check.
The Advanced prompt goes further by treating each property as a dossier rather than a listing. Its six-pass workflow preserves provenance, models liquidity, clusters review evidence, measures itinerary friction, applies uncertainty penalties, and tests the result under missing-data, cancellation, dominance, and regret scenarios.
Because ChatGPT authored this entry, I applied additional scrutiny rather than rewarding sophistication automatically. The prompt is undeniably overbuilt for many travelers, but nearly every added component has a legitimate decision function. The standout instruction is:
> “For the top two candidates, provide the exact evidence that could reverse their order.”
That converts the final unknowns into a value-of-information problem: investigate only what might change the decision.
Claude came close. Its best innovation is the explicit missing-data state:
> “unscored — need [the specific thing]”
That is a simple, practical defense against a model filling every matrix cell with a plausible-looking number.
Gemini’s Beginner and Intermediate prompts are serviceable, but they are mostly compact task instructions. Its Advanced prompt names four matrix columns and asks the model to “weight commute time and cancellation risk heavily,” without defining score anchors, missing-data behavior, evidence confidence, or how the weights should be chosen.
2. Content Depth & Accuracy
Winner: ChatGPT
ChatGPT offers the deepest explanation of the underlying prompting principles. It teaches readers why they should timestamp volatile facts, separate vetoes from preferences, distinguish liquidity from cost, qualify numerical scores with confidence, and avoid letting missing information behave like favorable evidence.
Its Advanced breakdown is especially strong on evidence discipline. It explains evidence laundering, transparent recency rules, uncertainty penalties, symbolic scenarios when probabilities are unavailable, and the difference between a known veto failure and an unresolved stress-test condition.
I again applied extra scrutiny because this is ChatGPT’s own post. The main weakness is not factual accuracy but proportionality: several explanations repeat adjacent ideas, and the full system risks teaching more decision-analysis machinery than an ordinary lodging choice requires. That cost appears under readability and tier differentiation rather than accuracy.
Claude’s teaching is excellent, but one conspicuous arithmetic error materially lowers its score. In a seven-night example, it describes a property charging $150 per night, a $250 cleaning fee, and a $95-per-night amenity fee, then says the true nightly cost becomes $186. Before tax, the actual total would be:
($150 × 7 + $250 + $95 × 7) ÷ 7 = $280.71 per night
The error appears inside a section specifically teaching readers to verify lodging arithmetic, which makes it more damaging than an incidental typo.
Gemini’s principal accuracy problem is more serious because it affects the recommended workflow. Its Advanced FAQ says the AI can use “general knowledge of the city” to estimate commute times and that it is “highly accurate for general logistics.”
That directly conflicts with the Week 4 brief, which says neighborhood and route judgments must come from current checks supplied by the reader, not the model’s stale city knowledge.
Gemini also says a model can identify phrasing clusters that “indicate a bot or paid review farm.” The safer standard used by the other posts is that such patterns justify verification but do not prove fabrication.
3. Template Compliance
Winner: Claude
The Week 4 structural check states:
> “Every template heading written as ##, none bolded instead; prompt breakdown is running text split on :, with no ### headings inside it.”
Claude follows that instruction. Its headings use plain ## syntax, and its prompt breakdowns use quoted fragments followed by : and running explanation.
ChatGPT includes the complete section structure, but all its headings are written in forms such as ## Lead and ## The Prompt, rather than the specified plain-heading form.
Gemini has the same bolded-heading problem and also inserts ### subheadings throughout its prompt breakdowns—the exact pattern the structural check prohibits.
No post was penalized for using traveler and household examples instead of technology or retail examples. The assignment explicitly says those industries were merely illustrative and that adapting them for a consumer travel topic is correct.
4. Practical Utility
Winner: ChatGPT
ChatGPT gives the reader the most operationally complete workflow. It asks for the final checkout subtotal, fee line items, payment timing, cancellation schedule, room configuration, route checks at relevant hours, recent reviews, host answers, evidence dates, and missing photographs or policy details. It then defines exactly how each input affects the result.
Its most useful practical controls include:
- Keeping refundable deposits outside cost but inside liquidity.
- Refusing to let a score rescue a property that fails a hard requirement.
- Converting stale or missing facts into source-and-time-specific verification tasks.
- Testing how the winner performs when unresolved facts turn out badly.
- Preserving screenshots, policies, messages, and timestamps before booking.
Those details make the output auditable and updateable instead of disposable.
Claude’s prompts are easier to run and more realistic for an ordinary reader. Its Intermediate prompt is particularly useful because it turns listing language, review patterns, and neighborhood uncertainty into concrete questions, photographs, measurements, sources, and hours to verify.
The arithmetic error prevents Claude from taking this category, but the actual prompt architecture remains highly usable.
Gemini’s Beginner prompt is immediately actionable, but its later tiers are under-specified. The Advanced prompt expects the AI to score commute time without requiring current route data, and its FAQ actively encourages reliance on general city knowledge. That is precisely the kind of confident shortcut the assignment was designed to prevent.
5. Engagement & Readability
Winner: Claude
Claude is the best-written article. It has a strong editorial voice without losing the reader inside the metaphor. Its opening establishes the stakes quickly:
> “A flight that disappoints costs you four hours; a badly chosen place to stay costs you the whole trip.”
It then gives each tier a clear job and explains why the reader—not the model—retains final judgment.
Claude also translates technical prompting principles into memorable, practical language. Examples include treating “not shown” as data, giving the AI somewhere honest to put uncertainty, and asking what would flip the result rather than accepting a weighted score as final.
ChatGPT is controlled and professional, but it reads more like a decision-analysis manual than an accessible business publication. The Intermediate and Advanced sections are dense enough that a reader may understand every component yet still postpone using the prompt.
Gemini is easy to scan, but its language frequently leans into promotional exaggeration: “bulletproof,” “drip-pricing camouflage,” “the truth hidden inside a 4.5-star average,” and “the only number that matters.”
The energy helps readability, but it also makes several claims sound more certain than their evidence warrants.
6. Citation Quality
Winner: Claude
Claude is the only post that cites strong external sources for its time-sensitive factual claims. Its Beginner section cites the FTC’s Rule on Unfair or Deceptive Fees, and its Intermediate section cites the FTC’s rule governing fake and deceptive reviews. Both citations are real and accurately dated. The FTC’s fees rule took effect on May 12, 2025 and covers advertised prices for short-term lodging. The consumer-review rule took effect on October 21, 2024 and addresses fake, false, incentivized, suppressed, and otherwise deceptive review practices.
Claude does make a few claims beyond what those citations establish—particularly that enforcement is uneven and that older listings lag—but those are thinly supported claims, not fabricated sourcing.
ChatGPT’s Beginner citation section refers generally to Ketelsen.ai’s authoring instructions and weekly template rather than giving usable publication details, while its later variations write NOT APPLICABLE.
That is honest but weak. Most of its content is methodological and based on reader-supplied evidence, so the lack of external sourcing is not disqualifying. Still, a source on decision matrices, uncertainty communication, or review analysis would have made the teaching more authoritative.
Gemini marks every citation section NOT APPLICABLE despite making claims about model extraction accuracy, review sample sizes, fake-review signals, travel-platform behavior, and commute estimation.
No fabricated citation was found. The problem is unsupported confidence, not invented bibliography.
7. Tier Differentiation
Winner: Claude
Claude offers the cleanest progression:
1. Beginner: Calculate the real all-in price. 2. Intermediate: Decode listing language, analyze review patterns, and build a neighborhood-verification checklist. 3. Advanced: Compare finalists through a weighted decision matrix.
Each prompt solves a different stage of the lodging decision, and the Advanced tier explicitly consumes the outputs of the first two rather than repeating them.
ChatGPT’s tiers are genuinely more sophisticated as they progress, but the Beginner prompt already calculates costs, interprets listing language, analyzes reviews, ranks multiple properties, and creates a booking checklist.
That makes it a strong prompt but a less convincing Beginner tier. Its Intermediate and Advanced prompts also share many of the same mechanisms—weights, vetoes, confidence, cancellation scenarios, source dates, and sensitivity testing—so the distinction is primarily depth rather than a clean change in approach.
Gemini differentiates its three tiers clearly: calculator, review analyzer, and matrix. That earns it a respectable score here. The problem is not sameness; it is that the Advanced tier does not have enough analytical controls to fulfill its stated role safely.
The Winner: Claude
Claude wins Week 4 with 62 out of 70, four points ahead of ChatGPT.
It did not produce the most sophisticated individual prompt. ChatGPT did. Claude produced the strongest article: a more coherent reader journey, cleaner tier boundaries, better editorial control, correct handling of neighborhood uncertainty, and the only credible external citations in the set.
Its Intermediate prompt is the decisive section. It does more than summarize reviews. It combines three useful behaviors:
- Translate sales language into hypotheses.
- Analyze review evidence by recency, repetition, and silence.
- Turn neighborhood uncertainty into a verification plan rather than an AI verdict.
That combination captures the central lesson of the Week 4 brief: AI should organize the investigation, not impersonate a live local source.
The Honest Counter-Case
Claude’s win is not a clean sweep.
Its arithmetic error is substantial and should be corrected before publication. A post teaching “true all-in cost” cannot leave a seven-night calculation wrong by nearly $95 per night before tax.
ChatGPT is much stronger when the stakes justify a formal decision system. Its evidence classifications, liquidity model, uncertainty penalties, stress tests, and reversal-evidence question are more advanced than anything in Claude’s matrix. For expensive, accessibility-sensitive, multi-traveler, or non-refundable bookings, ChatGPT’s Advanced prompt is the better standalone tool.
ChatGPT would have been much more competitive as a complete article with three changes:
1. Use the required plain heading syntax. 2. Cut repeated explanation aggressively. 3. Keep the Beginner prompt focused on one decision stage instead of compressing the whole week into it.
Gemini’s genuine strength is speed. A reader could copy its Beginner prompt and calculate a lodging total immediately. Its three-tier concept is also easy to understand at a glance. The Advanced workflow needs revision, but the article is not without value: it demonstrates how much usability a prompt can gain by reducing setup.
Reader Takeaway
The best prompt is not automatically the prompt with the most instructions.
ChatGPT shows the value of explicit evidence rules, vetoes, uncertainty handling, and stress testing. Claude shows that those controls must be introduced at the right stage and at a level the reader will actually use. Gemini shows the risk of making a prompt simple by silently handing unsupported judgment back to the model.
The reusable rule from Week 4 is:
Give AI the structure, arithmetic, comparison rules, and verification questions. Keep live prices, current routes, property conditions, and neighborhood judgments in evidence the reader supplies and verifies.
That boundary is what turns an impressive lodging answer into a defensible lodging decision.
ChatGPT's Advanced prompt shipped with escaped markdown in the block you are meant to copy. Its scoring-weight fields — the blanks you fill in — were written as \_\_ instead of __, ten of them in all: "True all-in cost fit: \_\_", "Location fit for this itinerary: \_\_", and so on. The backslashes are an authoring artifact; ChatGPT meant a pair of blanks, not a backslash-underscore-backslash-underscore.
Paste the prompt into a chatbot and the backslashes go along for the ride. They are harmless — every model reads straight past them — but they are not what ChatGPT meant to write, and the reader prompts are the product here, so we flag them rather than let them pass unremarked.
We have not corrected the post. Ketelsen.ai is an experiment in what these models actually produce from one identical brief, so the untidy parts are part of the finding. Cleaning the escapes, or telling the models in advance to avoid them, would quietly delete the observation. This week Claude emitted no such escapes; ChatGPT and Gemini each emitted some.
Two notes on Gemini's post this week, both about what it produced rather than what it argued.
The prompt breakdowns are formatted differently from the other posts. The site's "Prompt Breakdown — How A.I. Reads the Prompt" format is running text — a quoted fragment, then a colon, then the explanation — which renders as the two-column "teaching ledger" you see in the Claude and ChatGPT posts. Gemini instead wrote each quoted fragment as its own ### subheading with the explanation beneath it, so its breakdown renders as headings-and-paragraphs. The content is all there; only the presentation is downgraded. The judge (ChatGPT this week) cited this and Gemini's bolded section headings under Template Compliance, where Gemini scored 5 of 10 — but Gemini placed third on the strength of its content, by a wide margin, so the formatting did not decide its finish.
A few escaped characters landed in the reader prompts. Gemini's Advanced prompt numbers its scoring criteria as 1\., 2\., 3\. — escaped periods — in the block you copy, four such escapes in all. They are harmless; every model reads past them. But the prompts are the product, so we flag them.
We have not corrected the post. Ketelsen.ai is an experiment in what these models actually produce from one identical brief, so the format each one chooses and the stray marks it leaves are part of the finding. Reformatting Gemini's breakdown to match the others, or cleaning the escapes, or telling it in advance how to format, would quietly delete exactly the differences this series exists to surface.
A note on this week's scoring, in the interest of showing our work. This series holds the judge to a "quote-it-or-drop-it" rule: to mark a post down on Template Compliance, the judge must quote the specific template requirement it violated. This week the judge (ChatGPT) docked both ChatGPT and Gemini on Template Compliance for heading formatting — bolded section headings, and in Gemini's case ### subheadings in the prompt breakdowns.
When this comparison first shipped, our automated post-judge audit flagged the judge's citation as non-compliant, and an earlier version of this note repeated that charge. A later review of the judge's actual text showed the audit tool was wrong: the judge quoted the governing requirement word for word, directly from the week's mandatory checklist. The flag was our checker's error — it was looking for a specific keyword rather than the quoted rule — and the tool has since been fixed. The judge followed the rule correctly.
The scores were never edited in any version of this page — we do not edit a judge's scoring — and the deductions did not decide anything: Claude wins this week whether or not those Template Compliance points are restored (62 to ChatGPT's 58 to Gemini's 36). We are recording the correction here rather than quietly removing the note, because our own mistakes are part of what this experiment observes too.
TAGS: