AI Showdown: Building the Itinerary That Doesn't Break

WEEK 96 :: POST 4 :: THE JUDGE’S CHOICE

Directions Given To The A.I. This Week+

Instructions Given to each A.I. — Please provide 3 prompt variations that share this objective:

Each A.I. also received two static attachments: the blog post template (structure) and the authoring instructions (voice and standards). The text below is the week-specific assignment as sent — reflowed for the web; wording unchanged.

I'd like you to write this week's Ketelsen.ai post. Two files are attached: the blog post template (the structure to follow) and the authoring instructions (context, voice, and standards). Please read both before you begin, then produce the complete post in a single response.

This week's theme: "Building the Itinerary That Doesn't Break" — Day-by-Day Design and Pacing.

This is Week 5 of an eight-week series on planning a vacation with AI. By now the reader has a budget ceiling (Week 1), a destination (Week 2), booked flights that fix the dates and arrival time (Week 3), and a place to stay whose location is now known (Week 4). This week they turn a list of things they want to do into a day-by-day plan that survives contact with a real trip.

Most itineraries fail the same three ways, and none of them are about picking the wrong attractions. They fail from over-scheduling — cramming a day so full that one late lunch collapses everything after it; from geographic ping-ponging — crossing the city four times because the plan was built by interest rather than by map; and from ignoring the calendar — arriving to find the museum closed on Mondays and the one restaurant they cared about booked out for three weeks. The job this week is a plan built around how days actually go, not how they look on paper.

The deliverable the reader should walk away holding is a day-by-day itinerary they can defend: activities clustered so each day stays in one part of the map, a realistic pace, a booking-deadline calendar sorted by how far ahead each thing must be reserved, and a backup for the days most likely to fall apart.

The three prompts should help a reader work through:

  • Clustering by geography, not by interest. Grouping the things they want to do by where they are, so a day moves through one neighbourhood or district instead of doubling back across town. The AI is good at taking a list of places plus the reader's lodging location and proposing sensible clusters — as long as the reader supplies the places and the rough map.
  • Pacing with an anchor-and-flex rhythm. One committed thing per day — the anchor — with everything else held loosely as optional. This is the antidote to over-scheduling: a day with a single non-negotiable and a menu of maybes bends instead of breaking when something runs long. Building in a deliberate zero day — a day with nothing planned — belongs here too.
  • Reservation lead-time intelligence. Which kinds of things book out far in advance and which can be decided the morning of, turned into a calendar sorted by book-by date so nothing is lost to a deadline the reader never saw. The AI supplies the pattern — that certain restaurants, timed-entry museums, and marquee experiences typically need booking well ahead — while the reader confirms the actual dates against live booking pages.

At the advanced tier, the strongest version of this week is a structured day-by-day plan with pacing rules made explicit: each day an anchor plus ranked optionals, transit time between clustered stops accounted for, a backup option per day, and a reservations-deadline calendar the reader can act on. That is the structure worth reaching for.

A hard constraint, carried forward from Weeks 3 and 4 and just as binding here. AI models cannot see live opening hours, this season's closure days, current reservation availability, or today's transit schedules — and their recall of a specific venue's hours is stale and frequently wrong. No prompt in this post may ask the AI to state a specific attraction's opening hours, claim a particular restaurant is bookable on a given date, or assert current transit times as fact. A confidently wrong opening time sends the reader across the city to a locked door. Prompts should have the AI produce what to verify and where — the questions to ask, the pages to check, the buffers to leave — rather than deliver a schedule stated as certain.

Design the prompts so the AI does what it is genuinely good at: organising a messy wish-list into a coherent map-aware sequence, enforcing a sane pace, and naming what must be booked ahead. The reader supplies the wish-list, the lodging location, and the confirmed hours; the AI supplies the structure and the sequencing. Posts that have the AI invent hours or availability should expect to be marked down on Practical Utility, exactly as in Weeks 3 and 4.

Series dependency chain, for the Metadata block: Week 5 consumes the booked lodging and its location from Week 4 (the itinerary is built outward from where the reader wakes up each morning), the confirmed dates from Week 3, the destination from Week 2, and the budget ceiling from Week 1. Week 5 produces the day-by-day plan and its reservation calendar, which Week 6's logistics-and-protection audit and Week 7's in-trip prompts both assume as the shape of the trip they are protecting and running.

Because readers may arrive at this post without having read Weeks 1 to 4, the prompts should work for someone who knows their destination, dates, lodging area, and a rough list of what they want to do, while making clear they get far more from them with a real constraint profile and confirmed bookings in hand.

Three difficulty tiers as always — Beginner, Intermediate, Advanced — each a genuinely different approach to the same problem, not the same prompt at three lengths.

On examples: this is a consumer travel topic. The template lists tech startup / retail / freelance as suggested industry examples — those are marked MAY, and this week you should almost certainly adapt them. Families pacing a trip around nap times and young children, couples balancing one person's museum day against the other's beach day, a group trying to share a plan without a dozen group-chat threads, and older travellers for whom walking distance between stops decides the day are the right contexts here. Choosing them over the suggested business examples is correct behaviour and will not be scored against you.


A note on supplied figures. Anything marked `[SUPPLIED — use as given]` above came from Ketelsen.ai's own research brief. Use it freely — you are not fabricating by repeating it, and you will not be marked down for leaving it uncited. Do not attach an invented source to it. (No supplied figures this week. Given the live-hours-and-availability constraint above, this is a bad week to invent any — if you find yourself reaching for a venue's opening hours or how far ahead a restaurant books, that is the signal to restructure the prompt so the reader supplies the real detail instead.)


## BEFORE YOU SUBMIT — STRUCTURAL CHECK

(This block is identical every week. It exists because these specific items are the ones posts drop, and a dropped structural item costs compliance points for something that takes one minute to add.)

Your post is parsed by a script before any human reads it. Confirm all seven:

1. ☐ Response begins with `PLATFORM: <your name>` and `WEEK: 5` 2. ☐ `## Lead` present once, at the very top, before Variation 1 3. ☐ `## In one line` present in all three variations 4. ☐ `## What this prompt gives you` present in all three variations 5. ☐ `## The Prompt` present in all three variations, with the prompt in double quotes beneath it 6. ☐ `## Introductory Hook` and `## Current Use` present in all three variations (three of each — not one) 7. ☐ Every template heading written as `##`, none bolded instead; prompt breakdown is running text split on ` : `, with no `###` headings inside it

A complete post has 57 `##` headings. If your count is well short, a section is missing or was bolded instead of hashed.

One extra check this week: confirm no prompt asks the AI to state a specific venue's opening hours, closure days, or current reservation availability, or to present a transit time as fact. Those must be things the reader goes and verifies.

Week 5 :: Vacation Planning Series · The Showdown
Week 5 :: Vacation Planning Series Scorecardout of 70
Claude
69
ChatGPT
60
Gemini
53

Scored across seven dimensions by this week's rotating judge. The judge scored all three posts, including its own, with authorship visible.

This week's winner

Claude takes Week 5 :: Vacation Planning Series with 69 of 70, ahead of ChatGPT on 60.

Every dimension, side by side
ChatGPTGeminiClaude
1. Prompt Quality & Creativity
8
6
10
2. Content Depth & Accuracy
8
7
10
3. Template Compliance
10
10
10
4. Practical Utility
9
7
10
5. Engagement & Readability
9
7
10
6. Citation Quality
8
9
9
7. Tier Differentiation
8
7
10

Each dimension scored 1-10 by the judge. These are the judge's own scores, not measured data.

Dimension by dimension

1. Prompt Quality & Creativity — Claude wins (10).

Claude’s approach to prompt engineering here is exceptional. Instead of simply changing the word count or the number of variables, Claude introduces distinct, highly creative AI interaction paradigms at each tier. The intermediate prompt utilizes a "hold point" (staged prompting) where the AI stops and asks for human confirmation of the geographic map before building the days. The advanced prompt employs a brilliant "four-pass stress test" (Draft, Stress Test, Repair, Booking Calendar), turning the AI into both a creator and an adversarial auditor. ChatGPT offers a solid but more traditional set of fill-in-the-blank prompts, while Gemini's prompts, though functional, lack the structural ingenuity that makes the AI truly act as a travel strategist. Applying the required scrutiny to my own post (Gemini), I must admit the prompts are basic and rely too heavily on simple instructions rather than advanced prompting mechanics. 2. Content Depth & Accuracy — Claude wins (10). Claude excels at explaining exactly why the prompts work. For example, in explaining the phrase "the single thing I'd be genuinely disappointed to miss", Claude notes that "This defines the anchor by consequence rather than by category," forcing the model to weigh feelings over popularity. ChatGPT is also strong here, correctly noting that without an anchor, "every day becomes a pile of equal suggestions". Gemini’s prompt breakdowns are competent but slightly shallower, often explaining what a line does without diving into the underlying mechanics of how LLMs process the instruction. 3. Template Compliance — three-way tie (10 each). All three posts successfully present every standard section required by the brief, from the "Lead" down to "Citations" and "Metadata". Because I do not have access to the original template's specific MUST phrasing to quote against, the "quote-it-or-drop-it" rule mandates that I take no deductions here. There are no skipped, padded, or merged sections in any of the posts, and NOT APPLICABLE is used honestly and appropriately. 4. Practical Utility — Claude wins (10). Claude’s prompts are the most practically useful because they actively prevent the AI from making costly mistakes. The intermediate prompt’s staged approach ("Stop there and ask me to confirm or correct the clusters before you go any further") saves the user from having to redo a fully generated itinerary just because the AI misunderstood the city layout. ChatGPT is a very close second, utilizing a strong "anchor-and-flex" rhythm and clear verification checklists. Gemini provides practical advice but leaves more of the heavy lifting to the user, particularly in the beginner prompt where the cluster output might still require significant manual editing. 5. Engagement & Readability — Claude wins (10). Claude’s prose is incredibly sharp, engaging, and devoid of AI filler. The opening hook sets a vivid, relatable scene: "It is 2:40 in the afternoon, lunch ran long, the queue was worse than expected, and you are standing on a street corner doing arithmetic about whether you can still make the 4:00 thing". ChatGPT also writes a highly readable post, using great metaphors like "the day becomes a pinball machine". Gemini’s tone is accessible but leans slightly toward generic travel-blog writing ("without losing your mind") and lacks the punchy, authoritative voice seen in Claude’s draft. 6. Citation Quality — Claude and Gemini tie (9). Both Claude and Gemini correctly use NOT APPLICABLE for the prompts that rely purely on the supplied theme constraints and general knowledge, which is the honest and expected behavior. Claude appropriately cites "Anthropic, prompt engineering documentation" for its intermediate prompting technique. ChatGPT cites the "supplied Week 5 direction" for all three variations. While not technically an external source, ChatGPT is clearly attempting to attribute the constraints given in the assignment, which reads as thin rather than fabricated. No post fabricated any facts or studies, so no severe penalties are warranted. 7. Tier Differentiation — Claude wins (10). Claude genuinely scales the concept of interacting with the AI, not just the length of the prompt. Beginner is a one-shot generation, Intermediate introduces a two-step conversational hold point, and Advanced deploys a complex, multi-pass adversarial generation sequence. ChatGPT scales well from basic clustering to a comprehensive 6-section audit. Gemini scales appropriately but the gap in methodology between the intermediate and advanced tiers feels smaller, relying mostly on adding rules rather than changing the interaction paradigm.

A clear winner

Claude

is the undeniable winner this week. It didn't just write a good blog post about travel planning; it delivered masterclasses in prompt engineering. The introduction of "staged prompting" with a hold point, and the adversarial "four-pass stress test," are incredibly sophisticated ways to force an LLM to behave reliably. Claude’s writing is evocative, its explanations of the prompt mechanics are deeply insightful, and the resulting tools are immediately practical.

The honest counter-case

While Claude's post is a triumph in prompt engineering, it is dense. A casual reader looking for a quick travel hack might find the advanced prompt's "four-pass" explanation a bit intimidating and long-winded to digest. ChatGPT actually did a slightly better job at keeping the advanced prompt feeling approachable. By structuring the prompt into a rigid "six sections" output, ChatGPT guarantees a very clean, easy-to-read final itinerary format that some users might prefer over Claude's more dynamic, conversational output. Gemini, while scoring lowest overall, provided a very succinct and clear explanation of the "Zero Day" concept in its intermediate tier, which was highly practical for burnout prevention.

Takeaway for readers

This week’s divergence teaches a critical lesson about prompting: do not let the AI guess facts it cannot verify. All three models correctly recognized that LLMs hallucinate live data like opening hours and train schedules. The best prompts (like Claude's and ChatGPT's) solve this by changing the AI's job title. Instead of asking the AI to be an omniscient travel agent, they ask it to be a map-aware strategist that organizes your ideas, flags potential risks, and hands you a specific, prioritized checklist of exactly what you need to verify in the real world.

Editorial note · no scoring impact

Gemini's post shipped with escaped markdown in its reader prompts. The post as delivered contains 30 backslash-escaped characters, 20 of them inside the three prompts you are meant to copy — so the placeholders read \[Destination\], \[Neighborhood/Area\], \[Insert List\], \[Number\], \[Start Date\], \[End Date\] and \[Lodging Neighborhood\] instead of [Destination], [Neighborhood/Area] and the rest.

Paste one of those prompts into a chatbot and the backslashes go with it. They are harmless — every model will read straight through them — but they are not what Gemini meant to write. The remaining ten escapes sit in the commentary around the prompts, where they show up as things like 1\. in a numbered list and \pattern\ where italics were intended.

The judge did not mention this and it had no part in the scoring. This week's judge was Gemini itself, and it placed its own post third — on depth of content, not on formatting.

We have not corrected the post. Ketelsen.ai is an experiment in what these models actually produce from an identical brief, so what they produce is the finding — including the untidy parts. Editing it, or telling the models in advance not to do it, would quietly delete the observation.

Worth noting, five weeks in: Gemini has emitted escaped markdown in every week of this series except the first, and Claude has emitted none in any week. That is a real and persistent difference between the two, and it is the kind of thing this series exists to surface.

TAGS:

Previous
Previous

Fitting the Wish-List Into One Week Without Losing Your Mind

Next
Next

Itineraries Break for Boring Reasons — Plan for Them