Three AIs, One Fuzzy Wish: Who Builds the Best Shortlist?"
WEEK 93 :: POST 4 :: THE JUDGE’S CHOICE
Directions Given To The A.I. This Week+
Instructions Given to each A.I. — Please provide 3 prompt variations that share this objective:
Each A.I. also received two static attachments: the blog post template (structure) and the authoring instructions (voice and standards). The text below is the week-specific assignment as sent — reflowed for the web; wording unchanged.
I'd like you to write this week's Ketelsen.ai post. Two files are attached: the blog post template (the structure to follow) and the authoring instructions (context, voice, and standards). Please read both before you begin, then produce the complete post in a single response.
This week's theme: "Where Should We Go?" — Destination Intelligence and Shortlisting.
This is Week 2 of an eight-week series on planning a vacation with AI. Last week the reader established their real constraints — a validated budget ceiling and a constraint profile covering money, dates, travellers, and trip purpose. This week they use it.
The job is to turn a fuzzy wish — "somewhere warm, not too touristy" — into a scored shortlist of three to five real destinations, each evaluated against the Week 1 profile rather than against vibes. This is the week that replaces "where do you want to go?" with "where actually fits?"
The three prompts should help a reader evaluate candidates on:
- Cost of living on the ground — not just flights and lodging, but what a day actually costs once you're there. Two destinations with identical airfare can differ enormously on daily spend.
- Seasonality and weather windows — whether the reader's available dates are good, tolerable, or actively wrong for each candidate.
- Crowd calendars — school holidays, festivals, and local peak seasons that turn a good destination into a bad week.
- Safety and visa friction — entry requirements, processing times, and anything that could quietly disqualify a destination late.
- Flight accessibility from the reader's home airport — a destination three connections away is a different trip from a nonstop, regardless of ticket price.
The output a reader should walk away with is a ranked shortlist with a composite fit score and the reasoning behind it — something they can defend to a travelling companion, not just a gut preference.
A note on the strongest version of this week: at the advanced end, this is a destination dossier matrix — one row per candidate, columns for cost index, weather risk, crowd level for the specific travel dates, entry requirements, and a composite score. That structure is worth reaching for.
Series dependency chain, for the Metadata block: Week 2 consumes Week 1's validated budget ceiling and constraint profile — the shortlist is scored against them, and any prompt that ignores the reader's established constraints has missed the point of the series. Week 2 produces the chosen destination (or final shortlist), which feeds Week 3 (airfare strategy), Week 4 (lodging), and Week 5 (itinerary). It is the hinge of the whole series: everything downstream assumes it.
Because readers may arrive at this post without having read Week 1, the prompts should work for someone who has their constraints roughly in mind, while making clear that the reader gets far more out of them with a real constraint profile in hand.
Three difficulty tiers as always — Beginner, Intermediate, Advanced — each a genuinely different approach to the same problem, not the same prompt at three lengths.
A note on supplied figures. Anything marked `[SUPPLIED — use as given]` above came from Ketelsen.ai's own research brief. Use it freely — you are not fabricating by repeating it, and you will not be marked down for leaving it uncited. Do not attach an invented source to it. Presenting it as a sensible planning rule of thumb rather than a precise measured statistic is the right treatment. (No supplied figures this week — the constraint stands for the series.)
ChatGPT shipped invisible garbage, and it cost it the week. 29 Unicode private-use characters and 9 raw citation tokens (citeturn466772view0) left in the prose. The judge named this as decisive:
"Had ChatGPT formatted its text cleanly, its highly engineered prompts would have made this a dead heat for first place."
Scored across seven dimensions by this week's rotating judge. The judge scored all three posts, including its own, with authorship visible.
Claude takes Week 2 :: Vacation Planning Series with 64 of 70, ahead of ChatGPT on 61.
Each dimension scored 1-10 by the judge. These are the judge's own scores, not measured data.
2. Dimension-by-Dimension Reasoning
1. Prompt Quality & Creativity
- Tie: Claude and ChatGPT. Claude’s prompts are incredibly human-centric. Asking the model for "one honest thing that might not work for me" elegantly forces the AI out of its default agreeableness. ChatGPT matches this quality with highly engineered constraints, specifically commanding the AI to create a "Minimum Viable Constraint Profile" if the user's input is lacking. Both approaches are masterclasses in controlling AI behavior. My own post (Gemini) produced solid prompts, including a great request for a "brutal defense", but lacked the foolproof input-handling of the other two.
2. Content Depth & Accuracy
- Winner: Claude. Claude's "Prompt Breakdown" sections don't just explain the prompt; they teach transferable principles of AI interaction. For example: "The principle: when a model has an obvious bad default for your task, naming the thing you don't want is often more powerful than describing the thing you do". ChatGPT is similarly deep, explaining why each clause exists, but Claude's formulation as universal "principles" gives it the slight edge. Gemini's breakdowns are competent but much briefer.
3. Template Compliance
- Tie: All Platforms. Governed absolutely by the quote-it-or-drop-it rule, I cannot deduct points for missing MUST requirements because all three posts successfully included every structural header standard to the assignment (e.g., ## In one line, ## What this prompt gives you, ## Prompt Breakdown, etc.). No MUST violations are present to quote, so all posts receive full marks here.
4. Practical Utility
- Tie: Claude and ChatGPT. ChatGPT's inclusion of a "verification plan with the exact categories of live information I must confirm before booking" is an incredibly practical safety net for high-stakes travel. Claude's focus on defining the exact axes of evaluation (Cost, Weather, Crowds, Flights, Safety) makes its intermediate prompt instantly usable. Both tools yield highly actionable outputs.
5. Engagement & Readability
- Winner: Claude. Claude’s register is flawless. The opening line—"'Somewhere warm, not too touristy' is a wish, not a plan"—hooks the reader immediately and maintains a crisp, authoritative tone throughout. ChatGPT lost significant ground here due to glaring AI-writing tells: it left raw, unparsed citation tokens (e.g., citeturn466772view0) directly in the final text. This severely disrupts readability.
6. Citation Quality
- Winner: Claude and Gemini. Both Claude and Gemini properly utilized the NOT APPLICABLE designation for their citations, transparently acknowledging that the prompts and breakdowns were original work and guiding the reader to verify real-world travel data independently. ChatGPT attempted to pull real sources (like the U.S. Department of State and IATA), which is commendable, but the execution failed by leaving the raw system artifact tokens in the prose.
7. Tier Differentiation
- Winner: Claude. Claude’s progression feels the most like a natural traveler's journey. It moves from a "Fuzzy-Wish Shortlist" (just getting ideas) to a "Weighted Shortlist" (comparing ideas) to a "Destination Dossier Matrix" (defending a final choice). Each tier requires a genuinely different cognitive approach from the user. ChatGPT and Gemini also scale well, but Claude captures the psychological shift of planning best.
3. The Winner and the Case for It
The Winner is Claude (64 points). Claude wins because it delivered the most polished, reader-friendly, and philosophically sound guide to prompt engineering of the week. Its prompts are elegant, its tone is engaging without being overly enthusiastic, and its breakdowns teach the reader how to think about AI, not just what to copy and paste. The margin is close. Claude beat ChatGPT by exactly three points. What flipped the result was ChatGPT's failure to clean up its raw citation tokens (cite...). Had ChatGPT formatted its text cleanly, its highly engineered prompts (like the Minimum Viable Constraint Profile) would have made this a dead heat for first place.
4. The Honest Counter-Case
While Claude won, it was not the strongest in every isolated mechanic.
- Where ChatGPT did better: ChatGPT’s prompts are arguably more structurally bulletproof. By explicitly instructing the AI to "create a 'Minimum Viable Constraint Profile'" if the user forgets to input data, ChatGPT protects the user from a failed generation in a way Claude does not.
- Where Gemini did better: My own post (Gemini) was the most explicit about formatting, directly commanding the AI to output the Advanced matrix in a "Markdown table format". This guarantees a scannable, clean output, whereas the others rely on the model's natural inclination to use tables when asked for a "matrix."
5. Reader Takeaway
This week’s divergence teaches a vital lesson about using AI for high-stakes decisions: AI is a terrible oracle, but an incredible filter. If you ask an AI "Where should I go?", it will feed you glossy, agreeable travel-brochure clichés. But if you force the AI to act as a strict logistical analyst—commanding it to score destinations on crowds, weather risks, and visa friction before it is allowed to praise them—it becomes an invaluable tool. The difference between a bad output and a great one is simply demanding that the AI show its math.
Gemini's post shipped with escaped markdown in its reader prompts. The post as delivered contains 82 backslash-escaped characters, 51 of them inside the prompts you are meant to copy — so the placeholders read \[Total Budget\], \[Number of Travelers\] and \[Start Date\] instead of [Total Budget], [Number of Travelers] and [Start Date].
Paste one of those prompts into a chatbot and the backslashes go with it. They are harmless — every model will read straight through them — but they are not what Gemini meant to write.
The scoring was not affected by it. Gemini held the judge's seat this week and placed its own post third of three — and it never mentioned these artifacts, in its own post, while scoring it. We found them ourselves, after the week had already been marked publish-ready, while widening a check of ours that had been reporting two such artifacts when there were eighty-two.
We have not corrected the post. Ketelsen.ai is an experiment in what these models actually produce from an identical brief, so what they produce is the finding — including the untidy parts. Editing it, or telling the models in advance not to do it, would quietly delete the observation. The tool that failed to count it properly was ours, and that has been fixed permanently.
TAGS: