AI Showdown: Landing the Plane — One Clear Winner

WEEK 99 :: POST 4 :: THE JUDGE’S CHOICE

Directions Given To The A.I. This Week+

Instructions Given to each A.I. — Please provide 3 prompt variations that share this objective:

Each A.I. also received two static attachments: the blog post template (structure) and the authoring instructions (voice and standards). The text below is the week-specific assignment as sent — reflowed for the web; wording unchanged.

I'd like you to write this week's Ketelsen.ai post. Two files are attached: the blog post template (the structure to follow) and the authoring instructions (context, voice, and standards). Please read both before you begin, then produce the complete post in a single response.

This week's theme: "Landing the Plane" — Post-Trip Reconciliation and the Reusable System.

This is Week 8, the final week of an eight-week series on planning a vacation with AI. The trip is over. This week is about closing it out well and, more importantly, about keeping what the reader just built — turning eight weeks of one-off prompts into a personal, reusable system so the next trip takes a fraction of the effort. This is the payoff the whole series was pointing at.

There are two jobs. The first is reconciliation: a budget post-mortem comparing what was planned against what was actually spent and finding where the estimates were wrong; and the recovery work of disputing incorrect charges and pursuing any compensation the reader is owed — a delayed flight, a resort fee that was never disclosed, a charge that does not match what was agreed. The second job is the durable one: taking the workflow the reader has now lived through and compressing it into a template they can run again, so the knowledge does not evaporate the moment they unpack.

The deliverable the reader should walk away holding is twofold: a clear post-trip reconciliation — what was spent versus planned, what to dispute and how — and a reusable personal trip-planning template distilled from the eight-week process, ready to run for the next destination.

The three prompts should help a reader:

  • Run the budget post-mortem. Compare planned against actual from the reader's own records, surface where the plan was optimistic, and turn that into a sharper set of assumptions for next time — the AI structuring the comparison from numbers the reader supplies, not inventing what a trip "should" cost.
  • Pursue disputes and claims, methodically. Organise a charge dispute or a delay claim into who to contact, what evidence to attach, and what to ask for, in a calm and orderly sequence — while pointing the reader to the airline, card issuer, or platform to confirm what they are actually entitled to rather than asserting it.
  • Build the reusable template. Distil the eight-week workflow into a personal, repeatable planning system — the steps, the prompts worth keeping, and the reader's own hard-won preferences — so the next trip starts from a framework instead of a blank page. This is the series' real lesson: a good prompt, saved and adapted, becomes a system.

At the advanced tier, the strongest version of this week is a written, reusable planning template the reader can save and re-run — the whole series compressed into a sequence of steps and prompts tuned to how this traveller actually plans, plus a reconciliation summary that feeds next time's estimates. That structure is worth reaching for, and it is the natural landing point for a series about turning one-off AI help into a repeatable method.

A hard constraint, carried through to the last week. AI models cannot see the reader's actual receipts, a specific carrier's current compensation policy, or the rules that govern a particular claim, and these are jurisdiction- and date-specific. No prompt may ask the AI to confirm a specific compensation entitlement, promise a payout, adjudicate whether a charge is disputable, or state current claim rules as settled fact. The dispute and claim prompts should give the reader an organised process and the right questions — and send them to the airline, card issuer, or regulator to confirm — not a verdict on their case.

Design the prompts so the AI does what it is genuinely good at: structuring a reconciliation from supplied numbers, organising a dispute into an orderly sequence, and compressing a lived process into a reusable template. The reader supplies their records and their experience; the AI supplies structure, sequence, and the distilled system. Posts whose prompts have the AI assert entitlements or invent costs should expect to be marked down on Practical Utility and Content Accuracy.

Series dependency chain, for the Metadata block: Week 8 consumes the entire trip — the budget ceiling from Week 1 (to reconcile against), the bookings and protection work from Weeks 3, 4 and 6 (for disputes and claims), and the in-trip record from Week 7 (what actually happened and what it cost). Week 8 produces the reusable planning template — the series' closing artifact — which has no successor week because it is the thing designed to start the next trip's Week 1.

Because readers may arrive at this post having only just returned, without having read the earlier weeks, the prompts should work for someone who simply has their receipts and a sense of how the trip went, while making clear the reusable template is far richer when it is distilled from the full process.

Three difficulty tiers as always — Beginner, Intermediate, Advanced — each a genuinely different approach to the same problem, not the same prompt at three lengths.

On examples: this is a consumer travel topic. The template lists tech startup / retail / freelance as suggested industry examples — those are marked MAY, and this week you should almost certainly adapt them. A family reconciling a trip's spending against the plan, a traveller pursuing a delayed- flight claim, a couple disputing an undisclosed resort fee, and anyone building a personal template so the next trip is easier are the right contexts here. Choosing them over the suggested business examples is correct behaviour and will not be scored against you.


A note on supplied figures. Anything marked `[SUPPLIED — use as given]` above came from Ketelsen.ai's own research brief. Use it freely — you are not fabricating by repeating it, and you will not be marked down for leaving it uncited. Do not attach an invented source to it. (No supplied figures this week. Given the claims-and-entitlements constraint above, this is a bad week to invent any — if you find yourself reaching for a compensation amount or a typical trip cost, that is the signal to restructure the prompt so the reader supplies the real number and confirms entitlements at the source.)


## BEFORE YOU SUBMIT — STRUCTURAL CHECK

(This block is identical every week. It exists because these specific items are the ones posts drop, and a dropped structural item costs compliance points for something that takes one minute to add.)

Your post is parsed by a script before any human reads it. Confirm all seven:

1. ☐ Response begins with `PLATFORM: <your name>` and `WEEK: 8` 2. ☐ `## Lead` present once, at the very top, before Variation 1 3. ☐ `## In one line` present in all three variations 4. ☐ `## What this prompt gives you` present in all three variations 5. ☐ `## The Prompt` present in all three variations, with the prompt in double quotes beneath it 6. ☐ `## Introductory Hook` and `## Current Use` present in all three variations (three of each — not one) 7. ☐ Every template heading written as `##`, none bolded instead; prompt breakdown is running text split on ` : `, with no `###` headings inside it

A complete post has 57 `##` headings. If your count is well short, a section is missing or was bolded instead of hashed.

One extra check this week: confirm no prompt asks the AI to confirm a compensation entitlement, promise a payout, adjudicate whether a charge is disputable, or state current claim rules as settled fact. Those must be a process and questions the reader confirms with the airline, card issuer, or regulator.

Week 8 :: Vacation Planning Series · The Showdown
Week 8 :: Vacation Planning Series Scorecardout of 70
Claude
65
ChatGPT
60
Gemini
43

Scored across seven dimensions by this week's rotating judge. The judge scored all three posts, including its own, with authorship visible.

This week's winner

Claude takes Week 8 :: Vacation Planning Series with 65 of 70, ahead of ChatGPT on 60.

Every dimension, side by side
ChatGPTGeminiClaude
1. Prompt Quality & Creativity
9
6
10
2. Content Depth & Accuracy
9
6
9
3. Template Compliance
10
7
10
4. Practical Utility
10
7
9
5. Engagement & Readability
8
6
9
6. Citation Quality
5
4
8
7. Tier Differentiation
9
7
10

Each dimension scored 1-10 by the judge. These are the judge's own scores, not measured data.

Dimension by dimension

1. Prompt Quality & Creativity — Claude wins with 10; ChatGPT follows with 9; Gemini scores 6.

Claude built the most inventive prompt set, particularly at the advanced tier. Its Reusable Trip Playbook does not simply request a template. It creates a controlled four-phase interaction: “Work in four phases and stop for my input between each one,” followed by an interview, extraction pass, build phase, and adversarial stress test. The final instruction — “Now try to break what you built” — forces the model to identify which stage would fail under different destinations, companions, or planning windows before revising and versioning the playbook. That is a genuine system-design prompt rather than a long request for organized prose.

ChatGPT came close. Its advanced prompt uses five stages and includes an assumption ledger, an unresolved-items register, human-verification steps, completion criteria, saved handoff outputs, preference confidence labels, a portability audit, and a version-history entry. It is arguably the most comprehensive prompt in the set. Its weakness relative to Claude is that it behaves more like a detailed specification handed to the model all at once; Claude more deliberately controls the conversation through pauses, corrections, and a stress test.

Gemini’s prompts are usable but much more predictable. Its advanced instruction asks the AI to “distill this entire workflow into a sequential, repeatable master system” with customized prompts and bracketed placeholders. That is a sound objective, but it does not define an interview, correction loop, provenance rule, test procedure, or quality-control stage. The model is asked to produce a system, but not given much of a system for producing it.

2. Content Depth & Accuracy — ChatGPT and Claude tie at 9; Gemini scores 6.

ChatGPT and Claude both teach transferable prompting principles rather than merely paraphrasing their own prompts.

ChatGPT is strongest when explaining durable information architecture. Its advanced design separates destination-specific lessons from reusable traveler-specific lessons, labels assumptions KEEP, CHANGE, or TEST AGAIN, separates preferences into CONFIRMED, TENTATIVE, and TRIP-SPECIFIC, and requires source and last-updated fields. Those distinctions show an accurate understanding of the danger of converting one trip’s circumstances into permanent personal rules.

Claude reaches comparable depth through behavioral and conversational design. Its advanced prompt asks the model to infer what the traveler values “judged by what I chose rather than by what I say I prefer,” requires each extracted principle to be traceable to something the user actually said, and instructs the model to label inference separately from evidence. Its explanations of revealed preferences, provenance, bounded pushback, and staged self-critique are substantive and broadly reusable.

Gemini’s explanations are generally correct, but often stop after describing the immediate function of a sentence. For example, bracketed placeholders are called “the secret sauce” because they make a template reusable, but the post does not explore how the system should distinguish observations from inferences, prevent stale assumptions, test itself, or evolve over time.

There is no evidence that any post fabricated a factual statistic or assignment-supplied figure. Gemini’s lower score reflects shallower teaching, not a major factual failure.

3. Template Compliance — ChatGPT and Claude tie at 10; Gemini scores 7.

ChatGPT and Claude include the complete required section structure across all three variations. Their prompt breakdowns use running text with quoted fragments followed by : explanations, and neither substitutes third-level headings for that required format.

Gemini also includes the substantive sections, so this is not a gross omission. Its deduction rests on the attached Week 8 brief’s mandatory pre-submit checklist, which says:

> “Every template heading written as ##, none bolded instead; prompt breakdown is running text split on :, with no ### headings inside it.”

Gemini repeatedly formats headings as ## Lead, ## The Prompt, and similar bolded headings. It then divides the breakdown with ### headings such as ### "I have just returned from a vacation..." rather than using the required running-text construction.

That is a real, quotable structural miss, but it is not equivalent to omitting variations or whole required sections. A 7 reflects a complete post with a recurring mechanical-formatting defect.

4. Practical Utility — ChatGPT wins with 10; Claude scores 9; Gemini scores 7.

ChatGPT provides the most complete set of artifacts a reader could act on immediately.

Its beginner prompt accounts for refunds, shared expenses, cash, questionable charges, missing receipts, and incomplete amounts rather than limiting the exercise to a simple variance table. Its intermediate prompt produces a neutral case summary, evidence-linked timeline, evidence inventory, responsibility map, contact sequence, first-contact message, follow-up log, and escalation questions.

The advanced prompt then turns the entire series into an operational workflow. For every planning phase it requires an objective, inputs, decision questions, a copy-paste prompt, human-verification steps, completion criteria, and an output to save for the following phase. That explicit entry-work-control-exit-handoff structure gives the reader something closer to a personal operating manual than a retrospective essay.

Claude is nearly as useful. Its beginner prompt improves the budget by asking whether each major miss was a one-time event or a repeatable assumption. Its claim prompt asks the model to identify weak evidence, route questions to the correct organization, draft a factual request, and define what happens after silence. Its advanced playbook is excellent, although the interview and four-stage workflow require more time and sustained participation than ChatGPT’s more specification-driven approach.

Gemini’s prompts can all be used today, and its intermediate prompt has sensible boundaries. The main limitation is that the outputs are underspecified. The beginner prompt does not create an unresolved-items register. The intermediate prompt lacks a follow-up log or explicit escalation state. The advanced prompt assumes the reader can summarize eight weeks of work cleanly before the AI begins, rather than helping extract and validate that information.

5. Engagement & Readability — Claude wins with 9; ChatGPT scores 8; Gemini scores 6.

Claude has the most distinctive and memorable voice. Its opening puts “receipts, a boarding pass from a flight that left four hours late, and a card charge nobody in the house recognises” on the kitchen counter. That image establishes the entire week’s work without sounding like a table of contents.

The post repeatedly converts abstract prompting ideas into concrete human situations. “There is a particular kind of exhaustion that comes with being owed something” is an effective opening to the claim section because it identifies avoidance and administrative fatigue as the real user problem. Claude’s weakness is length. At more than ten thousand words, it sometimes explains a principle after the reader has already understood it, and several sections could be tightened without losing substance.

ChatGPT is more restrained and editorially disciplined. Its opening — “A trip is not finished when the suitcase is empty; it is finished when the money is reconciled, the loose ends are pursued, and the lessons are saved” — is clear, professional, and well matched to a mainstream business-publication register. Its prose is less vivid than Claude’s but easier to scan across a long post.

Gemini is the shortest and fastest to read, which is a genuine advantage. However, phrases such as “the real power of AI,” “the secret sauce,” “hard-won knowledge,” and “without the headache” give parts of the post a generic AI-explainer quality. The prose often announces importance rather than demonstrating it through specific detail.

6. Citation Quality — Claude wins with 8; ChatGPT scores 5; Gemini scores 4.

No post earns a fabrication penalty.

Claude is the only post that supplies meaningful authoritative destinations for the claim-sensitive material. It names the U.S. Department of Transportation’s Aviation Consumer Protection resources, the Consumer Financial Protection Bureau, and Regulation (EC) No 261/2004 in EUR-Lex. More importantly, it describes these as places where readers should verify their own cases, not as proof that a particular reader qualifies for compensation.

Claude still leaves many broader claims unsourced, so an 8 rather than a 9 or 10 is appropriate. The citations are concentrated in the intermediate variation, and some assertions about dispute sequencing and model behavior would benefit from support.

ChatGPT honestly uses NOT APPLICABLE rather than inventing references. That avoids the serious failure, but it leaves a long, research-like post with no supporting sources at all. Its intermediate and advanced sections discuss changing rules, authoritative verification, privacy, portability, and AI limitations; even a small source list of official consumer and travel authorities would have strengthened reader trust. The judge prompt explicitly treats honest NOT APPLICABLE as thin sourcing rather than dishonesty, which places this in the middle range rather than at the bottom.

Gemini also uses NOT APPLICABLE throughout. Its advanced tools section additionally recommends specific named model versions as “best suited” for the task without evidence or a date-qualified basis. That kind of product recommendation is time-sensitive and likely to age, making the absence of sourcing more consequential.

7. Tier Differentiation — Claude wins with 10; ChatGPT scores 9; Gemini scores 7.

Claude’s three tiers are genuinely different interaction designs.

The beginner prompt is a focused diagnostic conversation about financial variance. The intermediate prompt is a bounded case-building workflow with evidence, routing, communication, and escalation. The advanced prompt is a multi-turn interview, extraction, system-build, and stress-test process. The reader does not merely receive more output at each tier; the relationship between user and model changes.

ChatGPT also differentiates the tiers strongly: a closeout report, a recovery packet, and a complete personal operating system. Its advanced tier adds system controls, handoffs, preference states, quality control, and portability. It scores one point below Claude because ChatGPT’s three prompts retain a more similar specification-heavy style, while Claude changes both the scope and the conversational mechanism.

Gemini’s tiers solve different jobs, so this is not the same prompt repeated at three lengths. The limitation is methodological. Each tier remains primarily a one-shot instruction in which the reader provides material and the model returns an organized artifact. The advanced variation expands the scope substantially, but not the interaction design or validation discipline.

The winner

Claude wins Week 8 with 65 out of 70.

The five-point margin over ChatGPT is meaningful rather than cosmetic. Claude’s advantage comes from combining strong prompt architecture with unusually good explanatory writing. Its advanced prompt interviews before drafting, distinguishes evidence from inference, forces corrections between phases, specifies what must not be decided prematurely, stress-tests the finished playbook, and assigns it a version number and changelog. Those choices directly support the week’s central goal: converting a completed trip into a system that improves rather than a document that merely summarizes.

ChatGPT is a strong runner-up at 60. In a workflow-heavy setting, some readers may prefer it. Its operating-system prompt is more comprehensive, more explicit about verification and phase handoffs, and easier to audit as a formal specification.

Gemini finishes third at 43. Because Gemini holds the identified judge seat this week, that placement deserves to be stated plainly: the authorship labels did not create a close-call problem. Gemini’s post is materially thinner in prompt architecture, explanatory depth, sourcing, and structural execution. Its strengths are real, but they do not support a higher placement against these two competitors.

The honest counter-case

Claude did not win every trade-off.

ChatGPT produced the best immediately auditable workflow. Its advanced prompt explicitly requires human-verification steps, completion criteria, saved outputs, source fields, last-updated fields, portability checks, and an unresolved-items register. Claude’s playbook is more insightful and adaptive, but ChatGPT’s is closer to something an operations-minded reader could convert directly into a checklist or project document.

ChatGPT is also more concise than Claude while retaining substantial depth. Claude’s voice makes the post enjoyable, but the full article demands a long reading session, and some readers may abandon it before reaching the advanced prompt.

Gemini’s greatest strength is accessibility. Its beginner prompt is the easiest of the three to paste into a chat without preparation, and its advanced prompt communicates the essential idea of preserving reusable prompts with bracketed placeholders in a compact form. A reader intimidated by the larger systems in the competing posts might actually begin with Gemini’s version.

The problem is that ease of entry comes at the cost of safeguards and learning. Gemini gives the reader fewer intake questions, fewer explicit outputs, no correction stage, no stress test, and little support for separating durable preferences from one-trip circumstances.

What readers should take away

This week shows the difference between asking an AI for a document and designing a process that reliably produces one.

Gemini mostly specifies the desired destination: compare the budget, organize the claim, build the template. ChatGPT specifies the components and controls of the finished system. Claude goes one step further and specifies the conversation that must happen before the system can be trusted.

For one-time tasks, a concise destination-focused prompt may be enough. For reusable systems, the higher-value instructions are the less glamorous ones:

  • collect evidence before analysis; - stop for user correction between stages; - label inference separately from supplied facts; - preserve unknowns as visible blanks; - define what must not be decided yet; - require a stress test; - save provenance, versions, and change history.

That is the series’ real landing point. The best prompt is not merely one that generates a strong answer today. It is one that leaves behind a method the reader can inspect, correct, save, and run again.

TAGS:

Previous
Previous

Landing the Plane: Budget Post-Mortems and the Reusable System

Next
Next

Eight Weeks, One Reusable System: Landing the Plane