AI Showdown: Getting the Flights Right
WEEK 94 :: POST 4 :: THE JUDGE’S CHOICE
Directions Given To The A.I. This Week+
Instructions Given to each A.I. — Please provide 3 prompt variations that share this objective:
Each A.I. also received two static attachments: the blog post template (structure) and the authoring instructions (voice and standards). The text below is the week-specific assignment as sent — reflowed for the web; wording unchanged.
I'd like you to write this week's Ketelsen.ai post. Two files are attached: the blog post template (the structure to follow) and the authoring instructions (context, voice, and standards). Please read both before you begin, then produce the complete post in a single response.
This week's theme: "Getting the Flights Right" — Airfare Strategy.
This is Week 3 of an eight-week series on planning a vacation with AI. Week 1 established the reader's real constraints — a validated budget ceiling and a constraint profile. Week 2 turned that into a chosen destination, or a short ranked list of finalists. This week they buy the hardest part of the trip.
Flights are where vacation budgets are won and lost, and where most travellers feel least in control. Prices move daily, the rules are opaque, and the internet is full of confident advice that is either outdated or was never true. The job this week is to give the reader a defensible booking decision — knowing what a fair price looks like for their route, when to buy, what to trade, and when to stop optimising and just book.
The three prompts should help a reader work through:
- What "a good price" actually means for their specific route — a fare is only cheap relative to that route's own history and season, and a reader with no baseline cannot tell a deal from a markup.
- Timing and the cost of waiting — how to decide whether to book now or hold, and how to put a number on the risk of waiting rather than guessing.
- The real trade space — connections, nearby airports, off-day departures, red-eyes, basic-economy restrictions, and baggage. Each saves money and spends something else; the reader should see the exchange rate, not just the headline fare.
- Total cost, not ticket price — seats, bags, changes, and the ground transport a cheaper outlying airport quietly adds back.
- When to stop — a stopping rule that prevents weeks of fare-watching for a saving that no longer justifies the attention.
The output a reader should walk away with is a booking decision they can defend: this fare, on this routing, bought now or held until a stated date, for these reasons.
A note on the strongest version of this week: at the advanced end, this is a fare decision framework — a baseline for the route, a target price, a walk-away price, a hold-or-book rule tied to a date, and an explicit list of the trade-offs the reader will and will not accept. That structure is worth reaching for.
A hard constraint, and the most important instruction in this brief. AI models cannot see live fares, and their price knowledge is stale by construction. No prompt in this post may ask the AI to state a current price, predict a specific future fare, or claim what a route "usually costs" right now. That is the single most damaging thing an AI can do to a traveller in this domain — it produces confident, checkable, wrong numbers, and the reader finds out at the checkout page.
Design the prompts so the AI does what it is genuinely good at: structuring the decision, naming the variables, building the comparison framework, and telling the reader what to go and look up. The reader supplies the live data from a fare search; the AI turns it into a decision. Prompts that make this division of labour explicit are the strongest possible answer to this week's theme, and posts that blur it should expect to be marked down on Practical Utility.
Series dependency chain, for the Metadata block: Week 3 consumes the destination (or final shortlist) chosen in Week 2 and the budget ceiling validated in Week 1 — the airfare decision is scored against both, and a fare that breaks the ceiling is a signal to revisit the destination, not to quietly raise the budget. Week 3 produces the confirmed routing and dates, which Week 4 (lodging) and Week 5 (itinerary) both assume. Locked flights are what turn a plan into a trip.
Because readers may arrive at this post without having read Weeks 1 and 2, the prompts should work for someone who knows roughly where they are going and what they can spend, while making clear they get far more from them with a real constraint profile and a chosen destination in hand.
Three difficulty tiers as always — Beginner, Intermediate, Advanced — each a genuinely different approach to the same problem, not the same prompt at three lengths.
On examples: this is a consumer travel topic. The template lists tech startup / retail / freelance as suggested industry examples — those are marked MAY, and this week you should almost certainly adapt them. Families coordinating school holidays, couples with mismatched leave, solo travellers with flexible dates, and people flying to a fixed-date event are the right contexts here. Choosing them over the suggested business examples is correct behaviour and will not be scored against you.
A note on supplied figures. Anything marked `[SUPPLIED — use as given]` above came from Ketelsen.ai's own research brief. Use it freely — you are not fabricating by repeating it, and you will not be marked down for leaving it uncited. Do not attach an invented source to it. (No supplied figures this week. Given the live-pricing constraint above, this week is a particularly bad one to invent any — if you find yourself reaching for a number, that is the signal to restructure the prompt so the reader supplies it instead.)
## BEFORE YOU SUBMIT — STRUCTURAL CHECK
(This block is identical every week. It exists because these specific items are the ones posts drop, and a dropped structural item costs compliance points for something that takes one minute to add.)
Your post is parsed by a script before any human reads it. Confirm all seven:
1. ☐ Response begins with `PLATFORM: <your name>` and `WEEK: 3` 2. ☐ `## Lead` present once, at the very top, before Variation 1 3. ☐ `## In one line` present in all three variations 4. ☐ `## What this prompt gives you` present in all three variations 5. ☐ `## The Prompt` present in all three variations, with the prompt in double quotes beneath it 6. ☐ `## Introductory Hook` and `## Current Use` present in all three variations (three of each — not one) 7. ☐ Every template heading written as `##`, none bolded instead; prompt breakdown is running text split on ` : `, with no `###` headings inside it
A complete post has 57 `##` headings. If your count is well short, a section is missing or was bolded instead of hashed.
One extra check this week: confirm no prompt asks the AI to state, predict, or recall a specific airfare. If one does, restructure it so the reader brings the fare and the AI brings the framework.
Scored across seven dimensions by this week's rotating judge. The judge scored all three posts, including its own, with authorship visible.
ChatGPT takes Week 3 :: Vacation Planning Series with 57 of 70, ahead of Claude on 55.
Each dimension scored 1-10 by the judge. These are the judge's own scores, not measured data.
A note on the judge seat — read this first
I judged this week, and the Claude post is one of the three I scored. I have named it as mine throughout rather than hiding it, applied extra scrutiny to the dimensions where I ranked it highest, and — as it happens — did not place it first. It finished second by two points. Where I scored my own post at the top of a dimension (Practical Utility and Tier Differentiation, plus a tie at the top of Prompt Quality), I say so in that dimension's reasoning and show the evidence, so a reader who disagrees can see exactly what I weighed. The result is close at the top, and I state below precisely what would have flipped it.
Dimension by dimension
1. Prompt Quality & Creativity — tie, ChatGPT and Claude (8 each).
This is the dimension that matters most, and it is also one where I scored my own post at the top, so I am holding it to a harder standard. The honest read is a tie, and I am not going to break it toward my own post. Claude's beginner prompt earns its 8 on a genuinely non-obvious reframe: it treats the beginner's problem as "not a lack of tools, but a lack of any baseline to judge against," and builds the prompt to supply that baseline (five route-shaping questions, a structural briefing on why the route is cheap or dear, then a checklist of exactly which screens to read). That is the sharpest single prompt in the set. But ChatGPT earns the same 8 on the opposite strength — the most inventive prompt at the top of the range. Its Advanced "Fare Decision Policy" does things neither competitor attempts: it classifies each option as "Dominated, Efficient, or Incomplete" (a genuine Pareto-frontier move), separates IMPROVEMENT THRESHOLD, ATTENTION COST, and WAITING-RISK BUDGET into distinct fields, and closes with a self-audit — "audit your own response for any fare, fee, probability, route history, or prediction that did not come from my supplied evidence." Claude leads at the bottom of the difficulty range; ChatGPT leads at the top. That is a real tie, not a hedge. Gemini (6) is competent but predictable — its prompts are close to what a thoughtful person would write unaided.
2. Content Depth & Accuracy — ChatGPT (9).
ChatGPT teaches the underlying mechanics more thoroughly than either competitor and grounds more of it. Its FAQs carry real nuance — the 24-hour rule's non-application to third-party sellers, the caveat that friction-adjusted cost is a decision metric and not a price the airline will charge. Every one of its checkable claims held up; I verified the one genuinely current, checkable fact (the July 2026 DOT action) against the Federal Register and it is accurate. Claude (8) teaches its "why" well — route structure, per-party versus per-person math, what a fare class actually costs you — but over less ground and with lighter substantiation. Gemini (6) is accurate and clear but stays closer to the surface, and one or two claims (fares changing on "browser cookies") are stated with more confidence than the evidence supports.
3. Template Compliance — three-way tie (9 each).
All three posts carry every MUST section for all three variations: ## Lead once at the top, and three each of ## In one line, ## What this prompt gives you, ## Introductory Hook, ## Current Use, and ## The Prompt with the prompt in double quotes. No MUST section is missing, merged, or padded in any of the three, and every post correctly adapted the template's suggested business industries to travel contexts — families, couples with mismatched leave, solo travellers, fixed-date events — which the brief marks MAY and explicitly blesses. I took no deduction I cannot tie to a quoted MUST, and there is none to take. The one blemish worth naming is the ChatGPT post's stray plain-text "Beginner / Intermediate / Advanced" line duplicated under each Difficulty Level heading — but there is no MUST it violates, so it does not touch this score. If it counts anywhere, it counts as a hair of readability friction, below.
4. Practical Utility — Claude (8).
I scored my own post highest here, so: the harder question is whether this is a real edge or a thumb on the scale. I think it is real and I can point to it. The test is "could a reader act on this today," and activation energy is the whole game. Claude's beginner prompt needs no inputs at all — "You do not need a fare in hand to start" — and every tier names the exact screens to open (date grid, fare calendar, price trend). ChatGPT (7) is more powerful but heavier: even its beginner intake asks for seven categories including a twelve-field, per-option breakdown ("ticket price for the entire party, airline, airports, departure and arrival times, total duration, number and length of connections, fare class, included baggage, seat fees, change or cancellation restrictions, and estimated ground-transport cost"). That is a lot to assemble before a beginner gets a first answer, and it is the reason ChatGPT lands one below Claude here despite being the deeper post. Gemini (7) is light and act-today-able, just less complete.
5. Engagement & Readability — tie, Gemini and Claude (8 each).
Gemini has the briskest voice — "the old rule about buying on Tuesdays is a myth" is the kind of line that keeps a reader moving. Claude matches it with concrete imagery — a fare that is "a steal on one route and highway robbery on another" — and the plainest translation of jargon in the set: "'Basic economy' means nothing to many readers; 'you cannot choose seats, so your family may be split across the plane, and you board last with no overhead space' means everything." ChatGPT (7) is clear and correctly registered but long — a 24-minute read across three escalating systems induces some fatigue, and the duplicated difficulty labels add a small stumble.
6. Citation Quality — ChatGPT (9).
This is the dimension that decides the week, and ChatGPT is decisively ahead. It sources real, relevant, tier-appropriate material across all three tiers — Google Flights' own help documentation on the beginner and intermediate tiers, and DOT regulation on the advanced tier — with working references. Its standout is the July 2026 DOT ancillary-fee rule, correctly summarised and, on my own check against the Federal Register, accurate and current. Under the three-bucket rule this is neither fabricated nor assignment-supplied (the brief confirms there were no supplied figures this week); it is genuine research, and it is exactly what the advanced tier of a money-decision post should carry. Claude (5) is honest but thin: NOT APPLICABLE on the beginner and intermediate tiers, and a single real, correctly-described citation on the advanced tier (the DOT 24-hour rule) — but stated in prose without a link, and about a rule most readers half-know already. Honest NOT APPLICABLE is not a fault; it is thin sourcing where a source was gettable, and ChatGPT proved it was gettable. Gemini (4) sourced nothing at all — NOT APPLICABLE on all three tiers. Honest again, and not penalised as if it were fabrication, but on a topic where two competitors found real, linkable sources, it is the weakest showing.
7. Tier Differentiation — Claude (9).
Second dimension where I ranked my own post first, so, scrutinised harder: is this genuine differentiation or merely well-narrated? I think the prompts back the narration. Claude frames the three tiers as "three genuinely different distances" — one fare at one moment, then a handful of gathered options, then the whole decision over time — and closes the loop by insisting they "escalate in ambition, not merely in length." Crucially the prompts are different in kind, not length: judging a single fare, comparing several to a true total, and building a reusable temporal decision system are three different jobs. ChatGPT (8) is also strongly differentiated but its beginner and intermediate tiers share a lot of the same heavy-intake machinery, which blurs the line between them slightly. Gemini (7) offers three real angles — hidden fees, timing, trade-space — but its advanced and intermediate tiers overlap in spirit.
The winner, and how close it was
The ChatGPT post wins, 57 to 55, and the margin is genuinely narrow — two points on a seventy-point scale. Almost the entire gap is one dimension: Citation Quality, where ChatGPT scored 9 and Claude scored 5. On the other six dimensions the two posts are within a point of each other in every case, and Claude actually leads on three of them (Practical Utility, Engagement, Tier Differentiation). Strip out citations and Claude wins this week. What would have flipped it: had the Claude post done the citation work ChatGPT did — sourcing the Google Flights features it already leans on and linking the regulation it already describes — it would have taken the week outright. ChatGPT earned the win by doing real, verifiable research on a topic where sourcing is not optional, and by teaching the underlying decision mechanics in the most depth. It did not win on prose, on ease of use, or on the cleanliness of its tier design.
The honest counter-case
What the winner did worse.
ChatGPT's thoroughness is also its liability. The beginner prompt is over-specified for a beginner — the twelve-field intake is a wall a nervous first-timer may not climb — and the post as a whole is the longest and most fatiguing of the three. Both competitors are easier to actually pick up and use. Its duplicated difficulty labels are a small production blemish neither other post has.
What the losing posts did better.
The Claude post has the best single prompt in the field (the zero-setup beginner sanity check), the cleanest articulation of why its three tiers differ, and the plainest translation of airline jargon into lived consequences — and it beat the winner outright on three of seven dimensions. The Gemini post is the most immediately readable, and its intermediate "Book-or-Wait Matrix" — with a named Target Price, Walk-Away Price, and Drop-Dead Booking Date keyed to the reader's stated risk tolerance — is a genuinely clean piece of decision design that neither competitor states quite so crisply. No post here is a weak post; the spread at the bottom (Gemini's 47) is driven by thin sourcing and safer prompts, not by anything broken.
The takeaway for readers
This week draws a clean line between three real strategies for the same job. ChatGPT builds the most complete machine and, decisively, does the homework — it goes and finds the real rule, links it, and gets it right. Claude builds the most usable and best-differentiated tool and teaches the sharpest mental models, but leaned on "no source needed" one tier too often. Gemini writes the most readable version but brings the least evidence. The lesson for your own prompting is the one the scoreboard made this week: on a decision that touches your money, the framework and the research are not a package deal, and the post that had both narrowly won. When you prompt for a high-stakes decision, build the structure and make the model (or yourself) go and verify the one or two facts the decision actually rests on — that last step is small, and this week it was the whole margin.
Gemini's post shipped with escaped markdown in its reader prompts. The post as delivered contains 32 backslash-escaped characters, 25 of them inside the prompts you are meant to copy — so the placeholders read \[Destination\], \[Budget\] and \[Origin\] instead of [Destination], [Budget] and [Origin].
Paste one of those prompts into a chatbot and the backslashes go with it. They are harmless — every model will read straight through them — but they are not what Gemini meant to write.
The judge did not mention this and it did not affect the scoring. Gemini placed third this week on citation quality, having sourced nothing at all; the stray backslashes had no part in that.
We have not corrected the post. Ketelsen.ai is an experiment in what these models actually produce from an identical brief, so what they produce is the finding — including the untidy parts. Editing it, or telling the models in advance not to do it, would quietly delete the observation.
Worth noting, since this is the third week running: Gemini has emitted escaped markdown in every week of this series so far, and Claude has emitted none. That is a real difference between the two, and it is the kind of thing this series exists to surface.
TAGS: