Ketelsen.ai Weekly Comparison: Vacation Feasibility and Budget Architecture
WEEK 92 :: POST 4 :: THE JUDGE’S CHOICE
Directions Given To The A.I. This Week+
Instructions Given to each A.I. — Please provide 3 prompt variations that share this objective:
Each A.I. also received two static attachments: the blog post template (structure) and the authoring instructions (voice and standards). The text below is the week-specific assignment as sent — reflowed for the web; wording unchanged.
I'd like you to write this week's Ketelsen.ai post. Two files are attached: the blog post template (the structure to follow) and the authoring instructions (context, voice, and standards). Please read both before you begin, then produce the complete post in a single response.
This week's theme: "Should We Even Take This Trip?" — Feasibility and Budget Architecture.
This is Week 1 of an eight-week series on planning a vacation with AI. It is the series entry point, and its job is to stop the reader from doing what almost everyone does: picking a destination first and backing into the money afterward. Before any destination is chosen, the reader locks down their real constraints.
The three prompts you write should help a reader establish:
- A true budget ceiling, including the hidden-cost multiplier. People routinely underestimate total trip cost by 20-40% `[SUPPLIED — use as given]` once food, local transit, fees, tips, and "well, we're already here" spending are counted. A prompt that only asks about flights and hotels has already failed the reader.
- Available dates and PTO math — what time they can actually take, and what it costs them to take it.
- Traveler composition — kids, mobility needs, group size, solo travel, and how each changes the constraint set.
- Trip purpose — rest versus adventure versus culture versus a milestone celebration. Purpose determines what "worth it" means, and readers often cannot articulate it until asked directly.
The output a reader should walk away with is a validated trip budget ceiling and a constraint profile — a short, concrete artifact they can reuse. Every later week in the series references it.
Series dependency chain, for the Metadata block: Week 1 is the entry point and has no upstream prerequisite. Its budget ceiling and constraint profile feed Weeks 2 through 4 (destination shortlisting, airfare strategy, lodging), and its budget ceiling is referenced again in Week 7 for daily in-trip spend tracking.
A useful framing, if it helps: this is the vacation equivalent of asking "should I buy a car right now?" before walking onto a lot. The reader is being taught to define constraints before falling in love with options.
Three difficulty tiers as always — Beginner, Intermediate, Advanced — each a genuinely different approach to the same problem, not the same prompt at three lengths.
A note on supplied figures. Anything marked `[SUPPLIED — use as given]` above came from Ketelsen.ai's own research brief. Use it freely — you are not fabricating by repeating it, and you will not be marked down for leaving it uncited. Do not attach an invented source to it. Presenting it as a sensible planning rule of thumb rather than a precise measured statistic is the right treatment.
This comparison was re-scored after the original judging. The first version gave Post A a disqualifying 1 on Citation Quality for the "20–40% budget overrun" figure, calling it an invented statistic. That was wrong, and the fault was ours: Ketelsen.ai supplied that figure to all three AIs in the week's assignment brief, and the judge had not been shown the brief. It had no way to tell a supplied figure from an invented one.
The correction was made by the judge, not by us. ChatGPT re-scored Citation Quality from 1 to 3, which raised Post A's total from 36 to 38 out of 70, and rewrote that section in its own words — that rewrite is what appears below. This affected the score but not the outcome: Post C remains the winner. The scoring table and reasoning you are reading carry the judge's corrected figures throughout.
The pipeline defect has been fixed permanently: the assignment brief is now a required attachment to the judge, so a supplied figure is checkable rather than guessed at. That change is why no later week has repeated this.
Scored across seven dimensions by this week's rotating judge. Entries were reviewed blind — labelled only A, B and C.
Claude takes Week 1 :: Vacation Planning Series with 66 of 70, ahead of ChatGPT on 58.
Each dimension scored 1-10 by the judge. These are the judge's own scores, not measured data.
Fairness Note
I treated the three attachments as anonymous entries and judged only the writing on the page. I do not have a reliable basis to identify any post as my own, so I did not speculate about authorship.
For reference:
- Post A: “The Hard Boundary Blueprint,” “The Parameterized Constraint Engine,” and “The Multi-Variable Travel Constraint Matrix”
- Post B: “The Vacation Reality Check,” “The Constraint Matrix Builder,” and “The Trip Go/No-Go Decision Memo”
- Post C: “The Reality-Check Budget,” “The Constraint Profile Interview,” and “The Feasibility Model”
A Note on This Week's Scoring
This comparison was re-scored after publication review. The original version gave Post A a disqualifying 1 on Citation Quality for the "20-40% budget overrun" figure, describing it as an invented statistic. That was wrong: Ketelsen.ai supplied that figure to all three AIs in the week's assignment brief, and the judge had not been shown the brief. The score was corrected to 3 — the post's sourcing is genuinely thin, but nothing was fabricated.
The correction was made by the judge, not by us, and it did not change the ranking. We've since added the assignment brief to what the judge receives, so provenance is checkable rather than guessed at.
Dimension-by-Dimension Reasoning
1. Prompt Quality & Creativity
Winner: Post C
Post C has the strongest prompt design because each variation does a distinct job and gives the AI a clear operating model. The Beginner prompt is simple but not shallow: it asks the AI to “list every cost category a trip like this actually has,” then add “a hidden-cost cushion of 20 to 40 percent,” and finally “ask me the two most important questions I probably haven’t thought about yet.” That last move is especially useful because it turns the model from a calculator into a blind-spot finder.
The Advanced prompt is the best prompt in the set. Its instruction to build “three total-cost scenarios — Lean, Expected, and Blowout” is more useful than a single padded estimate, and its requirement to identify “the single change that would most improve feasibility” gives the reader a practical lever instead of a vague warning.
Post B is very close here. Its Advanced prompt has excellent controls: “Do not invent current travel prices,” “label it as user-provided, inferred estimate, or missing,” and a “Trip Feasibility Memo” with stress testing. That is strong prompt engineering. The reason Post C edges it out is compression. Post B’s Advanced prompt is powerful but heavy; Post C delivers nearly the same decision architecture with less friction.
Post A has some useful ideas, especially the “Sourcing Logic Block,” which could feed downstream destination shortlisting. But too much of the prompt relies on inflated language: “elite travel systems engineer,” “Teleological Intent,” “Constraint Vector Profile,” and “financial firewall.” Those phrases sound advanced, but they do not always make the prompt more usable.
2. Content Depth & Accuracy
Winner: Post B
Post B wins depth. It explains the underlying planning problem with patience and specificity. Its Intermediate prompt does not merely say “budget carefully”; it asks for “a conservative total ceiling, a comfortable target, and a hard stop number,” plus “workdays, PTO days, unpaid days, travel days, and recovery days.” That is a substantive model of feasibility, not just travel advice.
Post B’s Advanced section is especially thorough. It includes “debt, savings, or cash-flow boundaries,” “emotional pressure,” a “walk-away number,” purpose-fit scoring, and five explicit failure modes, including “PTO does not actually cover the trip” and “one traveler’s needs dominate the plan.” That is the deepest treatment of the actual decision.
Post C is also accurate and strong, especially where it explains why scenario ranges matter: “A single number implies a false precision no travel budget deserves.” Its content is tighter and often more memorable than Post B’s, but Post B covers more of the decision surface.
Post A is the weakest here because it makes confident claims without enough support. The phrase “20-40% budget overruns most travelers experience” is presented as a fact but not sourced. The Advanced explanation also calls its 30% buffer “aggressive, data-driven” without providing data. That weakens trust.
3. Template Compliance
Winner: Post B
Post B is the most complete template execution. Each tier includes the expected sections: hook, current use, prompt, breakdown, practical examples, creative uses, adaptability, pro tips, prerequisites, tags, tools, FAQ, follow-up prompts, citations, comparison, and metadata. It also gives multiple practical examples per tier and keeps the sections clearly separated.
Post C is also very compliant. It has all the major sections and a strong final comparison. Its structure is cleaner than Post B’s, but Post B is more exhaustive.
Post A includes the required sections, but the execution is thinner. Its “Practical Examples from Different Industries” sections generally include only two examples per variation, and one line contains a visible text corruption: “ปัจจัยs.” That kind of artifact matters in a publishing workflow because it signals that the post still needs editorial cleanup.
4. Practical Utility
Winner: Post C
Post C is the easiest post for a reader to use today. It turns fuzzy vacation planning into concrete actions: build a real total, ask what you missed, run an interview, then use a feasibility model when the stakes are higher. The Advanced example about the freelance consultant is particularly useful: “Two weeks off means roughly $9,000 in un-billed work,” which makes the “true cost” closer to $14,000 than the visible $5,000 budget. That is exactly the kind of reframing readers need.
Post B is also highly practical. Its “Go, Modify, Delay, or Do Not Take This Trip Yet” recommendation is clear, and its “Budget Ceiling,” “PTO and Time Cost,” and “Traveler Constraint Matrix” sections would produce a durable artifact. The only drawback is that the prompts can feel more like a full planning system than something a casual reader would actually run in one sitting.
Post A has practical elements, especially the basic “Travel Constraint Profile.” But the Advanced tier asks for “net financial liquid net worth allocated for leisure” and “an honest assessment of your current psychological state,” which may feel unnecessarily intrusive or formal for vacation planning. Its practical value is real, but less reader-friendly.
5. Engagement & Readability
Winner: Post C
Post C is the clear readability winner. It has the best voice: sharp, direct, and memorable without becoming snide. The Beginner hook lands immediately: “Get it wrong and you’re financing regret at 24% APR.” That line does more than entertain; it explains why the prompt matters.
The prompt breakdowns are also clean and teachable. “An unfenced helper is an over-eager helper” is a compact way to explain why scope control matters. “Garbage in isn’t just garbage out — it’s confident garbage out” is another strong line because it names an AI failure mode in plain language.
Post B is readable and humane, but it is long. Its tone is accessible, and lines like “AI can produce destination lists, hotel comparisons, sample itineraries, and packing guides in seconds, but speed is dangerous when the first assumption is wrong” are strong. Still, the density would test some readers.
Post A is the weakest on voice. It leans heavily on AI-writing tells: “bulletproof,” “elite,” “forensic,” “immutable,” “programmatic,” “friction vectors,” and “teleological focus.” Some of those terms may appeal to a systems-minded audience, but together they create distance from the reader.
6. Citation Quality
Winner: Post C
Post C is the strongest on citation quality because it makes a real sourcing effort and cites relevant conceptual sources. Its use of Kahneman’s Thinking, Fast and Slow for the planning fallacy and Flyvbjerg and Gardner’s How Big Things Get Done for reference-class forecasting supports the post’s underlying argument: people often underestimate cost, time, and complexity, so a better planning prompt should force buffers, scenarios, and outside-view thinking.
Post C is not perfect. It uses the “20–40%” hidden-cost cushion, but under the corrected review standard that figure is treated as assignment-supplied rather than independently claimed. No penalty is warranted for using it. The remaining weakness is that Post C could have strengthened its “current use” section with recent travel-industry sources on fees, dynamic pricing, or rising ancillary travel costs. Still, its citations are real, relevant, and tied to the post’s core reasoning.
Post B remains mid-low in this category. It lists NOT APPLICABLE for citations, which is not a disqualifying failure, but it is thin for a post making broad claims about AI planning behavior, hidden costs, feasibility risk, and travel decision-making. Its strongest citation-adjacent move is not a source but a prompt-design safeguard: “do not invent live prices” and “label them clearly as estimates.” That improves epistemic hygiene inside the prompt, but it does not replace external sourcing in the article itself.
Post A should not be treated as having fabricated the “20–40%” figure. That number was assignment-supplied, so the earlier fabrication finding was incorrect. The accurate criticism is thinner but still significant: Post A uses NOT APPLICABLE for citations across all three variations while making several confident claims about travel inflation, dynamic pricing, hidden-cost overruns, and the reliability of specific multipliers. Those claims may be directionally reasonable, but the post does not show readers where the confidence comes from. So Post A earns a low score for weak sourcing, not a disqualifying score for invention.
7. Tier Differentiation
Winner: Post C
Post C has the cleanest Beginner / Intermediate / Advanced ladder.
The Beginner prompt is a “Reality-Check Budget”: quick, low-friction, and focused on exposing the real total. The Intermediate prompt is a “Constraint Profile Interview”: slower, conversational, and designed to produce a reusable artifact. The Advanced prompt is a “Feasibility Model”: scenario-based, opportunity-cost-aware, and verdict-driven.
That is exactly what tiering should do. The tiers are not just longer versions of each other; they represent different operating modes.
Post B also differentiates well. Beginner is a guided intake, Intermediate is a constraint matrix, and Advanced is a decision memo. Its tiering is strong, but less elegant than Post C’s because the Intermediate and Advanced versions both become fairly large planning artifacts.
Post A differentiates the tiers more than a weak post would. The Beginner tier asks four basic questions, the Intermediate tier adds modules and a 1.25x multiplier, and the Advanced tier introduces phases and a “Sourcing Logic Block.” But some of the differentiation comes from added jargon and higher multipliers rather than genuinely different user workflows.
Winner: Post C
Post C is this week’s winner with 66 out of 70.
What separated it was not just polish. Post C had the strongest combination of prompt quality, reader usability, tier differentiation, and voice. It understood the core Week 1 idea — do not fall in love with a destination before defining feasibility — and turned that idea into three prompts that readers could actually use.
Its biggest strength is that every tier has a distinct job:
- Beginner: expose the real cost.
- Intermediate: build a reusable constraint profile.
- Advanced: model feasibility with scenarios, opportunity cost, worth-it criteria, and a clear verdict.
The Advanced prompt is especially strong because it avoids false precision. “Lean, Expected, and Blowout” is a better structure than one padded budget number. The “single change that would most improve feasibility” is also a standout feature because it gives the reader a next move, not just a score.
Post B was a serious competitor. It is deeper in places, and its Advanced “Trip Go/No-Go Decision Memo” may be the most complete individual framework. But Post C is more publishable, more memorable, and more likely to be used by a real reader without fatigue.
The Honest Counter-Case
The case against Post C is that it is not always the most exhaustive. Post B’s Advanced prompt captures more edge cases: debt boundaries, emotional pressure, provenance labels, walk-away numbers, stress tests, and traveler risk maps. A reader planning an expensive, complicated, group-heavy trip might get a more complete audit from Post B.
Post B also deserves credit for its seriousness. “If any number is uncertain, label it as user-provided, inferred estimate, or missing” is one of the best instructions in the whole set. It protects against fake precision and teaches a broader AI principle: separate known facts from assumptions.
Post A’s best contribution is the “Sourcing Logic Block.” That idea is genuinely useful for an eight-week travel-planning series because it creates a bridge from Week 1 feasibility into later destination filtering. Post A also has a clear structural progression from basic interview to parameterized profile to advanced matrix. Its problem is not that it lacks ideas; it is that the execution wraps those ideas in too much inflated language and too little sourcing.
Every post had at least one real strength:
- Post A had the strongest downstream-series idea with the “Sourcing Logic Block.”
- Post B had the deepest and most complete decision framework.
- Post C had the best balance of usefulness, clarity, voice, and prompt architecture.
Takeaway for Readers
This week’s comparison shows that the best AI prompt is not always the longest or most technical. The strongest prompt is the one that gives the model the right job, the right sequence, and the right stopping point.
Post A shows the risk of confusing advanced language with advanced usefulness. Post B shows the power of a complete framework, but also the cost of asking readers to carry too much structure. Post C shows the sweet spot: clear constraints, visible assumptions, scenario thinking, and a practical next action.
The bigger prompting lesson is simple: before asking AI to optimize a decision, ask it to define the boundaries of the decision. In vacation planning, that means budget, time, travelers, purpose, and the point where the trip stops being worth it. Once those are clear, every later AI task gets easier — and much less likely to produce a beautiful plan for the wrong trip.
Gemini's post shipped with two mechanical flaws, and we have published it exactly as delivered.
A corrupted word. One sentence reads "a strict budget ceiling thatปัจจัยs in the cost of lost billable hours". The intended word was almost certainly "factors"; what arrived was the Thai word for factor spliced into the middle of an English sentence. This is a token-level generation fault — the kind of seam that normally gets edited out before anyone sees it.
Escaped markdown in the reader prompts. 10 backslash-escaped characters, 8 of them inside the prompts you are meant to copy, so placeholders read \[like this\] rather than [like this].
Neither affected the scoring — the judge did not mention either one, and Post A's score was decided on sourcing.
We have not corrected any of it, though we very nearly did: an edited copy of this post was prepared in July 2026 with the corrupted word silently repaired, and it was never published once we settled the rule. Ketelsen.ai is an experiment in what these models actually produce from an identical brief. A generation fault is a real result. Cleaning it up would have made the models look more reliable than they are, which is the one thing this site must not do.
ChatGPT's post shipped with escaped markdown in its reader prompts: 10 backslash-escaped characters, all 10 inside the prompts you are meant to copy, so placeholders read \[like this\] instead of [like this].
Paste one of those prompts into a chatbot and the backslashes go with it. They are harmless — every model reads straight through them — but they are not what ChatGPT meant to write.
The judge never mentioned this and the scoring was not affected by it.
We have not corrected the post. Ketelsen.ai is an experiment in what these models actually produce from an identical brief, so what they produce is the finding, including the untidy parts.
TAGS: