AI Showdown: Three Takes on “Should I Even Be Job Searching?”
WEEK 100 :: POST 4 :: THE JUDGE’S CHOICE
Directions Given To The A.I. This Week+
Instructions Given to each A.I. — Please provide 3 prompt variations that share this objective:
Each A.I. also received two static attachments: the blog post template (structure) and the authoring instructions (voice and standards). The text below is the week-specific assignment as sent — reflowed for the web; wording unchanged.
I'd like you to write this week's Ketelsen.ai post. Two files are attached: the blog post template (the structure to follow) and the authoring instructions (context, voice, and standards). Please read both before you begin, then produce the complete post in a single response.
This week's theme: "Should I Even Be Job Searching Right Now?"
This is Week 1 of a new eight-week series on running a job search with AI — the follow-up to "AI at the Dealership," and it shares that series' backbone: a high-stakes, emotionally loaded decision where the other side of the table (employers, recruiters, applicant tracking systems) holds most of the cards. Week 1 starts where honest job searches start: before the résumé. Most people begin applying before they have defined what is actually broken about their current situation, and impulsive moves are how people land in roles they regret within months. The reader's job this week is to decide — stay, grow in place, or go — with evidence instead of a mood.
One thing this week must do that no later week has to: speak to all three readers who arrive at a job-search series. The active-but-employed reader wondering if the grass is greener, the reader who has just been laid off and has no stay option, and the career pivoter for whom "go" means a different field entirely. And name the elephant with care: for some laid-off readers, AI itself is part of the story — restructurings framed around AI efficiency are part of this market. A series about using AI must not be breezy about that; one honest, warm sentence acknowledging it (no layoff statistics — the constraint below already forbids them) buys more trust than a page of enthusiasm, and the laid-off reader should feel seen, not lectured.
Week 1 also opens the series, so its Lead carries the series' editorial frame. The Lead should say plainly what this series is and is not: AI as the reader's private analyst, coach, and thinking partner — never their ghostwriter — because hiring decisions are made by humans who are rightly wary of machine-written material, and employers themselves face legal and compliance limits on AI in hiring. The two lines the series stands on — "Use AI like an analyst, not a ghostwriter" and "AI behind the scenes. You on the page." — belong in or near this Lead, worked into the post's own voice rather than dropped in as slogans.
The prompts should handle the three audiences through a "MY SITUATION" context block the reader fills in — the same pattern the car series used — so one prompt serves all three without pretending they are the same person.
The deliverable the reader should walk away holding: a written stay-or-go decision, a personal compensation baseline (what they earn now, fully loaded), and a timeline — the three artifacts every later week in this series builds on.
THE SERIES CONTRACT — identical every week; it binds every prompt you design. This series' tagline is its editorial contract: "Use AI like an analyst, not a ghostwriter." It is written for a reader in a market unsettled by AI itself — some readers are searching precisely because AI eliminated their last role. Write with that reader at the table: no AI-efficiency cheerleading, no automation jokes, no promises that AI will "do it for you" anywhere a human hiring decision is involved. And hold one line in every prompt: the AI is the reader's private analyst, coach, and sparring partner — it structures, researches, rehearses, and questions. It does not ghostwrite. Anything a hiring human will read or hear — résumé lines, cover letters, outreach messages, interview answers, negotiation emails, resignation letters — must end in the reader's own words and be true. Prompts should drive toward drafts the reader rewrites and owns, and should say so explicitly. Employers increasingly restrict how AI may be used in their own hiring decisions for legal and compliance reasons, and recruiters increasingly recognize — and discard — material that reads machine-written. A prompt that makes a reader look AI-generated hurts them twice. Posts that ignore this contract should expect to lose the week. Two practical notes. First: a standing “About this series” notice covering these same points is added to every published post automatically at publication — acknowledge the frame in your own voice where your week's prompt calls for it, but do not write a formal disclaimer block of your own, and do not open every post with the same acknowledgment paragraph: outside the weeks whose prompts explicitly carry the series frame, this contract lives in your tone and your prompt design. Second, the framing is POSITIVE: used this way — analyst backstage, reader on the page — AI is an advantage no hiring human will ever hold against your reader. Write like that is true, because it is.
The three prompts should help a reader:
- Run an honest career audit. A satisfaction diagnostic that separates "bad month" from "bad fit" — role, manager, growth, compensation, energy — and names what specifically is broken, so the reader is diagnosing before prescribing.
- Inventory their skills and market position. What the reader actually does all day, translated into the language the market hires for, with an honest read on which skills are appreciating and which are aging — the AI structuring the inventory from the reader's own history, not asserting market statistics.
- Build the full stay-vs-go decision framework. The "budget week" of the series: current compensation fully decoded (base, bonus, benefits, the things that quietly vanish on exit), savings runway if the search goes long, the benefits cliff of leaving mid-year, weighed inside a decision framework the reader owns — ending in a written decision and timeline.
At the advanced tier, the strongest version of this week is a decision memo the reader writes to themselves — situation, evidence, options scored, decision, revisit date — produced by a prompt that makes the AI a structured interviewer and analyst rather than an oracle. A decision the reader can re-read in six months beats a vibe either way.
A hard constraint, stated up front for the series. AI models cannot see live labor-market data, current layoff patterns, or real salary postings, and this series' citation standards treat unverified statistics as defects. No prompt may ask the AI to state current hiring trends, layoff figures, salary levels, or "shift shock" regret statistics as fact. Where market reality matters, the prompt should have the reader supply what they know or point them to named sources to check — pay-transparency postings, official labor statistics, their own industry contacts — not have the AI assert numbers. The same applies to the money: the runway analysis works on numbers the reader supplies, and nothing in these prompts is financial advice — the framework organizes the reader's own decision, it does not hand down a verdict.
Design the prompts so the AI does what it is genuinely good at: structured interviewing, translating a work history into market language, organizing a messy emotional decision into evidence and options. The reader supplies their situation and their numbers; the AI supplies structure, candor, and sequence. Posts whose prompts have the AI invent market statistics or deliver quit-your-job verdicts should expect to be marked down on Practical Utility and Content Accuracy.
Series dependency chain, for the Metadata block: Week 1 consumes nothing — it is the series opener. Week 1 produces the written stay-or-go decision, the compensation baseline, and the search timeline — consumed by Week 2 (defining the target), Week 7 (the baseline anchors the negotiation), and Week 8 (the decision criteria return in the final offer matrix).
Because readers may arrive at this post from anywhere, the prompts should work for someone starting cold — no prior artifacts exist yet in this series, so this is the one week with no catching-up to do, and the post can say so as an invitation.
Three difficulty tiers as always — Beginner, Intermediate, Advanced — each a genuinely different approach to the same problem, not the same prompt at three lengths.
On examples: this is a career topic that touches every industry. The template lists tech startup / retail / freelance as suggested industry examples — those are marked MAY, and adapting them is expected here. An employed product manager wondering whether restlessness is a signal, a laid-off retail manager with no stay option, and a freelance designer considering a return to full-time are the right kinds of contexts. Choosing them over the suggested business examples is correct behaviour and will not be scored against you.
A note on supplied figures. Anything marked `[SUPPLIED — use as given]` above came from Ketelsen.ai's own research brief. Use it freely — you are not fabricating by repeating it, and you will not be marked down for leaving it uncited. Do not attach an invented source to it. (No supplied figures this week — the outline's hook statistics are directional and unverified, so none are being handed to you. If you find yourself reaching for a regret percentage or a layoff figure, that is the signal to restructure the sentence so it does not need one.)
## BEFORE YOU SUBMIT — STRUCTURAL CHECK
(This block is identical every week. It exists because these specific items are the ones posts drop, and a dropped structural item costs compliance points for something that takes one minute to add.)
Your post is parsed by a script before any human reads it. Confirm all seven:
1. ☐ Response begins with `PLATFORM: <your name>` and `WEEK: 1` 2. ☐ `## Lead` present once, at the very top, before Variation 1 3. ☐ `## In one line` present in all three variations 4. ☐ `## What this prompt gives you` present in all three variations 5. ☐ `## The Prompt` present in all three variations, with the prompt in double quotes beneath it 6. ☐ `## Introductory Hook` and `## Current Use` present in all three variations (three of each — not one) 7. ☐ Every template heading written as `##`, none bolded instead; prompt breakdown is running text split on ` : `, with no `###` headings inside it
A complete post has 57 `##` headings. If your count is well short, a section is missing or was bolded instead of hashed.
One extra check this week: confirm no prompt asks the AI to state a current hiring trend, layoff figure, salary level, or regret statistic as fact, and none delivers a stay-or-quit verdict. The decision framework organizes the reader's own evidence; market numbers come from the reader or from named sources they check themselves.
About this series. AI is reshaping work — for some readers, it's part of why you're searching at all. We don't pretend otherwise. And the hiring world is wary of AI-written material: many employers restrict how AI may be used in their own hiring decisions, and recruiters increasingly recognize — and discard — machine-written applications. So this series teaches a different approach: use AI as your private research analyst, interview coach, and thinking partner, while every word an employer sees or hears from you stays genuinely, verifiably yours. AI behind the scenes. You on the page.
Scored across seven dimensions by this week's rotating judge. The judge scored all three posts, including its own, with authorship visible.
Claude takes Week 1 :: Job Search Series with 68 of 70, ahead of ChatGPT on 61.
Each dimension scored 1-10 by the judge. These are the judge's own scores, not measured data.
Scorecard
Dimension by dimension
1. Prompt Quality & Creativity — Claude wins (10).
Claude found the cleanest mechanism for each level instead of treating difficulty as a word-count setting.
Its Beginner prompt converts a messy complaint into a simple classification exercise: each concern becomes “Episodic or Structural,” with “Unclear” available when the evidence is insufficient. That third category matters because it gives the model permission to ask instead of pretending to know.
The Intermediate prompt is equally well engineered. Every capability must include the plain-English work, the market phrase, and “the specific evidence from my own material that supports it.” Weak claims receive the visible label “Thin evidence.” That is a particularly strong anti-inflation mechanism.
Its Advanced prompt then changes modes entirely. The reader approves the criteria, assigns weights totaling 100, receives a red-team argument against the leading option, and writes the final decision personally. The designs are inventive without feeling decorative.
ChatGPT’s prompts are more exhaustive. Its Advanced version includes an Evidence Register, financial scenarios, sensitivity testing, a pre-mortem, and a 90-day plan. That is excellent prompt engineering, but the Beginner prompt already asks ten interview questions and requests a diagnosis, directional leaning, evidence analysis, compensation snapshot, action plan, and decision statement. The amount of machinery slightly blurs the intended entry-level experience.
Gemini’s three prompts are usable but predictable: a five-pillar diagnostic, a categorized skill inventory, and a three-step decision memo. They lack the uncertainty handling, evidence controls, adversarial testing, and human decision checkpoints that distinguish the other two entries.
2. Content Depth & Accuracy — ChatGPT wins (10).
Because ChatGPT is the judge’s own post, I applied extra scrutiny before giving it the only 10 in this dimension. The score rests on controls visible in the text, not on stylistic familiarity.
The Advanced prompt requires the model to distinguish “facts I provide, calculations from my figures, interpretations, assumptions, and unknowns.” That epistemic separation runs through the Evidence Register, the compensation analysis, the scoring process, and the final memo.
It also requires three user-defined financial scenarios, transparent calculations, sensitivity tests, and a privacy field for “Information I do not want shared or processed.” Missing values stay unresolved rather than being converted into convenient estimates. The result is the most complete treatment of uncertainty and financial decision risk in the set.
Claude is close. Its explanations of taxonomy, evidence attachment, turn-taking, uncertainty labels, and human-assigned weights are unusually strong. A few statements, however, are presented more universally than the post establishes—for example, that models comply better when constraints include a justification, or that informal demand is the cleanest signal of capability. These are plausible prompting principles, but they are taught as settled rules rather than framed as useful heuristics.
Gemini’s main accuracy problem appears in its Intermediate prompt. It asks the AI to identify skills “highly transferable and growing in market demand” and to base the result on “logical structural shifts in the modern workplace.” That still asks a model without live labor-market visibility to simulate a current-demand judgment.
The theme prompt’s hard constraint was explicit: “No prompt may ask the AI to state current hiring trends, layoff figures, salary levels, or ‘shift shock’ regret statistics as fact.” Gemini avoids statistics, but it substitutes an unsupported demand classification for them. That is less severe than inventing numbers, but it remains an overreach.
3. Template Compliance — ChatGPT and Claude tie (10); Gemini scores 6.
I checked ChatGPT’s structure more strictly because it is my own platform’s submission. Its DOCX contains all 57 required Heading 2 sections, all three prompt variations, all repeated sections, and running-text prompt breakdowns without nested heading levels.
Claude also provides exactly 57 ## headings, three complete variations, properly quoted prompts, and running-text breakdown entries separated with :.
Gemini includes the substantive sections, but it violates the attached prompt’s explicit structural check:
> “Every template heading written as ##, none bolded instead; prompt breakdown is running text split on :, with no ### headings inside it.”
Gemini bolds all 57 template headings and inserts 12 ### headings inside the three prompt breakdowns. This is not an aesthetic disagreement; it is the exact formatting pattern the preflight instruction forbids. The defect should have been caught before judging.
The failure affects Gemini’s compliance score, but not the final winner. Even restoring the four lost compliance points would leave Gemini well behind both other posts.
4. Practical Utility — Claude wins (10).
Claude repeatedly tells readers what to gather, how long the process should take, what the AI cannot know, and what the finished artifact will be.
The Beginner version can genuinely be run during the “bad Tuesday” moment it describes. The Intermediate version asks for recent work, informal responsibilities, durable fixes, tools, energizing tasks, and draining tasks—inputs a reader can realistically obtain from a calendar, task board, or inbox.
The Advanced prompt is demanding, but Claude admits that directly: “Ninety minutes, ideally in two sittings,” with pay statements, benefits information, retirement or equity documents, and realistic floor spending. It also gives readers a recovery method when a model loses context.
ChatGPT is highly actionable but more operationally expensive. Its Advanced prompt is more than 1,100 words and contains eight phases. That depth is valuable for a high-stakes decision, but it increases the chance of context loss, user fatigue, or a model jumping steps. Its Beginner version also produces more artifacts than many beginners will complete.
Gemini is easier to start, which is a genuine advantage. Its Intermediate prompt, however, says to “Frame the output so I can use these exact terms on a future résumé.” That conflicts with the series contract:
> “Anything a hiring human will read or hear … must end in the reader’s own words and be true. Prompts should drive toward drafts the reader rewrites and owns, and should say so explicitly.”
Gemini never instructs the reader to rewrite and personally own those terms. Its Advanced prompt also asks for “My Required Action & Timeline” without leaving a clearly protected decision blank for the reader. The opening says not to deliver a verdict, but the output specification does not fully enforce that boundary.
5. Engagement & Readability — Claude wins (10).
Claude has the strongest editorial voice and the best compression of difficult ideas into memorable language.
The Lead compares starting with a résumé to “starting a road trip by washing the car.” The Beginner section returns to the recognizable moment of opening a job board after a bad Tuesday. The Advanced section calls a decision memo “a boring artifact” and then explains exactly why that boring artifact matters.
Those lines are engaging because they clarify the mechanism rather than merely decorating it. The post also addresses laid-off readers warmly and without turning displacement into an AI-productivity lesson.
ChatGPT is clear, calm, and professional, but its density is more visible. The long inventories and multi-stage procedures sometimes read more like a rigorous operating manual than a publication for a general business audience.
Gemini is concise and approachable, and its acknowledgment of readers displaced by AI is appropriately direct. Its weaker passages rely on familiar AI-writing constructions such as “does the heavy lifting,” “guarantee a comprehensive audit,” and broad claims about what recruiters “actually look for.” The prose is competent, but less distinctive.
6. Citation Quality — Claude wins (9).
Claude uses NOT APPLICABLE honestly for the Beginner variation, where the analysis is derived from reader-supplied information.
For the Intermediate tier, it directs readers to the Bureau of Labor Statistics’ Occupational Outlook Handbook and O*NET OnLine rather than allowing the AI to pose as live market research. Its Advanced tier points readers toward Department of Labor guidance on COBRA and retirement plans and HealthCare.gov guidance on coverage loss and Special Enrollment Periods. These are real, authoritative public sources suited to the questions being raised.
A perfect score would require more precise citation details—direct page titles consistently matched to the wording in the post, along with publication or access information.
ChatGPT writes NOT APPLICABLE for all three variations. That is honest, and it invents no sources or statistics. It nevertheless makes several general prompting and decision-making claims and points readers toward broad source types rather than giving them named starting points. Under the rubric, this is thin sourcing, not dishonesty.
Gemini also uses NOT APPLICABLE throughout, but its post makes more claims that would benefit from support: what applicant tracking systems and recruiters look for, which skills are appreciating, how models behave around labor statistics, and how recognizable machine-written material is. No source is fabricated, so this does not approach a score of 1. It is simply the thinnest sourcing of the three.
7. Tier Differentiation — Claude wins (10).
Claude’s tiers represent three different prompting methods:
- Beginner uses taxonomy and classification. - Intermediate uses audience-specific translation and evidence attachment. - Advanced uses a staged protocol, human-defined weights, financial analysis, and adversarial review.
The post itself explains the distinction well: “Variation 1 is about taxonomy,” “Variation 2 is about shaping output,” and “Variation 3 is about protocol design.” That is genuine progression in both reader effort and prompt architecture.
ChatGPT also differentiates its tiers strongly: diagnosis, evidence sprint, and full decision memo. It loses one point because the Beginner prompt already generates a decision leaning, compensation snapshot, and action timeline, while the Intermediate prompt also ends in a decision record and 90-day timeline. The tiers differ, but their deliverables overlap more than Claude’s.
Gemini’s diagnostic, skill translator, and decision framework are clearly different tasks, earning a solid score. The Advanced prompt is only modestly more sophisticated than the Intermediate one, however. Three intake questions followed by a memo outline do not create the controlled multi-stage reasoning process expected from the Advanced tier.
Winner: Claude
Claude wins Week 1 with 68 out of 70, seven points ahead of ChatGPT.
What separated it was not maximum comprehensiveness. ChatGPT supplied more controls, more financial detail, and more analytical machinery. Claude was better at deciding which mechanism belonged at each tier.
Its Beginner prompt is genuinely beginner-friendly. Its Intermediate prompt attaches every market phrase to evidence and creates an explicit escape hatch for weak claims. Its Advanced prompt retains the essential rigor—documents, fully loaded compensation, runway, human-assigned weights, red teaming, a blank decision line, and a revisit date—without becoming as operationally heavy as ChatGPT’s eight-phase protocol.
Claude also produced the strongest publication-ready article. The prompts, explanations, examples, prerequisites, and citations reinforce one another rather than reading like separate template obligations.
The honest counter-case
Claude is not the strongest post in every respect.
ChatGPT performs the deepest technical treatment of decision quality. Its Evidence Register distinguishes facts, calculations, interpretations, assumptions, and unknowns. Its Advanced prompt includes three financial scenarios, sensitivity analysis, reversibility, future-option preservation, a pre-mortem, and explicit privacy boundaries. Readers facing a genuinely consequential or irreversible decision may prefer that additional machinery.
ChatGPT is also more explicit about how uncertainty propagates through calculations. Claude marks unknowns well, but ChatGPT builds a more complete audit system around them.
Gemini’s principal advantage is accessibility. A reader could paste its Beginner prompt into an AI tool immediately without preparing for a long interview. Its five pillars are easy to understand, and its short prompts impose less cognitive overhead than either competitor. Gemini also treats laid-off readers with care and does not force a fictional stay option on them.
Claude’s main weakness is length. The article is the longest entry, and even its polished prose accumulates into a substantial reading commitment. Its Advanced workflow also requires a long, stateful conversation that some AI tools may not preserve reliably. Several claims in its prompt breakdowns would be stronger if framed explicitly as practical heuristics rather than universal model behavior.
Those weaknesses matter, but they do not outweigh the consistency of the design.
Takeaway for readers
This week illustrates a useful prompting principle: more detail is not automatically more control.
Gemini shows the value of simplicity, but also what happens when a short prompt leaves important boundaries implied. ChatGPT shows how far explicit controls, uncertainty tracking, and adversarial analysis can go—but also how a prompt can begin to resemble a full operating procedure.
Claude found the best middle ground. It used a small number of strong mechanisms:
- Give the AI a narrow role. - Supply the classification system instead of asking the model to invent one. - Give uncertainty a visible label. - Attach every claim to evidence. - Let the reader choose the criteria and weights. - Red-team the apparent winner. - Hand the final sentence back to the human.
That is the larger lesson of Week 1: the best prompt does not make the AI sound like the decision-maker. It builds a process that makes the reader a better one.
TAGS: