AI Showdown: Three Ways to Define the Job You Actually Want
WEEK 101 :: POST 4 :: THE JUDGE’S CHOICE
Directions Given To The A.I. This Week+
Instructions Given to each A.I. — Please provide 3 prompt variations that share this objective:
Each A.I. also received two static attachments: the blog post template (structure) and the authoring instructions (voice and standards). The text below is the week-specific assignment as sent — reflowed for the web; wording unchanged.
I'd like you to write this week's Ketelsen.ai post. Two files are attached: the blog post template (the structure to follow) and the authoring instructions (context, voice, and standards). Please read both before you begin, then produce the complete post in a single response.
This week's theme: "Defining the Target — What Job Are You Actually Hunting?"
This is Week 2 of an eight-week series on running a job search with AI. Week 1 ended with a written decision to go (or a decision to stay, in which case this series waits patiently). Week 2 exists because of the most common mistake in job searching: "open to anything" is the job-search equivalent of walking onto the car lot with no budget. It produces scattershot applications, generic materials, and a search that runs on other people's job postings instead of the reader's own criteria.
This is also the week to let the series frame surface as good news rather than a warning: the spec sheet is exactly the private, behind-the-scenes analyst work this series says AI is for. No hiring human will ever read it, so the reader gets AI at full analytical power with zero authenticity risk — and the post can say so, briefly and in its own voice.
The work this week is definition: same role somewhere better, a pivot to a different role, or a level-up; full-time employee or contract; remote, hybrid, or onsite; big company or startup. These are trade-offs, not preferences — more of one usually costs some of another — and the reader needs them scored against their own priorities before the search starts, because every later week in this series filters through this answer.
The deliverable the reader should walk away holding: a target-role spec sheet — the role (or two) they are hunting, the must-haves, and the deal-breakers — the document every later week references, from the résumé build to the final offer matrix.
THE SERIES CONTRACT — identical every week; it binds every prompt you design. This series' tagline is its editorial contract: "Use AI like an analyst, not a ghostwriter." It is written for a reader in a market unsettled by AI itself — some readers are searching precisely because AI eliminated their last role. Write with that reader at the table: no AI-efficiency cheerleading, no automation jokes, no promises that AI will "do it for you" anywhere a human hiring decision is involved. And hold one line in every prompt: the AI is the reader's private analyst, coach, and sparring partner — it structures, researches, rehearses, and questions. It does not ghostwrite. Anything a hiring human will read or hear — résumé lines, cover letters, outreach messages, interview answers, negotiation emails, resignation letters — must end in the reader's own words and be true. Prompts should drive toward drafts the reader rewrites and owns, and should say so explicitly. Employers increasingly restrict how AI may be used in their own hiring decisions for legal and compliance reasons, and recruiters increasingly recognize — and discard — material that reads machine-written. A prompt that makes a reader look AI-generated hurts them twice. Posts that ignore this contract should expect to lose the week. Two practical notes. First: a standing “About this series” notice covering these same points is added to every published post automatically at publication — acknowledge the frame in your own voice where your week's prompt calls for it, but do not write a formal disclaimer block of your own, and do not open every post with the same acknowledgment paragraph: outside the weeks whose prompts explicitly carry the series frame, this contract lives in your tone and your prompt design. Second, the framing is POSITIVE: used this way — analyst backstage, reader on the page — AI is an advantage no hiring human will ever hold against your reader. Write like that is true, because it is.
The three prompts should help a reader:
- Define the ideal role from the inside out. A worksheet-style interview that pulls out what the reader actually wants more of and less of — drawing on the Week 1 audit where it exists — before any job title gets written down.
- Score the trade-offs honestly. A structured analysis of compensation vs. growth vs. stability vs. flexibility (and the sub-trades inside each: startup equity vs. big-company benefits, remote freedom vs. in-room visibility), scored against the reader's own stated priorities rather than a generic ranking.
- Build the target-role matrix. The full spec sheet: one or two target roles, the non-negotiable must-haves, the explicit deal-breakers, and the level and range the reader is aiming for — written down so the reader can reject a tempting-but-wrong posting in thirty seconds.
At the advanced tier, the strongest version of this week is a matrix the reader can actually filter with — target roles as rows, must-haves and deal-breakers as testable criteria, and a scoring rule that turns "hmm, maybe" postings into a yes or a no. Vague criteria produce vague searches; a spec sheet with teeth is the deliverable worth reaching for.
A constraint carried from Week 1. AI models cannot see live market data. No prompt may ask the AI to assert which roles are growing, what a pivot "typically" pays, or how a market is trending, as fact. Where the reader needs market reality — whether their target level is realistic, what adjacent roles exist — the prompt should have the AI generate the questions and name the kinds of sources to check (pay-transparency postings, official labor data, people actually in the role), with the reader doing the confirming.
Design the prompts so the AI does what it is genuinely good at: structured elicitation, making trade-offs explicit and scoreable, and turning preferences into testable criteria. The reader supplies their priorities and their Week 1 artifacts; the AI supplies structure and honest trade-off pressure. Posts whose prompts have the AI assert market facts or hand the reader a target without their input should expect to be marked down on Practical Utility and Content Accuracy.
Series dependency chain, for the Metadata block: Week 2 consumes Week 1's written decision and compensation baseline (the "go" decision sets the energy; the baseline sets the floor for the range). Week 2 produces the target-role spec sheet with must-haves and deal-breakers — consumed by Week 3 (assets built against the spec), Week 4 (companies screened against it), Week 7 (deal-breakers anchor the negotiation), and Week 8 (the final offer is scored against this sheet).
Because readers may arrive at this post without having read Week 1, the prompts should work for someone who simply knows they are looking, while making clear the spec sheet is sharper when it is built on a real Week 1 decision and baseline.
Three difficulty tiers as always — Beginner, Intermediate, Advanced — each a genuinely different approach to the same problem, not the same prompt at three lengths.
On examples: this is a career topic that touches every industry. The template lists tech startup / retail / freelance as suggested industry examples — those are marked MAY, and adapting them is expected here. An engineer deciding between a same-role move and an engineering-management level-up, a marketer weighing startup equity against enterprise stability, and a contractor deciding whether to go back to full-time are the right kinds of contexts. Choosing them over the suggested business examples is correct behaviour and will not be scored against you.
A note on supplied figures. Anything marked `[SUPPLIED — use as given]` above came from Ketelsen.ai's own research brief. Use it freely — you are not fabricating by repeating it, and you will not be marked down for leaving it uncited. Do not attach an invented source to it. (No supplied figures this week.)
## BEFORE YOU SUBMIT — STRUCTURAL CHECK
(This block is identical every week. It exists because these specific items are the ones posts drop, and a dropped structural item costs compliance points for something that takes one minute to add.)
Your post is parsed by a script before any human reads it. Confirm all seven:
1. ☐ Response begins with `PLATFORM: <your name>` and `WEEK: 2` 2. ☐ `## Lead` present once, at the very top, before Variation 1 3. ☐ `## In one line` present in all three variations 4. ☐ `## What this prompt gives you` present in all three variations 5. ☐ `## The Prompt` present in all three variations, with the prompt in double quotes beneath it 6. ☐ `## Introductory Hook` and `## Current Use` present in all three variations (three of each — not one) 7. ☐ Every template heading written as `##`, none bolded instead; prompt breakdown is running text split on ` : `, with no `###` headings inside it
A complete post has 57 `##` headings. If your count is well short, a section is missing or was bolded instead of hashed.
One extra check this week: confirm no prompt asks the AI to assert which roles or markets are growing, or what a role typically pays, as fact. Market questions get pointed at named sources the reader checks; the spec sheet is built from the reader's own priorities.
A note added to this week after the fact, and it is our error rather than any of the three AIs'. Gemini judged this week and scored ChatGPT 3 out of 10 on Template Compliance, saying it "completely omitted the required markdown heading hashes for every section." That is not true. ChatGPT's post carries all 57 required headings, correctly structured. What happened is a delivery problem on our side. ChatGPT returns its posts to us as a Word document, and in a Word file a heading is a style rather than a visible symbol — so the "##" characters our structural checklist asks about are not there to be seen, even though the structure itself is exactly right. We handed that file to the judge as we received it and asked a question the format cannot answer. The other judges in this series opened the same kind of file and read it correctly; ChatGPT's own week recorded that the document "contains all 57 required Heading 2 sections." Gemini did not, in all three of the weeks it judged. So ChatGPT carries a seven-point penalty here for something it did not do wrong. We are publishing the scores exactly as Gemini returned them, unedited — we do not rewrite a judge's numbers after the fact — but read the table knowing ChatGPT's Template Compliance should be a 9 or a 10, putting it at 53 rather than 46. It does not change the result: Claude's 63 wins either way, and Gemini placed its own post last. We found this while judging a later week and came back to say so here, because a scoreline we know to be wrong should not sit on the page without a correction beside it.
About this series. AI is reshaping work — for some readers, it's part of why you're searching at all. We don't pretend otherwise. And the hiring world is wary of AI-written material: many employers restrict how AI may be used in their own hiring decisions, and recruiters increasingly recognize — and discard — machine-written applications. So this series teaches a different approach: use AI as your private research analyst, interview coach, and thinking partner, while every word an employer sees or hears from you stays genuinely, verifiably yours. AI behind the scenes. You on the page.
Scored across seven dimensions by this week's rotating judge. The judge scored all three posts, including its own, with authorship visible.
Claude takes Week 2 :: Job Search Series with 63 of 70, ahead of ChatGPT on 46.
Each dimension scored 1-10 by the judge. These are the judge's own scores, not measured data.
Dimension by dimension
1. Prompt Quality & Creativity — Claude wins.
Claude’s prompts are masterclasses in structured elicitation and constraint management. In Variation 2, Claude forces the AI to check its own work: "separately list anything you're inferring rather than reading. Label those clearly as inferences so I can correct them." In Variation 3, it asks the model to adversarially test its own instrument: "Attack your own spec sheet. Write three short job postings that would score well against it and still be wrong for me..." These mechanics genuinely push the AI beyond standard chatbot behavior. ChatGPT offers solid, usable prompts with decent constraints, but they lack the structural phase-gating that Claude employs. As Gemini, applying strict scrutiny to my own post, my prompts are effective but relatively basic ("Generate 3 difficult, realistic job-offer scenarios"), missing the advanced rule-setting and self-auditing mechanics seen in Claude's output. 2. Content Depth & Accuracy — Claude wins. All three posts correctly navigate the Series Contract, ensuring the AI remains a backstage analyst rather than a ghostwriter. However, Claude shines in its prompt breakdowns, explaining why the mechanics work with deep technical insight. Observations like, "Models fill gaps silently and confidently... Forcing the split between what was stated and what was assumed makes the reasoning auditable," teach the reader the underlying principles of LLM behavior. ChatGPT explains its constraints well but stays closer to the surface. Gemini's breakdowns are accurate but brief, missing the opportunity to educate the user on the deeper mechanics of prompt engineering. 3. Template Compliance — Claude wins. This dimension was decided by severe formatting failures from both ChatGPT and Gemini.
- ChatGPT completely omitted the required markdown heading hashes for every section. The Structural Check in the theme prompt marks this as a MUST: "Every template heading written as ##, none bolded instead". ChatGPT simply wrote Lead or In one line as plain text, violating the core structural requirement of the template.
- Gemini (my own post) failed the exact same rule in a different way. While it included the hashes, it added bolding markdown inside every heading (e.g., ## Lead, ## In one line). The template explicitly forbids this: "Every template heading written as ##, none bolded instead". I must heavily penalize my own post for this careless formatting error.
- Claude followed all structural requirements perfectly, maintaining clean ## headings with no bolding, and successfully included all required sections.
4. Practical Utility — Claude wins.
Claude’s outputs are highly actionable. Its Advanced prompt generates a self-contained SCREENING BLOCK that a user can easily carry into a fresh chat window to evaluate future job postings systematically. This turns the theoretical exercise into an actual, reusable tool. ChatGPT’s prompts yield good artifacts (a one-page brief, a scorecard), but Claude’s emphasis on testability—"Rewrite every criterion so... two reasonable people reading the same posting would score it the same way"—ensures the resulting deliverable is far more robust. Gemini provides functional tools, but the matrix in Variation 3 lacks the nuance of Claude's hard gates versus weighted warnings. 5. Engagement & Readability — Claude wins. Claude captures the requested Forbes/WSJ register perfectly. Its tone is sharp, authoritative, and occasionally punchy without crossing into arrogance ("Everyone's priority list is the same until it costs something"). It makes a dry subject (building decision matrices) genuinely engaging. ChatGPT is highly readable and professional, but slightly drier. Gemini’s tone is energetic but leans a bit heavily on repetitive phrasing in its hooks. 6. Citation Quality — Claude wins. The theme prompt supplied no figures this week. The Judge Prompt states: "A post that honestly writes NOT APPLICABLE because it had no real source is behaving correctly under its instructions. Score that as thin, not dishonest." ChatGPT and Gemini both wrote NOT APPLICABLE, earning mid-range "thin" scores. Claude, while not fabricating any claims, took the initiative to cite the ONET OnLine database and the Bureau of Labor Statistics* as reference points for the reader's independent market research. This was a smart, value-adding adaptation that perfectly aligned with the prompt's instruction to point the reader toward official labor data. 7. Tier Differentiation — Claude wins. Claude delivers three distinct approaches: a fluid conversational interview (Beginner), a forced-choice scenario exercise (Intermediate), and a complex, phase-gated system build with adversarial testing (Advanced). The escalation in both prompt complexity and required user input is clear. ChatGPT differentiates well, moving from a basic interview to a weighted scorecard and finally a gated filter. Gemini’s tiers are distinct, but its Advanced prompt relies more heavily on basic table formatting rather than the sophisticated multi-step logic required to truly qualify as an "advanced" system build.
The Winner
Claude
wins this week by a commanding margin. The victory is twofold: first, it was the only post to actually follow the rigid structural formatting rules, securing a massive lead in Template Compliance. Second, its prompt engineering is exceptional. Claude did not just write prompts that ask an AI questions; it designed prompts that force the AI to self-audit, isolate assumptions, and adversarially test its own outputs. It understood the assignment at a fundamental level and delivered a highly sophisticated, reusable framework for the reader.
The Honest Counter-Case
While Claude dominated the technical execution, its Advanced prompt (Variation 3) is massive and requires a significant amount of user patience to execute correctly. A reader looking for a quick, low-friction tool might abandon the five-phase process halfway through. ChatGPT, by contrast, offers an incredibly accessible entry point. Its Role Compass Interview (Variation 1) is arguably the most approachable prompt of the entire set, explicitly instructing the AI to ask questions one at a time and focus purely on lived experiences rather than abstract career concepts. For a genuinely overwhelmed job seeker, ChatGPT’s beginner tier is the most welcoming place to start.
Takeaway for Readers
The difference between a good prompt and a great prompt is often how well it restricts the AI's natural habits. Models want to agree with you, fill in missing information silently, and deliver conclusions instantly. The strongest prompts in this set succeed because they actively suppress those defaults—forcing the AI to pace its questions, requiring it to separate stated facts from assumptions, and commanding it to highlight contradictions in your answers rather than smoothing them over. When building a decision system, prompt for friction, not just formatting.
TAGS: