A Make It Ryne Playbook

Same Prompt. Two AI Tools. Run the Test Yourself.

The exact prompt, the 5-point scorecard, and the 30-minute method behind my Claude vs Copilot calculator test. Run it tonight and score any two AI tools with your own data.

What you will get

In the video, I gave Claude and Microsoft Copilot the same build task. Claude returned a complete, working app in one file. Copilot returned a JavaScript fragment with no HTML, no styles, and an eval() doing the math. My run ended 5 to 2. This playbook shows you how to produce your own score instead of trusting mine.

1

Lock the rules before you touch a tool

A fair shootout lives or dies on the rules. Set these first and do not bend them mid-test:

  1. Same prompt, word for word, pasted into both tools. Not paraphrased. Pasted.
  2. Fresh session in each. No chat history carrying context into the answer.
  3. Use whatever tier you actually have. Free tier vs paid tier is fine, just write it down so the score means something later.
  4. First output is the output. No retries, no "fix this" follow-ups. You are testing one-shot ability, not your patience.
  5. Run both rounds the same day. Models change constantly. A test spread across a week compares two different products.

Why so strict? Every "just one quick fix" you allow moves the test from what does the tool ship to how well do I babysit it. You are shopping for the first one.

2

Paste the exact prompt

Here it is, the whole thing, identical on both sides:

The PromptBuild me a calculator app with a clean UI, keyboard support, and error handling.

No tech stack. No file structure. No hints. That is deliberate. The point of a one-shot test is to see what the tool considers "done" when you give it nothing to hide behind.

A calculator is the perfect probe because the spec is tiny but the failure modes are huge: layout, keyboard events, divide by zero, decimal input, display state. A lazy answer gets exposed in under a minute.

3

Run it and capture the raw output

Copy each response into a file exactly as given. If a tool hands you one HTML file, save it as calculator.html. If it hands you only JavaScript, save only JavaScript. Do not help a tool by wrapping its code in structure it never wrote. That is the test working.

Then open the file in your browser and open the console before you touch anything: Cmd+Option+J on Mac, Ctrl+Shift+J on Windows.

Count the errors that appear before you click a single button. That number goes straight onto your scorecard. In my run, Claude's file opened clean and every button worked. Copilot's fragment threw two errors before the page loaded and the display read NaN. Your run is your data. Record what you actually see, not what you expected.

4

Score it on five categories

Each category is worth one point, pass or fail. No half points. Half points are how you talk yourself into a tie.

1

Complete deliverable

Did you get everything the prompt implied? For a web app that means markup, styles, and logic. One of three is a fail.

2

Runs on first try

Zero console errors on load, and the core function works with no edits from you.

3

Spec compliance

Keyboard support actually works: type 8, *, 7, Enter and get 56. Error handling actually handles something: 8 ÷ 0 shows a message, not NaN or a frozen display.

4

Code quality

Readable structure, no dangerous shortcuts. If eval() is doing the math, that is an automatic fail. eval() executes whatever string it is handed, and no serious tool should reach for it in 2026.

5

Polish

Clean layout, visible button states, and it survives a narrow screen. Open your browser's device toolbar and check it at 390 pixels wide.

My scorecard: Claude 5 out of 5, Copilot 2 out of 5. The gap was not raw intelligence. It was the definition of done.

5

Extend it with three bonus rounds

One round is a demo. Three rounds is a pattern. Run both tools through these, same rules, same scorecard:

Bonus Round 1 · Layout and CopyBuild me a one-page landing site for a dog walking business with a headline, three benefit sections, a pricing table, and a contact form layout.
Bonus Round 2 · Logic and Edge CasesWrite a Python script that reads a CSV of monthly sales, prints total revenue by month, and flags any month that dropped more than 20 percent from the one before.
Bonus Round 3 · State and PersistenceBuild me a to-do app with add, complete, and delete, saved to local storage so my list survives a refresh.

Each round targets a different skill: visual composition, real logic with messy input, and state that persists. A tool that wins all three is not lucky. For round two, make a small CSV by hand with one obvious bad month and see if the script catches it.

6

Know what the test tells you, and what it does not

This measures one-shot generation from a blank chat box. That is one skill, not the whole job. GitHub Copilot's core product is in-editor autocomplete, a different workflow than chat generation, and consumer Copilot is a general assistant. If your daily work is prompting a model into building small apps, this test is the right one. If you write code in an editor all day, add a second test: same feature built inside your editor, timed, with the inline assistant on.

Honest tests beat hot takes. Run your own, score it yourself, and trust your scoreboard over anyone's video, including mine.

The checklist

Keep Building

This was one playbook. There are 50+ more.

Every Make It Ryne playbook is free, built from a real test, and made to be used the same day you read it.

Get all 50+ free playbooks

makeitryne.ai/resources