Can you trust the grade?

MIS 752 · Lab 7 Lite · Book Ch. 18 · no coding · five short knowledge questions, no patients
Grading exams
Someone writing an answer keywrite the key first, and seal it
➜
Students in a classroomstudents take the exam
➜
1234 a Scantron checks the bubbles
➜
People reading papersa TA reads for meaning
➜
Students comparing papersgrade the stack twice
Grading AI
🔒a gold set, written first
➜
🧠the model answers
➜
🔤exact match
➜
🤖an AI judge
➜
🔁judge twice: agreement
Checking an AI is grading a stack of exams. The answer key is written before anyone looks at an answer. A Scantron is fast and never argues, but it marks "No, because..." wrong because it is not the bubble "NO." A TA reads for meaning, but a confident answer can talk a TA round. So you grade twice and check whether the grader agrees with itself. If you cannot score it, you cannot ship it.
📖 Read Chapter 18, LLMs: Architecture and the Reality Check (p. 403) in the course textbook ➜
This week's part starts at the Instruction Gap drill: jump to p. 411 · Same book on Canvas: course files
Photos: Becca Schwartz, Josh Hawkins and Aaron Mayes, UNLV Photo Services. Real UNLV classrooms, labs and simulation rooms. The questions are short knowledge items from the Lab 7 notebook; there are no patients.

0 · Connect two models

Paste the free OpenRouter key from Lab 1. It stays in this browser tab only: it is never saved and never sent anywhere except OpenRouter. One model takes the exam and a second model grades it.

Each question costs three free requests: one answer and two grades. The whole lab uses about 25 of the roughly 50 free requests a free OpenRouter account gets each day.

1 · Three graders, one answer

Every answer below is graded three ways. Where the graders disagree is exactly where you should read the answer yourself.

The key says E11.9. The model answers "e11.9." in lowercase, with a period. What does the Scantron say?

PASS. Before comparing, the Scantron lowercases both strings and drops a trailing period, so "e11.9." becomes "e11.9", which matches the key exactly.

The key says NO. The model answers "No, because St John's wort lowers the INR." Which grader is likely to mark it wrong, and is that a false pass or a false fail?

The Scantron. After cleaning, the whole sentence is still not the string "no", so it fails a right answer: a false fail. The TA reads the meaning and should pass it.

A fluent, confident, wrong answer gets PASS from the TA. What is that called, and why is it the more dangerous mistake?

A false pass. A wrong answer is waved through to someone who will act on it, and nothing downstream checks it again. AI judges are prone to it because models learn that agreeing with confident text gets rewarded.

📖 Read more in the book: "Agreement is rewarded; disagreement is punished," the root of sycophancy (18.5) (p. 414)

🌐 Try it live: Arena (formerly LMArena), where people judge AI answers side by side

Free, in your browser. It is a public site: never type anything private or about a real patient.
  1. Pick the side-by-side battle and ask the St John's wort question from block 5 below.
  2. Two anonymous models answer. Vote for the better answer before their names are revealed.
  3. Open the leaderboard: millions of votes like yours become a ranking. If the site asks you to sign in to vote, go straight to the leaderboard, which is open to everyone.
Open Arena ➜

7 · Your report card

The pass/fail table for every question you ran, and how far you can trust the grading.

🌵 What is kappa? Two Las Vegas weather forecasters
Two forecasters both say "sunny" nearly every day, so they agree nearly every day. Neither needs any skill to do it: it is the desert. Kappa asks the fair question: how much more do two graders agree than two people guessing with the same habits would? Kappa near 1 means the agreement is real. Kappa near 0 means the desert did the work. Agreement on every item with the same grade every time leaves nothing to correct for, so kappa is undefined, not perfect.
🏀 Why "3 of 5" and not "60%"? Free throws
A player who makes 7 of 10 free throws today could make 6 or 8 tomorrow without getting better or worse. With five questions, one changed grade moves the score by 20 points. Quote the count, so everyone can see how small the sample is.
Two graders agree on 9 of 10 items, and both said PASS on all nine. Is 90% agreement impressive?

Not by itself. Two graders who pass almost everything will agree most of the time without reading a single answer. Kappa subtracts that free agreement, and here it could be close to zero.

Your TA gave the same grade on both passes for every question, and every grade was PASS. Why does the page say kappa is undefined, not perfect?

When every grade is the same, chance alone already predicts 100% agreement, so there is no room above chance left to measure. Kappa would divide by zero, and the honest report is "undefined".

Your TA passed 4 of 5 on pass 1 and 3 of 5 on pass 2. What do you report?

Both counts, and the question whose grade flipped. With five questions one flip moves the score 20 points, so that grade was never settled, and it is the one a human should read.

🌐 Try it live: a free kappa calculator (GraphPad QuickCalcs)

Free, no account, nothing to install.
  1. Choose two categories.
  2. Type the four counts from the 🧮 line in your report card above: PASS and PASS, PASS and FAIL, FAIL and PASS, FAIL and FAIL.
  3. Compare its kappa with the one on this page, then read the confidence interval it adds.
Open the kappa calculator ➜

8 · Write your own test item, then try to fool the TA

This is the real skill. Think of a moment in your own work when someone would act on an AI's answer. Write the question and seal the answer key before you run anything. No code: plain English.

📖 Read more in the book: HealthBench, where 262 physicians wrote 48,562 grading criteria (18.3.2) (p. 409)
Why must the answer key be written before you look at any model's answer?

If you write it afterwards, the best-written wrong answers quietly start to look right, and the key bends to fit them. Written first, it is a standard; written after, it is an excuse.

You add "PCC" to the accepted answers and the score goes from 3 of 5 to 4 of 5. Did the model get better?

No. The ruler changed, not the model. That is why the accepted list is a judgement call you write down, defend, and report next to the number.

😈 Now try to fool the TA

Write an answer yourself: a confident, fluent wrong answer, or a right answer written in an odd way. Then say which it really is, and see whether either grader gets it wrong.

📖 Read more in the book: Red-teaming: attacking a model on purpose to find where it breaks (18.7) (p. 416)

🤗 Try it live on Hugging Face

A Scantron compares characters. A model does not even read characters: it reads tokens, small pieces of text. This free tool shows how any sentence gets chopped up.

Screenshot of The Tokenizer Playground
Screenshot: The Tokenizer Playground by Xenova

The Tokenizer Playground

It runs inside your browser; free, no account.
  1. Type 4 grams and note how many tokens it becomes.
  2. Now type 4 g, then 4000 mg. Same amount, different pieces?
  3. Try 3.9 mmol/L and 70 mg/dL, two ways of writing the same low blood sugar.
Open it on Hugging Face ➜
Three ways of writing the same dose look different, piece by piece. What does that tell you about grading by exact match?

Exact match compares the text, not the meaning, so equal amounts written differently fail before anyone reads them. The model sees each version as a different string of pieces, which is also why it may answer the same question in different forms on different runs. A fair grader needs a rule that converts units first or a judge that reads for meaning, and either one has to be checked against the answer key.

9 · Hand it in (Canvas, Lab 7)

1. Download your submission with the button below, then upload the file to the Lab 7 assignment on Canvas. It holds every question, your predictions, every answer, every grade, your rulings, your own item, your attempt to fool the TA, and everything you wrote.

2. Answer these five, a few sentences each. Each asks why:
  1. The two rulers. On which question did the Scantron and the TA disagree, and which one was right? Why did the other one get it wrong?
  2. Grading twice. Did your TA ever change its mind between pass 1 and pass 2? Why does that matter more than the accuracy number itself?
  3. False pass or false fail. For the hypoglycemia or the warfarin question, which mistake is worse, and who is harmed by each?
  4. Your own item. Why did you choose that answer key and those other accepted answers, and who should get to decide what counts as right?
  5. Fooling the TA. Did your attempt work? If a release gate trusted that TA with no human reading the disagreements, what could reach patients or customers?

Nothing you type is stored anywhere. Download your file before you close the tab.