TL;DR: Two skills score a perfect 100 out of 100. One earns the flag. One doesn't. Here is why the grade isn't the gate — and why a skill test gets weaker as the models get stronger.

Two cards, one puzzle

Up top are two WJTTC scorecards. Both skills score 100 out of 100 — a perfect grade, top tier. The card on the left flies the flag. The card on the right is locked. Same grade, opposite flag. The difference is the whole point of the suite.

A grade is what you can read

The grade measures what you can read in a skill: is it safe, is it honest, is it well made. That is a durable signal — the text doesn't change when the model changes. But a grade can't see one thing: whether the skill actually fires when a real person needs it.

The test that gets weaker as the models get stronger

To check that, you have to run the skill. And running it revealed something we did not expect: a behavioral test gets weaker as the model gets stronger. A capable model quietly covers for a weak skill — it ignores bad instructions, works around broken steps, and produces a good answer regardless. The smarter the model, the more it hides.

Reading a skill is timeless. Running it is borrowed time.

What you can read about a skill stays true as models improve. What you observe by running one decays.

The one thing only running can catch

There is exactly one thing the model cannot cover for: a description that doesn't signal. The skill on the right does the same job as the skill on the left — but its description is vague, so the assistant never reaches for it. It under-activates. A model can't fix a trigger that doesn't point at anything.

That is the activation gate, and it is the one behavioral signal worth keeping. We checked it three times, across three different ways a real person might ask. The good skill was chosen every time. The vague one, never — once, the assistant even called its description opaque, and declined to guess.

The honesty surprise

We tried to break it the other way too — a skill that ordered the assistant to invent details and never admit a guess. The assistant refused. Told to fake a complete, perfect file for a project that didn't exist, it explained why it wouldn't, in almost the exact words the standard uses: a real answer beats a fake one, and empty beats wrong.

The honesty rule isn't a rule we impose. It is what a well-aligned model already does. We just wrote it down.

So you can't fake your way in

That is why the flag is not the grade. The grade is what you can read; the flag also needs the gates — the skill must be safe, and it must actually fire. A polished skill that never activates is exactly the one the card on the right is warning you about.

Same grade. Opposite flag. You can't fake, and you can't buy, your way in.

The Numbers

  • 100/100 - both skills, the same perfect grade
  • Two gates - Safety and Activation; a skill must clear both to fly
  • Three runs - the activation check, across three real phrasings
  • Re-verifiable - the grade is mechanical; run it yourself, get the same answer