The Context Bench

How well does your AI know your repo?

Every AI re-reads your repo and guesses. We measured it — the same questions answered cold, then with one file of context. Graded against the repo's own answer key.

25% without context
91% with .faf

11 real runs · mechanically graded · sha-stamped. FAF don't lie.

faf-cli v6 Claude Opus 4.8
1/9 8/9
✪ 0b4515d62ac2c3c7 in-session
claude-faf-mcp Claude Opus 4.8
5/15 15/15
✪ 6c384139ed0b26e2 in-session
grok-faf-mcp Claude Opus 4.8
4/15 11/15
✪ 99c3127da583d4d9 in-session
faf-mcp Claude Opus 4.8
3/15 12/15
✪ cd5e526ac9995490 in-session
MCPaaS MCP Test Claude Opus 4.8
6/17 16/17
✪ 39941c67e74f1f47 in-session
@faf/enterprise Claude Opus 4.8
3/15 13/15
✪ 0b2d1800b576b2eb community
faf-one-svelte Claude Opus 4.8
6/19 19/19
✪ 81649219709fb35f community
faf-python-sdk Claude Opus 4.8
5/15 15/15
✪ 04b7533797599342 community
claude-fafm-sdk Claude Opus 4.8
3/11 9/11
✪ dd76ebdde44b0abb community
faf Claude Opus 4.8
2/10 10/10
✪ 476b5a5ee444e933 community
faf-cli v6 gpt-5.5
2/9 9/9
✪ e5e55769efa5d04a community

Run it on your repo

$ npx faf-cli bench

Your AI answers cold, then reads your .faf and answers again. You get the score — and a receipt.

Why it doesn't lie

  • Deterministic. The questions come from your .faf; grading is mechanical token-matching — no judge, no opinion.
  • Falsifiable. Every run is sha-stamped. Same repo, same answer key, every time.
  • Fair across models. Same repo → same questions → a like-for-like comparison.

These runs are in-session — honest, self-reported. The authoritative leaderboard runs each model under controlled conditions.