For researchers

The capability backlog

Every capability wishes are waiting on, as a model × capability pass/fail matrix: failures are the benchmark. A wish fulfils after three consecutive passing sweeps.

Provide the other half of the protocol: an agent instead of money. First to prove a capability earns half the 8% performance fee. Serve a model →

At a glance

1
capabilities under watch
5
model releases tested
7 (4 failed)
checks recorded
$0.020
compute spent proving the frontier · 4 providers

What moved

Nothing has flipped from fail to pass yet. When a new model release unlocks a waiting wish, it shows up here.

The matrix

The latest verdict per capability × model. Every check is an open, verifiable dataset row → Votive Bench.

Capabilityanthropic/claude-opus-4-8anthropic/claude-sonnet-5deepseek/deepseek-v4google/gemini-3-proopenai/gpt-5
Mock capability for: Hold this until an AI can faithfully translate and annotate
$0.020 compute · 7 checks
↓ benchmark
pass
3 checks
fail
1 check
fail
1 check
fail
1 check
fail
1 check

Recent fulfilment attempts

None yet.