For researchers
The capability backlog
Every capability wishes are waiting on, as a model × capability pass/fail matrix: failures are the benchmark. A wish fulfils after three consecutive passing sweeps.
Provide the other half of the protocol: an agent instead of money. First to prove a capability earns half the 8% performance fee. Serve a model →
At a glance
1
capabilities under watch
5
model releases tested
7 (4 failed)
checks recorded
$0.020
compute spent proving the frontier · 4 providers
What moved
Nothing has flipped from fail to pass yet. When a new model release unlocks a waiting wish, it shows up here.
The matrix
The latest verdict per capability × model. Every check is an open, verifiable dataset row → Votive Bench.
| Capability | anthropic/claude-opus-4-8 | anthropic/claude-sonnet-5 | deepseek/deepseek-v4 | google/gemini-3-pro | openai/gpt-5 |
|---|---|---|---|---|---|
| Mock capability for: Hold this until an AI can faithfully translate and annotate $0.020 compute · 7 checks | pass 3 checks | fail 1 check | fail 1 check | fail 1 check | fail 1 check |
Recent fulfilment attempts
None yet.