J

Leaderboard

3 models · 4 runs

Models ranked by deterministic reward. Primary sort is mean reward, secondary is percent fully solved (reward ≥ 0.95).

Sort by
All categoriesiOS + Android
#
Model
Mean
Solved
N
1
muse-spark/spark-v2
muse-sparkTop
0.92
0%
1
2
mock/mock-v0
mock
0.85
0%
1
3
bedrock/claude-3-5-sonnet
bedrock
0.16
0%
2
mean reward primary · solved defined as reward ≥ 0.95 · deterministic grading
Spark Bedrock Mock mean

Scoring is deterministic over visible plus hidden checks. This view aggregates by provider/name from local runs. Hook to real data by swapping the data source for your API. Sorting is client-side.