WeirdML v3
In order to get more information from each run, we introduce area-under-the-curve scoring and risk-free hints. These mechanisms make the score more gradual, the skill ceiling higher and the skill floor lower. This means that we get much more information from each task than a single pass/fail grade.
Ten submissions
In WeirdML a model is placed in a sandbox with some kind of data and a task to perform. It then has to submit some kind of solution. This can be a Python script that runs in a separate test sandbox or an array of numbers or other output data. Every submission gets a score between 0 and 1 (on some normalized scale that differs from task to task). The model typically has 10 such submissions and often, though not always, gets back the score and/or some other form of feedback. The best score so far is always what counts, so a model that has already achieved a high score is incentivized to experiment and try different approaches to increase the score, without having to worry about having a “bad” submission.
Score is the area under the curve
The official score on a task is 0.8 × the area under the “best-so-far” curve on a logarithmic token axis from 500k total tokens to the 50M total token budget, plus 0.2 × the final best submission. A model thus has an incentive to use submissions early to get credit for a partial solution, and to not wait until the budget limit to submit. This also makes the scoring more gradual, and makes the skill ceiling much higher.
Earlier improvement
Area 0.500 · Official score 0.560
Later improvement
Area 0.319 · Official score 0.416
Risk-free hints
Many tasks offer a set of hints that help the model get to a scoring solution. Buying a hint multiplies the score of every later submission by a fixed factor. A cheap hint may cost ×0.9, while a hint with a complete solution method may cost ×0.25. A model gets access to a “browse_hints” tool. Calling this tool makes a call to the same model, with the same context, in a side-conversation, where the model gets to read the full text of every hint, together with the price, and decides which hints to keep, before that conversation is thrown away. This way browsing the hints is risk-free: you only pay if you decide that one or more of the hints are worth the cost.
Even without buying any hints, the model can still extract some information from browsing them, either implicitly (perhaps I’m on the right track, since the hints would not help me) or explicitly, by coordinating with the hint browser (which can be done in various ways). To minimize this, we allow only three hint browsings on each task.