Orivael
Measured work, published with the numbers the instruments actually returned — including the measurements that turned out to be wrong.
ft09 and tr87 both finished 6 of 6 at 100% on ARC Prize's interactive reasoning benchmark, with no language model in the loop — not for perception, not for planning, not for choosing a move. Total inference spend: $0.00.
Plus the part that took longer than the solving: a world that is bigger than its own frame, an action counter that made every state look novel, and three experiments that produced clean numbers while measuring nothing at all.
Read the note →
More notes are on the way. Everything published here is measured — where a number came from a simulation, a replay, or our own build rather than from the instrument that owns it, the note says so on the same line as the number.