The things we build to answer our own questions: experiments, research, and tools, shared in the open. We start with model benchmarks, run on the same work we ship for clients - currently 30 models across 87 tasks, topped by grok-composer-2.5-fast at 85.0% composite.
More experiments land here as we build them.