Evals.
Score cards Bench produced, published as they are recorded. A target is model, host, quant, serving config, sampling config and checkpoint together, because a score without that tuple is not a result.
Arena's score cards are published here, one per data environment per run, from two evaluations on one DGX Spark: Nemotron 3.5 Lightning 30B-A3B at NVFP4 on 19 August 2026, and Qwen3.8-27B on 21 August, 32 rollouts attempted in each environment. Every rollout in both was sampled with the model's reasoning mode switched off by a server default, which is now a required field on every row rather than a clause somebody remembered to type. They are an archive rather than a recipe — one host name served two different models two days apart. This page renders one committed results file and nothing else, so what is here is what the file holds rather than what a page managed to load.
A score this page no longer stands behind is removed rather than annotated: the raw figure is kept as scoreUnverified and the score itself is null. Three of the fourteen are void, and none for contamination — both grand-exchange rows, where the disabled reasoning mode rather than the model produced the number, and one drop-table-inference mean taken over two populations that share no behaviour. A score is the mean over rollouts attempted, counting an error as zero, never over the ones that came back. Nothing is averaged across hosts, because the host is part of the target.
Bench produces the score cards · Forge produces the artifacts they point at · Arena is where a gate becomes a gradient