Python · Apache-2.0

Bench

Should you run this model, on this box, at this quant, with these flags — and can you prove it later. A decision record for self-hosting, not an eval framework.

A target is model × host × quant × serving config × checkpoint, because none of it is incidental. Runs produce immutable score cards with per-sample outcomes; compare diffs two of them and refuses to pretend a confounded comparison is honest; sweep reduces a matrix to a frontier. Contamination probes void the quality score of any model that has seen the eval set. Public under Apache-2.0 at git.karti.ai/lumbridge-public/bench; the private eval set, the canary and the seeder stay behind, because publishing an eval destroys it.

Read about Lumbridge Bench