Gamesmark Bench

Measured comparisons of local vs paid AI — language models, image models, and the hardware they run on. Local models run on a Ryzen box with an Intel Arc Pro B70 and an RTX 3080; paid models via API.
Every LLM task uses --epochs: single draws proved to be noise (a 7B scored 5/5 then 0/1 on the same case). Image results are judged at the size assets actually ship at, against real shipped artwork.