Per-class token pricing, explained
Why input, cache reads, cache writes and output are priced apart, what each of the four models charges per token class, and one real request worked through.
What we learn running an OpenAI-compatible endpoint: how a request is priced, what the meter records, and the parts a rate card cannot tell you.
Why input, cache reads, cache writes and output are priced apart, what each of the four models charges per token class, and one real request worked through.