The Batch API

A discount for giving up latency, nothing else. Same model, same quality, results collected later.

Part of the Cheap and fast track on lAItest.

Roughly half price, for work you are willing to wait for.

The discount buys latency. It buys nothing else.

What batch means

You hand the provider a file of requests and collect the results later, instead of holding a connection open while a model types. The weights are the same, the quality is the same, and the price is roughly half at every vendor that offers it. Google sells the same discount twice under two names, Batch and Flex.

It stacks, and it sometimes unlocks capability

Batch pricing combines with prompt caching, so a cached batch job is cheap twice over. On the Claude API there is also a beta that raises the maximum output length far above the normal ceiling, and it is available only through the Batch API, only on some models, and not on Bedrock, Google Cloud or Foundry. Cheap and unusually large can be the same request.

List prices, before the batch discount

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Claude Opus 5 input$5 per 1Msource, verified .
Claude Opus 5 output$25 per 1Msource, verified .
Claude Sonnet 5 input$2 per 1Msource, verified .
Claude Sonnet 5 output$10 per 1Msource, verified . Batch work bills at about half of the standard rate.

Which of these should not go through a Batch API?

Answer: A chat reply a person is watching for. Batch trades latency for price. Anything a human is waiting on has to be served at standard rates. Anything whose deadline is measured in hours does not, and paying standard rates for it is a choice.

In one sentence

If nobody is waiting for the answer, you are overpaying at standard rates.