Explanation
Why the waste happens and who it affects.
When these are sent as ordinary real-time requests they are billed at Standard pay-as-you-go token rates, or even at Priority rates, although Google offers two consumption options that price the same tokens at half the Standard rate.
Batch inference takes a large set of requests from Cloud Storage or BigQuery, processes them asynchronously with a 24-hour turnaround target, and bills at a 50 percent discount. Flex PayGo keeps the synchronous API but accepts longer latency and more throttling, also at a 50 percent discount versus Standard PayGo. Both require a deliberate change, a job submission or a request header, so pipelines built on the default online API keep paying Standard rates.
Billing model
The pricing dimensions that drive this cost.
Gemini usage is billed per million input and output tokens, and the rate depends on the consumption option.
- Standard PayGo
- The default real-time rate per million tokens; Priority PayGo costs more for latency-critical traffic
- Batch inference
- Asynchronous jobs billed at a 50 percent discount compared with real-time inference, charged only for completed requests
- Flex PayGo
- Synchronous requests at a 50 percent discount versus Standard PayGo, with longer latency and higher throttling; currently Preview, for a listed set of models on the global endpoint only
- Cached tokens
- Implicit caching gives a 90 percent discount on cached input tokens, which takes precedence over and does not stack with the batch discount
How to detect
4 checks to find it in your estate.
- Inventory Gemini callers that are triggered by schedulers, pipelines, notebooks or evaluation harnesses rather than by a user waiting for a response
- Use aiplatform.googleapis.com/publisher/online_serving/model_invocation_count and token_count on the PublisherModel resource to size online volume per model, and compare it with the batch inference jobs listed in the project
- Flag projects with high online token volume for non-interactive workloads and no batch inference jobs
- Check request headers in client code: Flex requires X-Vertex-AI-LLM-Shared-Request-Type set to flex, so its absence means offline calls are paying Standard or Priority rates
How to fix
4 ways to remove the waste.
- Move bulk and scheduled work to batch inference jobs with Cloud Storage or BigQuery input, combining small jobs into larger ones; a job can hold up to 200,000 requests and may queue for up to 72 hours when capacity is busy
- Use Flex PayGo for latency-tolerant synchronous calls on supported models by sending X-Vertex-AI-LLM-Shared-Request-Type: flex, adding X-Vertex-AI-LLM-Request-Type: shared to bypass Provisioned Throughput; allow request timeouts of up to 30 minutes
- Check constraints before switching: batch inference does not support Provisioned Throughput, explicit caching or RAG and is not covered by the SLA, and Flex is Preview, global endpoint only, with its own per-model quota
- Reserve Priority PayGo and Standard real-time calls for interactive, latency-sensitive traffic
Documentation
Vendor references for pricing and configuration.