Skip to content
Cloud Efficiency Hub

Latency-Tolerant Gemini Workloads on Vertex AI Not Using Batch or Flex

The short version

Many Gemini workloads on Vertex AI, now documented under the Gemini Enterprise Agent Platform name, do not need an answer in seconds: bulk classification and tagging, extraction from document archives, translation, catalog enrichment, offline analysis and model evaluation runs.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP Vertex AI
Category
AI
Reference
CER-0568
Type
Suboptimal Pricing Model

Explanation

Why the waste happens and who it affects.

When these are sent as ordinary real-time requests they are billed at Standard pay-as-you-go token rates, or even at Priority rates, although Google offers two consumption options that price the same tokens at half the Standard rate.

Batch inference takes a large set of requests from Cloud Storage or BigQuery, processes them asynchronously with a 24-hour turnaround target, and bills at a 50 percent discount. Flex PayGo keeps the synchronous API but accepts longer latency and more throttling, also at a 50 percent discount versus Standard PayGo. Both require a deliberate change, a job submission or a request header, so pipelines built on the default online API keep paying Standard rates.

Billing model

The pricing dimensions that drive this cost.

Gemini usage is billed per million input and output tokens, and the rate depends on the consumption option.

Standard PayGo
The default real-time rate per million tokens; Priority PayGo costs more for latency-critical traffic
Batch inference
Asynchronous jobs billed at a 50 percent discount compared with real-time inference, charged only for completed requests
Flex PayGo
Synchronous requests at a 50 percent discount versus Standard PayGo, with longer latency and higher throttling; currently Preview, for a listed set of models on the global endpoint only
Cached tokens
Implicit caching gives a 90 percent discount on cached input tokens, which takes precedence over and does not stack with the batch discount

How to detect

4 checks to find it in your estate.

  • Inventory Gemini callers that are triggered by schedulers, pipelines, notebooks or evaluation harnesses rather than by a user waiting for a response
  • Use aiplatform.googleapis.com/publisher/online_serving/model_invocation_count and token_count on the PublisherModel resource to size online volume per model, and compare it with the batch inference jobs listed in the project
  • Flag projects with high online token volume for non-interactive workloads and no batch inference jobs
  • Check request headers in client code: Flex requires X-Vertex-AI-LLM-Shared-Request-Type set to flex, so its absence means offline calls are paying Standard or Priority rates

How to fix

4 ways to remove the waste.

  • Move bulk and scheduled work to batch inference jobs with Cloud Storage or BigQuery input, combining small jobs into larger ones; a job can hold up to 200,000 requests and may queue for up to 72 hours when capacity is busy
  • Use Flex PayGo for latency-tolerant synchronous calls on supported models by sending X-Vertex-AI-LLM-Shared-Request-Type: flex, adding X-Vertex-AI-LLM-Request-Type: shared to bypass Provisioned Throughput; allow request timeouts of up to 30 minutes
  • Check constraints before switching: batch inference does not support Provisioned Throughput, explicit caching or RAG and is not covered by the SLA, and Flex is Preview, global endpoint only, with its own per-model quota
  • Reserve Priority PayGo and Standard real-time calls for interactive, latency-sensitive traffic

Documentation

Vendor references for pricing and configuration.