# Latency-Tolerant Gemini Workloads on Vertex AI Not Using Batch or Flex

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/latency-tolerant-gemini-workloads-on-vertex-ai-not-using-batch-or-flex

Many Gemini workloads on Vertex AI, now documented under the Gemini Enterprise Agent Platform name, do not need an answer in seconds: bulk...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Many Gemini workloads on Vertex AI, now documented under the Gemini Enterprise Agent Platform name, do not need an answer in seconds: bulk classification and tagging, extraction from document archives, translation, catalog enrichment, offline analysis and model evaluation runs.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP Vertex AI](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-vertex-ai)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0568

Type

Suboptimal Pricing Model

## Explanation

Why the waste happens and who it affects.

When these are sent as ordinary real-time requests they are billed at Standard pay-as-you-go token rates, or even at Priority rates, although Google offers two consumption options that price the same tokens at half the Standard rate.

Batch inference takes a large set of requests from Cloud Storage or BigQuery, processes them asynchronously with a 24-hour turnaround target, and bills at a 50 percent discount. Flex PayGo keeps the synchronous API but accepts longer latency and more throttling, also at a 50 percent discount versus Standard PayGo. Both require a deliberate change, a job submission or a request header, so pipelines built on the default online API keep paying Standard rates.

## Billing model

The pricing dimensions that drive this cost.

Gemini usage is billed per million input and output tokens, and the rate depends on the consumption option.

Standard PayGo

The default real-time rate per million tokens; Priority PayGo costs more for latency-critical traffic

Batch inference

Asynchronous jobs billed at a 50 percent discount compared with real-time inference, charged only for completed requests

Flex PayGo

Synchronous requests at a 50 percent discount versus Standard PayGo, with longer latency and higher throttling; currently Preview, for a listed set of models on the global endpoint only

Cached tokens

Implicit caching gives a 90 percent discount on cached input tokens, which takes precedence over and does not stack with the batch discount

## How to detect

4 checks to find it in your estate.

- Inventory Gemini callers that are triggered by schedulers, pipelines, notebooks or evaluation harnesses rather than by a user waiting for a response

- Use aiplatform.googleapis.com/publisher/online\_serving/model\_invocation\_count and token\_count on the PublisherModel resource to size online volume per model, and compare it with the batch inference jobs listed in the project

- Flag projects with high online token volume for non-interactive workloads and no batch inference jobs

- Check request headers in client code: Flex requires X-Vertex-AI-LLM-Shared-Request-Type set to flex, so its absence means offline calls are paying Standard or Priority rates

## How to fix

4 ways to remove the waste.

- Move bulk and scheduled work to batch inference jobs with Cloud Storage or BigQuery input, combining small jobs into larger ones; a job can hold up to 200,000 requests and may queue for up to 72 hours when capacity is busy

- Use Flex PayGo for latency-tolerant synchronous calls on supported models by sending X-Vertex-AI-LLM-Shared-Request-Type: flex, adding X-Vertex-AI-LLM-Request-Type: shared to bypass Provisioned Throughput; allow request timeouts of up to 30 minutes

- Check constraints before switching: batch inference does not support Provisioned Throughput, explicit caching or RAG and is not covered by the SLA, and Flex is Preview, global endpoint only, with its own per-model quota

- Reserve Priority PayGo and Standard real-time calls for interactive, latency-sensitive traffic

## Documentation

Vendor references for pricing and configuration.

- [Batch inference with Gemini  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/capabilities/batch-inference)

- [Flex PayGo  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/flex-paygo)

- [Agent Platform Pricing  cloud.google.com](https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing)

- [Use Provisioned Throughput  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/provisioned-throughput/use-provisioned-throughput)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP Vertex AI  CER-0267

### [Using High-Cost Models for Low-Complexity Tasks in Vertex AI](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-models-for-low-complexity-tasks-bec92)

Vertex AI workloads often include low-complexity tasks such as classification, routing, keyword extraction, metadata parsing, document triage, or summarization of short and simple text. These operations do not require the advanced...

AI

- GCP Vertex AI  CER-0512

### [Underutilized Vertex AI Provisioned Throughput](https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-vertex-ai-provisioned-throughput)

Provisioned Throughput on Vertex AI, now documented under the Gemini Enterprise Agent Platform name, is a fixed-cost, fixed-term subscription that reserves throughput for a specific generative AI model in a specific region. It is bought in...

AI

- GCP Vertex AI  CER-0556

### [Vertex AI Workbench Instances With Idle Shutdown Disabled](https://www.pointfive.co/efficiency-hub/inefficiencies/vertex-ai-workbench-instances-with-idle-shutdown-disabled)

Vertex AI Workbench instances, now documented as Agent Platform Workbench under the Gemini Enterprise Agent Platform name, are notebook VMs that bill for CPU, memory and any attached GPUs whenever they are running. To help manage costs,...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

