# Idle Snowflake Cortex Search Services Without Serving Suspension

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/idle-snowflake-cortex-search-services-without-serving-suspension

A Cortex Search service keeps a search index available for low-latency hybrid (vector and keyword) retrieval, typically behind a RAG application or...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

A Cortex Search service keeps a search index available for low-latency hybrid (vector and keyword) retrieval, typically behind a RAG application or chatbot.

PointFive Research

Cloud cost research at PointFive

Snowflake service

[Snowflake Cortex AI](https://www.pointfive.co/efficiency-hub/cloud-services/snowflake-cortex-ai)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0519

Type

Idle or Unused Resource

## Explanation

Why the waste happens and who it affects.

Its serving layer is billed on the size of the indexed data for as long as serving is running, and Snowflake states this cost is incurred while the service is available to answer queries even if no queries are served. A service nobody calls therefore bills the same serving credits as one handling steady traffic.

The waste shows up with prototype and proof-of-concept RAG services that are never dropped, services whose calling application has been retired or pointed at a newer service, and services used only during business hours or in occasional batch evaluations. Serving is not auto-suspended by default, so unless AUTO\_SUSPEND is set or someone suspends serving, an idle service keeps billing indefinitely. The indexing side adds its own cost when source tables keep changing, because refreshes run on a warehouse and new or updated rows are re-embedded, even if no one reads the results.

## Billing model

The pricing dimensions that drive this cost.

Cortex Search bills several components separately; serving is the one that accrues with no activity.

Serving compute

Billed in AI Credits per GB per month of uncompressed indexed data (source text plus vector embeddings) for as long as serving is running, whether or not queries arrive; the Snowflake Service Consumption Table lists Cortex Search at 6.3 AI Credits per GB/mo of indexed data

Indexing warehouse compute

Standard warehouse credits consumed by the service's warehouse when a refresh detects changes in the base objects and rebuilds the index

EMBED\_TEXT tokens

AI Credits charged per token embedded each time a row in the source query is inserted or updated

Storage

Materialized source data and index structures stored in the account at the flat per-TB storage rate

## How to detect

5 checks to find it in your estate.

- Run DESCRIBE CORTEX SEARCH SERVICE for each service returned by SHOW CORTEX SEARCH SERVICES and list those whose serving\_state is RUNNING

- Query SNOWFLAKE.ACCOUNT\_USAGE.CORTEX\_SEARCH\_SERVING\_USAGE\_HISTORY (hourly serving credits per service) or CORTEX\_SEARCH\_DAILY\_USAGE\_HISTORY filtered to CONSUMPTION\_TYPE = 'SERVING' to find services with continuous serving credits

- Confirm whether those services receive requests: with REQUEST\_LOGGING enabled, each logged search request produces a row in the SNOWFLAKE.LOCAL.AI\_OBSERVABILITY\_EVENTS event table, readable through SNOWFLAKE.LOCAL.GET\_AI\_OBSERVABILITY\_EVENTS; otherwise check with the owning application team

- Flag services with steady serving credits and no requests over a representative period, and services whose traffic is limited to predictable windows such as business hours or scheduled batch runs

- Check whether AUTO\_SUSPEND is set on intermittently used services; it is off by default

## How to fix

5 ways to remove the waste.

- Drop services that no longer back any application with DROP CORTEX SEARCH SERVICE, after confirming no agent, app or pipeline still references them

- Suspend serving on services that are idle for known periods with ALTER CORTEX SEARCH SERVICE \<name\> SUSPEND SERVING, and RESUME SERVING before use; while serving is suspended the service cannot answer queries

- For intermittent traffic, set ALTER CORTEX SEARCH SERVICE \<name\> SET AUTO\_SUSPEND = \<seconds\> so serving is suspended after a period of query inactivity and resumed on the next query; the minimum is 1800 seconds and it affects serving only, not indexing

- Where freshness is not needed, suspend indexing (SUSPEND INDEXING) or lengthen TARGET\_LAG to reduce refresh warehouse credits and re-embedding

- Reduce the indexed data size, for example by narrowing the source query to the rows and columns actually searched, since serving is billed per GB of indexed data

## Documentation

Vendor references for pricing and configuration.

- [Understanding cost for Cortex Search Services  docs.snowflake.com](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-search/cortex-search-costs)

- [ALTER CORTEX SEARCH SERVICE  docs.snowflake.com](https://docs.snowflake.com/en/sql-reference/sql/alter-cortex-search)

- [DESCRIBE CORTEX SEARCH SERVICE  docs.snowflake.com](https://docs.snowflake.com/en/sql-reference/sql/desc-cortex-search)

- [CORTEX\_SEARCH\_SERVING\_USAGE\_HISTORY view  docs.snowflake.com](https://docs.snowflake.com/en/sql-reference/account-usage/cortex_search_serving_usage_history)

- [Monitor Cortex Search requests  docs.snowflake.com](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-search/cortex-search-monitor)

- [Snowflake Service Consumption Table  snowflake.com](https://www.snowflake.com/legal-files/CreditConsumptionTable.pdf)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Snowflake Cortex AI  CER-0518

### [Using High-Cost Models for Low-Complexity Tasks in Snowflake Cortex](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-models-for-low-complexity-tasks-in-snowflake-cortex)

AI\_COMPLETE lets each SQL call name any available model, from small open models to frontier models from Anthropic, OpenAI, Google and others, and the model choice sets the token rate. Teams often pick a frontier model once for a prototype...

AI

- AWS SageMaker  CER-0333

### [Idle SageMaker Notebook Instances Left Running Continuously](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-sagemaker-notebook-instances-left-running-continuously)

SageMaker notebook instances are billed continuously while in an active state - and critically, they do not automatically shut down when idle. Closing a browser tab, shutting down a Jupyter kernel, or simply walking away does not stop the...

AI

- AWS S3  CER-0330

### [Orphaned MLflow Training Artifacts and Model Checkpoints in Object Storage](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-mlflow-training-artifacts-and-model-checkpoints-in-object-storage)

Machine learning experimentation workflows - particularly those managed through experiment tracking platforms - generate large volumes of artifacts in object storage. Every training run produces model checkpoints, evaluation outputs,...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

