# Using High-Cost Models for Low-Complexity Tasks in Snowflake Cortex

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-models-for-low-complexity-tasks-in-snowflake-cortex

AI_COMPLETE lets each SQL call name any available model, from small open models to frontier models from Anthropic, OpenAI, Google and others, and the...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

AI\_COMPLETE lets each SQL call name any available model, from small open models to frontier models from Anthropic, OpenAI, Google and others, and the model choice sets the token rate.

PointFive Research

Cloud cost research at PointFive

Snowflake service

[Snowflake Cortex AI](https://www.pointfive.co/efficiency-hub/cloud-services/snowflake-cortex-ai)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0518

Type

Oversized Model Selection

## Explanation

Why the waste happens and who it affects.

Teams often pick a frontier model once for a prototype and then run it over entire tables for labeling, classification, extraction or short summaries that a smaller model handles as well. Because Cortex AI Functions are designed for throughput over large tables, a model choice that looks trivial per row multiplies across millions of rows.

The rate gap is wide. In Snowflake's Service Consumption Table effective September 28, 2026, AI\_COMPLETE with openai-gpt-6-astra is billed at 6.00 input and 30.00 output AI Credits per million tokens, claude-opus-5-5 at 2.40 and 12.00, claude-haiku-4-5 at 0.60 and 3.00, and llama3.1-8b at 0.132 for both. Snowflake's own model guide says that to achieve the best performance per credit, users should choose a model that matches the content size and complexity of the task. Task-specific functions such as AI\_CLASSIFY have their own flat rates, so they are not automatically cheaper than a small model through AI\_COMPLETE and should be compared too.

## Billing model

The pricing dimensions that drive this cost.

Cortex AI Functions run on Snowflake-managed compute and bill in AI Credits.

AI\_COMPLETE tokens

Billed per million input and output tokens at a rate set per model in the Service Consumption Table

Task-specific functions

Functions such as AI\_CLASSIFY, AI\_FILTER and AI\_SUMMARIZE have their own per-million-token rates and add a prompt, so billed tokens exceed the text supplied

Calling warehouse

The warehouse running the query still bills credits while the query runs, separately from token charges

## How to detect

4 checks to find it in your estate.

- Query SNOWFLAKE.ACCOUNT\_USAGE.CORTEX\_AI\_FUNCTIONS\_USAGE\_HISTORY grouped by FUNCTION\_NAME, MODEL\_NAME, USER\_ID and QUERY\_TAG to see which models drive AI credits

- Flag large batch AI\_COMPLETE jobs that use frontier models (for example gpt-6-astra, Opus or Fable class models) for labeling, classification, extraction or short summaries

- Use AI\_COUNT\_TOKENS on sample inputs to estimate tokens per row and project the cost difference between candidate models

- Check whether model access is unrestricted, meaning any role can call the most expensive models

## How to fix

5 ways to remove the waste.

- Benchmark smaller models on a labeled sample of the real table and switch the SQL to the cheapest model that meets the accuracy bar

- Compare task-specific functions such as AI\_CLASSIFY or AI\_EXTRACT against a small model through AI\_COMPLETE on cost per row, since their flat rates can be higher or lower

- Route by difficulty: run a small model first and send only low-confidence or failed rows to a larger model

- Restrict expensive models with model RBAC on objects in SNOWFLAKE.MODELS, which Snowflake recommends over the CORTEX\_MODELS\_ALLOWLIST parameter that is being deprecated

- Shorten prompts and constrain output length, since both input and output tokens are billed for generative functions

## Documentation

Vendor references for pricing and configuration.

- [Cost considerations for Cortex AI Functions  docs.snowflake.com](https://docs.snowflake.com/en/user-guide/snowflake-cortex/aisql-cost)

- [Snowflake Service Consumption Table  snowflake.com](https://www.snowflake.com/legal-files/CreditConsumptionTable.pdf)

- [Managing Cortex AI Function costs with Account Usage  docs.snowflake.com](https://docs.snowflake.com/en/user-guide/snowflake-cortex/ai-func-cost-management)

- [Privileges and model access for Cortex AI Functions  docs.snowflake.com](https://docs.snowflake.com/en/user-guide/snowflake-cortex/aisql-privileges-and-access)

- [Models and regional availability for Cortex AI Functions  docs.snowflake.com](https://docs.snowflake.com/en/user-guide/snowflake-cortex/aisql-regional-availability)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Snowflake Cortex AI  CER-0519

### [Idle Snowflake Cortex Search Services Without Serving Suspension](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-snowflake-cortex-search-services-without-serving-suspension)

A Cortex Search service keeps a search index available for low-latency hybrid (vector and keyword) retrieval, typically behind a RAG application or chatbot. Its serving layer is billed on the size of the indexed data for as long as serving...

AI

- GCP Vertex AI  CER-0267

### [Using High-Cost Models for Low-Complexity Tasks in Vertex AI](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-models-for-low-complexity-tasks-bec92)

Vertex AI workloads often include low-complexity tasks such as classification, routing, keyword extraction, metadata parsing, document triage, or summarization of short and simple text. These operations do not require the advanced...

AI

- AWS Bedrock  CER-0276

### [Using High-Cost Bedrock Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-bedrock-models-for-low-complexity-tasks-734ca)

Many Bedrock workloads involve low-complexity tasks such as tagging, classification, routing, entity extraction, keyword detection, document triage, or lightweight summarization. These tasks do not require the advanced reasoning or...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

