Skip to content
Cloud Efficiency Hub

Using High-Cost Models for Low-Complexity Tasks in Snowflake Cortex

The short version

AI_COMPLETE lets each SQL call name any available model, from small open models to frontier models from Anthropic, OpenAI, Google and others, and the model choice sets the token rate.

PointFive Research

Cloud cost research at PointFive

Snowflake service
Snowflake Cortex AI
Category
AI
Reference
CER-0518
Type
Oversized Model Selection

Explanation

Why the waste happens and who it affects.

Teams often pick a frontier model once for a prototype and then run it over entire tables for labeling, classification, extraction or short summaries that a smaller model handles as well. Because Cortex AI Functions are designed for throughput over large tables, a model choice that looks trivial per row multiplies across millions of rows.

The rate gap is wide. In Snowflake's Service Consumption Table effective September 28, 2026, AI_COMPLETE with openai-gpt-6-astra is billed at 6.00 input and 30.00 output AI Credits per million tokens, claude-opus-5-5 at 2.40 and 12.00, claude-haiku-4-5 at 0.60 and 3.00, and llama3.1-8b at 0.132 for both. Snowflake's own model guide says that to achieve the best performance per credit, users should choose a model that matches the content size and complexity of the task. Task-specific functions such as AI_CLASSIFY have their own flat rates, so they are not automatically cheaper than a small model through AI_COMPLETE and should be compared too.

Billing model

The pricing dimensions that drive this cost.

Cortex AI Functions run on Snowflake-managed compute and bill in AI Credits.

AI_COMPLETE tokens
Billed per million input and output tokens at a rate set per model in the Service Consumption Table
Task-specific functions
Functions such as AI_CLASSIFY, AI_FILTER and AI_SUMMARIZE have their own per-million-token rates and add a prompt, so billed tokens exceed the text supplied
Calling warehouse
The warehouse running the query still bills credits while the query runs, separately from token charges

How to detect

4 checks to find it in your estate.

  • Query SNOWFLAKE.ACCOUNT_USAGE.CORTEX_AI_FUNCTIONS_USAGE_HISTORY grouped by FUNCTION_NAME, MODEL_NAME, USER_ID and QUERY_TAG to see which models drive AI credits
  • Flag large batch AI_COMPLETE jobs that use frontier models (for example gpt-6-astra, Opus or Fable class models) for labeling, classification, extraction or short summaries
  • Use AI_COUNT_TOKENS on sample inputs to estimate tokens per row and project the cost difference between candidate models
  • Check whether model access is unrestricted, meaning any role can call the most expensive models

How to fix

5 ways to remove the waste.

  • Benchmark smaller models on a labeled sample of the real table and switch the SQL to the cheapest model that meets the accuracy bar
  • Compare task-specific functions such as AI_CLASSIFY or AI_EXTRACT against a small model through AI_COMPLETE on cost per row, since their flat rates can be higher or lower
  • Route by difficulty: run a small model first and send only low-confidence or failed rows to a larger model
  • Restrict expensive models with model RBAC on objects in SNOWFLAKE.MODELS, which Snowflake recommends over the CORTEX_MODELS_ALLOWLIST parameter that is being deprecated
  • Shorten prompts and constrain output length, since both input and output tokens are billed for generative functions

Documentation

Vendor references for pricing and configuration.