Explanation
Why the waste happens and who it affects.
Teams often pick a frontier model once for a prototype and then run it over entire tables for labeling, classification, extraction or short summaries that a smaller model handles as well. Because Cortex AI Functions are designed for throughput over large tables, a model choice that looks trivial per row multiplies across millions of rows.
The rate gap is wide. In Snowflake's Service Consumption Table effective September 28, 2026, AI_COMPLETE with openai-gpt-6-astra is billed at 6.00 input and 30.00 output AI Credits per million tokens, claude-opus-5-5 at 2.40 and 12.00, claude-haiku-4-5 at 0.60 and 3.00, and llama3.1-8b at 0.132 for both. Snowflake's own model guide says that to achieve the best performance per credit, users should choose a model that matches the content size and complexity of the task. Task-specific functions such as AI_CLASSIFY have their own flat rates, so they are not automatically cheaper than a small model through AI_COMPLETE and should be compared too.
Billing model
The pricing dimensions that drive this cost.
Cortex AI Functions run on Snowflake-managed compute and bill in AI Credits.
- AI_COMPLETE tokens
- Billed per million input and output tokens at a rate set per model in the Service Consumption Table
- Task-specific functions
- Functions such as AI_CLASSIFY, AI_FILTER and AI_SUMMARIZE have their own per-million-token rates and add a prompt, so billed tokens exceed the text supplied
- Calling warehouse
- The warehouse running the query still bills credits while the query runs, separately from token charges
How to detect
4 checks to find it in your estate.
- Query SNOWFLAKE.ACCOUNT_USAGE.CORTEX_AI_FUNCTIONS_USAGE_HISTORY grouped by FUNCTION_NAME, MODEL_NAME, USER_ID and QUERY_TAG to see which models drive AI credits
- Flag large batch AI_COMPLETE jobs that use frontier models (for example gpt-6-astra, Opus or Fable class models) for labeling, classification, extraction or short summaries
- Use AI_COUNT_TOKENS on sample inputs to estimate tokens per row and project the cost difference between candidate models
- Check whether model access is unrestricted, meaning any role can call the most expensive models
How to fix
5 ways to remove the waste.
- Benchmark smaller models on a labeled sample of the real table and switch the SQL to the cheapest model that meets the accuracy bar
- Compare task-specific functions such as AI_CLASSIFY or AI_EXTRACT against a small model through AI_COMPLETE on cost per row, since their flat rates can be higher or lower
- Route by difficulty: run a small model first and send only low-confidence or failed rows to a larger model
- Restrict expensive models with model RBAC on objects in SNOWFLAKE.MODELS, which Snowflake recommends over the CORTEX_MODELS_ALLOWLIST parameter that is being deprecated
- Shorten prompts and constrain output length, since both input and output tokens are billed for generative functions
Documentation
Vendor references for pricing and configuration.
- Cost considerations for Cortex AI Functionsdocs.snowflake.com
- Snowflake Service Consumption Tablesnowflake.com
- Managing Cortex AI Function costs with Account Usagedocs.snowflake.com
- Privileges and model access for Cortex AI Functionsdocs.snowflake.com
- Models and regional availability for Cortex AI Functionsdocs.snowflake.com