Many enterprises route every AI request, from a simple email summary to a complex reasoning task, through the most powerful and expensive model available. Snowflake is looking to fix that waste with dynamic model routing, a new capability inside its Cortex AI Gateway that matches each prompt to the most cost-effective model automatically.
A Traffic Controller for AI Models
Rather than sending every prompt to a heavy-duty flagship model, dynamic model routing evaluates each task in real time. Simple or repetitive jobs get directed to lighter, cheaper models, while genuinely complex reasoning problems are escalated to premium frontier models.
For enterprise teams, that means budgets are no longer spent on compute power that a task never actually needed, and developers don't have to hand-code routing rules for every possible request.
Measurable Cost and Quality Gains
The feature runs quietly in the background across Snowflake's CoCo and CoWork products, as well as connected third-party AI agents, balancing quality needs against cost the moment a prompt is sent.
Internal testing found AI agents using dynamic routing built data transformation pipelines with up to three times greater token efficiency than relying on frontier models alone, while holding output quality steady. Development teams also completed the same volume of pull requests using 25% fewer tokens.
The system also gives IT leaders regional and regulatory control, letting them restrict which models or geographic providers are accessible so compliance holds even as pricing or availability shifts.
More Model Choice, Same Governance
Smart routing only works if there's a rich pool of models to choose from. Snowflake is widening access to open models inside its secure data perimeter, adding options like DeepSeek-V4-Flash and GLM-5.3.
These aren't lightweight fallbacks. In Snowflake's own testing, DeepSeek-V4-Flash scored 74.4% on enterprise data engineering tasks, outperforming several proprietary rivals, while GLM-5.3 delivered strong results while using notably fewer tokens. Both join an existing library that includes models from Anthropic, OpenAI, Google, SpaceXAI, Meta and Mistral, so developers don't have to rebuild integrations every time a new model launches.
"the winning systems will be the ones that can adapt quickly"— BARIS GULTEKIN, VP OF AI, SNOWFLAKE
