Cortex Agent and Cortex Analyst evaluations: Default metric version and judge model changes (Pending)

Snowflake is changing evaluation defaults because the v1 judge model, claude-4-sonnet, reaches end-of-life on October 14, 2026. It has been in the legacy state since August 12, 2026, so accounts that didn’t use it before that date already can’t run v1 evaluations. The new judge models also have context windows of about 1 million tokens, which reduces failures on long agent traces.

Important

This change isn’t part of a behavior change bundle. It is expected to be enabled by default starting October 13, 2026. Dates are subject to change.

Behavior change

This change affects the following metrics when they omit the setting shown or set it to auto:

  • Cortex Agent system metrics (answer_correctness, logical_consistency, tool_selection_accuracy, and tool_execution_accuracy) and the Cortex Analyst sql_correctness metric: version
  • Cortex Agent custom metrics: model
Before the change:

System metrics default to v1, and custom metrics default to claude-4-sonnet.

After the change:

System metrics default to v3, and custom metrics default to claude-sonnet-4-6. For both, Snowflake uses openai-gpt-5.4 if claude-sonnet-4-6 isn’t allowed or available for your account.

Explicitly specified versions and models don’t change.

Impact and preparation

Scores from different metric versions aren’t directly comparable. Scores for metrics that use a judge model can change. tool_selection_accuracy doesn’t use a judge model, so its scores don’t change. Evaluation costs can also change because judge models have different credit rates; track them with the CORTEX_REST_API_USAGE_HISTORY view.

No action is required to adopt the new defaults. Before rollout:

  1. Test system metrics with version: "v3", and recalibrate thresholds in dashboards, alerts, and CI/CD checks.
  2. To control when you upgrade, set a system metric version, such as v2, or a custom metric model. Metrics pinned to v1 or claude-4-sonnet stop working at end-of-life.
  3. Confirm that claude-sonnet-4-6 or openai-gpt-5.4 is allowed for your account and the role that runs the evaluation, and is available through your cross-region inference configuration. See judge models and cross-region inference.

See also

Ref: 2442