Batch inference jobs¶
Note
Batch inference jobs require snowflake-ml-python version 2.0.0 or later. If you’re using
snowflake-ml-python 1.x, see
Batch inference jobs (snowflake-ml-python 1.x, deprecated).
Use Snowflake batch inference to run efficient, large-scale model inference on static or periodically updated datasets. The batch inference API runs on Snowpark Container Services (SPCS), which provides a distributed compute layer optimized for throughput and cost efficiency.
You start a batch inference job by calling ModelVersion.run_batch on a model version in the
Snowflake Model Registry. The job builds an
inference image, provisions an SPCS job on the compute pool you name, runs inference, writes results
to a stage, and then winds the compute down so you don’t keep paying for idle capacity.
If the model depends on packages from a private PyPI repository, set artifact_repository_map when
you log the model. Snowflake installs those packages when it builds the inference image. You don’t
pass the repository to run_batch. For an explanation and example, see
Use a private PyPI artifact repository.
To run batch inference directly in SQL, use EXECUTE INFERENCE JOB SERVICE.
When to use batch inference¶
Use the run_batch method for workloads that:
- Process images, audio, or video files using multimodal models with unstructured data.
- Execute inference over millions or billions of rows.
- Run inference as a discrete, asynchronous stage in a pipeline.
- Integrate inference as a step within an Airflow DAG or a Snowflake task.
Limitations¶
- The output stage must be a Snowflake internal stage.
- For multimodal use cases, only server-side encryption is supported.
partition_columnisn’t supported for Hugging Face pipeline models or for FUNCTION-type model methods.
Get started¶
Connect to the model registry¶
Connect to the Snowflake Model Registry and get a reference to the model version you want to run:
Run a batch inference job¶
Pass the input rows, a compute pool, and an output location. The call returns an MLJob handle:
By default run_batch submits the job asynchronously (async_=True) and returns right away. Pass
async_=False if you’d rather have the call block until the job finishes.
Job management¶
Use the ML Job APIs to list, inspect, cancel, and delete batch inference jobs:
Note
The result function in the ML Job APIs isn’t supported for batch inference jobs. Read the job’s
output from the output stage instead.
Specify inference data¶
Provide exactly one of the following:
X: a Snowpark DataFrame holding the input rows.input_stage_location: a stage path that already holds the input data.
Passing both, or neither, raises a ValueError.
DataFrame input¶
When you pass X, Snowflake materializes the rows as Parquet files in a reserved subdirectory of
output_spec.stage_location before the job starts, then reads them from there.
Stage input¶
If your input is already staged as Parquet, point input_stage_location at it. Snowflake reads
those files in place and doesn’t copy them:
Important
input_stage_location must not sit inside output_spec.stage_location. Keep the input and output
paths separate so the job doesn’t read its own output.
Unstructured input (multimodal)¶
For unstructured data, reference the files by their fully qualified stage paths in the input DataFrame. The job reads each file and passes its content to the model:
To list all files under a stage path as a DataFrame, use list_stage_files:
Stage support¶
Supported configurations for input:
- Internal stages: all types of internal stages are supported.
- External stages: Amazon S3 only, and the stage must use server-side encryption. Azure Blob Storage and Google Cloud Storage aren’t supported.
Input rows can reference different stages in the same DataFrame, mixing external and internal paths. Each path is resolved independently at read time.
External stages require a one-time admin setup: an S3 storage integration and IAM permissions on the
bucket. For details, see
CREATE STAGE and
Bulk loading from Amazon S3. The role running the batch inference job
must have USAGE on the external stage.
The output stage specified by OutputSpec(stage_location=...) must be an internal stage.
Convert files to a model-compatible format¶
Your model can accept file content in one of the following encodings:
FileEncoding.RAW_BYTESFileEncoding.BASE64FileEncoding.BASE64_DATA_URL
Use the column_handling field of InputSpec to tell the job which columns hold stage paths and
what encoding the model expects. Each entry is a ColumnHandlingOptions mapping with an
input_format and a convert_to value.
Output¶
OutputSpec.stage_location is a base path, not the final directory. Each job writes its results to
a per-job subdirectory named for the unqualified job name:
The unqualified job name is the trailing identifier of the fully qualified job.id, so you can
build the output path from the job handle:
Results are written as Parquet files. Set job_name on run_batch if you want a predictable output
directory instead of a server-generated one.
Save mode¶
OutputSpec.mode controls what happens when files already exist at the output location:
| Value | Behavior |
|---|---|
SaveMode.ERROR | Fail if the output directory already contains files. This is the default. |
SaveMode.OVERWRITE | Replace any existing files at the output directory. |
Pass model parameters¶
If the model’s signature includes parameters defined with
ParamSpec, pass values at inference
time through InputSpec.params. Any parameter you leave out uses its default from the signature.
Partitioned models¶
Run batch inference on a partitioned model by setting InputSpec.partition_column. Each partition
is processed independently, which is useful for models that train or predict per group.
run_batch raises a ValueError if you set partition_column for a Hugging Face pipeline model or
a FUNCTION-type method, or if the partition column collides with a column in the partitioned model’s
output signature.
For more information about partitioned models, see Partitioned models.
Configure the job¶
Job settings are split into three optional blocks that you pass to run_batch.
Resources¶
ResourcesSpec sets the per-node compute limits for the inference containers:
| Field | Description |
|---|---|
cpu_requests | CPU limit for CPU-based inference. Accepts an integer, a fraction, or a string. If omitted, the job tries to use all vCPUs on the node. |
memory_requests | Memory limit. Accepts an integer or a fraction, but requires a unit such as GiB or MiB. If omitted, the job tries to use all memory on the node. |
gpu_requests | GPU limit for GPU-based inference. Accepts an integer or a string. If omitted, the job runs on CPU. |
Inference¶
InferenceSpec tunes how the model is served inside the job:
| Field | Description |
|---|---|
num_workers | Number of workers that handle requests in parallel within one service instance. Determined automatically if omitted. |
max_batch_rows | Maximum number of rows processed in a single batch. Determined automatically if omitted. Larger values can improve throughput. |
engine_options | An EngineOptions instance that selects a custom inference engine. |
EngineOptions takes an engine value and an optional engine_args_override list of command-line
arguments for that engine. The engine value accepts an InferenceEngine enum member or a
case-insensitive string such as "vllm" or "python_generic".
Image build¶
ImageBuildSpec controls the container image build:
| Field | Description |
|---|---|
image_repo | Image repository for the inference image. Uses a default repository if omitted. |
force_rebuild | Whether to rebuild the image even if a matching one already exists. Defaults to False. |
Putting the blocks together¶
function_name selects the model function to call. If you omit it, Snowflake resolves it against
the model’s function list. Use mv.show_functions() to see the available functions.
Best practices¶
Use the sentinel file¶
A job can fail midway for many reasons, which can leave partial data in the output directory. To
mark completion, the job writes a _SUCCESS file to the output directory.
To avoid reading partial or incorrect output:
- Read output data only after the
_SUCCESSfile appears. - Start from an empty output directory.
- Use
OutputSpec(mode=SaveMode.ERROR)so the job fails instead of silently overwriting data.
Other recommendations¶
- Keep
input_stage_locationoutside the output base path. - Use a fully qualified
job_namewhen a downstream step needs to know the output path in advance. - Set
force_rebuild=Trueonly when you need it. Reusing a cached image makes jobs start faster.
Run batch inference in a Snowflake task DAG¶
Note
This feature requires the snowflake.core package.
Use BatchInferenceTask to run a batch inference job inside a
Snowflake task DAG. This is
useful for scheduled or recurring batch inference, for multi-step DAGs that include inference, and
for downstream tasks that read the inference output.
Construct BatchInferenceTask inside a with DAG(...) block, or pass dag= explicitly, and chain
it with other tasks using >>. The task differs from run_batch in a few ways:
- Supply the input with
queryorinput_stage_location, exactly one of the two. The task doesn’t accept a DataFrame. - The job runs synchronously, so the DAG step completes when the job does.
- The DAG must supply a warehouse, either through
DAG(warehouse=...)orwarehouse=on the task. Serverless DAGs aren’t supported. - Each run gets a server-generated job name and writes results under
<output_spec.stage_location>/<job_name>/, so repeated runs don’t overwrite each other.
Use fully qualified stage paths. An unqualified path resolves against the session namespace when the task is built, but against the task owner’s namespace when it runs.
The task publishes its output directory as the task return value, so a successor can read the path instead of hardcoding it:
For more information about Snowflake tasks, see Introduction to tasks.
Examples¶
Use a custom model¶
This example transcribes audio files from a stage with a custom model that wraps a Whisper pipeline.
Use a Hugging Face model¶
Supported Hugging Face tasks have their signatures inferred automatically.
Use a Hugging Face model with vLLM¶
Select vLLM through InferenceSpec(engine_options=EngineOptions(engine=...)).
Text generation¶
Image text to text¶
Troubleshooting¶
Get job logs¶
Get metrics¶
To get metrics for a batch inference job, use one of the following approaches, depending on whether the job still exists.
If the job hasn’t been deleted, use the SPCS_GET_METRICS function, which returns container metrics for the job’s underlying SPCS service:
If the job has been deleted, query your event table directly. The event table retains historical metrics even after the service is dropped:
Replace <JOB_NAME> with the job_name you passed to run_batch, or with the server-generated name
if you didn’t specify one. You can get the generated name from job.id.
Common errors¶
| Symptom | Likely cause |
|---|---|
ValueError about X and input_stage_location | You passed both or neither. Pass exactly one. |
| Job fails because output files already exist | The output directory isn’t empty and mode is SaveMode.ERROR. Use a new directory or SaveMode.OVERWRITE. |
ValueError about partition_column | The model is a Hugging Face pipeline or a FUNCTION-type method, or the partition column collides with the model’s output signature. |
| Output looks truncated | The job may not have finished. Check for the _SUCCESS file before reading results. |