Start Building →
Guide

CI/CD for ML Pipelines: What's Different, and How SageMaker, Kubeflow, and Vertex AI Compare

Vertex AI Pipelines runs on the Kubeflow SDK, not a separate engine. Here is what actually changes when CI/CD meets machine learning.

Author photo
packet.ai Team
September 30, 2026

Standard CI/CD is built around testing changes to code against known inputs and expected behavior. A model depends on its training data and the data distribution it eventually sees in production, neither of which is captured by versioning code alone. This guide covers what genuinely changes when cicd for machine learning gets applied in practice, and how that maps onto the three platforms most teams actually use: SageMaker Pipelines, Kubeflow Pipelines, and Vertex AI Pipelines.

Key takeaways

  • Mlops pipelines often add a concern beyond standard CI/CD's continuous integration and deployment: continuous training (CT), where retraining is triggered by new data, detected drift, performance changes, or a fixed schedule, since a model degrades as real-world data drifts away from its training distribution
  • ML pipelines need to version data and trained model artifacts alongside code; reliably reproducing a model can depend on the dataset version, parameters, random seeds, and execution environment, not just the training code itself
  • Vertex ai pipeline supports pipelines authored with the kubeflow pipelines SDK, providing a managed, serverless execution environment rather than requiring teams to operate the underlying Kubernetes infrastructure themselves; the KFP SDK defines the pipeline, and Vertex AI provides the execution engine and Google Cloud integrations around it
  • Sagemaker pipeline takes a different architectural approach from both: SageMaker manages the pipeline orchestration infrastructure for you, and pipeline steps become managed SageMaker jobs rather than Kubernetes workloads, trading some portability for less operational overhead
  • ML CI adds a layer of checks beyond standard software tests: validating that incoming data matches the expected schema, and gating deployment on model evaluation metrics clearing a defined threshold, since model quality is measured on a continuous scale rather than a single pass/fail assertion

Why CI/CD for Machine Learning Isn't Just CI/CD

Standard CI/CD is built around testing changes to code against known inputs and expected behavior. Machine learning adds dependencies that aren't captured by code alone. The training dataset affects the resulting model, the trained model itself becomes an artifact that needs to be tracked, and its quality is usually evaluated through metrics rather than a simple pass/fail assertion the way testing a function's return value works. A model can also degrade after deployment as the data it sees in production shifts away from what it was trained on, something that doesn't have a direct equivalent in traditional software, which doesn't erode on its own purely from the passage of time.

This is why mlops pipelines often add a concern beyond the two that define standard CI/CD. Continuous integration and continuous deployment remain, but continuous training (CT) gets added specifically to handle retraining, typically triggered by new data arriving, by detected drift in the data or the model's performance, or on a fixed schedule, rather than purely by a code change the way CI is normally triggered.

What Actually Gets Versioned, and Why It's Harder

In standard CI/CD, versioning code is sufficient because code fully determines behavior. In an mlops pipeline, code alone doesn't reliably reproduce a model; the relevant dataset version, model artifact, parameters, random seeds, and execution environment may all matter, since a model trained on a slightly different dataset version, or even a different library version, is for reproducibility purposes a different model even with identical training code. A dvc pipeline is one common way teams solve the data side of this: tools like DVC extend Git-style version control to large datasets that don't fit well into a standard code repository, giving data the same kind of change tracking code already gets.

Skipping this is one of the most common, specific failure modes in ML CI/CD setups: a team versions the model but not the exact dataset it was trained on, and loses the ability to reproduce that model later when the underlying data has since changed or been deleted. This isn't a hypothetical edge case; it's a direct consequence of treating ML pipelines as standard software pipelines with a training step bolted on, rather than designing the versioning strategy around what ML specifically requires.

⚡ Vertex AI Pipelines uses Kubeflow's pipeline model

A useful distinction that comparisons often miss: Vertex AI Pipelines can execute pipelines authored with the Kubeflow Pipelines SDK, including KFP 2.x definitions, in a managed serverless environment, confirmed directly in Google's own architecture documentation. The KFP SDK defines the pipeline; Vertex AI provides the managed execution engine, scheduling, and Google Cloud integrations around it, including automatic lineage tracking through Vertex ML Metadata. That makes self-managed Kubeflow and Vertex AI related but not interchangeable deployment choices: self-managed KFP runs on Kubernetes infrastructure you operate yourself, while Vertex AI handles that underlying infrastructure for you.

SageMaker, Kubeflow, and Vertex AI: The Three Platforms Compared

PlatformPipeline executionInfrastructureMain tradeoff
Kubeflow PipelinesKubernetes workloadsSelf-managed or managed KubernetesMore control, more infrastructure responsibility
SageMaker PipelinesManaged SageMaker jobs and workflow stepsAWS-managedLess infrastructure work, stronger AWS coupling
Vertex AI PipelinesManaged execution of KFP/TFX pipeline definitionsGoogle-managedLess infrastructure work, stronger GCP integration

Kubeflow Pipelines runs ML workflows as native Kubernetes workloads, giving teams direct control over the underlying cluster at the cost of managing that cluster themselves, a meaningful operational burden that's the clearest tradeoff against its portability. Its artifact and metadata model, tracking executions, lineage, and artifacts across pipeline runs, is also a meaningful part of why it's useful for reproducibility specifically, not just an orchestration convenience. Sagemaker pipeline, also called aws sagemaker pipeline or amazon sagemaker pipelines depending on which naming convention a given team uses, takes a genuinely different architectural approach: SageMaker manages the pipeline orchestration infrastructure for you, and individual pipeline steps become managed SageMaker jobs rather than raw Kubernetes workloads, which means less infrastructure to operate directly, in exchange for tighter coupling to AWS specifically. Vertex ai pipeline, as covered above, provides managed execution for pipelines authored with Kubeflow Pipelines' own SDK, which positions it as the serverless, hands-off option specifically for teams already working within Google Cloud's ecosystem. For many teams, the practical decision among these three is driven as much by existing infrastructure and cloud commitments as by differences in pipeline capability, alongside factors like Kubernetes expertise, portability requirements, and existing data platform choices.

Testing and Quality Gates Look Different for ML

ML CI has to validate things a standard software CI pipeline doesn't check on its own: whether incoming data matches the schema the model expects, and whether the model's evaluation metrics clear a defined quality threshold before it's allowed to proceed toward deployment. The usual software tests, unit tests, integration tests, schema validation, still matter and don't go away. ML adds another layer of checks on top because model quality is measured through metrics rather than a single assertion about whether a function returned the expected value.

A pipeline that trains a new model version but doesn't gate on these metrics before deployment is missing the check that actually prevents a worse-performing model from reaching production, which is a meaningfully different risk than a software CI pipeline's typical concern of a build simply failing to compile. Mlops continuous training, or machine learning continuous training more broadly, specifically depends on this gate existing, since retraining automatically on a schedule or a drift trigger without a quality check in between means a bad retraining run can replace a working model with a worse one, silently.

Why Training Steps Often Need Different Infrastructure Than the Rest of the Pipeline

Model training and hyperparameter optimization typically take meaningfully longer than most software builds, and often require specific hardware, GPUs, that a standard CI runner generally doesn't provide. This is one practical reason teams often separate expensive training workloads from the CI system that handles code changes and testing, rather than running every retraining job inside the same CI infrastructure used for standard code changes: cost and runtime control, not just architectural preference.

This separation also means the GPU-provisioning side of an ML pipeline is a genuinely distinct infrastructure concern from the orchestration layer itself, whichever of the three platforms is handling that orchestration. The pipeline tool schedules and tracks the training step; something else actually needs to provide the GPU capacity that step runs on.

Where GPU Capacity Fits Into This Picture

The orchestrator decides when and how a training job runs; the GPU layer determines whether that job actually has the compute, memory, and runtime environment it needs. Whichever orchestration platform a pipeline uses, the training steps inside it still need GPU capacity to run, and that capacity is a separate decision from which pipeline tool manages the workflow. packet.ai provides GPU cloud infrastructure that a training step from any of these pipeline tools can run on, independent of which specific orchestrator is managing the surrounding workflow.

Request a quote to size the right GPU configuration for your training workloads.

Sources and Further Reading

Frequently asked questions

ML adds dependencies and evaluation requirements that standard software CI/CD doesn't normally have to manage. The training dataset, model artifact, and evaluation metrics all matter alongside the code, and deployed models can change in quality as production data shifts. This is why ML pipelines often add data validation, model evaluation, and continuous training alongside traditional CI/CD.
Not identical, but closely related. Vertex AI Pipelines can execute pipelines authored with the Kubeflow Pipelines SDK, including KFP 2.x definitions, in a managed, serverless environment. The KFP SDK defines the pipeline; Vertex AI provides the managed execution engine and Google Cloud integrations around it, per Google's own documentation.
SageMaker Pipelines is a fully managed orchestration service where SageMaker manages the underlying infrastructure and pipeline steps run as managed SageMaker jobs. Self-managed Kubeflow runs pipelines as native Kubernetes workloads on a cluster the team manages itself, trading SageMaker's lower operational overhead for Kubeflow's greater portability and infrastructure control.
Cost and runtime control. Training and hyperparameter optimization typically take much longer than standard software builds and often need GPU hardware that standard CI runners don't provide. Separating the training pipeline from CI lets teams control when and how expensive GPU time gets used, rather than running every retraining job on every pipeline trigger.

Last reviewed: September 30, 2026. MLOps platform capabilities and pricing change frequently; verify current features and costs against the official documentation for SageMaker, Kubeflow, and Vertex AI before committing to one for production use.

Waste less compute.

Same models. Same API. Fraction of the cost. Start free — no credit card required.

Start Building →

More from the blog