Standard CI/CD is built around testing changes to code against known inputs and expected behavior. A model depends on its training data and the data distribution it eventually sees in production, neither of which is captured by versioning code alone. This guide covers what genuinely changes when cicd for machine learning gets applied in practice, and how that maps onto the three platforms most teams actually use: SageMaker Pipelines, Kubeflow Pipelines, and Vertex AI Pipelines.
Key takeaways
Standard CI/CD is built around testing changes to code against known inputs and expected behavior. Machine learning adds dependencies that aren't captured by code alone. The training dataset affects the resulting model, the trained model itself becomes an artifact that needs to be tracked, and its quality is usually evaluated through metrics rather than a simple pass/fail assertion the way testing a function's return value works. A model can also degrade after deployment as the data it sees in production shifts away from what it was trained on, something that doesn't have a direct equivalent in traditional software, which doesn't erode on its own purely from the passage of time.
This is why mlops pipelines often add a concern beyond the two that define standard CI/CD. Continuous integration and continuous deployment remain, but continuous training (CT) gets added specifically to handle retraining, typically triggered by new data arriving, by detected drift in the data or the model's performance, or on a fixed schedule, rather than purely by a code change the way CI is normally triggered.
In standard CI/CD, versioning code is sufficient because code fully determines behavior. In an mlops pipeline, code alone doesn't reliably reproduce a model; the relevant dataset version, model artifact, parameters, random seeds, and execution environment may all matter, since a model trained on a slightly different dataset version, or even a different library version, is for reproducibility purposes a different model even with identical training code. A dvc pipeline is one common way teams solve the data side of this: tools like DVC extend Git-style version control to large datasets that don't fit well into a standard code repository, giving data the same kind of change tracking code already gets.
Skipping this is one of the most common, specific failure modes in ML CI/CD setups: a team versions the model but not the exact dataset it was trained on, and loses the ability to reproduce that model later when the underlying data has since changed or been deleted. This isn't a hypothetical edge case; it's a direct consequence of treating ML pipelines as standard software pipelines with a training step bolted on, rather than designing the versioning strategy around what ML specifically requires.
⚡ Vertex AI Pipelines uses Kubeflow's pipeline model
A useful distinction that comparisons often miss: Vertex AI Pipelines can execute pipelines authored with the Kubeflow Pipelines SDK, including KFP 2.x definitions, in a managed serverless environment, confirmed directly in Google's own architecture documentation. The KFP SDK defines the pipeline; Vertex AI provides the managed execution engine, scheduling, and Google Cloud integrations around it, including automatic lineage tracking through Vertex ML Metadata. That makes self-managed Kubeflow and Vertex AI related but not interchangeable deployment choices: self-managed KFP runs on Kubernetes infrastructure you operate yourself, while Vertex AI handles that underlying infrastructure for you.
Kubeflow Pipelines runs ML workflows as native Kubernetes workloads, giving teams direct control over the underlying cluster at the cost of managing that cluster themselves, a meaningful operational burden that's the clearest tradeoff against its portability. Its artifact and metadata model, tracking executions, lineage, and artifacts across pipeline runs, is also a meaningful part of why it's useful for reproducibility specifically, not just an orchestration convenience. Sagemaker pipeline, also called aws sagemaker pipeline or amazon sagemaker pipelines depending on which naming convention a given team uses, takes a genuinely different architectural approach: SageMaker manages the pipeline orchestration infrastructure for you, and individual pipeline steps become managed SageMaker jobs rather than raw Kubernetes workloads, which means less infrastructure to operate directly, in exchange for tighter coupling to AWS specifically. Vertex ai pipeline, as covered above, provides managed execution for pipelines authored with Kubeflow Pipelines' own SDK, which positions it as the serverless, hands-off option specifically for teams already working within Google Cloud's ecosystem. For many teams, the practical decision among these three is driven as much by existing infrastructure and cloud commitments as by differences in pipeline capability, alongside factors like Kubernetes expertise, portability requirements, and existing data platform choices.
ML CI has to validate things a standard software CI pipeline doesn't check on its own: whether incoming data matches the schema the model expects, and whether the model's evaluation metrics clear a defined quality threshold before it's allowed to proceed toward deployment. The usual software tests, unit tests, integration tests, schema validation, still matter and don't go away. ML adds another layer of checks on top because model quality is measured through metrics rather than a single assertion about whether a function returned the expected value.
A pipeline that trains a new model version but doesn't gate on these metrics before deployment is missing the check that actually prevents a worse-performing model from reaching production, which is a meaningfully different risk than a software CI pipeline's typical concern of a build simply failing to compile. Mlops continuous training, or machine learning continuous training more broadly, specifically depends on this gate existing, since retraining automatically on a schedule or a drift trigger without a quality check in between means a bad retraining run can replace a working model with a worse one, silently.
Model training and hyperparameter optimization typically take meaningfully longer than most software builds, and often require specific hardware, GPUs, that a standard CI runner generally doesn't provide. This is one practical reason teams often separate expensive training workloads from the CI system that handles code changes and testing, rather than running every retraining job inside the same CI infrastructure used for standard code changes: cost and runtime control, not just architectural preference.
This separation also means the GPU-provisioning side of an ML pipeline is a genuinely distinct infrastructure concern from the orchestration layer itself, whichever of the three platforms is handling that orchestration. The pipeline tool schedules and tracks the training step; something else actually needs to provide the GPU capacity that step runs on.
The orchestrator decides when and how a training job runs; the GPU layer determines whether that job actually has the compute, memory, and runtime environment it needs. Whichever orchestration platform a pipeline uses, the training steps inside it still need GPU capacity to run, and that capacity is a separate decision from which pipeline tool manages the workflow. packet.ai provides GPU cloud infrastructure that a training step from any of these pipeline tools can run on, independent of which specific orchestrator is managing the surrounding workflow.
Request a quote to size the right GPU configuration for your training workloads.
Last reviewed: September 30, 2026. MLOps platform capabilities and pricing change frequently; verify current features and costs against the official documentation for SageMaker, Kubeflow, and Vertex AI before committing to one for production use.
Same models. Same API. Fraction of the cost. Start free — no credit card required.
Start Building →