Case Studies | Retail

From on-prem GPUs to elastic AWS training

SageMaker Training and Tuning Migration

About

A Retail business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas, each about turning a fixed on-premises setup into a scalable, well-governed AWS training pipeline.

Elastic GPU compute

The client's training and tuning were bound to two RTX 3090 cards. The migration needed to provide comparable and greater GPU capacity on demand, from a single-GPU instance for standard runs through to a multi-GPU instance for heavier hyperparameter tuning, so experiments were no longer queued behind fixed hardware.

Faithful reproduction of the existing pipeline

The move had to run the client's own custom forecasting repository and full Python dependency stack, including a private statsforecast dependency, and produce results comparable to the on-premises environment. A migration that changed the answers would be no migration at all.

Secure, repeatable, self-serviceable infrastructure

Everything had to be provisioned as code so the environment could be redeployed reliably, with IAM, VPC and S3 access configured to least privilege, and monitoring in place so the client could see cost and job health at a glance and run it without us.

Solution

We delivered the migration as a sequence of milestones, each with a clear deliverable and acceptance point, so the client could see progress and validate results at every stage.

17 days

Engagement to a production-ready SageMaker pipeline

96 GB

Multi-GPU VRAM on demand for tuning (4x NVIDIA A10G)

100%

Infrastructure provisioned as code via CloudFormation

By the numbers:

  • 17 days - Engagement to a production-ready SageMaker pipeline
  • 96 GB - Multi-GPU VRAM on demand for tuning (4x NVIDIA A10G)
  • 100% - Infrastructure provisioned as code via CloudFormation
Changes

All three focus areas were delivered and accepted. The client's demand-forecasting training and tuning now run on Amazon SageMaker, reproduced faithfully from their existing repository, on infrastructure that is entirely defined as code and monitored end to end.

  • Elastic GPU computeThe client moved from two fixed RTX 3090 cards to on-demand SageMaker GPU tiers, with a multi-GPU ml.g5.12xlarge (4x NVIDIA A10G, 96 GB VRAM) for hyperparameter tuning and single-GPU ml.g5.xlarge for standard runs.
  • Faithful reproductionTheir custom forecasting repository and full Python dependency stack, including the private statsforecast dependency, run inside SageMaker, with tuning results validated as comparable to the on-premises environment.
  • Repeatable infrastructureThe entire environment, SageMaker Studio, S3, VPC, IAM and CloudWatch, deploys from a single CloudFormation template with least-privilege access throughout.
  • Full observabilityCloudWatch dashboards and alarms give the client a live view of training runtime, cost and job failures across the pipeline.
  • HandoverThe client received a runbook, reference architecture and a knowledge-transfer session, and signed off that they can operate, maintain and scale the pipeline independently.

With training no longer bound to the hardware on the desk, the client can scale experiments up and down as their models demand, and the GPU spend on SageMaker also builds toward the NVIDIA Inception Programme credits the team is pursuing. The pipeline is no longer capped by fixed on-premises GPUs, but set up to support whatever the client takes on next.

AWS Stack

Amazon SageMaker

For managed, on-demand GPU training and hyperparameter tuning that scales with the workload.

Amazon S3

For durable storage of training data, model artefacts, tuning outputs and logs.

AWS CloudFormation

For provisioning the entire environment as repeatable infrastructure as code.

Amazon CloudWatch

For dashboards, alarms and logging across training jobs and infrastructure health.

Amazon VPC

And AWS IAM for private networking and least-privilege access control around the training environment.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.