Case Studies | SaaS B2B

A reproducible computer-vision platform on AWS

Computer Vision MLOps Platform

About

A SaaS B2B business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The engagement had four focus areas, each captured as a concrete success criterion.

An unified, cost-aware inference API

Both new models needed to sit behind one API that accepts images from Amazon S3 and returns JSON with masks and overlays, with all preprocessing moved inside the model containers so that training and inference behave identically. Given a modest and cost-sensitive traffic profile, inference had to be asynchronous rather than always-on.

Reproducible training with full lineage

Every training run and every inference result had to record the model package, the dataset manifest, the container digest and the code revision, so that retraining from a manifest reproduces equivalent metrics and any result can be audited or compared for a client.

Cost efficiency by design

With around 150 images a day expected in the year ahead, the platform had to lean on Spot training, asynchronous and batch inference, and S3 lifecycle policies rather than expensive standing infrastructure.

Operability and drift control

Once live, the team needed dashboards, alarms and data-drift monitoring so on-call engineers could diagnose issues from the console alone and retraining could be triggered when the incoming data shifts.

Solution

The work was structured as three phases across roughly twelve weeks, each with its own acceptance criteria, taking the platform from an usable minimum through reproducible pipelines to a fully operated and cost-controlled service.

12 wk

Three-phase build to an operated platform

2

New computer-vision models productionised

150/day

Target image throughput the design was sized for

By the numbers:

  • 12 wk - Three-phase build to an operated platform
  • 2 - New computer-vision models productionised
  • 150/day - Target image throughput the design was sized for
Changes

The platform was accepted against its success criteria: an image submitted through the API returns asynchronous outputs in S3 tagged with the model version and dataset manifest, retraining from a manifest reproduces equivalent metrics, and only approved models reach production. The client moved from two locally run experiments to an operated, auditable service. The figures below are the design targets the platform was built and validated against.

  • One cost-aware APIServes both the turf and weed models through asynchronous SageMaker inference, with preprocessing inside the containers so there is no training-to-inference drift.
  • Reproducible, versioned pipelinesRecord the model package, dataset manifest, container digest and code revision for every run, so results are auditable and retraining is deterministic.
  • Cost efficiency by designUses Spot training, asynchronous and batch inference and S3 lifecycle policies to Glacier IR, keeping cost per thousand images inside an agreed envelope.
  • Operable and monitoredGives the team CloudWatch dashboards and alarms plus SageMaker Model Monitor drift detection, so issues are diagnosable from the console and drift can trigger retraining.
  • HandoverDelivered the API contract, a golden test set, a "add a new model" guide, incident runbooks and infrastructure-as-code so the client can extend the platform independently.

With the foundation in place, the client can add new models through the documented path, lean on drift-triggered retraining to keep quality high, and scale comfortably beyond today's throughput as its imagery and customer base grow.

AWS Stack

Amazon SageMaker

For asynchronous and batch inference, Spot training, the Model Registry and Model Monitor that make the machine-learning lifecycle reproducible and observable.

Amazon S3

For the versioned data lake holding raw, processed, synthetic and inference data with lifecycle policies for cost control.

AWS Lambda

For the API router and post-processing that tie the inference flow together.

Amazon API Gateway

For a single, secure entry point to both models.

Amazon DynamoDB

For tracking inference jobs and their status.

Amazon ECR

For the versioned container images that carry preprocessing and model code.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.