A Middle East financial-services AI firm
AWS-validated reference — full details available to AWS Partner Validation or on request. (Customer name held on file with VeUP.)
VeUP migrated the customer financial-services AI workload off another cloud to AWS, onto Amazon SageMaker with Capacity Blocks for ML for predictable GPU cost and AWS Graviton for cheap inference, under a multi-account FinOps structure with CUR visibility that put the GPU-heavy spend under control.
The challenge
the customer, a Dubai-based AI-driven financial-services firm, ran its GPU/ML workload on a different cloud provider and needed to move to AWS to scale on AWS's AI/ML ecosystem and unlock AWS credits/EDP — while bringing a GPU-heavy spend profile under disciplined management. Uncontrolled on-demand GPU is the most expensive way to run a bursty ML workload, and a single-account billing posture gave no scoped, non-root cost visibility. A Marketplace-vs-direct spend mis-attribution further distorted the cost picture.
The solution
A cloud-to-AWS migration across all phases (assess->mobilize->migrate) per the the customer Migration SOW, landing the AI workload on Amazon SageMaker. GPU capacity for training/inference is reserved through Capacity Blocks for ML (predictable GPU cost vs uncontrolled on-demand); inference runs on AWS Graviton for better price-performance. A multi-account payer/child structure with managed AWS credits/EDP, Amazon Cost Explorer, and a CUR-ingest cost platform gives the firm scoped (non-root) cost visibility and a monthly optimization cadence. VeUP runs the engagement as managed billing / FinOps.
Production outcomes
| KPI | Result |
|---|---|
| Production outcomes | FinServ AI workload migrated from another cloud to AWS and running in production under a live resell relationship (~$360K). FinOps normalized a Marketplace-vs-direct spend mis-attribution that had distorted the cost picture, managed AWS credit/EDP coverage across on-demand GPU/inference usage, and shifted the compute mix toward Capacity Blocks for ML (predictable GPU cost) and AWS Graviton (price-performant inference) — reducing the effective run cost of the AI workload, tracked in CUR dashboards. |
| Engagement window | 2024-12-13 (Resell Customer Live); migration kickoff 2024-12-16 → 2025-01-30 (AWS Referral Customer Live); migration target 2025-01-31; ongoing managed billing/FinOps |
| Cost / TCO posture | Cost modeling is the heart of the engagement: VeUP runs managed billing / FinOps on a monthly cadence, modeling and optimizing the AI workload's cost via Capacity Blocks for ML (predictable GPU pricing vs on-demand), AWS Graviton (price-performant inference), managed AWS credits/EDP coverage, and a CUR-ingest cost platform + Cost Explorer for visibility across a multi-account payer/child structure. |
| Lessons & continuation | For a GPU-heavy FinServ AI workload, the controlling cost levers are reserving GPU via Capacity Blocks for ML (not on-demand) and shifting inference to Graviton; a multi-account payer/child + CUR structure is the precondition for scoped, non-root cost discipline; correcting a Marketplace-vs-direct spend mis-attribution before optimizing avoids modeling against a distorted baseline. |
Amazon SageMaker · Capacity Blocks for ML · AWS Graviton · AWS Cost Explorer · AWS Organizations (payer/child) · AWS credits / EDP
Architecture
AWS Well-Architected view of the GPU-heavy FinServ AI workload — from the source-cloud, on-demand-GPU, single-account baseline through to the production AWS Organizations payer/child structure, multi-AZ SageMaker training and Graviton inference, and the security and FinOps rails that engineered the cost profile down.