VeUP
VeUP Frontier Practice

Every engagement lands on Bedrock. The model has to earn it.

Claude, GPT, open weights and first-party Nova all arrive through one governed channel — and which one ships is settled by benchmark on the customer's own labelled data, workload by workload. Pick a provider to see the proof behind it.

Every engagement lands here
Amazon Bedrock
one governed model channel — metered, guardrailed, swappable
46
customer accounts in the frontier estate
198
linked AWS accounts underneath them
29
where VeUP is the AWS payer of record
9
published case studies that name their models
What counts as proof here

A published case study carries the production claim — named services, measured outcomes, acceptance criteria. Around it sits the working estate: accounts where a model family is under evaluation or actively deployed.

The exhibit

Five candidates. One labelled dataset. The open model won.

Multimodal content moderation at platform scale for a global consumer social platform. This is what model-agnostic looks like when it is real: VeUP's own partner model was in the run and did not win it.

The bench, as run
  • Qwen 3 VL
    Open-weight
    Selected
  • Meta Llama 4 Maverick
    Open-weight
    Evaluated
  • Meta Llama Guard 4
    Open-weight
    Evaluated
  • Claude Opus
    Anthropic
    Evaluated
  • Amazon Nova Premier
    Amazon
    Evaluated
  • Amazon Rekognition
    Amazon
    Incumbent

One labelled dataset built from the platform's own content. Every candidate invoked through Amazon Bedrock, scored on the same images, under the same prompts, in a Phase 1 evaluation deliverable — with Amazon Rekognition scored as the incumbent it had to beat — which it did on recall, accuracy, F1 and unit cost.

What shipped
54% → 99.58%
NSFW accuracy — prompt engineering alone, no fine-tuning
$0.43
per 1,000 images vs $1.00 on Rekognition
11/11
SOW acceptance criteria passed
Read the case study
Bedrock wins. Bedrock is really the path, really clearly.
CEO of the platform (anonymized; on-call quote permission, replay-verifiable)
The method

How the model actually gets chosen.

01

Land on Bedrock

One governed model channel inside the customer's own AWS account. Every call metered, tagged and budget-enforced; guardrails and observability wired before the first prompt ships.

02

Build the labelled set

From the customer's own data, not a public leaderboard. On the moderation engagement that meant a labelled image set; on the PM SaaS it meant 12,000 real production prompts.

03

Score every candidate identically

Same data, same prompts, same harness — frontier labs, open weights and first-party models side by side. The result is a deliverable the customer keeps, not a slide.

04

Ship the winner, qualify the alternate

The winner takes the workload; a second model is qualified as the tested fallback. Because everything runs through one channel, changing your mind later is a config change with a receipt.