www.appliedcompute.com/case-studies/nvidia

AUGUST 5, 2026

Post-Training Collaboration with NVIDIA

Applied Compute supports a thriving open model ecosystem, and we've spent the past several months utilizing NVIDIA Nemotron.

About the company
NVIDIA designs GPUs and the CUDA software stack used for AI training and inference. Nemotron is its family of open models. Visit Site
Industry
Technology

Applied Compute deploys custom Nemotron models in production for our frontier enterprise customers in software engineering, financial services, logistics, and AI-native companies.

We are collaborating on multiple projects that are in service of better open-source models and systems around those models.

De-risking mainline RL runs

The Applied Compute Agent Cloud, AC2, enables the full iteration cycle across evals, data curation, training experiments, and inference in a single platform. This surfaces instability, reward hacking, and recipe problems early, when they still cost a fraction of what they would at full scale.

Production workload statistics

Production workload statistics. Applied Compute runs a large, varied production footprint and shares statistics from it such as number of turns and number of output tokens per turn.

Our public inference benchmark is a preview of this work: production agentic workload profiles (agentic coding, code QA, and office work) drawn from multi-turn deployments and released with an open-source harness for replaying them against inference engines. The profiles come from recorded traces, not fixed input and output lengths. Agentic coding averages about 20 tool turns per trace, with office work averaging 41 turns, and code QA running as high as 200, with 200 to 300 assistant tokens per turn against prompts that start near 10k and grow with every tool result. Replaying them on DeepSeek R1, vLLM and SGLang track each other closely but both lose throughput as concurrency rises, because KV evictions drag the eligible cache hit rate down.

An Ongoing Initiative

Applied Compute runs Nemotron across multiple frontier-enterprise deployments today, and our post-training and inference platform AC2 is built for the multi-turn, agentic workloads Nemotron targets.

As each new Nemotron version ships to customers, this loop is what makes it better, proven on the real workloads people run in production rather than on benchmarks alone.