Media and Entertainment

Pinterest Brings Conversational AI to Visual Discovery With NVIDIA Blackwell and Dynamo

Image courtesy of Pinterest

Objective

Pinterest is one of the world’s largest visual discovery platforms, helping 640 million monthly active users find inspiration for home design, fashion, recipes, travel, and more. Because every experience on Pinterest is grounded in images, its AI must reason over both visual and language content simultaneously—in real time, at internet scale. After nearly five years of collaboration with NVIDIA spanning recommendation systems and generative AI, Pinterest built its vision-language model (VLM) serving stack on NVIDIA Blackwell, enabling a new class of multimodal AI experiences—from the conversational Pinterest Assistant to multimodal reranking, content safety, and signal generation—across a fleet of 14,000 NVIDIA GPUs.

Customer

Pinterest

Partner

Inferact - vLLM

Use Case

Accelerated Computing Tools & Techniques

Key Takeaways

Scaling Purchase Intent

  • 14,000 NVIDIA GPUs—including Blackwell, Hopper, and previous-generation NVIDIA architectures—help Pinterest turn 80 billion+ monthly searches, 16 billion+ boards, and 96%+ unbranded text searches into high-value signals for AI-powered discovery and shopping. 

Lower AI Costs, Better Ad Performance 

  • Open models post-trained on NVIDIA GPUs using Pinterest data run at less than 8% of previous AI transaction costs, while AI creative optimization delivered a 6% average click-through rate lift. 

Faster Pinterest Assistant Inference 

Turning AI-Powered Visual Discovery Into Business Momentum

Pinterest is a visual discovery platform. It uses AI to help people move from inspiration to action, and then helps advertisers reach those users at the right moment.

When a user browses Pins (a bookmark of an image or video), refines a search, or asks Pinterest Assistant to help plan a room, the underlying AI is doing something far more demanding than answering a text question: It must process images—sometimes thousands of them in a single request—understand their content, relate the images to the user’s intent, and generate a response, all within the latency window that keeps the experience feeling real time.

On the commercial side, those same AI signals help Pinterest improve ad targeting, bidding, creative optimization, and measurement. Pinterest’s proprietary Taste Graph draws on more than 80 billion monthly searches, plus more than 16 billion boards created on the platform. And because more than 96% of text-based searches are unbranded, Pinterest can help advertisers reach people with clear intent before they have chosen a brand or product.

Pinterest’s AI-led strategy is driving measurable business momentum, with the platform reaching 640 million monthly active users, extending a multi-quarter streak of record users and double-digit user growth, and making Gen Z its largest and fastest-growing cohort.

Pinterest

Pinterest’s new AI tools for the discovery era: from ads to personalized shopping

Building on Blackwell to Accelerate AI Inference and Ad Performance

For years, Pinterest ran this kind of intelligence through a mix of recommendation systems and large language models. But as the company expanded into vision-language models (VLMs), a new class of serving challenges emerged. Unlike text-only LLM workloads, VLM inference is prefill-heavy: Encoding image content is computationally expensive, request payloads grow dramatically when they carry multiple images, and multi-turn conversations accumulate visual context across turns, compounding KV cache pressure with every exchange. In production, a single Pinterest request may include thousands of images, each requiring preprocessing, tokenization, and encoding before the language model can act on it.

Serving this class of workload required more than adding GPU capacity. It required disaggregating the distinct computational phases of inference, routing requests intelligently based on visual context, and building custom support for Pinterest’s proprietary image embeddings. And it required a serving AI infrastructure that was flexible enough to expand across not one product surface, but many—recommendation rerankers, content safety systems, OCR pipelines, and conversational agents—without each team rebuilding the same foundation from scratch.

Pinterest built its platform  around two foundational choices: NVIDIA Blackwell B200 GPUs as the compute layer, and NVIDIA Dynamo as the serving orchestration framework.

Pinterest standardized on B200 GPUs for their industry-leading inference economics. Blackwell’s architectural advances in low-precision data formats, memory bandwidth, and transformer engine made it the right foundation for workloads where prefill is the dominant cost. In preliminary benchmarking for Pinterest Assistant, B200 instances delivered more than 2x latency improvement over Hopper-generation hardware.

For orchestration, Pinterest selected NVIDIA Dynamo for its inference-engine-agnostic design, native compatibility with Pinterest’s service discovery infrastructure, and a high-performance router. Equally important was the close working relationship Pinterest built with the NVIDIA Dynamo team, who provided hands-on support throughout development. The result is an end-to-end generative AI serving platform centered on Dynamo, with vLLM as the inference engine that Pinterest’s teams across the company can now build on through a common Chat Completions API.

That platform approach supported Pinterest’s cost and monetization agenda. Pinterest achieves inference transaction costs at less than 8% of comparable closed proprietary models by post-training open models on NVIDIA GPUs using its unique data, giving the company headroom to expand AI-powered experiences. On the advertising side, Pinterest’s AI creative optimization tools are already translating into performance gains: Smart Assembly, a Pinterest Performance+ capability that automatically builds and serves the best-performing ad from advertiser-uploaded images, delivered a 6% average click-through-rate improvement in early alpha testing.

Find the right gift on Pinterest with visual search features

Building Pinterest's VLM Serving Stack on NVIDIA Dynamo

How NVIDIA Dynamo Powers Pinterest Assistant's Real-Time Visual AI

Pinterest

Pinterest Assistant is Pinterest’s AI-powered, visual-first shopping and discovery collaborator. It is designed to help users find what they want, even when they do not have the exact words for it, whether they are looking for an outfit, refreshing a room, planning an event, or shopping for products that fit their personal style.

The experience relies on vision-language models that can understand both images and text, making it possible to interpret a user’s style, visual context, and intent in real time. Serving those models at Pinterest scale requires fast image encoding, efficient memory management, and intelligent routing across GPU infrastructure. One of the biggest breakthroughs comes from Pinterest’s internal image encoder, which lets Pinterest serve precomputed visual embeddings through NVIDIA Dynamo and vLLM instead of processing raw pixels for every request, helping accelerate end-to-end latency by up to 44x.

With NVIDIA Dynamo, Pinterest Assistant can process 25x more visual context per request while keeping the experience responsive, with projection embeddings delivering up to 369x faster time to first token and 44x faster end-to-end latency. The outcome is AI that feels less like a search box and more like a personal stylist, interior decorator, or shopping collaborator, helping users move from inspiration to action faster.

Pinterest Scales Multimodal AI Across Every Product Surface with NVIDIA

Pinterest is adopting NVIDIA Dynamo end to end, including AI Configurator—a tool that can simulate more than 10,000 deployment configurations in seconds—to find optimal parallelism settings and hardware allocations with significantly less manual tuning and DynoSim. Paired with Dynamo Planner, an autoscaler designed for VLM and LLM workloads, Pinterest will be able to dynamically adjust compute in response to real-time traffic patterns, balancing cost efficiency with latency SLAs as demand fluctuates.

Looking further ahead, Pinterest is closely tracking NVIDIA’s Vera Rubin platform roadmap for the next generation of inference capability—a signal of confidence in a partnership that has grown alongside Pinterest’s AI ambitions for nearly five years. From early recommendation systems to a 14,000-GPU fleet powering LLMs, RecSys, and now production VLM serving, the foundation is in place for every Pinterest experience—search, discovery, conversation—to run on real-time multimodal intelligence.

With NVIDIA accelerated computing as the platform and NVIDIA Dynamo as the serving layer, Pinterest is building toward a future where the line between seeing something and understanding it—for both users and AI—disappears entirely.

“Building the next generation of AI-powered discovery means investing in infrastructure that can keep up with the scale and complexity of Pinterest,” said Kartik Paramasivam, Chief Architect at Pinterest. “Our collaboration with NVIDIA helps us deliver faster, smarter and more personalized experiences for the hundreds of millions of people who use Pinterest.”

Kartik Paramasivam
Chief Architect at Pinterest

Scale and serve AI inference fast with NVIDIA Dynamo.

Related Customer Stories