Higher Education / Research

StudyFetch Cuts Inference Costs ~10x With NVIDIA, Expanding Access to AI Learning for Millions of Learners

Objective

StudyFetch is building AI-powered learning infrastructure for more than 7 million students, educators, and lifelong learners. Its platform turns live lectures, notes, and course materials into personalized AI tutoring, study plans, flashcards, practice tests, and course-building curriculum grounded in each learner’s real coursework.

As demand grew into hundreds of thousands of lectures every month, the cost and accuracy ceilings of off-the-shelf cloud transcription threatened the company’s core promise: keeping capable AI affordable for the people who need it most. By migrating its transcription pipeline to NVIDIA Riva with the Parakeet ASR model, packaging it in NVIDIA NIM™ microservices, and adopting NVIDIA Nemotron™ open models with distillation across its stack, StudyFetch cut its largest inference workload’s cost by roughly 10x—unlocking voice tutoring, real-time personalization, and a new agentic learning platform built to lower barriers to AI for millions more students and lifelong learners. This partnership is now expanding into dedicated infrastructure planning, with StudyFetch exploring three 8-way NVIDIA B300 systems to support a shift of conversational tutoring and course-generation workloads from closed models to open models, giving the team more predictable capacity and lower inference costs as usage grows.

Customer

StudyFetch

Partner

AWS · Google Cloud

Use Case

Generative AI / LLMs Speech AI,
Inference,
Workforce AI,
Learning

Key Takeaways

Inference Efficiency

  • ~10x reduction in cost on the largest inference workload by replacing managed cloud transcription with NVIDIA Riva + Parakeet in NVIDIA NIM containers

Personalized Learning

  • Hundreds of thousands of college lectures transcribed in real time every month, with the same orchestration layer now scaling NVIDIA-aligned learning to lifelong learners and the next-generation workforce

Scalable Learning

  • 100M+ learning interactions personalized in real time through StudyFetch’s Learn Engine on NVIDIA H100 and L40S GPUs, with Honen launching AI literacy programs reaching 250,000 K–12 students
  • StudyFetch now evaluating three eight-way NVIDIA B300 systems to expand open-model tutoring and course generation workloads with predictable capacity and lower inference costs

Keeping Personal AI Tutoring Affordable for Millions of Learners

StudyFetch is an AI-native learning workspace used by more than 7 million learners—primarily in higher education and postgraduate programs—who come to the platform for AI tutoring, practice, and study materials grounded in their actual coursework. Live lecture transcription forms the foundation of the experience: It is how the platform captures what each learner is taught and turns that input into a personal AI tutor, flashcards, practice tests, and study plans aligned with a real curriculum.

That foundation is also where the company hit its hardest constraints. Running live transcription on a major cloud provider’s managed service at hundreds of thousands of college lectures every month, StudyFetch hit two walls at once.

The first was cost. Live transcription was already a six-figure monthly line item, and the unit economics were getting worse with every new learner. For a company whose mission is to make great education accessible to anyone, every dollar of model cost compounds on margins that must remain low enough for learners to afford the product. Frontier-model inference at internet scale simply does not pencil out on traditional cloud economics.

The second was accuracy in the environments in which learners actually learn. Real college lectures are noisy—full of domain-specific terminology in subjects like organic chemistry, anatomy, statistics, and programming—and often delivered by professors with a wide range of accents and pacing. Off-the-shelf transcription was not tuned for that reality, and the quality gap showed up downstream in every product surface that depended on the transcript: the AI tutor, the flashcards, and the study guides.

The whole edtech category is wrestling with the same tension. Learners need increasingly capable AI to keep up with how fast the world is changing, but serving frontier models at scale does not pencil out under traditional cloud inference economics. StudyFetch needed an inference stack that could deliver capable, accurate AI at a price its learners could afford—without trading away accuracy or deployment flexibility.

Tractian

Powering Live Transcription and Personalization With NVIDIA Riva and Nemotron

Through the NVIDIA Inception team, StudyFetch mapped its highest-cost inference workloads to NVIDIA’s speech AI and open-model stack. As the technical partnership deepened, its scope expanded—from infrastructure optimization to aligning StudyFetch’s growing platform with NVIDIA’s broader mission of scalable, accessible AI across institutions, workforce programs, and learning communities. At the core of the new architecture:

  • NVIDIA Riva, a GPU-accelerated speech AI service running the Parakeet ASR model, produces live, high-accuracy transcripts of lectures in real time—including the noisy classrooms and domain-heavy vocabulary that broke off-the-shelf services.
  • NVIDIA NIM microservices—prebuilt, GPU-optimized inference containers—package and deploy the Riva pipeline across both AWS and Google Cloud, letting a small team run production speech AI without a dedicated MLOps function.
  • NVIDIA Nemotron, a family of open large language models, is integrated into StudyFetch’s API platform alongside leading frontier models. Because Nemotron is open and fine-tunable, StudyFetch runs distillation pipelines served through NVIDIA NIM, compressing larger models into smaller, lower-latency deployables tuned for the platform’s specific tasks.
  • As StudyFetch shifts more conversational tutoring and course-generation workloads from closed models to open models, the team is evaluating dedicated NVIDIA B300 infrastructure to support predictable capacity, lower inference costs, and continued Nemotron evaluation at production scale.
  • StudyFetch’s Learn Engine—its Educational Context Layer—runs on NVIDIA H100 and L40S GPUs, routing each learner’s requests across more than 100 million learning interactions and improving continuously as more learners use the platform.
  • Honen, StudyFetch’s new agentic learning platform, is built on NVIDIA GPUs from day one and extends the personalized learning model into AI literacy and workforce education. Learners can launch GPU-backed Jupyter notebooks inside NVIDIA-aligned courses, giving them hands-on experience with AI workflows. NVIDIA Training and NVIDIA Deep Learning Institute content is being integrated, helping learners build familiarity with NVIDIA tools, accelerated computing, and practical AI development environments.

Cutting Inference Costs ~10x While Unlocking Voice Tutoring and Real-Time Personalization

Migrating live lecture transcription onto NVIDIA Riva and Parakeet running on NVIDIA L40S GPUs in NVIDIA NIM containers delivered roughly a 10x reduction in cost on the company’s largest inference workload. Nemotron distillation on NVIDIA H100s is on track to deliver similar substantial reductions across the rest of the inference stack. The impact is visible inside the product itself. Where the previous service noticeably lagged, learners now watch their lectures being transcribed in real time as the professor speaks. Hundreds of thousands of lectures move through the pipeline every month with accuracy that holds up on real classroom audio. NVIDIA NIM packaging means a small team can run all of this without standing up a separate MLOps function. The cost savings paid for capabilities StudyFetch could not previously justify on the same budget:

  • Multimodal, Voice-Based AI Tutoring: NVIDIA Riva text-to-speech (TTS) pipelines for the AI Tutor are in development, targeted for summer rollout ahead of the next school year—opening voice-led learning to learners who need accessible, conversational modalities.
  • Real-Time Personalization at Scale: The Learn Engine on NVIDIA H100 and L40S GPUs continuously optimizes 100M+ learning interactions across the learner base.
  • Deployable in Any Environment: The NVIDIA Riva + NVIDIA NIM + Nemotron stack runs equally well in cloud, single-tenant, and regulated environments, opening Honen to enterprise, government, and education customers.
  • Wider AI Literacy and Ecosystem Reach: Cost reductions allow StudyFetch to ship new learner-facing capabilities, lowering barriers into AI for students, working professionals, and reskilling cohorts.

The cost savings translate directly into more capability in more learners’ hands—and made the case for something larger than transcription.

“StudyFetch exists to make great education accessible to anyone. NVIDIA is what makes that economically possible. Partnering with their team on Riva, Nemotron, and the rest of the stack did more than cut our costs—it let us put more capability in the hands of more students, so they can learn the things they need to learn at a price they can actually afford.”

Ryan Trattner
Co-Founder & CTO, StudyFetch

Scaling Personal Education to the Workforce and the Next Generation

With transcription, personalization, and Honen all running on a shared NVIDIA foundation, StudyFetch is expanding the partnership across several fronts:

  • NVIDIA Riva TTS rolls out for the AI Tutor’s voice pipelines this summer, ahead of the next school year, and Nemotron distillation is approaching peak rollout across the broader inference stack as the company accumulates more curriculum-specific usage data.
  • StudyFetch is evaluating three 8-way NVIDIA B300 systems to support its shift from closed-model dependencies toward open-model conversational tutoring and course generation, with Nemotron evaluations expected to inform future deployment decisions.
  • Honen is launching with NVIDIA, reaching 250,000 K–12 students through AI literacy programs delivered alongside NVIDIA Training and NVIDIA Deep Learning Institute content, with broader higher education and workforce deployments planned.
  • StudyFetch is exploring on-premises and single-tenant Honen deployments for regulated industries and government education customers, using distilled Nemotron models tuned for those environments.
  • The team is evaluating NVIDIA-accelerated vision models to train nursing students and allied health technicians on medical imaging—extending the platform’s reach into clinical and allied health education.

The StudyFetch and NVIDIA partnership shows how a production AI infrastructure win can become a model for broader access. By reducing inference costs, improving real-time learning experiences, and connecting learners to NVIDIA-aligned tools and content, StudyFetch is helping make AI literacy and workforce learning scalable across classrooms, campuses, communities, and the workforce at large.  

Related Customer Stories