NVIDIA has announced a major expansion of NVIDIA AI for Media, its software stack for broadcast, sports and streaming production. Timed to IBC 2026, held in Amsterdam from September 11 to 14, the update covers generative frame interpolation, detection of AI-generated footage, and lip-synced dubbing for multiple languages. What ties it together is that each piece is designed to slot into existing broadcast plants rather than replace them.

Smoother slow motion without swapping cameras

The headline addition is Video Frame Generation (VFG). It uses generative AI to create new frames between the original frames of a clip, raising frame rates by 2x or 4x while holding temporal consistency so fast-moving subjects do not fall apart.

Ross Video has built VFG into its Rio Replay platform to produce AI-assisted slow motion for sports coverage. The integration currently supports 6x slow-motion generation, with work underway toward 8x. In other words, a replay team can get fluid motion without a super-high-frame-rate camera capturing every frame, which changes the order in which broadcasters need to spend on hardware.

Two other tools raise the quality of the underlying video. Video Super Resolution (VSR) upscales footage while reducing noise, blur and compression artifacts, and the new release adds 10-bit support along with streaming modes that let developers choose between real-time speed and higher image quality. TrueHDR converts standard-dynamic-range video to HDR in real time, reaching up to roughly 2,000 nits while preserving local contrast. VSR, VFG and TrueHDR can run in a single video-effects pipeline, which makes reworking an existing content library for streaming a practical option.

AI-video detection reaches 99.3 percent accuracy

For newsrooms, the Synthetic Video Detector (SVD) is the piece worth watching. The NIM microservice returns a probability that a given clip is authentic or AI-generated, and was first shown at SIGGRAPH earlier this year. Since that release, accuracy has reached 99.3 percent for text-to-video content and 97.7 percent for image-to-video content, with the largest gains on the harder image-to-video cases.

Dalet has wrapped SVD into a cloud-hosted verification workflow so editorial teams can submit footage and review the resulting scores and metadata inside a Dalet interface. TwelveLabs moved Compliance by TwelveLabs, the first application built on its video intelligence platform, to general availability and folded in SVD to add frame-level authenticity signals to regional compliance screening. Wowza, whose streaming engine powers more than 35,000 video deployments across over 170 countries, will distribute SVD through its Video Intelligence Framework, with deployment options spanning on-premises, edge, cloud, hybrid and fully air-gapped environments.

Ten languages from one stream, at one-eighth the video

The localization updates are equally practical. LipSync reshapes mouth movement in a video to match a translated audio track, and the new release improves handling of partially obscured faces while better preserving teeth, lip and facial textures. Active Speaker Detection, which identifies who is talking in a multi-person scene, no longer requires speaker diarization for multiple audio tracks, adds voice activity detection, and supports deployment over a gRPC interface.

NDI, a Vizrt company, has put that combination to work. It demonstrated real-time translation, lip-synced dubbing and regional language adaptation running inside existing broadcast workflows. In internal testing of a 10-language workflow, the approach carried 8 times less video than the traditional method of building a separate stream per language. If the assumption that each additional language means more bandwidth and more infrastructure no longer holds, the effect on regional stations and smaller streaming operators could be substantial.

NVIDIA has also brought its Content Localization technologies, covering captions, translated audio, dubbing, synchronized video and localized graphics, into the Holoscan for Media developer toolkit. AI-Media, CAMB.AI, Chyron and Panjaya each handle a specific stage of that chain. Holoscan for Media itself gains an integration with Media Exchange Layer (MXL), an open way for software-based media functions to exchange live video, audio and data across a distributed environment.

A recipe for training sport-specific models

The other pillar is NVIDIA Sports Intelligence Playbooks, structured frameworks that let leagues, media companies and vendors fine-tune NVIDIA open models on their own footage and annotations. The playbooks span data preparation, fine-tuning, inference, evaluation, optimization and deployment.

Early results are easy to read. Evaluated on previously unseen footage, multiple-choice accuracy rose from roughly 53 percent to 94 percent, and open-ended evaluation from roughly 5.7 percent to 66 percent. That is a direct answer to the familiar problem of general-purpose models misreading sport-specific rules and tactics. Machina Sports is integrating the playbooks with its sports-native data and agent infrastructure so that rights holders can run private, deployable intelligence on their own material.

Summary

This expansion is less a flashy product launch than a coordinated refresh of parts that drop into equipment broadcasters already own. The numbers are concrete: 6x slow motion from generated frames, 99.3 percent accuracy in spotting AI-generated video, and a 10-language workflow that moves one-eighth the video. AI in production has moved past the demo stage and into conversations about bandwidth bills and staffing, which is the most straightforward way to read the announcement.