The one who tied the knot must untie it.

Jensen Huang, who has frequently been impersonated in AI-generated videos across the internet, has decided to take matters into his own hands.


Recently, NVIDIA announced the launch of its Synthetic Video Detector (SVD) NIM. Billed as a tool that analyzes videos frame by frame, it identifies AI-generated content by detecting statistical and frequency-domain traces left by generative models, boasting a detection accuracy of up to 92%.


image.png

(Image source: NVIDIA)


It’s hardly a surprise, given how low the barrier has become for creating fake images and videos.


Take a video I came across not long ago, for example. The premise was simple: a cat knocks over a vase, then bolts away faster than lightning. The dog beside it freezes on the spot, pawing frantically as if trying to explain, “It wasn’t me, I swear it wasn’t me.”


GIF 2026-7-24 10-25-43.gif

(Image source: Douyin)


Pretty funny, right? It would be even better if it weren’t an AI-generated video.


What makes things trickier is that fake videos used to give themselves away with obvious visual flaws—mismatched lip sync, wrong number of fingers, and so on. But today’s video generation models are advancing at breakneck speed, and much of the AI content they produce no longer leaves such glaring loopholes. It’s no wonder my elderly family members fall for them every single day.


Since ordinary people struggle to tell the difference, it’s high time dedicated detection tools stepped up.


NVIDIA Joins the “Using AI to Detect AI Videos” Movement


For a quick primer: NVIDIA unveiled NVIDIA Inference Microservices, or NIM for short, in March 2024. Essentially, NIM packages models, inference engines, API interfaces and runtime environments into a single deployment tool, giving developers a fast track to accessing AI services.


The newly released Synthetic Video Detector NIM can be understood as an AI video detection service packaged under NVIDIA’s NIM standard.


Since it is primarily designed for on-premises deployment, online access currently requires applying for NVIDIA’s media testing program. Local deployment, meanwhile, demands a Linux system and hardware compatible with NVENC/NVDEC video codecs—making it rather inaccessible for average users.


image.png

(Image source: NVIDIA)


That said, not being able to try it ourselves doesn’t stop us from diving into the available data and demos.


According to official documentation, once a video is uploaded, SVD NIM first splits the H.264-encoded MP4 file into individual frames. These frames are then analyzed by vision models built on DINOv2 and DINOv3, which assign a score to each frame and aggregate the results to indicate how likely the entire video is to be AI-generated.


Crucially, it does not rely on conventional telltale signs like six fingers, character clipping, or objects vanishing out of nowhere.


Per NVIDIA’s official statement, SVD NIM identifies underlying statistical features, frequency anomalies, and pixel patterns uncommon in real camera footage—artifacts left behind during the generation and denoising process. It maintains a high detection rate even when videos are compressed and re-encoded.


image.png

(Image source: NVIDIA)


Looking at an official test case, the video in question shows none of the visual glitches or inconsistencies typical of early AI videos. Yet SVD NIM returned an overall synthetic score of 99%.


To be fair, the video does look fake at a glance—the camera movement is just a little too smooth.


GIF 2026-7-27 10-47-46.gif

(Image source: NVIDIA)


In my view, NVIDIA’s approach is far more reliable than fixating on superficial visual flaws.


After all, if you showed most people the clip below with no context, I’d wager at least 80% wouldn’t guess it was 100% AI-generated. That’s how impressive the video generation capabilities of models like Seedance and Kling have become.


GIF 2026-7-27 11-07-36.gif

(Image source: Bilibili)


NVIDIA reports that its internal test set includes over 4,000 videos, on which SVD NIM achieves a detection success rate of 85.64%. For uncompressed videos, accuracy climbs as high as 92%.


Of course, 92% accuracy doesn’t mean it gets every call right.


While omitted from promotional materials, NVIDIA notes in its technical documentation that this high accuracy is achieved using a relatively strict detection threshold, which may produce a high number of false positives in real-world environments. The tool is therefore better suited for initial screening on video platforms, with suspicious footage referred to human reviewers for further verification.


What’s more, this accuracy figure is based on H.264-encoded MP4 videos; any transcoding or compression will reduce accuracy. Jensen Huang’s fine print is certainly professional, to say the least.


Real-World Anti-AI Test: Zhuque Performs Decently, Anti-Fraud Center Proves Unstable


For now, NVIDIA’s SVD NIM remains out of reach for most everyday users.


China does have similar consumer-facing detection tools that work straight in the browser. One is Tencent Zhuque AI Detection Assistant; the other is the AI Content Verifier feature inside the National Anti-Fraud Center app. Both support text, image and video detection, and offer a limited number of free checks per day.


To see how well “AI catching AI” actually works, I put together a mixed test set:


several images generated directly by mainstream models, plus real photos taken on a smartphone; a carefully selected AI-generated video paired with an authentic vlog clip. I also shared some of the files via WeChat and took screenshots to test how they hold up against secondary compression.


Round one: AI-generated video


GIF 2026-7-27 11-33-11.gif

(Image source: Bilibili)


I think this video is quite well made. Its slightly blurry resolution, subtly shaky simulated camera work, and the character’s movements and expressions blend together to give it a real vintage film feel. The creator hand-drew extensive storyboards for the model to reference, fully leveraging Seedance’s production capabilities.


And Tencent Zhuque was the first to flub this test.


Screen Capture 2026-07-27 113125.png

(Image source: LeiTech)


For comparison, I borrowed a Zhuque account from a friend who had media access. SVD NIM, by contrast, returned an overall score of 93%, indicating a high probability that the video is AI-synthesized, with frame-by-frame anomalies marked on a coordinate graph.


That said, the detection speed is hardly fast—it took two minutes to analyze a 13-second video.


image.png

(Image source: LeiTech)


Even more frustrating was the Anti-Fraud Center’s content detector: I tried for half an hour straight and couldn’t get the file to upload at all, so I had to give up.


Round two: real video


GIF 2026-7-27 15-07-05.gif

(Image source: LeiTech)


Naturally, I didn’t pick a simple real video. I chose a vlog clip that many video generation models love to imitate. Shot with a good camera, it has smooth, natural transitions and textured street lighting—exactly the kind of footage that’s hard to tell apart from AI at a glance.


Guess what? Tencent Zhuque struck out again.


Screen Capture 2026-07-27 112712.png

(Image source: LeiTech)


NVIDIA’s tool fared considerably better. While it still returned a 34% synthetic score, this is merely a reference value, meaning the model found little evidence of synthetic content.


Screen Capture 2026-07-27 112726.png

(Image source: LeiTech)


As for the Anti-Fraud Center app, I still couldn’t get the upload to work. Elderly users who actually need this feature would be in for a rough time.


Next up: image testing. I won’t go into detail, but the set includes real photos on the same subject, plus two AI-generated images made with Doubao and Image 2, both edited and compressed to remove watermarks.


New Project.jpg

(Image source: LeiTech—from left to right: Doubao, Image2, real photo)


I was just about to praise Zhuque for correctly identifying the AI images this time, when it went and labeled a real photo as AI-generated too.


image.png

(Image source: LeiTech)


Conversely, the content detector that failed to upload videos multiple times actually got all three image tests right. It seems the Anti-Fraud Center does have some algorithmic advantages after all.


New Project (1).jpg

(Image source: LeiTech)


One final gripe: neither service offers many free detection attempts. They might not even be enough if you actually need to verify something.


AI Detection Is Weak—but Better Than Nothing


After this round of testing, my confidence in AI detection tools can be summed up as: better than nothing, but only just.


For video detection, NVIDIA’s tool performs passably, while Tencent Zhuque is thoroughly unreliable—even giving the opposite result. Image detection works somewhat better, but Zhuque still manages to mislabel real photos as fake. The Anti-Fraud Center’s detector, meanwhile, is almost always overloaded, making uploads impossible.


If the algorithm itself can’t make up its mind, people shouldn’t be rushing to post detection screenshots in comment sections as definitive proof.


In my opinion, rather than trying to guess whether an image has “AI vibes” after the fact, embedding watermarks and provenance records directly at the point of generation is clearly a more reliable approach.


Google already uses SynthID to add invisible markers to generated content, while Adobe and other companies are pushing the C2PA standard to track the full lifecycle of files from creation through modification. China is following suit: the Cyberspace Administration of China has explicitly mandated labeling for AI-generated content, and Tencent, Alibaba and ByteDance are actively ramping up efforts in detection, digital signatures and platform-level notifications.


6f5e992c-2b2f-4e6b-ad09-632ffb5dcf3e.png

(Image source: LeiTech)


The problem is that this ecosystem is not yet fully connected.


If I deploy an open-source model, the content it generates will naturally carry no watermark. Forward an image or video a few times, and its metadata may disappear entirely. As for whether platforms can recognize each other’s watermarks—given how guarded domestic tech firms are with one another—that’s another open question.


Until all these links are properly connected, ordinary users will simply have to stay vigilant.



ChinaJoy 2026, themed Roaming with AI, will kick off with great fanfare on July 31.


A total of 900 pan-entertainment exhibitors, including Tencent, NetEase, Sony, Qualcomm, MACHENIKE and Qingxian, will join forces to present a grand feast of global entertainment.

Powered by AI, how will hardware brands spanning terminals, peripherals, robotics, displays and chips cross paths with the gaming content industry to deliver new entertainment experiences?


Led by founder and editor-in-chief Luo Chao, Leitech’s ChinaJoy 2026 reporting team will descend on Shanghai for comprehensive on-site coverage. Stay tuned for more.

29672c82277ced8ca662ed2066d41085.jpg