AI Agents, AI Video

Tavus Unveils Griffin, Claims Video Turing Test Milestone

Tavus has introduced Griffin, a real-time conversational AI system that the company says has passed a video Turing test. Its headline finding: 48% of participants in a short live-call study believed they had spoken to a human.

Announced on 1 October 2026, Griffin is not a general product release. The Griffin-Lite research preview is restricted to selected testers while Tavus develops additional safety and disclosure measures.

For people evaluating AI assistants, the distinction matters. Looking and sounding human is one achievement. Providing dependable help, handling interruptions and making clear that the user is talking to software are separate requirements.

What is Tavus Griffin?

Griffin is what Tavus calls a Human Interaction Model. The system is designed to keep processing what it hears and sees while generating speech, expressions and movement, rather than treating listening and responding as separate stages.

The company describes this as full-duplex interaction: both sides can contribute at the same time. An assistant can, in principle, acknowledge someone mid-sentence, respond to an interruption or adjust its behaviour as the exchange develops.

The practical attraction is not simply a more animated face. Consider someone explaining a problem, pausing to find the right word, then correcting themselves. A useful assistant should distinguish those moments from an invitation to start speaking. Otherwise, the person still has to adapt their conversation to the software.

What the 48% result actually measures

Tavus’s published Griffin research account says 26 of 54 participants considered Griffin-Lite human after a one-minute call. Its previous system received that judgement from one of 41 participants, or 2.4%.

Participants expected to meet another participant and discuss what they were looking forward to that year. AI involvement was raised only at the survey’s end; participants were then told their partner was artificial.

This was Tavus’s study, not an independent replication.

The result concerns a brief interaction under specific conditions. It does not establish performance in longer calls, under deliberate scrutiny or on specialist questions. Tavus’s “Turing test” interpretation should not be mistaken for proof of dependable reasoning.

NVIDIA’s benchmark provides a separate check

NVIDIA’s VideoFDB leaderboard lists Griffin-Lite first among published non-human entries on both its generation and perception tracks. The benchmark uses a language-model judge and a 0-5 scoring rubric.

MeasureGriffin-LiteHuman reference
Generation overall score3.83 / 53.92 / 5
Perception overall score3.73 / 54.20 / 5
Generation median latency1,892 ms900 ms
Perception median latency2,232 ms1,400 ms
NVIDIA VideoFDB results, checked on 2 October 2026. The two tracks assess different tasks.

Griffin-Lite leads the AI entries, but remains below the human reference on both overall scores and responds more slowly by these median measurements.

These are behavioural benchmark scores, not the percentage of people deceived. Neither result establishes reliable performance on a specific business task.

What would make this useful beyond a convincing demo?

For DIY AI, the useful next comparison would hold the task constant and test text, voice and live video. Does seeing the assistant help someone finish the job with fewer explanations? Or does it add a persuasive face without improving the answer?

A support scenario could test that distinction directly. Give each interface the same product information and the same problem. Check whether the assistant identifies the relevant detail, asks an appropriate follow-up and reaches a correct resolution. Score the answer separately from how natural the conversation feels.

Then introduce awkward moments: pause before finishing a sentence, correct an earlier detail or interrupt an incorrect explanation. The question is whether the assistant preserves the user’s intended meaning, not simply whether its animation continues smoothly.

Disclosure should also be part of that evaluation. Repeat the task with participants who know from the outset that the assistant is AI. An interface that remains comfortable and useful without being mistaken for a person would offer a clearer basis for deployment.

Can customers use Griffin yet?

Not through a general customer release. Tavus describes a trusted-tester preview and further disclosure work before wider availability. The reviewed launch materials do not specify Griffin pricing or a firm public-release date.

Existing Tavus subscriptions should therefore not be treated as a promise of Griffin access. A team considering a future pilot still needs confirmed access terms, usage costs and a way to assess the complete experience rather than a selected demonstration.

Creators making a finished presenter video are solving a different problem. Our HeyGen and ElevenLabs workflow covers prerecorded production, where narration and footage can be reviewed before publishing. Live interaction removes that opportunity to edit each response before the audience sees it.

DIY AI’s take: the important question is not whether an AI can persuade someone that it is human for a minute. It is whether a clearly identified AI can use conversational timing and visual context to provide better help. Griffin gives that question a concrete system to investigate, but the announcement is not yet a buying recommendation.

DIY AI has not independently tested Griffin. This report separates Tavus’s study claims from NVIDIA’s published benchmark results.

Written by Steven Jones

AI Tools Reviewer and Technical Analyst

Steven Jones is a technology analyst specialising in artificial intelligence, machine learning workflows, and emerging automation tools.

At DIY AI, he focuses on clear, practical guidance for people comparing AI tools in the real world. His work covers text generation, image generation, video tools, data platforms, developer-focused AI products, and the automation workflows that connect them.

Steven's reviews are built around hands-on testing, practical benchmarks, and transparent scoring rather than vendor claims. He looks closely at where each tool performs well, where it falls short, and what those trade-offs mean for creators, teams, and businesses trying to make sensible AI adoption decisions.

He has a particular interest in safety, reliability, output quality, performance metrics, and dataset quality. When he is not reviewing the latest AI model updates, he experiments with prompt engineering techniques and contributes to DIY AI ongoing work on fair, explainable scoring frameworks for AI tools.

Back to AI News