Summary
- Tavus reports that 48% of participants (26 out of 54) believed they were interacting with a real person during a one-minute video call with Griffin-Lite, compared to only 1 out of 41 for its previous technology.
- Griffin-Lite achieved the highest score on Nvidia's VideoFDB benchmark, receiving 3.83 out of 5 for its generation quality, surpassing the next-best system at 2.80 and a human reference at 3.92.
- The findings stem from Tavus' proprietary research, where participants were informed they would meet another individual, with Griffin-Lite currently restricted to select trusted testers as the company enhances safety protocols.
The AI startup Tavus has introduced its latest model, Griffin, which managed to convince 48% of individuals in a live video call that they were speaking with a real human being.
Launched on October 1, the company describes Griffin as the first Human Interaction Model, designed to comprehend and engage in face-to-face conversations, focusing on both verbal cues and non-verbal expressions like gestures and pauses. Its predecessor only achieved a 2.4% success rate on the same assessment.
Myriad: How high will Nvidia go? Click to make your prediction.During the test, Griffin-Lite interacted with 54 individuals, of whom 26 expressed belief that they were conversing with a real person. In contrast, the previous model engaged 41 individuals but only garnered one believer.
Participants were led to believe they were paired with another person for a one-minute discussion about their expectations for the upcoming year. Only after the conversation were they asked if they had considered that their partner might not be real.
This setup deviates from the traditional Turing test, introduced by mathematician Alan Turing in 1950, which involves a judge determining which of two hidden entities—one human, one machine—is which. In this case, no one was informed that a bot could be involved.
The results were compiled from Tavus' own research, with participants recruited via an independent research platform. A community note on X has pointed out that these findings lack independent verification and do not adhere to standard testing protocols. Tavus indicated that participants who became suspicious typically did so within the first 20 seconds.
AI firms have been pursuing advancements in this area for some time. A study from UC San Diego found that OpenAI's GPT-4.5 was perceived as human in 73% of conversations when instructed to act as an introverted, internet-savvy young person, although this was a text-only evaluation.
Griffin adds a visual and auditory component in real-time interactions.
On NVIDIA's VideoFDB benchmark, which assesses live audio and video dialogue, Tavus claims Griffin-Lite achieved the top rank. It received a score of 3.83 for the naturalness and expressiveness of its responses, while the next-best system scored 2.80, and the human reference achieved 3.92.
In terms of perception, which evaluates a model's understanding of its audio-visual input, Griffin-Lite scored 3.73 compared to the best baseline's 3.44, while the human reference reached 4.20. Tavus states that NVIDIA conducted this evaluation independently.
Griffin is also full-duplex, meaning it can listen, observe, and respond simultaneously, akin to a phone conversation rather than a walkie-talkie. In a Tavus demonstration, it guided a user through solving a Rubik's cube, adapting its responses based on the user's actions and pausing when they needed time to think. The average audio-to-video delay is 0.43 seconds on NVIDIA H100 chips, which Tavus claims is half the delay of the next fastest approach.
This technology holds relevance beyond the tech industry, particularly as scammers have begun utilizing video calls. For instance, in January, North Korean-linked hackers employed deepfake technology during Zoom or Teams calls to impersonate trusted contacts, leading to security breaches attributed to BlueNoroff, a subsidiary of the Lazarus Group. Victims were manipulated into installing malware disguised as audio fixes.
David Liberman, co-creator of Gonka, a decentralized network for AI computing, noted in a report that visual and audio evidence can no longer be relied upon as definitive proof of authenticity, even before this technology reached its current level of sophistication.
Organizations are already adapting their security measures. In 2025, Kraken flagged a suspected North Korean job applicant after its security team posed unexpected questions, such as requesting government identification and local restaurant names, which the candidate struggled to answer.
Upon testing Tavus’ models, the results were somewhat underwhelming. Further investigation revealed that Griffin-Lite is currently not available for general use, as it is limited to select trusted testers in a research preview. The company is focused on developing disclosure features and collaborating with AI safety organizations.
Tavus emphasizes that Griffin requires additional safety measures before it can be released to the public.
In November 2025, Tavus secured a $40 million Series B funding round, led by CRV. The previous system, which achieved a 2.4% success rate, combined three distinct models—one for visuals, one for dialogue, and one for perception.
Those interested in becoming trusted testers for Griffin-Lite can apply through a form on the Tavus website.
