Griffin is Tavus's new model for live, face-to-face AI interaction. Its October 1, 2026 announcement centers on two results: 48% of participants in a short video-call study thought they had spoken to a human, and Griffin-Lite placed first on NVIDIA's full-duplex video benchmark.
Those claims deserve a close look. A convincing face, well-timed interruption and useful answer are separate achievements. The published evidence is strongest on conversational behavior; it leaves important questions about production access, cost and performance on real tasks unanswered.
This article examines Griffin's specifications, architecture, benchmarks, availability and limitations. Sources were checked on October 1, 2026. Kingy has not received Griffin access or conducted an original live-call test. Statistical intervals, score comparisons and cost examples below are our calculations.
What Griffin changes in a live conversation
Tavus describes Griffin as a Human Interaction Model, or HIM. Its technical release describes continuous conversational control over streaming speech and video. Incoming audio and video inform sub-second decisions about speaking, waiting, yielding and expression. Source: Griffin research release.
Consider an assistant helping someone identify a broken component over video. The person starts explaining, pauses to turn the component toward the camera, then continues. An assistant that answers during that pause can interrupt the explanation. One that waits through every silence can leave the person wondering whether it heard anything. A visible acknowledgment can help without taking over the conversation.
That example illustrates the problem Griffin targets. Speech recognition supplies words, but the decision to speak also depends on what the other person is doing. When the assistant speaks, the participant may react before the sentence ends. Conversation therefore requires ongoing decisions throughout an exchange.
Full-duplex means both sides can send and receive at the same time. For an AI interface, the useful question is whether new input changes its behavior while it is already responding. Simultaneous media transport alone would not demonstrate that ability.
Our interpretation is that Griffin's contribution concerns coordination: the relationship between perception, response timing, vocal delivery and visible behavior. The benchmarks later in this article help assess that coordination. They provide much less information about whether Griffin can solve a technical problem, follow a complicated company policy or retrieve an accurate answer from business records.
Griffin-Lite technical specifications
The published technical figures describe Griffin-Lite, the preview variant. They should not be treated as specifications for a future, more capable Griffin release.
| Component | Published specification |
|---|---|
| Video output | 720p, 25 fps |
| Video chunk | 320 ms; eight frames per latent |
| Video generation | Autoregressive diffusion; three steps per latent |
| Visual reference | Single image; full-scene generation |
| Speech generator | Autoregressive diffusion transformer, VDiT |
| Voice reference | Approximately 10 seconds |
| Tavec audio codec | 48 kHz; 40 continuous values/frame; 100 frames/second |
| Audio decoding | Fully causal; packets as small as 10 ms |
Source: Tavus's Griffin technical specifications. These are vendor-reported figures.
The video arithmetic is consistent: eight frames at 25 frames per second cover 0.32 seconds. A 320 ms video chunk describes the amount of video represented by each chunk. It does not by itself establish the time between a participant finishing a question and hearing an answer.
Likewise, a 10 ms audio packet is a unit of streamed output. It is not a claim that the system understands a new request and answers within 10 ms. An audio codec's 48 kHz sample rate also says little about semantic accuracy, accent coverage or voice consistency across a long call.
The distinction between video output rate and visual understanding matters too. A rendered face at 25 fps does not establish that incoming participant video is analyzed at that rate. We found no complete public Griffin specification for input sampling, network bandwidth or supported client devices.
A single-image reference can simplify likeness setup, but it leaves practical questions to test: stability when the character turns, hand and finger artifacts, background consistency, and recovery after repeated interruptions. Output resolution and frame rate cannot answer those questions alone.
Architecture and the specifications still missing
The release describes a conversational engine and an audiovisual generation engine. “Unified” is therefore best read as coordinated operation; the public description does not establish one neural network or a single set of weights. Source: Griffin architecture.
For developers, the distinction matters because the integration contract determines what they can control. A system may coordinate face and voice effectively while exposing only a narrow application interface. Another may allow extensive customization but require the customer to manage more of the timing. The launch evidence does not yet establish Griffin's production tradeoff.
The current public materials leave the following specifications unresolved:
| Missing specification | Why it matters |
|---|---|
| Parameter count and model sizes | Hardware and deployment comparisons |
| Context-window limit | Capacity for long instructions and conversation history |
| Complete language and accent matrix | Suitability for the actual audience |
| Customer API and versioning contract | Integration effort and upgrade control |
| Downloadable weights and license | Whether local deployment is possible |
| GPU configuration and memory requirements | Whether the component results can be reproduced |
| Detailed training-data inventory | Provenance and evaluation-overlap review |
| Production SLA and regional coverage | Reliability and deployment planning |
These are gaps in the public release reviewed for this article, rather than proof that Tavus lacks the corresponding internal information. A prospective customer should obtain Griffin-specific answers before designing around assumptions borrowed from another Tavus model.
We also found no published Griffin results for coding, mathematics, factual accuracy or business-task completion. A model that responds naturally to a confused expression still needs to explain the subject correctly. Improved communication can make an answer easier to receive without making the answer more reliable.
The 48% human-identification study
Tavus reports 26/54 human identifications for Griffin-Lite versus 1/41 for its earlier stack. Independently recruited participants expected another person in a one-minute call; AI questions and disclosure followed. Source: study methodology.
| System | Participants identifying a human | Observed rate | Calculated 95% Wilson interval |
|---|---|---|---|
| Griffin-Lite | 26 of 54 | 48.1% | 35.4%–61.1% |
| Phoenix-4.5 + Sparrow-2 + Raven-1 | 1 of 41 | 2.4% | 0.4%–12.6% |
The observed improvement is approximately 45.7 percentage points. Our intervals use a simple binomial model and the reported counts. They describe sampling uncertainty, without accounting for participant expectations, recruitment effects or differences between a study call and a customer deployment.
The result supports a specific conclusion: under this short-call protocol, many participants did not identify Griffin-Lite as AI. That is useful evidence of visual and conversational realism.
The experiment also gave participants a reason to expect a human. They were not assigned to investigate whether the partner was artificial or to stress-test the system. A one-minute conversation offers fewer opportunities to expose inconsistency than a long support call involving documents, corrections and difficult questions.
Recruiting people through an independent platform does not make the study an independent replication. Tavus ran and reported this evaluation. Independent testing with different participants, disclosed AI identity and longer sessions would help establish how broadly the finding applies.
The “passes the video Turing test” phrase is Tavus's characterization of this study. It should not be read as evidence of consciousness, general human-level reasoning or trustworthy advice. The reported identification rate also leaves 28 of 54 participants outside the group that identified a human.
For a product team, realism and disclosure should be measured together. A user who knows the assistant is AI may still find it easy to talk to. That is a useful outcome to test directly, because a disclosed production experience differs from this experimental setup.
NVIDIA VideoFDB benchmark results
NVIDIA's VideoFDB leaderboard independently lists Griffin Lite's generation overall score as 3.83 and its perception score as 3.73. The tables below show selected published comparisons checked on October 1.
Generation: the model's speech and visible response
| System | Overall / 5 | TOR alignment | Median latency |
|---|---|---|---|
| Human reference | 3.92 | 78% | 900 ms |
| Tavus Griffin Lite | 3.83 | 62.8% | 1,892 ms |
| Gemini 2.5 + Anam | 2.80 | 44% | 2,840 ms |
| Gemini 2.5 + Keyframe | 2.39 | 31% | 3,520 ms |
Perception: the response to incoming conversational cues
| System and input mode | Overall / 5 | TOR alignment | Median latency |
|---|---|---|---|
| Human reference | 4.20 | 90% | 1,400 ms |
| Tavus Griffin Lite, audio + video | 3.73 | 73.8% | 2,232 ms |
| MiniCPM-o 4.5, audio only | 3.44 | 72% | 920 ms |
| MiniCPM-o 4.5, audio + video | 3.40 | 73% | 720 ms |
| Gemini 2.5 Flash Native, audio + video | 3.17 | 72% | 3,160 ms |
| OpenAI gpt-realtime, audio only | 2.97 | 67% | 4,440 ms |
| OpenAI gpt-realtime, audio + video | 2.75 | 72% | 5,400 ms |
Source for both tables: NVIDIA VideoFDB. TOR denotes takeover-rate alignment, a conversational timing measure.
The VideoFDB paper describes 237 clips from two-person calls across 11 conversational dynamics. A language-model judge scores each response on a 0–5 rubric. Perception examines fluency, conversational flow and visual grounding; generation examines fluency, dyadic affect and nonverbal-cue appropriateness.
Our comparison of the listed scores puts Griffin first among nonhuman entries on both tracks. The generation gap to Gemini 2.5 + Anam is 1.03 points. Dividing that gap by Anam's 2.80 score gives a calculated 36.8% increase, which explains the announcement's rounded 37% claim.
That percentage describes a gain in the overall rubric score. It is not a percentage of correct answers, a measure of customer satisfaction or a 37% reduction in latency. The rubric's scale also does not establish that each percentage increase corresponds to an equal increase in practical value.
Griffin's generation score is 0.09 below the human reference, while its perception score is 0.47 below. Those calculated gaps are useful descriptions of this evaluation. They do not establish that the systems are statistically equivalent: the headline leaderboard does not provide enough uncertainty information to support that claim.
Timing gives a different view. Griffin's reported generation median is about 2.1 times the human reference; its perception median is about 1.6 times the reference. MiniCPM-o's audio/video perception run has substantially lower latency than Griffin's despite scoring lower overall. The benchmark therefore shows a strong quality result alongside remaining timing differences.
Input modes need attention. The 3.44 MiniCPM-o result is audio only; its audio/video result is 3.40. Comparing Griffin's audio/video score with both is informative, provided the inputs remain labeled. Anam appears as one particular Gemini-driven combination, so the table cannot rank every configuration that Anam offers.
These results are relevant when natural conversational behavior matters. They are less useful for selecting an assistant whose main job is answering factual questions or completing transactions. That selection requires additional task evaluations.
Latency and visual quality: what was measured
Tavus reports Griffin-Lite's video-component latency at 0.43 seconds on H100s, first place on DOVER/FID/THEval, and second on LSE-C at 7.27. Source: component evaluation.
The component test and NVIDIA's conversation test answer different questions:
| Number | What it describes | What it cannot establish alone |
|---|---|---|
| 320 ms | Video duration represented by a generated chunk | Complete question-to-answer delay |
| 0.43 seconds | Vendor-reported audio-to-video component latency | Full application response time |
| 1,892 ms | VideoFDB generation median | Tail latency for customer calls |
| 2,232 ms | VideoFDB perception median | Performance across every network and device |
Sources: Griffin release and NVIDIA leaderboard. The interpretations in the last column are ours.
A component can render quickly after receiving speech while the full system still spends time deciding what to say. Network transport, room startup and client playback introduce other stages. Combining measurements with different starting points into one “response time” would conceal those stages.
The published H100 result also leaves a reproducibility question. A GPU family name does not specify GPU count, memory use, batching, model precision or competing workloads. It cannot serve as a reliable local-hardware requirement or a calculation of commercial inference cost.
Visual quality and lip alignment measure different aspects of output. The component ranking is useful evidence, but it does not supply an artifact rate for every camera angle, reference image or conversation length. A deployment test should count visible failures and recovery time, including during interruptions and silence.
Median latency describes the middle of a distribution. For a live interface, repeated long delays may matter more than the median. Griffin-specific p95 and p99 measurements, measured over realistic calls and network conditions, remain useful information to request.
Griffin costs and the available Tavus prices
We found no public Griffin commercial price, per-minute rate, tester fee or guaranteed subscription entitlement in the launch materials. Griffin's cost is currently undisclosed. That prevents a defensible cost-per-call or return-on-investment estimate for Griffin itself.
Tavus does publish prices for its existing Conversational Video Interface. Its current model documentation distinguishes rendering, perception and conversation-flow models. The visible pricing page names Phoenix-4.5, Raven-1 and Sparrow-2 in that available stack. The following prices are platform context, not a Griffin rate card.
| Current CVI plan | Fee/month | Included minutes | Extra minute | Concurrent sessions |
|---|---|---|---|---|
| Free | $0 | 20 | None listed | 1 |
| Starter | $22 | 60 | No PAYG | 1 |
| Builder | $59 | 175 | $0.35 | 3 |
| Growth | $397 | 1,300 | $0.31 | 10 |
| Business | $975 | 4,000 | $0.26 | 15 |
| Enterprise | Custom | Custom | Negotiated | Custom |
Source: Tavus's current visible plan comparison, checked in a browser. Older indexed blocks on the page contain different names and allowances.
For perspective, our simple CVI estimate uses:
Monthly cost = fee + max(0, billable minutes − included minutes) × extra-minute rate
| Billable CVI minutes/month | Builder estimate | Growth estimate | Business estimate |
|---|---|---|---|
| 1,000 | $347.75 | $397.00 | $975.00 |
| 5,000 | $1,747.75 | $1,544.00 | $1,235.00 |
| 10,000 | $3,497.75 | $3,094.00 | $2,535.00 |
These calculations exclude tax, recording, external providers and discounts. They cannot be used to forecast Griffin's bill. Their purpose is to show the existing commercial baseline that a future Griffin price could be compared with.
Current CVI billing has a 30-second minimum per conversation and rounds to six-second increments. Created rooms can be billed before a participant joins. Growth and Business list recording at $0.03/minute. Source: Tavus pricing and billing FAQ.
At that minimum, 1,000 calls ending within ten seconds imply at least 500 billable minutes. A deployment should also budget for abandoned rooms, human escalation and integration work. Whether Griffin adopts the same billing mechanics remains unannounced.
When pricing arrives, compare cost per completed task as well as cost per minute. A more natural conversation could improve completion, or it could lengthen calls without helping resolve the issue. Measure both effects before assuming that a better benchmark score offsets a higher rate.
Griffin availability and how access works
Griffin-Lite is restricted to select trusted testers. Tavus says customer access awaits safety work and gives no firm general-release date. Source: Griffin release status.
The public launch is therefore an announcement and research preview. We found no published self-service Griffin endpoint, downloadable checkpoint or customer rollout schedule. A Tavus account or paid CVI plan should not be assumed to unlock it.
Readers seeking evaluation access should use the official Griffin page and Tavus's current access process. An access request does not guarantee acceptance, production permission or a commercial service level.
The public CVI developer quickstart documents an existing integration path using API-created conversations and embedded video. That is useful for evaluating Tavus's available platform, but it does not establish Griffin API compatibility. Source: CVI quickstart.
Before investing in integration, obtain explicit answers about the version offered, permitted use, migration path, regional availability and whether any prototype interface may change. Teams with a near-term launch deadline need an available model and a supported contract; a research preview has different planning assumptions.
Griffin limitations, disclosure and privacy
Griffin's clearest limitation today is access. Its other limitations require care because the public evidence covers a narrower range than a production product would need.
The human-identification study is short and vendor-run. VideoFDB evaluates conversational cues using a particular dataset and rubric. Neither evaluation establishes accurate medical advice, reliable account actions, prompt-injection resistance or correct handling of every customer policy. Those capabilities require their own evidence.
Long conversations remain a substantial evaluation gap. A production trial should look for identity drift, repeated gestures, voice changes, forgotten instructions and inconsistent answers after corrections. We found no public Griffin failure-rate table spanning session lengths, accents, difficult lighting and poor connections.
Language coverage is also unresolved for Griffin. The existing CVI documentation describes its supported languages and configuration choices, but those are specifications for the available platform. Griffin needs its own coverage and quality measurements before a team can assume parity. Source: CVI language documentation.
Disclosure matters because a convincing interface can change how much trust a user places in an answer. Tavus's current acceptable-use policy requires clear AI disclosure and explicit, informed consent for a replicated likeness. It also requires qualified involvement for tailored professional advice. A study conducted without initial disclosure should not be treated as a template for a customer experience.
Griffin-specific data terms should be checked separately. Tavus's privacy policy says information processed on behalf of business customers may be governed by their agreements. A general website policy therefore does not establish the precise retention, training-use or deletion terms of a Griffin trial.
The current pricing matrix advertises enterprise security and retention features, but that does not establish their inclusion in Griffin's preview. Ask which agreement covers live audio/video, reference images, voice samples, transcripts and any recordings. Include subprocessors, storage regions, deletion timing and training use in that review.
There is a practical control question too: what happens when a participant asks the model to stop, withdraws camera access or requests a human? A convincing face makes the experience smoother only if users can still understand and control it. Those behaviors deserve explicit acceptance tests.
Where Griffin could be useful, and what to test
Griffin is most interesting for applications where seeing and responding during a conversation could improve the task. The following are proposed evaluation scenarios, not claims of demonstrated production performance.
For a visual support assistant, ask users to show an unfamiliar object, rotate it, interrupt an explanation and correct a mistaken identification. Measure correct diagnosis, successful resolution, interruption recovery and the number of human handoffs. Include objects outside the assistant's knowledge so that admitting uncertainty is part of the test.
For tutoring, compare the same lesson through Griffin and the existing interface. Measure comprehension after the lesson, incorrect explanations, repeated clarification requests and the learner's ability to interrupt. A nod or encouraging expression is useful only when the teaching produces an accurate understanding.
For interview or sales practice, test whether the model follows the assigned role, changes its response when challenged and provides consistent feedback afterward. Disclose that the counterpart is AI. Measure whether the practice helps the participant perform the task, alongside how natural the conversation feels.
Use the same instructions and tasks across systems wherever possible. Record how often each succeeds, how long it takes, how often users interrupt successfully, and the full distribution of response delays. Include participants with different accents, devices and connection quality.
Once access and pricing are available, calculate the total cost of successful sessions, including failed attempts and human escalation. That will make it possible to judge whether Griffin's measured conversational advantages justify its integration effort and commercial terms.
