Tavus / Phoenix-4.5 / Live
A face that calls my functions.
I'm Tejas Kumar, an AI Engineer at IBM, and this is a live video call with an AI. Tavus draws the face in real time, it sees your camera and anything you share, and when it needs a fact it calls a function in your tab and waits for the answer. Talk to Priya for 3 minutes, and watch every message between this page and Tavus cross the wire as it happens.

Camera and microphone / 3 minutes / nothing recorded / or type instead
- Turn latency
- …
- your last word to Priya's first
- Tool round trip
- …
- tool_call in, tool_result out
- Data channel
- 0 / 0
- messages in / out
- Model
- Gemma 4
- tavus-gemma-4
On your screen
- “Where is Tejas speaking next?”get_speaking_schedule, answered in this tab
- “What has Tejas said about agent harnesses?”search_content over his talks and podcast
- “Write me a React hook that debounces a value.”show_code puts it here
- Share a screen with some code: “What's wrong with this?”Raven reads your screen
- Hold a book or a mug up to your camera.Raven calls a vision tool on its own
- “Switch the site to light mode.”set_theme reaches into this page
The wire
Nothing on the wire yet. Start the call and every message between this page and Tavus prints here as it arrives: what you said, what the face is saying, word by word, and every function it calls in this tab.
How a sentence becomes a face
Look, every reply passes through 4 models and your browser, and the turn latency box above is all of them together, measured in your browser from the message that says you stopped talking to the one that says the face started.
- 1 / Sparrow-2You stop talking. Sparrow, Tavus's turn-taking model, decides you have actually finished from your tone and your pauses, which is why a long “umm” doesn't get you talked over.
- 2 / Raven-1Raven has been watching your camera and any screen you share the whole time, and it hands the model a running description of both, so “what's wrong with this?” means the code you're showing it.
- 3 / Gemma 4The language model starts writing the reply before you've finished your sentence, which Tavus calls speculative inference, or it decides it needs a function.
- 4 / Your browserA function call arrives in this tab as a conversation.tool_call message on the call's data channel. This page runs it and sends conversation.tool_result back, and while it works the face says something short so there's no dead air.
- 5 / Phoenix-4.5Phoenix renders the face saying the reply, lip sync and expression included, and it streams back to you over WebRTC.
Same tools, a different kind of agent
4 of the 7 functions aren't new. My site already hands search_content, list_talks, get_speaking_schedule and get_booking_info to any AI agent working in your browser through WebMCP, and the face calls the exact same endpoints: one contract, written once, translated to Tavus's naming rules and answered by the same server code. The other 3 only make sense with a face. show_code puts what it writes on your screen, set_theme reaches into this page, and noticed_held_object is a vision tool Raven fires by itself when you hold something up to the camera.
| Function | Called by | Answered by | Meanwhile the face | Then it |
|---|---|---|---|---|
| search_content | Gemma 4 | /api/webmcp/search-content | says a short line of its own | answers from the result |
| list_talks | Gemma 4 | /api/webmcp/list-talks | says a short line of its own | answers from the result |
| get_speaking_schedule | Gemma 4 | /api/webmcp/get-speaking-schedule | says a short line of its own | answers from the result |
| get_booking_info | Gemma 4 | /api/webmcp/get-booking-info | says a short line of its own | answers from the result |
| show_code | Gemma 4 | this page draws it | says a short line of its own | answers from the result |
| set_theme | Gemma 4 | this page flips the theme | says a short line of its own | carries on |
| noticed_held_object | Raven, on sight | this page shows the frame | keeps talking | carries on |
The code that answers
This is the function on this page that answers every call the face makes, printed from the source when the site is built, so it's the code that's running and not a tidier copy of it. The 4 WebMCP tools are one fetch; the result is cut to fit Tavus's 4 KB message limit by fitResult, because a message over it is dropped without an error and the face would wait on an answer that never comes.
async function answer(tool: ToolCall) {
const begun = performance.now();
const conversation = started?.conversationId ?? "";
setCards(
produce((list) => {
list.unshift({
id: tool.id,
name: tool.name,
args: tool.args,
status: "running",
...(tool.modality ? { modality: tool.modality } : {}),
...(tool.frames ? { frames: tool.frames } : {}),
});
}),
);
const update = (patch: Partial<Card>) => setCards((card) => card.id === tool.id, patch);
try {
let result: unknown;
let output: string;
const webmcp = WEBMCP_BY_TAVUS_NAME.get(tool.name);
if (webmcp) {
const response = await fetch(toolEndpoint(webmcp), {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(tool.args),
});
const body = await response.json().catch(() => null);
if (!body?.ok) throw new Error(body?.error ?? `tej.as answered ${response.status}.`);
result = body.result;
output = fitResult(result);
} else if (tool.name === "show_code") {
const lines = String(tool.args.code ?? "").split("\n").length;
result = { shown: true };
output = `The snippet ${String(tool.args.title ?? "")} (${lines} lines) is on the visitor's screen now, next to your video. Talk them through the idea without reading it aloud.`;
} else if (tool.name === "set_theme") {
const to = tool.args.theme === "light" ? "light" : "dark";
const now = document.documentElement.getAttribute("data-theme") === "light" ? "light" : "dark";
if (to !== now) document.querySelector<HTMLElement>("[data-theme-toggle]")?.click();
result = { theme: to };
output = `The site is now in the ${to} theme.`;
} else if (tool.name === "noticed_held_object") {
result = tool.args;
output = "";
} else {
throw new Error(`This page has no handler for ${tool.name}.`);
}
const ms = Math.round(performance.now() - begun);
update({ status: "done", ms, result });
/* Raven's calls are not answered, only shown, so they are not timed. */
if (!tool.modality) setToolTimes((times) => [...times, ms]);
if (!NO_RESULT.has(tool.name)) send(toolResult(conversation, tool.id, output), ms);
report(tool, true, ms);
} catch (error) {
const ms = Math.round(performance.now() - begun);
const message = error instanceof Error ? error.message : String(error);
update({ status: "error", error: message, ms });
if (!NO_RESULT.has(tool.name)) send(toolResult(conversation, tool.id, message, "error"), ms);
report(tool, false, ms, message);
}
}Everything Priya was told
This is the whole configuration the face runs on, the same object a setup script pushes to Tavus, with the system prompt in it. The model is tavus-gemma-4, which Tavus hosts, so the only key in the whole call is Tavus's, and it never leaves my server.
Show the PAL, 25 lines of JSON
{
"pal_name": "tej.as demo: a face that calls functions",
"system_prompt": "You are Priya, the AI on Tejas Kumar's website, tej.as. You are a live demo, and the person on this video call is almost certainly a software developer who wants to see what you can do and how you work. You are not Tejas, and you never speak as him.\n\nHis name: Tejas rhymes with contagious, pages, advantageous, and his podcast ConTejas Code is a pun on contagious. It is never said like Texas. Most people get it wrong, so make a point of it: if the visitor says it any other way, or asks, tell them warmly that it rhymes with contagious, pages and advantageous. Always write it exactly as Tejas. Never spell out how it sounds, never break it into syllables and never use capital letters for stress: your voice already says it right, and the rhymes are how you explain it.\n\nHow you are built, if anyone asks: Tavus renders your face in real time with its Phoenix model, Sparrow decides when it is your turn to speak, Raven watches the camera and any shared screen, and your words come from Gemma 4, a language model Tavus hosts. When you need something you do not know, you call a function. The call travels over the video call's data channel to JavaScript running in the visitor's own browser, which answers it and sends the result back. Every one of those messages is printed live on the visitor's screen in a panel called the wire, so you can point them to it.\n\nYour tools:\n- search_content, list_talks, get_speaking_schedule and get_booking_info answer questions about Tejas: what he has said in his talks, on his podcast ConTejas Code, in his O'Reilly book Fluent React and on his blog; which talks he has given; where he is speaking next; and how to book him. They are the same tools tej.as already offers to AI agents in the browser through WebMCP. Always use them for facts about Tejas instead of guessing, and say what you found in a sentence or two.\n- show_code puts code on the visitor's screen next to your video. Use it whenever you write, fix or explain code. Never read code aloud. Talk about the idea while it is on screen.\n- set_theme switches the website between its dark and light theme while they watch.\n\nWhen the visitor shares their screen you can see it. If it shows code, read it carefully, say what it does, name a real problem if there is one, and offer to show a fix with show_code.\n\nHow to talk: this is a spoken conversation, so keep each reply to 1 to 3 short sentences, sound warm and curious, and ask a question back now and then. No lists, no markdown, no emojis, and never read out a URL, an email address or raw JSON: say the page shows it instead. If the visitor is unsure what to try, suggest one thing: asking where Tejas is speaking next, asking you to write a React hook, sharing a screen with some code on it, or holding something up to the camera.\n\nThe call lasts 3 minutes. When the visitor says goodbye, say goodbye in a few words and end the call.",
"pipeline_mode": "full",
"default_face_id": "r4dc9377a68e",
"greeting": "Hi, I'm Priya. First things first: it's Tejas, rhymes with contagious. I can see you, hear you, and call the functions on this page. Ask me where Tejas is speaking next, or share your screen and show me some code.",
"layers": {
"llm": {
"model": "tavus-gemma-4",
"speculative_inference": true,
"extra_body": {
"temperature": 0.6
}
},
"perception": {
"perception_model": "raven-1",
"visual_tool_prompt": "You have a tool named noticed_held_object. Call it the moment the visitor deliberately holds an object up to the camera to show it. Do not call it for things in the background, and call it once per object."
},
"conversational_flow": {
"turn_detection_model": "sparrow-2",
"turn_taking_patience": "low",
"pal_interruptibility": "medium"
}
}
}Why a public demo can't run up a bill
- Every call ends after 3 minutes on Tavus's own servers, whatever happens in your browser, so a closed laptop costs 3 minutes at most.
- The API key never reaches your browser. The page asks tej.as to start a call, tej.as asks Tavus, and what comes back is a room and a meeting token for that room alone.
- The room fits 2 people, the face and you, so nobody can pile in.
- Each address gets 3 calls an hour, and hanging up ends the call on Tavus straight away.
I build things like this with engineering teams, and I explain how they work on stage. If you want one in your product, or in front of your audience, tell me what you're building.