Architecture

Your phone shouldn't limit your bud.

Companion apps use flat illustrations because a phone can't render a film-quality human. So we don't ask it to. Unreal runs on a cloud GPU, and we stream the finished frames to you.

A voice turn, end to end

What happens between you speaking and your bud answering.

01 · You speak

The browser catches your voice and turns it into text, then sends it on with your bud's identity and your history attached.

02 · The bud thinks

The reply comes back in that character's voice: their cadence, their temperature, what they remember about you. Then their own ElevenLabs voice speaks it.

03 · The face answers

Unreal drives the lip sync off that audio on the GPU, and the finished frames come down to your phone over WebRTC, the same pipe a video call uses. You see a face, not a waveform.

Language

How one voice speaks thirty-two languages.

Nobody re-recorded anything. Each of the three pieces already in that loop turned out to do the work on its own.

01 · The mouth

The voice model infers the language from the text it's handed and speaks it in the same voice. Send it Spanish and you get Spanish, in that character's accent and cadence, with no new voice recording. That is normally the blocking cost of a feature like this, and it was already paid for.

02 · The mind

The character isn't translated into the language. It's told to be the same person speaking it natively, with that language's own idiom and humour, because a literal translation of a personality reads like a subtitle.

03 · The face

Lip sync is driven off the audio itself rather than a table of English phonemes, so it tracks whatever is being spoken. The face doesn't need to know the language either.

The honest half. Listening is the constraint, not speaking. Speech recognition runs on your own phone, so accuracy varies by language in a way the voice never does — strongest across European languages, solid for Japanese, Korean and Mandarin, and still catching up on some of the smaller ones.
Fidelity

Nothing gets turned down.

The render happens on a datacentre GPU instead of in your pocket, so nothing gets stripped out to fit. Lumen bounces real light around the room. Substrate makes skin behave like skin, with the light sinking in and coming back out. The hair is drawn strand by strand, the clothes move on Chaos cloth physics, and the lip sync is driven straight off the audio as it plays. Your phone does none of that work.

We tried the other path first. Two weeks of packaging Unreal natively for Android told us what the empty list of competitors already had: nobody ships that, because it doesn't work.

A MetaHuman bud rendered in Unreal Engine
Memory

It carries what you told it.

Your bud remembers what you were worried about last month, what you're working towards, who matters to you. All of it is encrypted with a passphrase you choose, on your phone, so we can't read it and neither can anyone who gets hold of our database. Standard keeps a rolling 30 days. Heyabud+ keeps the lot.

Kitana on the talk stage
The trade we made. Cloud rendering is what buys you a face like that, and it's also why there's no free tier. Every session is a GPU. We'd rather charge honestly for something real than give away something thin.