The Panel

Character: identity, face & voice

The tabs that decide who the character is, what it looks like, how it sounds and how it thinks.

Opening a character gives you a numbered rail of tabs down the left. This page covers the ones that define the character itself. The rest — documents, memory, tools, phone, publishing — are covered in Knowledge, memory & tools and Faces, phone & publishing.

You do not need to fill in every tab. A character works with a name, a background, one language, a model and a voice. Everything else has a sensible default you can revisit later.

1 · Identity — who it is

Identity tab with name, background persona, conversation style, example dialogues, languages and interaction mode
Identity. The Background field is the single most influential setting in the entire panel.
FieldWhat it does
NameWhat the character calls itself, and how it appears everywhere in the panel.
Background / PersonaWho this character is, in your own words. Three to six plain sentences: their role, where they work, who they help, what they care about. This text is handed to the AI before every single reply.
Conversation StyleOne line on how they speak — formal or casual, brisk or patient, long answers or short.
Example DialoguesOptional, and more powerful than it looks. One example per line, written as the character would actually answer. Models imitate examples far more reliably than they follow written instructions.
Cover PhotoThe thumbnail on the character card. Keep it under 500 KB.
LanguagesThe languages the character is permitted to answer in. Leave at least one selected.
Interaction ModeHow a conversation starts. Wake word is the one with a real effect on the server: the character ignores everything it hears until someone says the phrase.

Writing a persona that works

Write it as a description of a person, not a specification of a robot. "You are a helpful assistant" produces a generic answer machine; "You have run the front desk of this hotel for eleven years, you know every regular by name, and you would rather solve a problem than explain a policy" produces a character.

Keep facts out of the persona. Prices, opening hours, product specifications and policy text belong in the Knowledge Base — that way they are searched only when relevant, instead of being re-read on every single turn.

Wake word

Choosing Wake word reveals four extra settings. The character then listens continuously but only responds after hearing one of your phrases — the right choice for a noisy showroom or a device left running all day.

Leaving the phrase list empty switches the gate off entirely and the character answers everything — deliberate, so a half-configured kiosk stays usable rather than falling silent, but surprising if you expected it to be listening for a phrase.

2 · Avatar Studio — what it looks like

Avatar Studio showing avatar type selection, model source and the blend shape mapping editor
Avatar Studio. The mapping editor appears once the character has a 3D model to map.

This tab answers two questions: which 3D model is this character, and which shape in that model corresponds to which facial movement.

Avatar typeWhen to use it
MetaHuman (UE)Unreal Engine. Nothing to map — the facial performance is generated for MetaHuman's own rig on the server.
Ready Player Me / AvaturnWeb avatars. Both arrive with a complete ready-made mapping.
Reallusion CC4Character Creator exports. A partial mapping is provided; the rest is manual.
Custom GLBYour own model. Upload the .glb so the panel can read its shape names.
  1. Pick from the catalogue if one is offered — a ready-made avatar arrives fully mapped and you can skip the rest of this tab.
  2. Otherwise choose the type that matches where your model came from.
  3. Upload the model rather than pasting a URL. Uploading lets the panel read the list of shapes inside the file, which means the dropdowns are filled in for you and typos become impossible. Maximum 100 MB.
  4. Press Apply preset if one exists for your type, or Auto-detect to match names automatically.
  5. Check the Jaw and Mouth groups. Any name the model does not actually contain is shown in red.
  6. Test in the Playground and raise Global Gain if the mouth movement looks too subtle.
Switching avatar type does not replace an existing mapping — the old names stay and the face goes still. After changing type, always press Apply preset explicitly.

Unreal-side behaviour switches

The bottom of the tab carries switches that the Unreal plugin reads:

5 · Language & Voice — how it sounds

Language and Voice tab with TTS provider cards, API key field, voice selection and prosody sliders
Language & Voice. Provider first, then key, then voice, then model.
Voice providerKey neededNotes
Chatterbox (local)NoRuns on the IAMX servers and is free on every plan. Includes free voice cloning. Start here.
Piper (local)NoAlso free, with a wide range of ready-made voices across many languages.
ElevenLabsYesThe best quality and the best cloning. Your own account, your own bill.
OpenAI TTSYesGood and inexpensive. There is no voice list to pick from — open Enter voice ID manually and type one of alloy, echo, fable, onyx, nova, shimmer.
MiniMaxYesStrong on Turkish and Chinese, priced below ElevenLabs.

Voice cloning, free

With Chatterbox selected you can clone a voice at no cost on any plan: upload 5–15 seconds of clean speech, name it, and press Upload & Clone. The new voice is selected automatically and stays private to your workspace.

The model matters more than the voice

For ElevenLabs, the TTS Model selector is the difference between a character that answers instantly and one that feels laggy:

ModelDelay before speech startsUse for
Flash v2.5~75 msLive kiosks and NPCs — recommended
Turbo v2.5~250 msA good compromise
Multilingual v2~500 msRecorded content where quality beats speed
In conversation, half a second of silence before every reply is very noticeable. Flash is the right default for anything interactive.

Expressiveness

Four sliders shape delivery: Stability (low is dynamic and emotional, high is even and predictable), Similarity (how tightly a cloned voice is followed), Style Intensity (0 is neutral, 1 is theatrical) and Speaker Boost. The defaults are good; raise Style only if the voice sounds flat.

These are a starting point, not a fixed setting — the emotion engine nudges expressiveness up and down as the conversation goes, so a character genuinely sounds different when the conversation turns tense.

If the character uses real-time voice

When Core AI is set to Gemini Live or OpenAI Realtime, this tab changes completely: there is no separate speech engine, because the AI produces the audio itself. You simply pick a voice from the grid, filtered by All / Female / Male.

Voice names are not interchangeable between the two worlds. After switching between a real-time provider and a normal one, come back to this tab and pick the voice again — otherwise the character is left holding a voice name its new engine does not recognise.

7 · Personality

Personality tab with six preset cards and five Big Five trait sliders
Personality. The six presets are the practical control here.

Six ready-made temperaments — Friendly, Professional, Creative, Witty, Empathetic and Custom — set five traits at once. Pick the one closest to your intent: Friendly for a greeter, Professional for banking or support, Empathetic for care and health contexts.

Personality shapes the general temperament. For a specific voice — a catchphrase, a habit, a way of deflecting a question — write it into the Background and the Example Dialogues on the Identity tab. Those two fields do far more work than any slider.

8 · Core AI — the brain

Core AI tab with provider cards, model dropdown, temperature and max tokens sliders
Core AI. Choosing a provider resets the model to that provider's first entry.
ProviderCharacter of it
OpenAIThe default. Widest model range, strongest tool calling.
AnthropicClaude. Careful, accurate, good with long context.
Google GeminiMultimodal, with a genuinely usable free tier.
Grok (xAI)Looser, more conversational register with live web knowledge.
OpenAI Realtime / Gemini LiveSpeech straight in and out of the model. The lowest possible latency.
Ollama (local)A model on your own server. Nothing leaves your building.

Choosing a model

Start with a mini or flash class model. They are fast, they are cheap, and for a kiosk or an NPC the quality difference is far smaller than the latency difference. Move up to a flagship model only when you can point at answers that were not good enough.

SettingGuidance
TemperatureHow much the character varies its wording. 0.3–0.4 for banking and support; 0.7 for a general assistant; 0.9+ for comedy and entertainment.
Max TokensThe reply length ceiling. 500–800 suits spoken conversation — long answers are tiring to listen to.
Before switching to real-time voice, know the trade. Real-time gives you the fastest, most natural interruption handling available. In exchange the model answers from its own audio pipeline, so a real-time character does not draw on the Knowledge Base, long-term memory or the emotion engine the way a normal one does — and it consumes the monthly allowance 6 to 22 times faster. Excellent for a busy front desk; the wrong choice for a document-heavy support agent.

9 · Guardrails — the rules

Guardrails tab with the enable toggle, rules textarea and example rule buttons
Guardrails. The example buttons append a ready-made rule to the list.

Guardrails are the things the character must never do, written in plain language, one rule per line. They are delivered to the AI with the strongest possible framing — they survive a visitor insisting, role-playing, or claiming to be testing.

Guardrails are the one behavioural setting that still applies when a character is running in real-time voice mode. Whatever else changes, your rules travel with the character.

10 · Emotion

Emotion tab with the eight Plutchik emotion sliders and the emotion engine settings
Emotion. Eight axes on top, three engine settings underneath.

IAMX tracks a live emotional state for each person the character talks to. After every message the visitor sends, a small classifier nudges eight axes — joy, sadness, trust, disgust, fear, anger, surprise, anticipation — and that state feeds both the character's tone and its facial expression.

SettingWhat it controls
Decay RateHow quickly the mood drains back to neutral. The default settles in about three minutes. For a public kiosk raise it to around 0.0200 (roughly 45 seconds) so each new visitor meets a fresh character rather than inheriting the last person's argument.
Classifier EnabledWhether the character reacts emotionally at all. It runs alongside the reply and adds no delay. Leave it on.
Transition SmoothingHow long the face takes to move between expressions. 1500 ms feels natural; lower it towards 600 ms if the face seems to lag behind the voice.
On a shared kiosk, a decay rate of zero means moods never fade — an insult from an hour ago is still colouring the tone for the next person in the queue. Always give a public device a decay rate.

Where to go next