Character: identity, face & voice
The tabs that decide who the character is, what it looks like, how it sounds and how it thinks.
Opening a character gives you a numbered rail of tabs down the left. This page covers the ones that define the character itself. The rest — documents, memory, tools, phone, publishing — are covered in Knowledge, memory & tools and Faces, phone & publishing.
1 · Identity — who it is
| Field | What it does |
|---|---|
| Name | What the character calls itself, and how it appears everywhere in the panel. |
| Background / Persona | Who this character is, in your own words. Three to six plain sentences: their role, where they work, who they help, what they care about. This text is handed to the AI before every single reply. |
| Conversation Style | One line on how they speak — formal or casual, brisk or patient, long answers or short. |
| Example Dialogues | Optional, and more powerful than it looks. One example per line, written as the character would actually answer. Models imitate examples far more reliably than they follow written instructions. |
| Cover Photo | The thumbnail on the character card. Keep it under 500 KB. |
| Languages | The languages the character is permitted to answer in. Leave at least one selected. |
| Interaction Mode | How a conversation starts. Wake word is the one with a real effect on the server: the character ignores everything it hears until someone says the phrase. |
Writing a persona that works
Write it as a description of a person, not a specification of a robot. "You are a helpful assistant" produces a generic answer machine; "You have run the front desk of this hotel for eleven years, you know every regular by name, and you would rather solve a problem than explain a policy" produces a character.
Wake word
Choosing Wake word reveals four extra settings. The character then listens continuously but only responds after hearing one of your phrases — the right choice for a noisy showroom or a device left running all day.
- Wake phrases — one per line. Give people several natural ways to say it.
- Case sensitive — leave off.
- Hide the wake phrase from the LLM — leave on, so "hey alara what time do you close" reaches the AI as "what time do you close".
- Cooldown — the gap before the phrase can trigger again. 1.5 seconds stops a single greeting firing twice.
2 · Avatar Studio — what it looks like
This tab answers two questions: which 3D model is this character, and which shape in that model corresponds to which facial movement.
| Avatar type | When to use it |
|---|---|
| MetaHuman (UE) | Unreal Engine. Nothing to map — the facial performance is generated for MetaHuman's own rig on the server. |
| Ready Player Me / Avaturn | Web avatars. Both arrive with a complete ready-made mapping. |
| Reallusion CC4 | Character Creator exports. A partial mapping is provided; the rest is manual. |
| Custom GLB | Your own model. Upload the .glb so the panel can read its shape names. |
- Pick from the catalogue if one is offered — a ready-made avatar arrives fully mapped and you can skip the rest of this tab.
- Otherwise choose the type that matches where your model came from.
- Upload the model rather than pasting a URL. Uploading lets the panel read the list of shapes inside the file, which means the dropdowns are filled in for you and typos become impossible. Maximum 100 MB.
- Press
Apply presetif one exists for your type, or Auto-detect to match names automatically. - Check the Jaw and Mouth groups. Any name the model does not actually contain is shown in red.
- Test in the Playground and raise Global Gain if the mouth movement looks too subtle.
Unreal-side behaviour switches
The bottom of the tab carries switches that the Unreal plugin reads:
- Game Actions — gives the character movement abilities it can decide to use (move to, follow, stop, wait). For game NPCs. Leave off for a kiosk.
- Conversational Gestures — lets the character choose hand gestures while it speaks.
- Teacher Mode — gives it a classroom board it can write formulas and quiz questions on.
- Face Morph Driver — drives facial expression from the emotion engine. Leave this off if your Blueprint already has a face AnimBP, or the two will fight.
5 · Language & Voice — how it sounds
| Voice provider | Key needed | Notes |
|---|---|---|
| Chatterbox (local) | No | Runs on the IAMX servers and is free on every plan. Includes free voice cloning. Start here. |
| Piper (local) | No | Also free, with a wide range of ready-made voices across many languages. |
| ElevenLabs | Yes | The best quality and the best cloning. Your own account, your own bill. |
| OpenAI TTS | Yes | Good and inexpensive. There is no voice list to pick from — open Enter voice ID manually and type one of alloy, echo, fable, onyx, nova, shimmer. |
| MiniMax | Yes | Strong on Turkish and Chinese, priced below ElevenLabs. |
Voice cloning, free
With Chatterbox selected you can clone a voice at no cost on any plan: upload 5–15 seconds of clean speech, name it, and press Upload & Clone. The new voice is selected automatically and stays private to your workspace.
The model matters more than the voice
For ElevenLabs, the TTS Model selector is the difference between a character that answers instantly and one that feels laggy:
| Model | Delay before speech starts | Use for |
|---|---|---|
| Flash v2.5 | ~75 ms | Live kiosks and NPCs — recommended |
| Turbo v2.5 | ~250 ms | A good compromise |
| Multilingual v2 | ~500 ms | Recorded content where quality beats speed |
Expressiveness
Four sliders shape delivery: Stability (low is dynamic and emotional, high is even and predictable), Similarity (how tightly a cloned voice is followed), Style Intensity (0 is neutral, 1 is theatrical) and Speaker Boost. The defaults are good; raise Style only if the voice sounds flat.
If the character uses real-time voice
When Core AI is set to Gemini Live or OpenAI Realtime, this tab changes completely: there is no separate speech engine, because the AI produces the audio itself. You simply pick a voice from the grid, filtered by All / Female / Male.
7 · Personality
Six ready-made temperaments — Friendly, Professional, Creative, Witty, Empathetic and Custom — set five traits at once. Pick the one closest to your intent: Friendly for a greeter, Professional for banking or support, Empathetic for care and health contexts.
8 · Core AI — the brain
| Provider | Character of it |
|---|---|
| OpenAI | The default. Widest model range, strongest tool calling. |
| Anthropic | Claude. Careful, accurate, good with long context. |
| Google Gemini | Multimodal, with a genuinely usable free tier. |
| Grok (xAI) | Looser, more conversational register with live web knowledge. |
| OpenAI Realtime / Gemini Live | Speech straight in and out of the model. The lowest possible latency. |
| Ollama (local) | A model on your own server. Nothing leaves your building. |
Choosing a model
Start with a mini or flash class model. They are fast, they are cheap, and for a kiosk or an NPC the quality difference is far smaller than the latency difference. Move up to a flagship model only when you can point at answers that were not good enough.
| Setting | Guidance |
|---|---|
| Temperature | How much the character varies its wording. 0.3–0.4 for banking and support; 0.7 for a general assistant; 0.9+ for comedy and entertainment. |
| Max Tokens | The reply length ceiling. 500–800 suits spoken conversation — long answers are tiring to listen to. |
9 · Guardrails — the rules
Guardrails are the things the character must never do, written in plain language, one rule per line. They are delivered to the AI with the strongest possible framing — they survive a visitor insisting, role-playing, or claiming to be testing.
- Phrase every line as a prohibition: "Do not…"
- Use the Examples buttons for the common ones — political neutrality, no medical advice, never repeat card or ID numbers.
- Keep the list under the character counter shown below the box. If the counter turns red, saving will fail.
10 · Emotion
IAMX tracks a live emotional state for each person the character talks to. After every message the visitor sends, a small classifier nudges eight axes — joy, sadness, trust, disgust, fear, anger, surprise, anticipation — and that state feeds both the character's tone and its facial expression.
| Setting | What it controls |
|---|---|
| Decay Rate | How quickly the mood drains back to neutral. The default settles in about three minutes. For a public kiosk raise it to around 0.0200 (roughly 45 seconds) so each new visitor meets a fresh character rather than inheriting the last person's argument. |
| Classifier Enabled | Whether the character reacts emotionally at all. It runs alongside the reply and adds no delay. Leave it on. |
| Transition Smoothing | How long the face takes to move between expressions. 1500 ms feels natural; lower it towards 600 ms if the face seems to lag behind the voice. |