Embodied systems · Explore this field ↗ · Field map · 2 min read
A taxonomy of embodied and 3D chatbots
A rendered character qualifies only when conversation and embodied behavior form one responsive loop.
The inclusion test
A 3D-powered chatbot combines meaningful AI conversation with a realtime character or environment. The embodiment should respond dynamically through lip-sync, expression, gaze, gesture, posture, movement, or world action. A pre-recorded talking head, a text bot beside a decorative model, or a game NPC with only fixed dialogue does not meet that test. Borderline systems should disclose which signals are genuinely driven by conversation.
Useful categories
The field includes virtual assistants, companions, customer-service agents, education and training characters, health-navigation guides, game and entertainment characters, brand ambassadors, spatial-computing agents, and multi-user world participants. The same platform may support several categories, but the service obligations differ. A training character needs scenario control and assessment; a support agent needs policy accuracy and handoff; a companion raises dependency and age concerns.
Do not collapse adjacent fields
Digital humans describe representation and rendering, not necessarily conversation. VTubers are usually human-performed. Virtual influencers are media identities. Voice assistants may have no body. General NPC systems may optimize narrative play rather than reliable service. The distinctions matter because evidence, rights, costs, and safety controls do not transfer automatically.
Research boundary
This guide recovered a detailed owner-supplied 3D chatbot brief and the Character Current research edition. Product pricing and capability move quickly, so the durable contribution here is the inclusion test and evaluation framework. Vendor claims remain claims until a current implementation is tested; this build does not invent hands-on trials.
A classification checklist
To classify a candidate, record whether conversation is generated or scripted; whether the character is 3D, 2D, or video; which embodied channels respond to live state; whether the user can interrupt; whether the agent can act in a world; and whether a human performs any layer. Then name the use context and consequence. A visually impressive demo may still be a scripted presenter. A simple stylized body with responsive gaze and action may be genuinely embodied. Publish borderline cases with the missing criterion instead of forcing them into the directory.
Primary reading
Sources and limits
These links support the architecture, policy, or product behavior discussed above. Vendor documentation describes vendor features; it is not independent proof of performance. Current details should be rechecked before a production decision.