01 · SENSE
MoLing
MoLing
MoShun AI Lab's interaction entry point — an exploration of how text, voice, vision, context and tool use can share a single space.
What this is
MoLing is not a chat box but an interaction field you can see. Voice, text and visual signals enter one shared context, and the system’s state — listening, retrieving, answering — is expressed visually rather than hidden behind a spinner.
The problem
Most AI interfaces compress rich capability into a single input box: you cannot see what the system is doing, or where an answer came from. When interaction is reduced to a stream of text, the value of voice, vision and context is wasted.
Core capabilities
- 01
Multimodal input, one context
Text and voice share one session and one knowledge scope; switching input mode never drops the thread.
- 02
Visible system state
Listening, thinking, retrieving and answering each map to a distinct visual state, so the process itself is readable.
- 03
Bounded knowledge scope
Answers are drawn from a configured knowledge scope rather than unbounded open generation.
- 04
Explicit device permission
Microphone and camera activate only when the user turns them on, and can be switched off at any moment.
How it works
- 01
Enter
The interface wakes, states that it is in integration mode, and shows which inputs are available.
- 02
Input
The user opens with text or voice; the visual state responds to input intensity.
- 03
Understand
Intent is parsed and, where needed, retrieved against the configured knowledge scope.
- 04
Respond
The answer returns, context is retained, and the user can follow up or switch modality.
Where it sits across the capability domains
- 01 · SENSESense & InteractPrimary domain
- 02 · UNDERSTANDKnowledge & JudgementAlso touches
- 03 · CREATEContent & GenerationNot directly involved
- 04 · ACTAutomation & ExecutionNot directly involved
- 05 · ORCHESTRATEEnterprise OrchestrationNot directly involved
Use cases
- 01
Brand narration on a website
Moves a visitor from static reading into a brand experience they can talk to.
- 02
Customer enquiry
Answers questions on capability, use cases and partnership quickly.
- 03
On-site exhibition displays
Suits immersive large-format screens, product launches and client reception spaces.
Enterprise system connections
- Enterprise knowledge base
- Product and documentation sources
- Web and exhibition front-ends
- Support ticketing
Data, permission and deployment
Text and live voice connect to Mo-Voice with a one-time short-lived token; camera frames are never read, displayed, stored or uploaded. Microphone audio is sent only for the active session while the user has enabled it, and stops immediately when switched off. Knowledge scope and session retention are configured per engagement.
What is open today
MoLing on this site runs in integration mode: text and live voice are connected to Mo-Voice, while the camera/vision channel remains unconnected. Full capability for enterprise scenarios is delivered as a project.
Apply for access
Want to know what this does in your own context?
Tell us the scenario and the systems already in place. We will start by judging whether it is worth doing at all, then talk about how.
Openly explorable, still evolving, and not representative of the final product.
