MoLing
Openly explorable, still evolving, and not representative of the final product.
MoShun AI Lab's interaction entry point — an exploration of how text, voice, vision, context and tool use can share a single space.
MoLing is not a chat box but an interaction field you can see. Voice, text and visual signals enter one shared context, and the system’s state — listening, retrieving, answering — is expressed visually rather than hidden behind a spinner.
Most AI interfaces compress rich capability into a single input box: you cannot see what the system is doing, or where an answer came from. When interaction is reduced to a stream of text, the value of voice, vision and context is wasted.
The field
SENSEIntegration mode: text and live voice share one Mo-Voice MoLing session; the camera/vision channel is still not connected.
Inputs & connections
All seven entries remain visible; availability and permission use are stated before and after each action.
Text is available. Every other entry explicitly shows whether it is a permission preview, experience-only or coming later.
Waiting for input. No device or file is being read.
State sequence
- 01Idle
- 02Listening
- 03Understanding
- 04Calling
- 05Generating
- 06Complete
- 07Failed
Device permissions
These turn on only when you enable them. Nothing is requested automatically.
While enabled, microphone audio is converted to PCM16 and streamed to Mo-Voice; capture stops immediately when switched off.
No frame is ever displayed anywhere; the video track is released the instant permission is granted, and nothing is read or stored.
Camera frames are never read, displayed or uploaded; the microphone streams current-session audio to Mo-Voice only while enabled and stops immediately when switched off.
- MoLing
Hello — I am MoLing.
Text and live voice are connected to Mo-Voice. Ask a question directly, or enable the microphone for a continuous conversation.
Suggested questions
Core capabilities
- 01
Multimodal input, one context
Text and voice share one session and one knowledge scope; switching input mode never drops the thread.
- 02
Visible system state
Listening, thinking, retrieving and answering each map to a distinct visual state, so the process itself is readable.
- 03
Bounded knowledge scope
Answers are drawn from a configured knowledge scope rather than unbounded open generation.
- 04
Explicit device permission
Microphone and camera activate only when the user turns them on, and can be switched off at any moment.
What is open today
MoLing on this site runs in integration mode: text and live voice are connected to Mo-Voice, while the camera/vision channel remains unconnected. Full capability for enterprise scenarios is delivered as a project.
Data, permission and deployment
Camera frames are never read, displayed or uploaded; the microphone streams current-session audio to Mo-Voice only while enabled and stops immediately when switched off.
Text and live voice connect to Mo-Voice with a one-time short-lived token; camera frames are never read, displayed, stored or uploaded. Microphone audio is sent only for the active session while the user has enabled it, and stops immediately when switched off. Knowledge scope and session retention are configured per engagement.
