Skip to main content
01SENSE · Sense & Interact

MoLing

Experience Preview

Openly explorable, still evolving, and not representative of the final product.

MoShun AI Lab's interaction entry point — an exploration of how text, voice, vision, context and tool use can share a single space.

Positioning

MoLing is not a chat box but an interaction field you can see. Voice, text and visual signals enter one shared context, and the system’s state — listening, retrieving, answering — is expressed visually rather than hidden behind a spinner.

Most AI interfaces compress rich capability into a single input box: you cannot see what the system is doing, or where an answer came from. When interaction is reduced to a stream of text, the value of voice, vision and context is wasted.

The field

SENSE
Integration mode

Integration mode: text and live voice share one Mo-Voice MoLing session; the camera/vision channel is still not connected.

Inputs & connections

All seven entries remain visible; availability and permission use are stated before and after each action.

07 / entries

Text is available. Every other entry explicitly shows whether it is a permission preview, experience-only or coming later.

Idle

Waiting for input. No device or file is being read.

State sequence

  1. 01Idle
  2. 02Listening
  3. 03Understanding
  4. 04Calling
  5. 05Generating
  6. 06Complete
  7. 07Failed

Device permissions

These turn on only when you enable them. Nothing is requested automatically.

MicrophoneOff

While enabled, microphone audio is converted to PCM16 and streamed to Mo-Voice; capture stops immediately when switched off.

CameraOff

No frame is ever displayed anywhere; the video track is released the instant permission is granted, and nothing is read or stored.

Camera frames are never read, displayed or uploaded; the microphone streams current-session audio to Mo-Voice only while enabled and stops immediately when switched off.

  1. MoLing

    Hello — I am MoLing.

    Text and live voice are connected to Mo-Voice. Ask a question directly, or enable the microphone for a continuous conversation.

Suggested questions

Core capabilities

  1. 01

    Multimodal input, one context

    Text and voice share one session and one knowledge scope; switching input mode never drops the thread.

  2. 02

    Visible system state

    Listening, thinking, retrieving and answering each map to a distinct visual state, so the process itself is readable.

  3. 03

    Bounded knowledge scope

    Answers are drawn from a configured knowledge scope rather than unbounded open generation.

  4. 04

    Explicit device permission

    Microphone and camera activate only when the user turns them on, and can be switched off at any moment.

What is open today

MoLing on this site runs in integration mode: text and live voice are connected to Mo-Voice, while the camera/vision channel remains unconnected. Full capability for enterprise scenarios is delivered as a project.

Data, permission and deployment

Camera frames are never read, displayed or uploaded; the microphone streams current-session audio to Mo-Voice only while enabled and stops immediately when switched off.

Text and live voice connect to Mo-Voice with a one-time short-lived token; camera frames are never read, displayed, stored or uploaded. Microphone audio is sent only for the active session while the user has enabled it, and stops immediately when switched off. Knowledge scope and session retention are configured per engagement.