Skip to content

Feature: Add SenseVoice for local voice input #61

Description

@LauraGPT

Important

Correction: The earlier comparative claims below are withdrawn; they were not established by a matched benchmark. SenseVoiceSmall supports Mandarin Chinese, Cantonese, English, Japanese, and Korean. It can emit language, emotion, and audio-event tags, but speaker diarization requires a separate model or pipeline (for example CAM++) and is not a built-in SenseVoice result. Runtime, timestamps, punctuation, and performance depend on the selected model, interface, hardware, and audio. FunASR and SenseVoice repository source code is MIT; model weights follow each model card. Please evaluate the exact integration on this project's workload.

Hi! VisionClaw is impressive — real-time AI for smart glasses with voice + vision.

For the voice input/ASR component, SenseVoice could improve responsiveness:

Why SenseVoice for smart glasses?

  • 5x faster than Whisper — critical for real-time wearable interaction
  • 234M params — lightweight enough for edge/phone processing
  • Non-autoregressive — constant-time decoding, instant results
  • Emotion detection — understand user's mood from voice
  • Audio event detection — context awareness (music, traffic, crowds)
  • 50+ languages — auto-detection

Edge deployment options

Links

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions