Skip to content

About

Software for RPI5 based smart speaker with reSpeaker microphone

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AdaAssistant

A Raspberry Pi 5 smart speaker runtime with wake-word detection, utterance recording, speech-to-text, stateful command routing, local actions, stenography, and spoken responses.

Ada is designed as a single-user assistant. Only one domain owns the conversation at a time. When the wake word is heard, active output is interrupted and the next utterance is processed against the current domain state.

Stack

  • Audio: PipeWire, WirePlumber, pw-cat, Pulse/PipeWire tools such as pactl
  • Wake/VAD: openWakeWord wake-word detection, WebRTC VAD
  • Primary STT: Vosk WebSocket ASR
  • Enhanced STT: OpenAI-compatible speech recognition API, enabled by default
  • Routing: RapidFuzz intent scoring plus FSM-aware domain routing
  • Actions: volume control, Yandex Music playback, stenography, LLM chat, timers, alarms, stopwatches, current time
  • TTS: RHVoice
  • Runtime: Python 3.11+, asyncio, FastAPI, Redis pub/sub, systemd user services

Architecture

The project is split into two runtime services plus local infrastructure.

  • speech_service owns the microphone path. It reads audio from PipeWire, detects the wake word, plays the recording cue, records utterance WAV files, runs VAD, handles stenography stop detection, and publishes speech events to Redis.
  • ada-service owns assistant state and domain execution. It receives speech_service events through SpeechEventAdapter, turns them into internal AssistantMessage values, routes messages with RoutingService, executes domains through AssistantDispatcher, persists state through ContextStore, and renders results through OutputDomain.
  • vosk provides fast local recognition for normal commands and for detecting the end of stenography recordings.
  • redis is the event bus between speech capture and assistant orchestration.
  • PipeWire and WirePlumber provide speaker playback, microphone capture, and AEC reference routing.

Current repo architecture scheme:

Audio / OS layer
  PipeWire + WirePlumber
    |        |
    |        +--> speaker output, AEC reference, cue/TTS/music playback
    |
speech_service
  wake word -> cue -> VAD recording -> WAV files -> Redis speech events
    |
    v
Redis: ada:speech:events
    |
    v
ada-service
  Orchestrator
    |  app construction, Redis lifetime, speech-event runtime shell
    |
    +--> SpeechEventAdapter
    |     external Redis payloads -> AssistantMessage
    |     wake/new-utterance cancellation and generation control
    |
    +--> SpeechInputDomain
    |     utterance WAV -> primary STT -> user_text message
    |
    +--> RoutingService
    |     DomainRouter scores
    |     enhanced-STT route recovery before any LLM path
    |     LLM route fallback using prompts.yml on enhanced text only
    |
    +--> AssistantDispatcher
    |     selected domain.run(message, context_view)
    |     sequential next_messages dispatch
    |
    +--> ContextStore
    |     AssistantContext mutation and persistence
    |
    +--> TimeService + progressive LangChain tool catalog
    |     system timezone from /etc/timezone
    |     initial LLM schema: time_help only
    |     TimerModule / AlarmModule / StopwatchModule
    |     persistent alarms.json; in-memory timers and stopwatches
    |
    +--> OutputDomain
          Redis result publish
          TTS / output cancellation / follow-up listening

Domains
  system:    SpeechInputDomain, OutputDomain
  global:    VolumeDomain, ListeningModeDomain
  exclusive: MusicDomain, StenographyDomain, ChatDomain

Shared utilities and adapters
  utils/llm.py       low-level LLM calls and response parsing
  prompts.yml        route and service prompt templates
  services/*         unified STT service, TTS, runtime assembly
  routes/*           HTTP API adapters

Inside ada-service, the current responsibility split is:

  • Orchestrator is the FastAPI construction target and speech-event runtime shell. It wires the runtime components, owns shutdown/Redis connection lifetime, and coordinates the speech-event pipeline. Domain-specific behavior is expected to live outside it.
  • SpeechEventAdapter subscribes to ada:speech:events, parses external payloads, owns single-flight wake/utterance cancellation, and emits internal messages.
  • SpeechInputDomain owns command STT. It converts utterance_ready messages with WAV paths into user_text messages.
  • RoutingService owns route selection. It wraps deterministic DomainRouter scoring, enhanced-STT route recovery, and LLM route fallback over eligible domains only. LLM-backed routing is allowed only after the message has been recovered through enhanced STT.
  • AssistantDispatcher routes one message, runs the selected domain, applies the returned result through ContextStore, and dispatches returned next_messages sequentially.
  • ContextStore owns mutation and best-effort persistence of AssistantContext.
  • TimeService is the composition root and thin public facade for time capabilities. TimerModule, AlarmModule, and StopwatchModule own their operations and state; only AlarmModule persists data. The LLM initially sees only time_help, which discloses and binds one requested tool group.
  • OutputDomain owns result publishing, TTS, output cancellation, interruptible media halt, and follow-up listening decisions.

Business domains own their own FSM state and execution rules: MusicDomain, VolumeDomain, StenographyDomain, ListeningModeDomain, and ChatDomain. System domains provide runtime capabilities: SpeechInputDomain and OutputDomain. Domain state is stored in AssistantContext and saved under the configured output directory, so recent domain memory can survive a service restart.

The command path is:

  1. speech_service publishes a Redis event.
  2. SpeechEventAdapter converts it to an internal message and handles superseding/cancellation.
  3. SpeechInputDomain performs STT when the event contains an utterance WAV.
  4. RoutingService selects the eligible domain using deterministic scores, enhanced STT recovery, and LLM fallback when needed.
  5. AssistantDispatcher executes the selected domain and asks ContextStore to apply the result.
  6. OutputDomain publishes the result and performs TTS/follow-up handling; due timers and alarms enter this same output path.

The public HTTP API is for health checks, status, TTS, STT, and speech-service control. Domain actions are not public HTTP endpoints; they run through Redis speech events and domain routing.

Setup

Prepare .env from example.env, set the required microphone, speaker, and inference values, then setup the assistant with two commands:

  1. sudo ./scripts/setup.sh .env
  2. ./scripts/init.sh

setup.sh installs the system packages, Python environments, code, models, audio units, and systemd user services. It also generates /etc/ada/ada.env from the source .env and runs an Ada import/config smoke check. Do not edit /etc/ada/ada.env by hand as the source of truth.

init.sh reloads and starts the user services, then verifies speech_service and ada-service through /health and /ready. If either service does not become ready, init.sh exits with an error instead of reporting success.

Services

Service Purpose Default endpoint
ada-combine-sink.service PipeWire speaker + AEC reference routing setup n/a
vosk.service Vosk STT WebSocket server ws://127.0.0.1:2700
speech_service.service Wake detection, VAD recording, Redis speech events, REST control API http://127.0.0.1:2710
ada-service.service API, Redis subscriber, STT/domain/output runtime http://127.0.0.1:8000

Startup order is PipeWire/WirePlumber, ada-combine-sink, vosk, speech_service, then ada-service.

Configuration

Edit .env before running setup. The most important values are:

  • A trained openWakeWord .onnx wake-word model under models/openwakeword/keywords/ (see the README.md there) plus WAKE_THRESHOLD for wake-word detection.
  • MICROPHONE_SOURCE_NAME, MICROPHONE_SINK_NAME, and SPEAKER_SINK_NAME for PipeWire routing (SPEAKER_SINK_NAME is the main playback sink; MICROPHONE_SINK_NAME is the AEC reference playback sink).
  • AEC_LOOPBACK_LATENCY_MSEC, MUSIC_OUTPUT_VOLUME_PERCENT, and CUE_PEAK to fine-tune AEC/reference timing and loudness balance.
  • YANDEX_MUSIC_TOKEN if music playback should use Yandex Music.
  • INFERENCE_IP, LLM_PORT, and SPEECH_RECOGNITION_PORT for local OpenAI-compatible inference services.
  • ENHANCED_STT_MODEL and ENHANCED_STT_TIMEOUT_SEC for the enhanced recognition path.
  • STENOGRAPHY_OUTPUT_PATH and ORCHESTRATOR_OUT_DIR for generated files and context persistence.
  • LLM_MAX_TOOL_ROUNDS limits the model/tool loop used by time tools.
  • /etc/timezone supplies the system timezone; persistent alarms are stored in ORCHESTRATOR_OUT_DIR/alarms.json.

Enhanced STT is enabled by default. The setup script derives the LLM and speech-recognition base URLs from the inference host and ports, so the inference IP should not be duplicated in multiple URL variables.

Runtime Behavior

The normal command path is wake word, cue, recording, Vosk recognition, domain routing, domain execution, context update, Redis result publish, then output playback/TTS.

Playback policy is quality-first: music/TTS/cues target the direct speaker sink, while ada-combine-sink.service creates a PipeWire loopback from speaker monitor to the AEC reference sink.

Vosk text is allowed only for deterministic local command routing. If the selected path needs an LLM, Ada first upgrades the utterance with enhanced STT. Chat, LLM route classification, and LLM parameter enrichment all run from enhanced-STT text only; if enhanced STT is disabled or returns no text, Ada asks the user to repeat instead of sending Vosk text to the model.

On wake, interruptible output playback is stopped before processing the new utterance. Music keeps queue and resume memory, stenography keeps its session paths, and chat keeps its history. Global volume commands can run without stealing the active domain.

Time tools

Time operations are exposed to ChatDomain as typed LangChain tools. They use enhanced STT and the LLM path, so Ada only confirms a change after the tool has returned a successful result.

Example request Operation
«Который час?» Return the current time using /etc/timezone
«Поставь таймер на десять минут» Create an in-memory countdown timer
«Поставь будильник на шесть утра» Create a persistent one-time alarm at the nearest future 06:00
«Будильник по будням на 7:30» Create a persistent recurring alarm with repeat=mon..fri
«Отключи будильник “Рабочие дни”» Set active=false without deleting the alarm
«Запусти секундомер» Create and start an in-memory stopwatch

time_help keeps the small-model context compact by loading one interface group:

Category Interfaces
clock current date and time
timer set, list, cancel
alarm current time; set, list, enable, disable, delete
stopwatch start, read, list, stop, reset

Time behavior:

  • The model initially receives only the time_help schema. That command discloses and binds one group (clock, timer, alarm, or stopwatch) for the remainder of the tool-call loop, keeping the small-model context compact. An exact hidden time-tool name is also accepted as a compatibility fallback; this does not expose the other tool schemas to the model.
  • HH:MM[:SS] for a one-time alarm means the nearest future occurrence: later today when possible, otherwise tomorrow. Future ISO 8601 date-times keep their exact date; a past date-time rolls forward to the nearest occurrence of its local clock time.
  • One-time alarms have repeat=[]. Recurring alarms contain canonical weekday values: mon, tue, wed, thu, fri, sat, sun.
  • active=false keeps an alarm in storage but excludes it from the scheduler.
  • The LLM supplies a short message; a due timer or alarm publishes one result and sends that exact message to TTS once.
  • Only alarms survive service restarts, in ORCHESTRATOR_OUT_DIR/alarms.json. Timers and stopwatches intentionally exist only in process memory.
  • Missed alarms are silent at startup. A missed one-time alarm becomes inactive; a recurring alarm advances to its next future occurrence.

Stenography uses Vosk while recording only to detect stop phrases. After recording ends, Ada transcribes the final WAV with enhanced STT, asks the LLM for a short title and Markdown summary, and uploads that note to Nextcloud under NEXTCLOUD_BASE_PATH/NEXTCLOUD_STENOGRAPHY_DIR as {title}-{timestamp}.md. The WAV stays local for STT; raw transcripts are not uploaded.

Logs And Checks

Use systemctl --user status speech_service.service --no-pager and systemctl --user status ada-service.service --no-pager to check service state.

Use journalctl --user-unit speech_service.service -f for wake-word, cue, VAD, and recording logs.

Use journalctl --user-unit ada-service.service -f for STT, routing, enhanced STT, domain execution, assistant context, time scheduling, and TTS logs. Time events use names such as time_service_started, time_alarm_created, time_alarm_fired, and time_notification_publish_failed; progressive tool loading logs llm_tools_disclosed and llm_hidden_tool_resolved.

Useful Ada endpoints are /health, /ready, /metrics, /api/status, /api/stt/recognize, and /api/tts/speak. /api/status includes the current assistant context and time_service, including its timezone, scheduler_running, and counts for timers, alarms, active alarms, and stopwatches.

The local time-feature verification command is:

./venv/bin/python -m pytest -q \
  tests/test_setup_scripts.py tests/test_time_service.py \
  tests/test_time_tools.py tests/test_llm_workflow.py \
  tests/test_prompt_config.py tests/test_time_notifications.py

Hardware and end-to-end audio checks are intentionally performed only after deployment to the target Raspberry Pi.

References

About

Software for RPI5 based smart speaker with reSpeaker microphone

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages