A Raspberry Pi 5 smart speaker runtime with wake-word detection, utterance recording, speech-to-text, stateful command routing, local actions, stenography, and spoken responses.
Ada is designed as a single-user assistant. Only one domain owns the conversation at a time. When the wake word is heard, active output is interrupted and the next utterance is processed against the current domain state.
- Audio: PipeWire, WirePlumber,
pw-cat, Pulse/PipeWire tools such aspactl - Wake/VAD: openWakeWord wake-word detection, WebRTC VAD
- Primary STT: Vosk WebSocket ASR
- Enhanced STT: OpenAI-compatible speech recognition API, enabled by default
- Routing: RapidFuzz intent scoring plus FSM-aware domain routing
- Actions: volume control, Yandex Music playback, stenography, LLM chat, timers, alarms, stopwatches, current time
- TTS: RHVoice
- Runtime: Python 3.11+, asyncio, FastAPI, Redis pub/sub, systemd user services
The project is split into two runtime services plus local infrastructure.
speech_serviceowns the microphone path. It reads audio from PipeWire, detects the wake word, plays the recording cue, records utterance WAV files, runs VAD, handles stenography stop detection, and publishes speech events to Redis.ada-serviceowns assistant state and domain execution. It receivesspeech_serviceevents throughSpeechEventAdapter, turns them into internalAssistantMessagevalues, routes messages withRoutingService, executes domains throughAssistantDispatcher, persists state throughContextStore, and renders results throughOutputDomain.voskprovides fast local recognition for normal commands and for detecting the end of stenography recordings.redisis the event bus between speech capture and assistant orchestration.- PipeWire and WirePlumber provide speaker playback, microphone capture, and AEC reference routing.
Current repo architecture scheme:
Audio / OS layer
PipeWire + WirePlumber
| |
| +--> speaker output, AEC reference, cue/TTS/music playback
|
speech_service
wake word -> cue -> VAD recording -> WAV files -> Redis speech events
|
v
Redis: ada:speech:events
|
v
ada-service
Orchestrator
| app construction, Redis lifetime, speech-event runtime shell
|
+--> SpeechEventAdapter
| external Redis payloads -> AssistantMessage
| wake/new-utterance cancellation and generation control
|
+--> SpeechInputDomain
| utterance WAV -> primary STT -> user_text message
|
+--> RoutingService
| DomainRouter scores
| enhanced-STT route recovery before any LLM path
| LLM route fallback using prompts.yml on enhanced text only
|
+--> AssistantDispatcher
| selected domain.run(message, context_view)
| sequential next_messages dispatch
|
+--> ContextStore
| AssistantContext mutation and persistence
|
+--> TimeService + progressive LangChain tool catalog
| system timezone from /etc/timezone
| initial LLM schema: time_help only
| TimerModule / AlarmModule / StopwatchModule
| persistent alarms.json; in-memory timers and stopwatches
|
+--> OutputDomain
Redis result publish
TTS / output cancellation / follow-up listening
Domains
system: SpeechInputDomain, OutputDomain
global: VolumeDomain, ListeningModeDomain
exclusive: MusicDomain, StenographyDomain, ChatDomain
Shared utilities and adapters
utils/llm.py low-level LLM calls and response parsing
prompts.yml route and service prompt templates
services/* unified STT service, TTS, runtime assembly
routes/* HTTP API adapters
Inside ada-service, the current responsibility split is:
Orchestratoris the FastAPI construction target and speech-event runtime shell. It wires the runtime components, owns shutdown/Redis connection lifetime, and coordinates the speech-event pipeline. Domain-specific behavior is expected to live outside it.SpeechEventAdaptersubscribes toada:speech:events, parses external payloads, owns single-flight wake/utterance cancellation, and emits internal messages.SpeechInputDomainowns command STT. It convertsutterance_readymessages with WAV paths intouser_textmessages.RoutingServiceowns route selection. It wraps deterministicDomainRouterscoring, enhanced-STT route recovery, and LLM route fallback over eligible domains only. LLM-backed routing is allowed only after the message has been recovered through enhanced STT.AssistantDispatcherroutes one message, runs the selected domain, applies the returned result throughContextStore, and dispatches returnednext_messagessequentially.ContextStoreowns mutation and best-effort persistence ofAssistantContext.TimeServiceis the composition root and thin public facade for time capabilities.TimerModule,AlarmModule, andStopwatchModuleown their operations and state; onlyAlarmModulepersists data. The LLM initially sees onlytime_help, which discloses and binds one requested tool group.OutputDomainowns result publishing, TTS, output cancellation, interruptible media halt, and follow-up listening decisions.
Business domains own their own FSM state and execution rules: MusicDomain, VolumeDomain, StenographyDomain, ListeningModeDomain, and ChatDomain. System domains provide runtime capabilities: SpeechInputDomain and OutputDomain. Domain state is stored in AssistantContext and saved under the configured output directory, so recent domain memory can survive a service restart.
The command path is:
speech_servicepublishes a Redis event.SpeechEventAdapterconverts it to an internal message and handles superseding/cancellation.SpeechInputDomainperforms STT when the event contains an utterance WAV.RoutingServiceselects the eligible domain using deterministic scores, enhanced STT recovery, and LLM fallback when needed.AssistantDispatcherexecutes the selected domain and asksContextStoreto apply the result.OutputDomainpublishes the result and performs TTS/follow-up handling; due timers and alarms enter this same output path.
The public HTTP API is for health checks, status, TTS, STT, and speech-service control. Domain actions are not public HTTP endpoints; they run through Redis speech events and domain routing.
Prepare .env from example.env, set the required microphone, speaker, and inference values, then setup the assistant with two commands:
sudo ./scripts/setup.sh .env./scripts/init.sh
setup.sh installs the system packages, Python environments, code, models, audio units, and systemd user services. It also generates /etc/ada/ada.env from the source .env and runs an Ada import/config smoke check. Do not edit /etc/ada/ada.env by hand as the source of truth.
init.sh reloads and starts the user services, then verifies speech_service and ada-service through /health and /ready. If either service does not become ready, init.sh exits with an error instead of reporting success.
| Service | Purpose | Default endpoint |
|---|---|---|
ada-combine-sink.service |
PipeWire speaker + AEC reference routing setup | n/a |
vosk.service |
Vosk STT WebSocket server | ws://127.0.0.1:2700 |
speech_service.service |
Wake detection, VAD recording, Redis speech events, REST control API | http://127.0.0.1:2710 |
ada-service.service |
API, Redis subscriber, STT/domain/output runtime | http://127.0.0.1:8000 |
Startup order is PipeWire/WirePlumber, ada-combine-sink, vosk, speech_service, then ada-service.
Edit .env before running setup. The most important values are:
- A trained openWakeWord
.onnxwake-word model undermodels/openwakeword/keywords/(see theREADME.mdthere) plusWAKE_THRESHOLDfor wake-word detection. MICROPHONE_SOURCE_NAME,MICROPHONE_SINK_NAME, andSPEAKER_SINK_NAMEfor PipeWire routing (SPEAKER_SINK_NAMEis the main playback sink;MICROPHONE_SINK_NAMEis the AEC reference playback sink).AEC_LOOPBACK_LATENCY_MSEC,MUSIC_OUTPUT_VOLUME_PERCENT, andCUE_PEAKto fine-tune AEC/reference timing and loudness balance.YANDEX_MUSIC_TOKENif music playback should use Yandex Music.INFERENCE_IP,LLM_PORT, andSPEECH_RECOGNITION_PORTfor local OpenAI-compatible inference services.ENHANCED_STT_MODELandENHANCED_STT_TIMEOUT_SECfor the enhanced recognition path.STENOGRAPHY_OUTPUT_PATHandORCHESTRATOR_OUT_DIRfor generated files and context persistence.LLM_MAX_TOOL_ROUNDSlimits the model/tool loop used by time tools./etc/timezonesupplies the system timezone; persistent alarms are stored inORCHESTRATOR_OUT_DIR/alarms.json.
Enhanced STT is enabled by default. The setup script derives the LLM and speech-recognition base URLs from the inference host and ports, so the inference IP should not be duplicated in multiple URL variables.
The normal command path is wake word, cue, recording, Vosk recognition, domain routing, domain execution, context update, Redis result publish, then output playback/TTS.
Playback policy is quality-first: music/TTS/cues target the direct speaker sink, while ada-combine-sink.service creates a PipeWire loopback from speaker monitor to the AEC reference sink.
Vosk text is allowed only for deterministic local command routing. If the selected path needs an LLM, Ada first upgrades the utterance with enhanced STT. Chat, LLM route classification, and LLM parameter enrichment all run from enhanced-STT text only; if enhanced STT is disabled or returns no text, Ada asks the user to repeat instead of sending Vosk text to the model.
On wake, interruptible output playback is stopped before processing the new utterance. Music keeps queue and resume memory, stenography keeps its session paths, and chat keeps its history. Global volume commands can run without stealing the active domain.
Time operations are exposed to ChatDomain as typed LangChain tools. They use
enhanced STT and the LLM path, so Ada only confirms a change after the tool has
returned a successful result.
| Example request | Operation |
|---|---|
| «Который час?» | Return the current time using /etc/timezone |
| «Поставь таймер на десять минут» | Create an in-memory countdown timer |
| «Поставь будильник на шесть утра» | Create a persistent one-time alarm at the nearest future 06:00 |
| «Будильник по будням на 7:30» | Create a persistent recurring alarm with repeat=mon..fri |
| «Отключи будильник “Рабочие дни”» | Set active=false without deleting the alarm |
| «Запусти секундомер» | Create and start an in-memory stopwatch |
time_help keeps the small-model context compact by loading one interface group:
| Category | Interfaces |
|---|---|
clock |
current date and time |
timer |
set, list, cancel |
alarm |
current time; set, list, enable, disable, delete |
stopwatch |
start, read, list, stop, reset |
Time behavior:
- The model initially receives only the
time_helpschema. That command discloses and binds one group (clock,timer,alarm, orstopwatch) for the remainder of the tool-call loop, keeping the small-model context compact. An exact hidden time-tool name is also accepted as a compatibility fallback; this does not expose the other tool schemas to the model. HH:MM[:SS]for a one-time alarm means the nearest future occurrence: later today when possible, otherwise tomorrow. Future ISO 8601 date-times keep their exact date; a past date-time rolls forward to the nearest occurrence of its local clock time.- One-time alarms have
repeat=[]. Recurring alarms contain canonical weekday values:mon,tue,wed,thu,fri,sat,sun. active=falsekeeps an alarm in storage but excludes it from the scheduler.- The LLM supplies a short
message; a due timer or alarm publishes one result and sends that exact message to TTS once. - Only alarms survive service restarts, in
ORCHESTRATOR_OUT_DIR/alarms.json. Timers and stopwatches intentionally exist only in process memory. - Missed alarms are silent at startup. A missed one-time alarm becomes inactive; a recurring alarm advances to its next future occurrence.
Stenography uses Vosk while recording only to detect stop phrases. After recording ends, Ada transcribes the final WAV with enhanced STT, asks the LLM for a short title and Markdown summary, and uploads that note to Nextcloud under NEXTCLOUD_BASE_PATH/NEXTCLOUD_STENOGRAPHY_DIR as {title}-{timestamp}.md. The WAV stays local for STT; raw transcripts are not uploaded.
Use systemctl --user status speech_service.service --no-pager and systemctl --user status ada-service.service --no-pager to check service state.
Use journalctl --user-unit speech_service.service -f for wake-word, cue, VAD, and recording logs.
Use journalctl --user-unit ada-service.service -f for STT, routing, enhanced STT, domain execution, assistant context, time scheduling, and TTS logs. Time events use names such as time_service_started, time_alarm_created, time_alarm_fired, and time_notification_publish_failed; progressive tool loading logs llm_tools_disclosed and llm_hidden_tool_resolved.
Useful Ada endpoints are /health, /ready, /metrics, /api/status, /api/stt/recognize, and /api/tts/speak. /api/status includes the current assistant context and time_service, including its timezone, scheduler_running, and counts for timers, alarms, active alarms, and stopwatches.
The local time-feature verification command is:
./venv/bin/python -m pytest -q \
tests/test_setup_scripts.py tests/test_time_service.py \
tests/test_time_tools.py tests/test_llm_workflow.py \
tests/test_prompt_config.py tests/test_time_notifications.pyHardware and end-to-end audio checks are intentionally performed only after deployment to the target Raspberry Pi.
- openWakeWord: https://github.com/dscripka/openWakeWord
- WebRTC VAD: https://github.com/wiseman/py-webrtcvad
- Vosk: https://alphacephei.com/vosk/
- PipeWire: https://pipewire.org/