What it can do

Voice mode

Talk to it, and have it answer out loud.

Talk in the Chat sidebar is a spoken conversation, not dictation. It listens continuously, works out when you have finished, answers out loud, and lets you cut in.

It is beta, and that is honest. It wants a decent GPU and a headset. Over laptop speakers it will occasionally answer itself.

Draggy's voice settings: which model answers out loud, a choice between the system voice and the Natural one, which voice, and its speed.

The orb

What it doesWhat it means
Grows with your voiceListening
A dot with a rotating arcThinking, or searching
Pulsing steadilySpeaking
An outlineMicrophone muted

Keys

KeyAction
EscEnd the conversation
MMute the microphone
CShow the transcript
SpaceSkip the rest of the reply

Voices

System uses the voices your operating system has: instant, and they sound like your operating system. Natural is a neural voice, about 90 MB on first use, English only, and considerably better. Speed runs from 0.9x to 1.25x.

Everything spoken stays local. Recognition runs with Whisper on your machine, and audio is never uploaded or written to disk. Web searches are the only thing that leaves.

When it goes wrong

  • It interrupts itself. Its own voice is reaching the microphone. Use headphones.
  • It is slow. Watch the millisecond counter. Over 1500 ms means the talk model is too big.
  • "Basic detection". The neural voice detector did not load; background noise fools the simpler one.