Voice is not the ideal interface

It is completely linear, a single stream.

Voice is not the ideal interface:

  • It has the same fallbacks as text back and forth.
  • It is completely linear, a single stream. Repeatability is extremely high friction.

Although to steelman, "voice" is a two part claim. This is voice in as input which is low input, but there is voice out. Be careful not to conflate the two.

Voice input

  • Low friction, often faster for most people.
  • Very broad, can't easily input specific characters.
  • Some privacy concerns around being in public.

Voice output

  • Maybe works for some people, not so much for me.
  • "Hands free" where it can be done remotely.

Be careful not to conflate

  • People conflate the capabilities improvements of models being more "forgiving" with inputs and being able to infer more from prompts with "voice".
  • A lot of the benefits of voice is that people also just talk more in their prompts than they otherwise would if they had typed them.

The human eye: instantaneous and effortless movement, high bandwidth and capacity for parallel processing, intrinsic pattern recognition and correlation, a macro/micro duality that can skim a whole page or focus on the tiniest detail. Meanwhile, a graphic sidesteps human shortcomings: the one-dimensional, uncontrollable auditory system, the relatively sluggish motor system, the mind's limited capacity to comprehend hidden mechanisms.

(Bret Victor, Magic Ink)

An update

Maybe this should make me update that voice may actually be ok. It is extremely low friction to the user because they can just bumble about! BUT can the user bumble about with other input forms instead? I do this with text with longwinded prompts, random words, attaching #tags also helping with model performance. What would low friction natural gesture input look like? Is this what I imagine MCQs to be like? Or the Google guy with the knob to "tune" parameters. Can you just be more fuzzy with inputs as the models get better? Is there a way I can give people the ability to rapid fire input, without entering text? Simple yes no maybe next until they get what they want?

See Learning, creating, communicating, Induced fit and Will we gesture at all?.

Induced fit

Typing is really quite a drag. People like to click choose!

Learning, creating, communicating

Mechanical metaphors are extra noise.

Shared layers

AR glasses would mean you just render the image in shared space together.

The body as input

Hands, everyone has them. Can be very intuitive and natural.