Skip to content

Local speech-to-text option (Whisper/on-device) #7

Description

@BraedenBDev

Web Speech API sends audio to Google's servers for transcription. For privacy-sensitive workflows — especially when narrating proprietary UI or internal tools — we need a local alternative that keeps audio on-device.

Options investigated:

  • Whisper.cpp via WebAssembly — community WASM builds exist. Would run entirely in the extension's offscreen document or a worker. Latency is higher than streaming Web Speech API but acceptable for our use case (short narration clips, not real-time captions).
  • Chrome's on-device speech recognition — behind chrome://flags/#enable-on-device-speech-recognition. Not yet stable or available on all platforms, but worth tracking.

Implementation notes:

  • The transcription interface in the sidepanel already receives text chunks — swapping the backend should be transparent to the rest of the pipeline.
  • Model download size (~40MB for tiny.en) needs a first-run UX.
  • This would make PointDev fully offline-capable when combined with the existing local-only capture pipeline.

Related to NLnet milestone M4 (privacy and offline support).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1-criticalMust-have for NLnet submission or core functionalityenhancementNew feature or requestprivacyPrivacy and data handling

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions