diff --git a/.github/workflows/pages.yml b/.github/workflows/pages.yml new file mode 100644 index 00000000..65dad439 --- /dev/null +++ b/.github/workflows/pages.yml @@ -0,0 +1,31 @@ +name: Deploy GitHub Pages + +on: + push: + branches: [main] + paths: [docs/**] + workflow_dispatch: + +permissions: + contents: read + pages: write + id-token: write + +concurrency: + group: pages + cancel-in-progress: false + +jobs: + deploy: + environment: + name: github-pages + url: ${{ steps.deployment.outputs.page_url }} + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: actions/configure-pages@v5 + - uses: actions/upload-pages-artifact@v3 + with: + path: docs + - id: deployment + uses: actions/deploy-pages@v4 diff --git a/Docs/index.html b/Docs/index.html new file mode 100644 index 00000000..72e6fd29 --- /dev/null +++ b/Docs/index.html @@ -0,0 +1,1103 @@ + + +
+ + ++ AI-powered voice input that lives in your menu bar. + Hold a key, speak naturally — your words appear instantly, polished and precise. +
+
+ Hold your configured key — Fn, Ctrl, Shift, or Option. Recording begins instantly. A subtle overlay confirms you're live.
+Talk as you would to a person. WhisperKit runs entirely on your device — no audio ever leaves your Mac.
+Release the key. Your speech is transcribed, optionally refined by a local LLM, and inserted directly into your active app.
+Screen OCR captures what's visible. The AI uses this to fix homophones and align output with your actual context.
+Pick the mode that matches your task — from raw dictation to AI-powered command execution.
+Raw transcription with the lowest possible latency. What you say is exactly what gets typed.
+ Lowest Latency +An LLM cleans up filler words, fixes grammar, and structures your speech into polished text.
+ AI-Refined +Speak a command. The AI reads your screen context via OCR and generates a contextually relevant response — summarize, translate, reply.
+ Screen-Aware + AI-Generated ++ "All three modes use your personal Edit Rules and Language Style Presets — concise, formal, casual, or fully custom prompts." +
+Built for power users who care about privacy, performance, and control over their tools.
+Apple Speech (built-in), WhisperKit (offline Whisper), or Doubao ASR (cloud). Pick what fits your workflow.
+WhisperKit + MLX Qwen runs entirely on-device. Audio and text never leave your machine.
+MLX-powered Qwen2.5 and Qwen3 models for on-device smart formatting and correction.
+OpenAI, Claude, Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax — your key, your model.
+Configure any modifier key with long-press, double-tap, or single-tap trigger modes.
+ScreenCaptureKit + Vision captures on-screen text to help the LLM correct homophones in context.
+Recent inputs are injected as LLM context, improving accuracy for continuous dictation sessions.
+Personal text replacement rules applied automatically to every output. Your voice, your vocabulary.
+Full history with raw vs. processed comparison, word counts, and configurable retention period.
+Download the latest release. Drag to Applications. Hold Fn. Start speaking.
+ +xattr -cr /Applications/OpenType.app
+
+ Run this once after install if macOS shows a security prompt
+