Skip to content

Repository files navigation

PawnAI Recorder

Audio recording toolkit with a Python CLI (desktop/Linux) and a Jetpack Compose Android client. Both record in timed chunks, upload to S3-compatible storage, and can publish PawnQueue jobs for transcription/diarization.

Features

Python CLI

  • Real-time audio level monitoring with visual dB meter
  • Support for multiple audio devices and drivers (PulseAudio, ALSA, JACK, USB)
  • Configurable audio gain/amplification
  • Automatic audio segmentation into chunks
  • Input gain control
  • Interactive device selection
  • Local JSONL recording log – per-session and per-chunk metadata (device, duration, S3 status) persisted alongside audio files

Android client

  • Foreground-service recording with VU meter
  • Timed chunking (default 120 s) plus force-upload of the current chunk
  • FLAC (default, libFLAC via NDK) or WAV encoding
  • In-app settings for recording, S3, and PawnQueue (same contract as .pawnai-recorder.yml)
  • Kotlin PawnQueue publishing for transcribe-diarize / analyze / sync-siyuan

See android/README.md for build and NDK details.

Installation

From source

pip install -e .

After installation, you can run the pawnai-recorder command from anywhere:

pawnai-recorder --help

Development installation

For development with optional tools (testing, linting, type checking):

pip install -e ".[dev]"

Usage

List available audio devices

pawnai-recorder list-devices

Filter by driver type:

pawnai-recorder list-devices --driver pulse

Show debug output from audio libraries (ALSA, JACK warnings):

pawnai-recorder list-devices --verbose

The status command prints system and device information. When an S3 configuration is present it also verifies that the configured bucket is reachable and reports the result.

pawnai-recorder status

Show debug output from audio libraries (ALSA, JACK warnings):

pawnai-recorder status --verbose

Filter by driver type:

pawnai-recorder list-devices --driver pulse

Record audio

Start recording with interactive device selection:

pawnai-recorder record

Show debug output (useful for troubleshooting audio issues):

pawnai-recorder record --verbose

Record for a specific duration (in seconds):

pawnai-recorder record --duration 60

Specify device ID directly:

pawnai-recorder record --device-id 0

Apply input gain:

pawnai-recorder record --gain 2.0  # 2x amplification (+6dB)

Specify custom output directory:

pawnai-recorder record --output ./my-recordings/

Timestamp / filename formatting

Every session and chunk filename is derived from a configurable timestamp template. Two options control the output:

Option Default Description
--timestamp-format {ts} Format string; supports {ts} and {device_id} placeholders
--datetime-format %y%m%d%H%M%S strftime pattern applied to {ts}

Placeholders

Placeholder Example value Notes
{ts} 231015143022 Shaped by --datetime-format
{device_id} 3 Numeric device ID; default when not set

Examples

Embed the device ID in every filename:

pawnai-recorder record --timestamp-format '{ts}_dev{device_id}'
# → audio/231015143022_dev3_01.flac

Human-readable date with device tag:

pawnai-recorder record \
  --datetime-format '%Y-%m-%dT%H%M%S' \
  --timestamp-format '{ts}_dev{device_id}'
# → audio/2023-10-15T143022_dev3_01.flac

Date-only prefix (group by day):

pawnai-recorder record \
  --datetime-format '%Y%m%d' \
  --timestamp-format '{ts}_{device_id}'
# → audio/20231015_3_01.flac

Both options can also be set permanently in .pawnai-recorder.yml (see Configuration below).

Combined example

pawnai-recorder record --device-id 0 --duration 30 --output ./recordings/ --gain 1.5

Per-chunk diarization

queue.transcribe_diarize.mode in .pawnai-recorder.yml is end_of_session by default: one transcribe-diarize message is published when the take stops, with every chunk path. per_chunk publishes that message immediately after each chunk upload, which is the right choice for short chunks. The CLI flag overrides the file:

pawnai-recorder record --device-id 0 --chunk-size 30 --diarize-mode per_chunk

While a take is open, force the current buffer out (upload, and diarize when the mode is per_chunk) from the session window (u) or the status-bar menu Force upload chunk. That matches the Android client's force-upload action.

Session window, notes, and the status bar

On a terminal, record opens a window (install the extra: pip install -e ".[ui]"). It shows the level, the chunk list, and a note field you can type in while recording. Enter attaches the note to the open take. The note is written to the JSONL log and, on the next transcribe-diarize publish, sent as an annotations entry. Pawn does not render those entries yet; see docs/PAWNAI_ANNOTATIONS_CONTRACT.md.

Keys: s start/stop, u force-upload, c screenshot, q quit. Stopping returns to idle so you can start another take. Quit, and --duration, exit the process.

--plain keeps the Rich level meter for scripts and non-interactive runs.

pawnai-recorder record --device-id 0 --plain --no-tray

A Linux status-bar icon is shown for the same process (a StatusNotifierItem). It uses the desktop icon theme: play while idle, pause while recording. Left click toggles recording. The right-click menu has Start, Stop, Force upload chunk, Take screenshot, and Quit. KDE, GNOME (with AppIndicator support), and waybar show it. --no-tray skips the icon. There is no Windows or macOS tray.

Screen capture

Capture stays off unless you name a screen. --screenshot-output is a Wayland output name (DP-1) or a monitor index starting at 0. --screenshot-every 30 repeats; omit it and capture only when you ask (window key c or the tray item).

pawnai-recorder record --device-id 0 --screenshot-output DP-1 --screenshot-every 60

Backend choice:

  • Wayland with grim on PATH (preferred; a numeric index is mapped with swaymsg when that exists)
  • Wayland without grim: xdg-desktop-portal Screenshot (the portal does not pick a monitor; the first call may ask permission)
  • X11: the mss package, by monitor index

PNGs are saved next to the audio, uploaded with the same S3 layout, and listed on the transcribe-diarize payload as screenshots (region is reserved and currently null). A crop/area comes later.

S3-compatible upload organization

The status command now checks for S3 availability when configuration is provided and will surface an error or warning if the bucket cannot be contacted.

Saved files are uploaded automatically when s3 is configured in .pawnai-recorder.yml. Use --no-upload to bypass upload for a specific recording run.

Default key layout:

  • Without conversation ID: timestamp/filename
  • With conversation ID: conversation_id/timestamp/filename

Example with conversation ID:

pawnai-recorder record --device-id 0 --duration 60 --conversation-id conv-123

Bypass upload for one run:

pawnai-recorder record --duration 60 --no-upload

Copy the example config and edit credentials:

cp .pawnai-recorder.yml.example .pawnai-recorder.yml

Then configure .pawnai-recorder.yml in project root:

recording:
  rate: 16000
  chunk: 16000
  channel: 1
  chunk_size: 120
  file_extension: "flac"
  output_dir: "audio/"
  sample_width: 2
  # Timestamp / filename formatting
  # Placeholders: {ts} (datetime string), {device_id} (numeric device ID)
  timestamp_format: "{ts}"            # e.g. "{ts}_dev{device_id}" to embed device
  datetime_format: "%y%m%d%H%M%S"    # strftime pattern for {ts}

s3:
  bucket: "your-bucket-name"
  endpoint_url: "https://your-s3-compatible-endpoint"
  access_key: "your-access-key"
  secret_key: "your-secret-key"
  region: "us-east-1"      # optional
  prefix: "conversations"  # optional
  verify_ssl: true          # optional
  path_style: true          # optional

log:
  file: "recordings.jsonl"  # filename relative to output_dir (optional)

If upload fails, recording continues and local files are kept.

Local recording log

Every record run appends structured entries to a JSON Lines file (recordings.jsonl) inside the output directory. Record types written per session:

type event When written Key fields
session start Immediately after the stream opens session_id, conversation_id, device_id, device_name, sample_rate, channels, format, started_at
chunk — After each chunk file is saved chunk_index, file_path, duration_sec, s3_object_key, s3_uploaded, started_at
note — When you submit a note during the take id, at, text
screenshot — When a screen capture is saved id, at, file_path, output, s3_uri
session end After all chunks have been saved total_duration_sec, chunk_count, ended_at

Example entries

{"type":"session","event":"start","session_id":"260223143022","conversation_id":"mtg-01","device_id":3,"device_name":"USB PnP Audio Device","sample_rate":16000,"channels":1,"format":"flac","started_at":"2026-02-23T14:30:22"}
{"type":"chunk","session_id":"260223143022","chunk_index":1,"file_path":"audio/260223143022_01.flac","started_at":"2026-02-23T14:30:22","duration_sec":600.0,"s3_object_key":"conversations/mtg-01/260223143022/260223143022_01.flac","s3_uploaded":true}
{"type":"session","event":"end","session_id":"260223143022","ended_at":"2026-02-23T14:40:22","total_duration_sec":600.0,"chunk_count":1}

The log file is created automatically; no configuration is required.

Override the log filename for a single run:

pawnai-recorder record --log-file custom_log.jsonl

Inspect the log with standard tools:

# Pretty-print all entries
python -c "import json,sys; [print(json.dumps(json.loads(l), indent=2)) for l in open('audio/recordings.jsonl')]"

# List only session-start records (one per run)
grep '"event":"start"' audio/recordings.jsonl | python -m json.tool

# Show all chunks that failed S3 upload
grep '"type":"chunk"' audio/recordings.jsonl | python -c "
import json,sys
for line in sys.stdin:
    r = json.loads(line)
    if not r.get('s3_uploaded'):
        print(r['file_path'], r['started_at'])
"

Configure the log filename permanently in .pawnai-recorder.yml (see Configuration).

You can also run the application as a Python module:

python -m pawnai_recorder record

Android client

The Android app under android/ is a Jetpack Compose client that mirrors the CLI pipeline: record → chunk → S3 upload → optional PawnQueue jobs.

Build

cd android
./gradlew :template:assembleDebug
./gradlew :template:test

Requires Android SDK (local.properties → sdk.dir=...). FLAC encoding needs the NDK and CMake; see android/README.md.

Configure

Open Settings in the app and set recording, S3, and queue fields to match .pawnai-recorder.yml (bucket, endpoint, credentials, topic, job options). Choose file format flac or wav in the same screen.

Full feature list and native FLAC notes: android/README.md.

Project Structure

pawnai-recorder/
├── pawnai_recorder/
│   ├── __init__.py
│   ├── __main__.py
│   ├── cli/
│   │   ├── commands.py
│   │   ├── session_ui.py     # Textual session window
│   │   └── utils.py
│   ├── core/
│   │   ├── config.py
│   │   ├── jobs.py           # transcribe-diarize payload and attachment deltas
│   │   ├── log.py            # JSONL recording log
│   │   ├── recording.py
│   │   ├── session.py        # start/stop, notes, screenshots
│   │   ├── s3_upload.py
│   │   ├── storage.py
│   │   └── processing.py
│   ├── desktop/
│   │   ├── capture.py        # grim / portal / mss
│   │   └── tray.py           # status-bar icon
│   ├── server/               # loopback API for Obsidian (`serve`)
│   └── utils/
├── obsidian-plugin/
│   └── pawn-recorder/        # desktop Obsidian client for `serve`
├── packaging/
│   └── pawnai-recorder.service
├── android/                  # Jetpack Compose recording client
│   ├── README.md
│   └── template/             # App module (UI, service, FLAC JNI)
├── audio/                    # Default output directory for recordings
│   └── recordings.jsonl      # Auto-created recording log
├── docs/
├── tests/
├── .pawnai-recorder.yml.example
├── pyproject.toml
├── setup.py
├── requirements.txt
└── README.md                 # This file

Local service

pawnai-recorder serve keeps one recording session open and exposes it on 127.0.0.1 (default port 8765) for the Obsidian plugin. The record command is unchanged. Audio files stay in the configured output_dir, outside the vault.

pawnai-recorder serve --config ~/.config/pawnai-recorder/config.yml --bind 127.0.0.1:8765

A relative output_dir is resolved next to that config file. Optional api.token in the YAML is checked as Authorization: Bearer. GET /health does not require the token. The other routes are GET /status, GET /devices, GET /sinks, GET /sessions, GET /sessions/{id}, and POST /sessions/start, /stop, /flush, /notes, /screenshot.

This has to be a systemd user service. A system unit runs as root and cannot see your Pulse/PipeWire devices or the desktop session. Install packaging/pawnai-recorder.service:

mkdir -p ~/.config/pawnai-recorder ~/.config/systemd/user
cp packaging/pawnai-recorder.service ~/.config/systemd/user/
# Copy .pawnai-recorder.yml.example and point output_dir at a directory outside the vault.
systemctl --user daemon-reload
systemctl --user enable --now pawnai-recorder.service

If pawnai-recorder is not on the unit's PATH (a virtualenv, for example), set ExecStart to the absolute path of the executable.

Obsidian plugin

obsidian-plugin/pawn-recorder is a desktop-only plugin (id: pawn-recorder) that talks to the service. It does not replace the Pawn chat plugin. make dist writes obsidian-plugin/pawn-recorder/pawn-recorder.zip. Unzip that into <vault>/.obsidian/plugins/ so Obsidian sees the pawn-recorder folder.

The plugin writes recorder-session: <session id> on the active note when you start a session for that note, or link the note to the current one. Chunk files and upload state stay on the service. Timestamped annotations ("add selection as session note") are the existing recorder notes, separate from that link.

A Sessions tab inside the Pawn sidebar can call this same localhost API later. Recording stays in this process, not in pawn-server.

Configuration

All project metadata and dependencies are configured in pyproject.toml using the modern Python packaging standard (PEP 621). The setup.py file is minimal and only needed for backward compatibility.

Development tools configuration

The pyproject.toml includes configurations for:

  • black: Code formatting
  • ruff: Fast Python linter
  • mypy: Static type checking
  • pytest: Testing framework with coverage reporting

Requirements

  • Python 3.8+
  • PyAudio 0.2.14+
  • Loguru 0.7.2+
  • Typer 0.9+
  • NumPy 1.21+

Optional session window, status-bar icon, and X11 screenshots:

pip install -e ".[ui]"

That extra installs Textual, pystray, Pillow, mss, and jeepney. Wayland capture prefers the grim binary. The portal fallback needs a session bus.

Development Dependencies (optional)

  • pytest 7.0+
  • black 23.0+
  • ruff 0.1.0+
  • mypy 1.0+

License

MIT

About

A recorder for pawnai tool

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages