Seisami progress tracker

by Mr Lawrence

August 29, 2025

Seisami Tasks (Progress Tracker)

Input Layer

  1. ✅ Detect when fn key is pressed (via C + cgo).

  2. ✅ Start/stop audio recording when fn pressed/released.

  3. Handle microphone permissions (macOS/Linux/Windows separately).

  4. Add configurable hotkey (not just fn).

Audio Layer

  1. ✅ Record & save audio as .wav.

  2. ✅ Stream audio frames into encoder.

  3. Improve latency (smaller buffer size, realtime flush).

  4. Allow cancel/discard recording (not just stop).

  5. Manage file cleanup (avoid piling up recordings).

Transcription & Tooluse

  1. ✅ Integrate OpenAI Whisper via go-openai (full transcription).

  2. ✅ Add duration guard (skip too-short recordings, currently less < 3s).

  3. ✅ Handle transcription errors gracefully and emit transcription:error events.

  4. Design a simple command schema (e.g. {"action": "move", "task": "Fix bug", "to": "tomorrow"}).

  5. Implement mapping: when a phrase matches a command → trigger the corresponding tool function.

  6. Emit structured events to frontend (task:create, task:move, task:update).

  7. Fallback mode: if command not understood, show raw transcript for manual handling.

Backend / Data Layer

  1. ✅ Create simple task model (id, title, description, status, due_date).

  2. ✅ Store tasks locally, SQLite

  3. ✅ Event emitters for task:created, task:updated, task:moved.

Frontend Integration

  1. Show recording waveform indicator.

  2. ✅ Kanban Board Setup

  3. ✅ Connect board to db

  4. ✅ Show live transcript bubble while recording.

  5. ✅ Update board UI when transcription is processed.

  6. ✅ Add fallback UI (manual add/edit task).

Later / Nice-to-Have

  1. Allow sync/export (JSON → Trello, Jira, Notion API).

  2. Multi-project support.

  3. Voice commands for settings ("make this private", "show today’s tasks").

  4. Offline transcription (Whisper.cpp integration).

  5. Auto-summarize daily spoken log into "what I did today".

  6. Integration (gmail, github, jira, trello & more). e.g emailing, being able to tell seisami to send emails for you.

Something I wanna do for seisami, making it agnostic as possible

  1. i’ll be shipping it with portaudio

  2. Audio transcriptions can be in 3 ways

    • You can use the my cloud for transcribing audios

    • I’ll give you the option to use whisper.cpp i.e you’ll just select the path to whisper binary & the model you’ve installed to use

    • Give you the option to input your own openai api key which will be stored locally & can be used for transcriptions