Seisami Tasks (Progress Tracker)
Input Layer
✅ Detect when
fnkey is pressed (via C + cgo).✅ Start/stop audio recording when
fnpressed/released.Handle microphone permissions (macOS/Linux/Windows separately).
Add configurable hotkey (not just fn).
Audio Layer
✅ Record & save audio as
.wav.✅ Stream audio frames into encoder.
Improve latency (smaller buffer size, realtime flush).
Allow cancel/discard recording (not just stop).
Manage file cleanup (avoid piling up recordings).
Transcription & Tooluse
✅ Integrate OpenAI Whisper via
go-openai(full transcription).✅ Add duration guard (skip too-short recordings, currently less < 3s).
✅ Handle transcription errors gracefully and emit
transcription:errorevents.Design a simple command schema (e.g.
{"action": "move", "task": "Fix bug", "to": "tomorrow"}).Implement mapping: when a phrase matches a command → trigger the corresponding tool function.
Emit structured events to frontend (
task:create,task:move,task:update).Fallback mode: if command not understood, show raw transcript for manual handling.
Backend / Data Layer
✅ Create simple task model (
id,title,description,status,due_date).✅ Store tasks locally, SQLite
✅ Event emitters for
task:created,task:updated,task:moved.
Frontend Integration
Show recording waveform indicator.
✅ Kanban Board Setup
✅ Connect board to db
✅ Show live transcript bubble while recording.
✅ Update board UI when transcription is processed.
✅ Add fallback UI (manual add/edit task).
Later / Nice-to-Have
Allow sync/export (JSON → Trello, Jira, Notion API).
Multi-project support.
Voice commands for settings ("make this private", "show today’s tasks").
Offline transcription (Whisper.cpp integration).
Auto-summarize daily spoken log into "what I did today".
Integration (gmail, github, jira, trello & more). e.g emailing, being able to tell seisami to send emails for you.
Something I wanna do for seisami, making it agnostic as possible
i’ll be shipping it with portaudio
Audio transcriptions can be in 3 ways
You can use the my cloud for transcribing audios
I’ll give you the option to use whisper.cpp i.e you’ll just select the path to whisper binary & the model you’ve installed to use
Give you the option to input your own openai api key which will be stored locally & can be used for transcriptions