Set up the Mac
Your Mac is the brain. It transcribes each session, works out who spoke, and builds a searchable index, all locally on Apple Silicon.
Clone to the expected path
The scripts use recordings/ and chroma/ at the repo root. Clone to ~/Workspace/transcriber to match the Pi's default MAC_RECORDINGS_DIR.
git clone git@github.com:ample/pocket-protector.git \ ~/Workspace/transcriber
Somewhere else? Set MAC_RECORDINGS_DIR in the Pi's .env to your clone's recordings/ path.
Tools and models
ffmpeg stitches the chunks together. Ollama runs the embedding model and the model that answers. Everything else comes from the Python requirements.
brew install ffmpeg ollama ollama pull nomic-embed-text ollama pull llama3.1:8b python3 -m venv env source env/bin/activate pip install -r mac/requirements.txt
Unlock the speaker model
pyannote is a gated model on Hugging Face. You download it once with your account, and after that it runs offline.
Accept the terms on speaker-diarization-3.1 and segmentation-3.0.
Create a read token at huggingface.co/settings/tokens.
Run huggingface-cli login (or export HF_TOKEN) before the first run.
Start the processor
A daemon that checks recordings/ every 5 seconds. Marker files are its state machine, so a crash just means it picks up where it left off.
source env/bin/activate python mac/processor.py
TRANSCRIBED, then INDEXED.Failures retry automatically. They're logged and tried again on the next poll. To reprocess a session, delete its marker files.
