Set up free local Whisper on a Mac with Python ALREADY INSTALLED. Illustrated saved-run explanation, not live app footage. Synthetic narration and original sample. Tested: macOS arm64, Python 3.12.14, a new venv and new model cache. Python installation, other platforms and hardware are not covered.
1. In Terminal, create a NEW folder and enter it:
mkdir whisper-demo
cd whisper-demo
2. Replace /path/to/python3 with your installed Python executable:
/path/to/python3 -m venv .venv
.venv/bin/python -m pip --isolated install --index-url https://pypi.org/simple --no-cache-dir --only-binary=:all: torch==2.14.0 transformers==5.17.0 numpy==2.5.3 scipy==1.18.1 soundfile==0.14.0
3. Copy only the code between the DOWNLOAD markers into a plain-text file named download_whisper.py in this folder. Preserve indentation.
BEGIN DOWNLOAD
"""Download one pinned public Whisper model into a new local cache."""
from pathlib import Path
import os
os.environ['HF_HUB_DISABLE_IMPLICIT_TOKEN']='1'
os.environ['HF_HUB_DISABLE_TELEMETRY']='1'
for key in ['HF_HUB_OFFLINE','TRANSFORMERS_OFFLINE','HF_DATASETS_OFFLINE']:
os.environ.pop(key,None)
cache=Path('whisper-cache').resolve()
os.environ['HF_HOME']=str(cache/'home')
os.environ['HF_XET_CACHE']=str(cache/'xet')
from huggingface_hub import snapshot_download
if cache.exists() and any(cache.iterdir()):
raise SystemExit('Choose a new empty whisper-cache folder; refusing existing cache.')
cache.mkdir(exist_ok=True)
folder=snapshot_download('openai/whisper-tiny.en',
revision='87c7102498dcde7456f24cfd30239ca606ed9063',
cache_dir=str(cache),token=False,
allow_patterns=['*.json','merges.txt','vocab.json','model.safetensors','normalizer.json'])
print('Use this folder as --model-dir:')
print(folder)
END DOWNLOAD
4. While online, run:
.venv/bin/python download_whisper.py
Copy the printed snapshot path. Setup downloads packages and model files. The helper requires a new empty whisper-cache folder; use a new working folder if needed.
5. Copy only the code between the TRANSCRIBE markers into local_transcribe.py in the same folder:
BEGIN TRANSCRIBE
import argparse,math,os,sys
from pathlib import Path
for k in ['HF_HUB_OFFLINE','TRANSFORMERS_OFFLINE','HF_HUB_DISABLE_TELEMETRY']:
os.environ[k]='1'
def offline(event,args):
if event in {'socket.connect','socket.getaddrinfo','socket.sendto'}:
raise RuntimeError('Network disabled')
sys.addaudithook(offline)
p=argparse.ArgumentParser()
p.add_argument('input',type=Path)
p.add_argument('output',type=Path)
p.add_argument('--model-dir',required=True)
a=p.parse_args()
if not a.input.is_file():p.error('Input file missing')
if a.output.suffix.lower()!='.txt':p.error('Choose a .txt output')
if a.output.exists():p.error('Output exists; choose a new filename')
if not Path(a.model_dir).is_dir():p.error('Cached model folder missing')
import numpy as np,soundfile as sf,torch
from scipy.signal import resample_poly
from transformers import WhisperFeatureExtractor,WhisperTokenizerFast,WhisperForConditionalGeneration
try:
x,sr=sf.read(a.input,dtype='float64');info=sf.info(a.input)
except (RuntimeError,ValueError) as e:p.error('Unreadable audio: '+str(e))
if info.format!='WAV' or x.ndim!=1:p.error('Use mono WAV')
if not len(x) or not np.all(np.isfinite(x)):p.error('Empty or invalid audio')
if np.max(np.abs(x))==0:p.error('Silent audio')
if len(x) not in range(1,30*sr+1):p.error('Maximum 30 seconds')
torch.set_num_threads(4)
f=WhisperFeatureExtractor.from_pretrained(a.model_dir,local_files_only=True)
t=WhisperTokenizerFast.from_pretrained(a.model_dir,local_files_only=True)
m=WhisperForConditionalGeneration.from_pretrained(a.model_dir,local_files_only=True,use_safetensors=True).eval()
g=math.gcd(sr,16000)
x=resample_poly(x,16000//g,sr//g).astype('float32')
i=f(x,sampling_rate=16000,return_tensors='pt',return_attention_mask=True)
with torch.inference_mode():
ids=m.generate(i.input_features,attention_mask=i.attention_mask,max_new_tokens=440,do_sample=False,num_beams=1)
s=t.batch_decode(ids,skip_special_tokens=True)[0].strip()
if not s:raise RuntimeError('Empty transcript')
with a.output.open('x',encoding='utf-8') as out:out.write(s+'\n')
print(s)
END TRANSCRIBE
6. Put your own non-silent English mono WAV, at most 30 seconds, here as input.wav. Replace the quoted path with the snapshot path printed in step 4:
.venv/bin/python local_transcribe.py input.wav my-transcript.txt --model-dir "/printed/snapshot/path"
Choose a new output filename; its parent folder must exist. Open my-transcript.txt and listen against your source.
Observed output: The workshop has 12 seats, bring a notebook.
The original says twelve and uses two sentences. Check meaning, names, dates and numbers; ASR can make mistakes. The input is unchanged. Inference uses local model files and a Python network guard, not a system-wide privacy guarantee. No speed or general accuracy claim.
Model: https://huggingface.co/openai/whisper-tiny.en