Back to directory
Research & analysis · 2026-09-14

Local only for what is deterministic: embeddings, speech and TTS

After testing every model that fits a 128GB Strix Halo, this user keeps local strictly to embeddings, text-to-speech and speech-to-text, and relies on cloud for reasoning.

This person tested every local model they could fit on a 128GB Strix Halo and concluded that none competes with GLM 5.3 Flash, which costs roughly five cents per task. Their position is blunt: for work or serious consultation, local 27B to 70B models are hobby-grade, and pretending otherwise is hopeful. They keep local strictly for what they consider safe and deterministic.

The local components are F5 for text-to-speech, Parakeet for speech-to-text with a large custom correction dictionary, and small embedding models for a local RAG over Obsidian vaults. They trust embeddings locally because an embedding model is a single deterministic pass with no sampling, so the same text always yields the same vector and it is cheap to run. The RAG lookup itself is done with cloud models against those locally built embeddings. They do not expect hardware prices to fall for at least two years.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 1

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-14Local only for what is deterministic: embeddings, speech and TTS

After testing every model that fits a 128GB Strix Halo, this user keeps local strictly to embeddings, text-to-speech and speech-to-text, and relies on cloud for reasoning.

Research & analysis1 machineModels unspecifiedRuntime unspecified

This is the currently published snapshot.