Back to directory
Business automation · 2026-09-14

Running sales work on GLM 5.3 Flash across two DGX Sparks, 200M tokens a week

A sales professional dropped every frontier subscription and runs work agents on two DGX Sparks, averaging about 200 million tokens per week.

This person uses GLM 5.3 Flash with Hermes for work in sales, running on two DGX Sparks, and reports decent token speed. They average about 200 million tokens per week just on work tasks and have cancelled all frontier subscriptions, saying it has worked flawlessly.

Their opinion is that if you want the best results after walking away from closed frontier models, you need at least 256GB of VRAM, which is roughly where they feel quality starts to improve.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 1

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-14Running sales work on GLM 5.3 Flash across two DGX Sparks, 200M tokens a week

A sales professional dropped every frontier subscription and runs work agents on two DGX Sparks, averaging about 200 million tokens per week.

Business automation2 machinesGLM 5.3 FlashRuntime unspecified

This is the currently published snapshot.