An RTX 5090 runs recurring agents for project maintenance, site contributions and personalized news monitoring.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A legal user has Claude build task plans while a local Qwen 27B executes them, with anonymisation, OCR for difficult documents and a RAG pipeline using nomic embeddings.
Three GPUs spread across machines run Qwen 3.8 27B through a custom stateful agent layer to build a personal finance and portfolio tracker without shipping documents to the cloud.
An office replaced roughly $2,000 a month of API consumption with two RTX 5090 machines running Qwen 3.8 27B under vLLM, one for helpdesk and one for ops.
A developer with 20 years of experience runs three concurrent coding sessions at about 150 tok/s each on a single RTX 5090 and has been cloud-free for three months.
A Python collector pulls read-only configs, firewall rules, and network health data every day and feeds it into a local RAG pipeline, so the advisor always knows the current state of the homelab without ever being able to touch it.