Back to directory
Personal assistant · 2026-09-28

My personal JARVIS

Personal life/project management agent with persistent memory, reminders, smart home duties (timers lights etc), computer use.

Personal assistant originally made for managing a project, started as a persistent memory experiment, expanded into a full on home voice assistant.

Smart device control (currently limited to Kasa but if I ever care to mess with that again it would not be hard to expand it to others. Auto discovered by harness upon setup.

Speakers are interesting. Each speaker satellite is a raspberry pi zero 2 w with a speaker and microphone in a 3D printed box. No standard speaker/mic setup, each one is different. When they turn on, they automatically connect to the hub and open an audio stream over TCP. The hub streams audio back over that same connection. Open wake word is run on each stream, and for each one that gets a hit (gated on oww threshold+consecutive wake audio frames), we arbitrate. It's not perfect. Certain spots in the house will consistently trigger the wrong speaker, and then transcription is garbled and I hear Paul Bettany in the other room either controlling the lights, because it guessed what I wanted, or expressing confusion.

Model maintained memory, one of two modes, second mode I'm still working out. First mode, agent updates memory deliberately. Works well enough, atomic facts, automatically dated, model puts them in groups. Second mode has a separate inference pass with special memory management prompt per interaction, and only memory management tools, usually run later if on the same model, or in parallel running on a second, smaller model. Needs more work, better prompting, better algorithm. Memory bloats fast when doing it that way.

Recall is pretty simple. Memories are embedded by a small embedding model, user message has an embedding made from it and all are compared to the user message and injected based on vector similarity. It's loose, more gets through than should, but enough that is relevant makes it in without issue that the model either just knows what I'm on about, or knows where to look to get more info for memories that didn't make it but should have.

Intended to be run in smaller chats. Smart speaker chats reset after 1 minute of inactivity. Chat length limited by token count. Intended use is smaller chats with model constantly updating memory so nothing is ever actually lost. Tools are extremely limited. Instead has a series of python modules, and bash (or just python if bash is spooky). Has access to multiple machines on my tailscale. Favorite party trick is have it turn on my PC and get Overwatch and Discord going for me. Of course when my PC is off it falls back to the gemini API so that first bit is non-local. Models started when PC is started.

Model is made aware of known limitations in it's harness, and is able to appropriately reason about them, fill in the blanks where the harness has failed. Using Bash and/or python allows it to perform several tool calls in one tool turn, and is especially effective when subsequent tool calls would depend on the results of prior ones. For example, if I ask it to turn on the lights in the living room, with standard tools it would first list the devices in the living room, then turn on each light that states it's location is in the living room. With bash/python, the list devices function returns an array of devices and their local IPs, so it can take that result and for each light with 'living room' in it's name, turn it on. The agent runs on code snippets.

Notable lack of security. Anyone with access to the speakers can basically ask the model to do anything on my PC, or on it's own box, or on any of the pis running the speakers. Testing speech vector security gating. I can semi-reliably get it to recognize when it's me talking. Needs more work. Recently moved from hosting the actual harness on a raspberry pi 5 to an intel n150 powered mini PC, so I have a lot more compute available to me now. Exploring more complicated options. Uncertain if speech recognition will be enough however; exploring other options. Model does have strong egress protection, and any attempt to reach out into the greater web (curl and the like) marks the chat as tainted, and the model cannot continue executing tools. The intention is to send out agents to do that, and report back. The fear that this mechanism guards against is web search and website text containing prompt injection attacks. Even if the model actively acknowledges it in it's reasoning, it is now in the context and, more importantly, in the attention mechanism. When the model reports back to the main instance, it is effectively filtered and far less likely to launder instructions from the webpage.

One major observation I've made since I've started this project: thinking/CoT length are impacted by the harness in a way that is very, very interesting. Every other instance I have with this model, outside of this harness. Different harnesses, straight llama-server or llama-cli, whatever, it thinks probably a bit too much. More than it needs to. I've learned that what chain of thought is is a form of augmenting the prompt, to create information where there isn't information. Fill in gaps. My harness has a length system prompt. The model is told in great detail what it is, what it's doing, who I am, what it needs to know about this interaction, what it's capabilities are, what they aren't, and how to do what I need it to do. For most tasks, chain of thought is sub 200 tokens before a tool call starts being made, often sub 100. Even when it doesn't know exactly what I'm talking about, it knows enough about me that it knows what kind of information to search for or the kind of task it needs to complete even with ambiguity. Prefill on a cold prompt just absolutely kills TTFT, but when hot, it will often start doing exactly the same task in a similarly capable but less wordy/complete harness later than in my personal harness.

Everything in my harness is wildly incompatible with any other standard harness or system. No MCP connectivity or anything.

PROVENANCE

Submitted through the vram.wiki contributor dashboard · reviewed before publication.

Published by @Mrinohk

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-28My personal JARVIS

Personal life/project management agent with persistent memory, reminders, smart home duties (timers lights etc), computer use.

Personal assistant1 machineQwen 3.6 35B-A3Bllama.cpp

This is the currently published snapshot.