← all posts

The Rubber Duck Can Use Tools Now

Paloma has weaker model weights and a growing body of experience. Each interaction leaves the assistant better prepared for the next one.

I used to hear "Hey babe, how do I..." whenever my wife needed technical help. Sometimes she wanted an answer. Other times she needed a rubber duck while she worked an idea into something concrete.

Asking me also meant waiting for me. I may be deep inside a project and difficult to pull away, and she sometimes hesitates to interrupt. Her task would sit until I became available.

Over the last couple of weeks, that exchange has changed. Sometimes she asks Paloma first. Sometimes the same problem reaches both of us and I find myself racing it: can I finish before Paloma does?

A smaller model that gets better

Paloma is the local AI assistant I set up for her. Its model is much smaller than the largest cloud models, yet it handles a growing share of the questions that previously came to me. My wife uses it to work through ideas, and Paloma can continue into the technical work that follows. One day, she needed something done while we were away from home. She sent a quick message to Paloma, and the task was finished when she returned.

Tool access lets a rubber-duck conversation continue into action.

I can solve more of these problems than Paloma can, though my attention is often elsewhere. Paloma is available when the question occurs and can resume earlier work without waiting for me to restore the context. For everyday use, that availability matters as much as raw capability.

The assistant learns outside the model weights

Paloma has learned how my wife prefers to communicate. Phrases she dislikes have stopped appearing, frequently used resources are easier to find, and each conversation begins with more shared context.

The model weights remain fixed. Its harness accumulates experience, retrieves what applies to the current request, and exposes the relevant tools. The model starts from what earlier interactions established.

Tool use follows the same loop. Plain-language context describes the goal. The system explores the tool, observes the outcome, and retains a procedure that worked. Later requests begin from that evidence.

Paloma remains weaker at general reasoning while becoming steadily better at the work my wife asks it to do. That experience is specific to her work and remains available for inspection.

Inside the AI Harness covers those mechanics in more detail. Paloma shows their effect on a regular user.

Guardrails make tool use practical

An assistant that writes scripts can cause damage, especially for a nontechnical user who may lack the context to recognize a dangerous command. Paloma's harness limits what it can reach and stops when an action moves outside its authority. Escalation is part of the workflow. Those constraints let my wife use the capability without learning the operational risks behind every request.

The question became a race

Paloma's usefulness has created a small competition. My wife can give us the same problem, and I see whether I can solve it first.

I miss being the rubber duck sometimes. Those conversations gave me a window into her ideas, and I liked helping her move them forward. I also know how hard I can be to interrupt when a project has my full attention. Her work should be able to continue during those periods.

The race is fun. Her work can continue regardless of my availability, and Paloma keeps learning how to help her. That small change says more than a model benchmark.

The one-line version

A smaller model can become a better personal assistant when its harness preserves what each interaction teaches it.