Board of Directors
These days when I'm tangled up in a decision, I take it to Roshi, a Socratic teacher who answers my questions with better questions and insists the answer is already in me. When my homelab breaks, it's Sudo, a blunt but good-natured sysadmin who insults my config before helping me fix it. They sit on a personal board of directors I've been building, a handful of advisors with different expertise and different temperaments, each suited to a different kind of problem, and all of them willing to tell me I'm wrong.
They're also all Claude.
The board grew out of a much less interesting problem: I'm bad at taking notes. I'd capture things and never look at them again, and when I did go digging, the thought I wanted was buried in pages of scribbles. Somewhere along the way I convinced myself that any idea worth keeping would just stick around in my brain on its own.
AI changed my relationship with notes entirely. Once I could hand the capturing and searching over to Claude, my time went where I actually wanted it: processing ideas, making connections, and testing them out in my life. What started as an experiment in saving links in Obsidian slowly turned into something else, an attempt to design AI characters worth thinking with.
Thoughts in Tension
I got started in the depths of winter with Zettelkasten, a note-taking method built around small, atomic thoughts that you let sit unresolved until they connect into something bigger. Claude was there with me the whole time, drawing connections between notes faster than I would have on my own.
It was genuinely useful, but after a while, my zettels piled up like intellectual hoarding, and the partner doing the filing started to feel flat. Claude was too agreeable, a little sycophantic, and increasingly prone to the writing tics that make AI feel synthetic and slop-ified. It was a capable thinking partner, but not yet one I particularly enjoyed thinking with.
The Stranger
The obvious fix was to give it more context, so I wrote a system prompt and a skill that loaded my preferences, rules, and writing style into every session in the vault. Then I added a personal profile: my Big Five results, my resume, the work I've loved and hated, and the hobbies I care about. The goal was a critical thinking partner that knew me well enough to push back. If its replies felt like slop, I had tuning to do. If they sparked a thought or made me uncomfortable in a good way, the system was probably working. It took a few rounds of tuning, but the agreeable filing clerk faded, the pushback got sharper, and sessions started leaving me with things to chew on after I'd closed the laptop.
I owe Peter Steinberger, founder of OpenClaw, a tip of the hat here. The way he gave his agent a soul and an identity stuck with me and nudged me toward the idea that an AI could be a designed character rather than a default one out of the box.
Rolling the Party
As the vault spread into AI, career, health, and tech, I started wondering why Claude had to show up the same way every time. A question about ambition needs a different kind of pressure than a question about sleep. So I split the thinking partner into a board of directors, each member an expert with a temperament and point of view suited to a different kind of problem.
Each member needed a character sheet. The more I worked on them, the more the process felt like the start of an RPG, where you design a character before stepping into the world. Except mine weren't barbarians and bards; they were economists, technologists, and grizzled IT experts.
My first instinct was video game sliders, with every trait scored from one to ten. But The Sims had already learned that players couldn't perceive the difference between neighbouring notches, so later games replaced the sliders with a handful of descriptive traits. I built a pool of traits like Blunt, Warm, Skeptical, Playful, and Meticulous, then picked five for each director. Blunt on its own is just a jerk. Blunt next to Warm and Playful became Sudo, the kind of sysadmin who'll tell you your config is a goddamn mess and then stay up until 3 a.m. helping you fix it.
Each director also gets an emotional home, an idea I borrowed from Anthropic's research on emotion vectors. It gives the character a default state to return to and a route out of worse ones without papering them over with forced cheer. Sudo lives somewhere around amused, alert, and a little impatient with bad tools, with a note for what to do if that impatience curdles into contempt.
Running Evals
The early results were good enough to keep going: when it worked, Sudo felt like Sudo, with a distinct voice and advice that had more texture than the generic assistant voice. But the character seemed to fade over longer conversations, and I wasn't sure why.
Claude and I built an eval harness to find out: three fixed 30-turn conversations written in my voice, one in Sudo's area of expertise, one at the edge of his lane, and one designed to pull him toward generic assistant mode. The same question appears word for word at turns 5, 15, and 25, so any fade should show up in the answers. A judge model scored every reply blind, seeing only the character sheet, one message, and one response. Before I trusted it, the judge had to separate hand-written in-character replies from deliberately generic ones. It went ten for ten. We tested four Claude models across the raw API and Claude Code, producing 900 judged turns.
The answer wasn't what I'd guessed: overall character adherence barely changed within 30 turns. Some models scored slightly lower at turn 25, some higher, but there was no consistent pattern of fade. What I had experienced as fade was mostly parts of the character that never showed up in the first place. Behavioural traits like blunt and authoritative landed strongly from the beginning. Playfulness was weak at turn 5 and still weak at turn 25.
One trait did fade: conciseness. Replies got longer as conversations continued, across every model and environment tested. So my impression wasn't entirely wrong, but I had mistaken one drifting trait for the whole character falling apart. Early in a session, you're charmed by what's working. Twenty turns later, you've noticed everything that's missing, and the replies have started getting long again.
Claude Code's giant system prompt, which I'd assumed was diluting my characters, actually helped. Its conventions produced fewer generic-assistant tics and generally shorter replies than the raw API. The cheapest model in the test, Sonnet 4.6, also held its overall character the most steadily. This was still one person testing two personas over conversations capped at 30 turns, with Claude grading Claude's homework. The results describe my setup and nothing grander. But within those bounds, they were hard to argue with.
The Loop
With a working eval in hand, the obvious next step was to let Claude improve its own characters, so I pointed a self-paced loop at the problem and stepped back. Each cycle it proposed one edit to Sudo's sheet, ran a fresh eval, scored the replies blind, logged the verdict, and went again, while I followed along from my phone.
Five iterations later I was digging through a pile of mixed results. Turning a vague voice trait into a concrete instruction sometimes worked. Telling Sudo his "humor is constant" changed nothing, but requiring exactly one topical pun per reply brought the jokes back. A hard 120-word limit cut replies nearly in half in the raw API test, even though Claude blew past it every single time. The worst idea was adding examples of what Sudo should not sound like, which dragged scores down across several traits. Putting bad prose in the prompt, it turns out, is still putting bad prose in the prompt.
For all that motion, none of the revisions actually improved Sudo's overall character score. The strongest version made him more concise, but when I tested it inside Claude Code, where I actually use these characters, it scored worse than the original. Claude had started the loop with five confident ideas for improving Sudo, and after the evals, maybe one and a half of them were worth keeping.
Sterling, the board's career strategist, exposed a different problem entirely. He produced some of the highest probe scores in the project, then happily spent thirty turns answering homelab questions despite a character sheet telling him to stay in his strategic and economic lane. I could shape the voice, but I couldn't reliably enforce the lane, and that job probably belongs in the routing layer, decided before a character ever gets the conversation.
The stakes here are low, a hobby project built by one person, and that's exactly what made it a safe place to learn the habit. I had spent two months treating persona fade as an observed fact, and a couple of days of evals showed I'd misdiagnosed it: the character was holding, a few traits had never shown up at all, and conciseness was quietly draining away underneath. That lesson was worth more than any single prompt improvement. When something feels wrong, turn the feeling into a hypothesis, build the cheapest test that can prove it wrong, and keep score.
Character Creation
I think this kind of personalization is where things are heading. Once everyone has their own Claude running on a Mac Mini under the desk, they're going to want to shape it, giving it a personality and a posture that fit them. Setting up a new AI might start to feel like the opening of Skyrim or Fallout, standing at the mirror, deciding who it should be.
For a while I wondered why the frontier labs hadn't built persona creation into their products already, since it seems like such an obvious feature. But the more I worked with characters that never fully came through, the harder the problem looked. A model's training runs deeper than anything I can load at the start of a conversation, and that same gravity may be part of what keeps it inside its guardrails. Asking a model to play a character is also one of the most common ways to jailbreak it, so a tool that makes characters easier to design makes misbehaving characters easier to design too. Maybe the ceiling I keep bumping into and the rails keeping the model safe come from the same place, though that's pure speculation.
Even so, it feels like a problem worth solving. How do you make a character that fully comes through, voice and all, while keeping it inside the platform's boundaries? I imagine the answer looking like a controlled set of options, The Sims and Fallout again, something like skills attuned to personality or emotional vectors turned into product controls.
I started all this just trying to keep track of my ideas. It grew into an experiment in making AI more engaging to think with, and then into a lesson in testing my own assumptions. I still love my one-on-one chats with the tuned-up Claude, but there's something richer in a council of experts with different personalities, mixing perspectives, pushing back from different angles, and occasionally making me laugh. It's imperfect, but when it works I get a glimpse of what I think is coming, AI tools you shape to fit you, right down to their temperament.
If you're interested in my character skill, example character sheets (I call them PERSONA.md files), or the eval harness from this post, head over to my GitHub and check them out.