Scarlett: a personal AI assistant that stays honest
Scarlett reads your inbox, calendar, and Slack, then runs your day like a quiet chief of staff: a morning brief, drafts in your voice, protected time, research you can act on. I led product design end to end, from the jobs-to-be-done scope to the trust architecture to the shipped surfaces.
Overview
Personal AI assistants are the fastest-moving product surface in software, and most of them ship theater: capability claims the backend can't honor, meters that measure nothing. Scarlett took the opposite bet. One rule governed every screen: no pixel promises what the backend can't do. Drafts wait for you to press send. Research compares flights and never books one. The calendar defends your evenings and asks first.
I was the principal product designer on this from the start, which on a small team means the scope, the trust architecture, and the screens were one job. This page walks through all three.
Scoped by jobs
The scope came from jobs to be done. I interviewed the way I always do, listening for the hire: what would you fire your current tools from, what do you keep dropping, who are you afraid of disappointing. Six jobs survived the cut.
- JOB 01Stay on top of what needs me
- JOB 02Know the people who matter
- JOB 03Answer in my voice
- JOB 04Keep my calendar sane
- JOB 05Brief me before I walk in
- JOB 06Do my research legwork
Every skill in the catalog had to name its job, and anything that couldn't got cut. Assistants die of feature sprawl. The nav shows the discipline: seven destinations, each one a job you can say out loud.
Two products in one
The framing that shaped everything else: an assistant is two products wearing one face. The split organized the build, and it organized the design.
The app
what you can point at- morning brief
- inbox triage
- drafts
- people
- calendar
- research
The context layer
what makes it worth using- memory
- voice model
- retrieval
- learned patterns
- provenance
- audit log
Every screen is a view over the memory. So every claim on every screen traces to a source you can name: the insight panel ties an email to tomorrow's meeting, a draft cites the thread and the last reply, a to-do quotes the promise you made. Nothing arrives from nowhere.
The honesty rule
Write access to someone's inbox is earned twice. Platforms gate it behind verification, and people gate it behind trust that arrives slowly and leaves fast. So I drew the line as a verb law and held every surface to it.
Scarlett's verbs
draft · remind · watch · propose · compare
Your verbs
send · book · confirm · approve · decide
The rule pays for itself twice over. Users trust an assistant that shows its limits, and the preview-then-send loop is also the learning loop: every edit you make before approving teaches the voice model what you actually sound like. A constraint and a data flywheel in the same gesture.
It also turned out to be the strongest art director in the room. Fake sends, decorative confidence meters, and capability theater were unavailable by rule, and the design got more distinctive because of it.
Trust surfaces are product surfaces
Four architecture decisions became flagship UX, and I designed them as product, never as a settings page.
The Context Inspector. Everything Scarlett has learned about you, every fact, preference, and pattern, is inspectable, editable, and deletable. If the promise is "it knows your life," the counter-promise has to be "and you can see exactly what it knows, and change it."
A visible voice model. Scarlett learns tone from your sent mail and then shows you what it learned, as settings you can adjust rather than a black box. Feedback loops back in. "Make it a bit more concise" changes the next draft, and you can see why.
Bounded actions with an audit trail. Anything Scarlett initiates passes a confirmation gate. Harmless operations run free; anything that touches your world asks first; every action lands in a log you can read and reverse. Actions in the UI read as proposals until you approve them.
Ingested content carries no authority. An email can contain text that tries to give your assistant orders. Scarlett treats everything it reads as material to reason about, and the design reinforces it: third-party content is never presented in Scarlett's own voice, and her voice is reserved for what she actually concluded.
Home: the day, quietly in order
Home answers the question you'd ask a chief of staff at 8am: what does today need from me? The greeting makes a promise and keeps it small: three things before noon, I've got the rest. Around the central orb, five bubble nodes carry the whole surface. Each one holds its count, one sentence of why it matters, and exactly one action. Two decisions are waiting, about four minutes, start with Ava. Taylor's been waiting on your feedback, reply to Taylor. Your Acme reply is ready, see the draft.
The columns bow around the circle because straight edges against a circle read as templates, and the canvas metaphor stays home's alone. My favorite detail is the smallest one: the status line under the composer. Watching your inbox and calendar, quiet until 9pm. What it's doing, and when it will leave you alone. It proposes. You decide.
Inbox: a hierarchy of attention
Triage sorts by what it costs you: what needs you now, what waits on you, what can disappear into a digest. The tier counts stay visible, because curation you can't inspect is hiding. Scarlett shows what it set aside and lets you disagree, and the connective work happens unprompted: a thread gets tied to the meeting it affects, a reply gets tied to the person waiting on it.
Drafts: your voice, with receipts
Every draft carries its provenance: the thread it answers, the reason it was suggested, when you last replied. Steering it takes plain words. You annotate the draft the way you'd brief a colleague, shorter, warmer, ask for the first slot, and the rewrite honors the note. Send sits at the end of that chain, after the context, after your edits. Drafts stay private to you until you act.
The draft body is deliberately plain. People fear pasting an assistant's flourishes into a corporate thread, so Scarlett's own serif voice stays in her margins and the payload looks like you wrote it. Because by the time you press send, you mostly did.
Research: compare, never book
Research is comparison with a spine. You set the budget and the preferences; Scarlett scores each option against criteria you can see and change, and the recommendation explains itself in reasons you can check: within budget, fastest total time, the airline you already fly. Then the law holds. Scarlett compares. Booking stays yours.
How it was designed
The build ran through the Figma MCP, agent-driven with human art direction. I directed, three model families argued, and the founder corrected live. Gemini reviewed interaction details against screenshots; its serif-payload catch shipped. GPT concept boards were raw material to edit against. Claude ran the build. Cross-family review catches what self-review misses, and I run the same stack on code.
The design file underneath the shipped surfaces runs deep: full state machines, superseded directions kept as an archive of changed minds, a persona reveal that took four rounds to earn. A few studies from it:
What the founder's red pen taught the system
The highest-value inputs were corrections, and each one became a law instead of a patch.
- "What are you doing? We already have a live file." Read the file before building. The working canvas is ground truth, and the person you're designing with edits mid-session.
- "People won't be able to read them." Letterspaced caps got a floor: 13px minimum, tracked, secondary ink. Sixty-one labels changed in one pass because it became a rule, never a fix.
- "This looks like AI slop." Floating white cards with no reason to exist got stripped. The hierarchy that replaced them is better: a box now means editable or transient, and everything else earns structure from typography.
- Metaphors need jurisdictions. One clever idea stretched over every surface is how "innovative" AI products die. Each surface got the metaphor its job could carry, and no more.
- Placeholder everything. Real colleagues and real deadlines were swapped for a coherent demonstration cast without breaking one narrative thread across screens.
Reference intake, the same discipline. Concept boards and a founder-admired email product fed the work through one filter: list what the reference does, keep what the backend supports and the brand survives. Thread context and privacy lines came in. Voice-quality meters and magazine-image inboxes stayed out, because meters that measure nothing are theater and there was no image data to show.
Verdict
"Capability-honest" sounds like a limitation and behaves like a style. What distinguishes Scarlett isn't one screen; it's that every surface obeys the same small constitution: every claim traceable, one accent spent on urgency, boxes only when they mean something, and no pixel promising what the backend can't deliver. The most copied AI patterns of this cycle were unavailable here by rule, and the product is more distinctive for it.
What I'm watching after launch: whether protected time survives contact with real calendars, and how fast the voice profile converges once feedback starts flowing. Both are measurable, and both will tell me if the trust design actually earned trust.
The verb law and provenance rules here are the applied form of Designing AI Behavior. The same stance runs through Articos, where the product refuses to ship a report it cannot defend. If you're building an assistant and arguing about how much it should do without asking, I've had that argument for real. Email me.
// Screens are from the launch build. Names, companies, and threads on them are a demonstration cast, not real user data.