Jobs-to-be-Done for LLM product features
A capability is not a job. Write the job as one solution-free line, break it into steps, and decide which single step a model may touch. That call survives the model being swapped.
A capability is not a job. Most LLM features ship because the model can do the thing, and then they underperform, because “summarize this” names the tool and says nothing about the work a person was trying to finish. Write the job as one solution-free line, break it into steps, and decide which single step a model may touch. That call survives the model underneath it being swapped.
01Capability-led roadmaps were leaking before the models arrived
Pendo looked at feature usage across 615 of its subscriptions and reported in February 2019 that 80% of features in the average software product are rarely or never used. Their estimate of the cloud R&D behind those features: $29.5 billion. No language model anywhere in that data. Feature-led roadmaps were already burning most of their money before anyone shipped a chat box, and the models did not cause that. They made it cheap to build the next thing nobody hired.
You have probably seen the newer number. MIT’s Project NANDA put out “The GenAI Divide” in July 2025, and the press squeezed one line of it into a headline: 95% of enterprise GenAI pilots return nothing. Handle that one carefully. Kevin Werbach at Wharton read the report several times and could not find where 95% comes from, the nearest figure in it is much narrower, and “successful” there means a marked and sustained productivity or P&L impact, so “unsuccessful” does not mean zero. Scott Raynovich went further and called the whole thing unfounded. Take the trend, leave the decimal. Companies are buying a lot of model and converting a thin slice into anything a CFO will sign.
02Write the job so the technology can’t get into the line
Outcome-Driven Innovation gives you a grammar for this, and it is stricter than most teams expect. A job is a verb, an object of that verb, and a clarifier saying when. Nothing else. Tony Ulwick published the method in Harvard Business Review in January 2002, and the rule that makes it work is that the wording stays solution-free.
Here is the test. You can run it on your roadmap this afternoon. Read the job out loud, and if a technology appears in it, you wrote a capability wearing a job’s clothes. “Let advisors chat with the estimate” fails. So does anything with “AI-powered” in it.
Then break it into steps. Lance Bettencourt and Ulwick published the universal job map in 2008: define, locate, prepare, confirm, execute, monitor, modify, conclude. Eight steps, and every job you care about splits into them. Your AI call lives inside that split. You are never choosing whether to add AI to a product, you are choosing which step of a known job a model may touch and which step a person keeps.
At Strategyn, the firm behind Jobs-to-be-Done and ODI, I helped build the tooling for this. One product I shaped scores a customer need line and suggests fixes against ODI rules. Writing that scorer showed me how many needs people submit that are really feature requests with the verb moved to the front.
03One job at AutoLeap, from line to shipped feature
I was Principal Product Designer at AutoLeap, which makes shop-management software for auto-repair shops across North America. Service writing was the sore spot. Every new job started with a blank page, and the service advisor pieced together customer concerns, inspection findings, services and parts by hand. Slow. Error-prone. It pulled them away from the customer standing right there.
Written with no technology in it, the job goes like this: turn what a customer says is wrong with their car into a priced repair order they will approve.
Now walk the map. Locating the inputs and preparing a draft are two steps. Confirming that draft and closing the order are two more. The model got the first pair, the human kept the second, and that split was the design rather than an afterthought. Our AI Receptionist answers calls the shop would have missed and hands what the customer said into the estimate. Advisors can also type or speak plain commands into the estimate screen, “oil change, replace rear wiper, squeaking brakes,” and the model separates those into distinct jobs with parts and labor attached. The advisor reviews, edits, approves. Always. A shop’s name is on the invoice and the model does not get to sign it.
So an advisor opens a nearly complete order instead of an empty one. Across the four initiatives I led there, my résumé figures are +20% ACV, +15 NPS, and −23% churn. Those belong to the whole body of work rather than to service writing alone, and I am not going to split them after the fact to tell a cleaner story.
What the job line really bought was the shape of the argument in the room. Nobody had to debate whether a chat box was a good idea. We debated which map step was underserved, and evidence can settle that.
04What the job map could not tell me
Opportunity scoring in ODI has two axes: how important a step is, and how satisfied people are with it now. No slot for how sure the model is. Every line of a draft carries some chance of being wrong, which puts whoever reviews it into a job the framework never described.
A blank page is honestly empty. You know exactly how much work sits in front of you. A draft that is right most of the time is a different animal, because reading it costs something, and when the errors look plausible rather than obvious, reading it can cost more than starting clean. No outcome line I ever wrote at Strategyn had a term for that. Shipping surfaced it, and what we settled on was structural: review and approval stay with the advisor, and a wrong line has to be easy to spot and cheap to kill.
Take that away if you take nothing else. Jobs-to-be-Done tells you where to put the model. What a wrong answer costs at that spot is a number you only get by handing someone a draft they distrust and watching what they do.
05Why the job outlives the model
Models get replaced every few months now. Job lines do not, and that gap is the whole practical case for this framework in AI work. Service advisors have needed to turn a complaint into an approvable estimate for as long as repair shops have existed, and they will still need it when whatever you shipped this quarter is obsolete twice over.
Real usage points the same way. In September 2025 a team including Aaron Chatterji and David Deming published “How People Use ChatGPT” through NBER, drawing on roughly 700 million users and 18 billion messages a week. Practical guidance, seeking information, and writing cover close to 80% of chats. Asking beats doing, 49% to 40%. Read that list again. Those are jobs in the plainest language available, and not one of them names a capability. People showed up with work in hand and hired the model for whichever step of it hurt most.
Same loop I keep coming back to: research to product to code, run by whoever is holding the evidence. The job line is what makes the loop portable. Swap the model, keep the line, re-measure the same step. I wrote about the model-churn half of that in Building with LLMs.
06Where to start this week
Take the last AI feature your team shipped. Write down the job it serves, verb and object and context, with no technology in the wording. If you cannot finish it, you already have your answer. Then mark which map step the model touched and go look at your dashboard. Are you measuring that step, or counting prompts sent?
Prompts sent is theatre. The step is the thing.
One caveat before you commit. Jobs-to-be-Done is not a single framework, and the camp I trained in is the one that says jobs hold still. Alan Klement argues close to the opposite, that a job is a story a person tells about their own change, and if he is right my portability claim gets much weaker. I worked at Strategyn. Read that bias into everything above, and the branches below will point you at both sides.
What would you need to see before you let a model write the first draft of something your customer has to sign? For me it came down to how a wrong line gets caught, and I wrote up how the whole thing ran in the AutoLeap case study.
07References
- Turn Customer Input into Innovation
- The Customer-Centered Innovation Map
- Know Your Customers’ “Jobs to Be Done”
- The 2019 Feature Adoption Report
- The GenAI Divide: State of AI in Business 2025
- Why We Don’t Believe MIT NANDA’s Weird AI Study
- How People Use ChatGPT
- Know the Two, Very, Different Interpretations of Jobs to be Done
- Alan Klement’s War On Jobs-To-Be-Done
- Marketing Myopia
08Further study
Branch by what you are trying to do next.
If you are writing your first job statements.
- Turn Customer Input into Innovation — the grammar and the reason for it; the solution-free rule is the one everyone breaks first.
- The Customer-Centered Innovation Map — the eight-step breakdown; use it as the worksheet for deciding which step your model touches.
- Know Your Customers’ “Jobs to Be Done” — the accessible version; the milkshake story is there because the competitor set changes the moment you define the job properly.
If you are explaining to a sceptical exec why the AI feature underperformed.
- The 2019 Feature Adoption Report — the cleanest pre-AI proof that feature-led shipping wastes most of its money; predating the argument makes it harder to wave away.
- The GenAI Divide — where the 95% number comes from; read it yourself before you quote it.
- Why We Don’t Believe MIT NANDA’s Weird AI Study — the rebuttal; bring both to the meeting, because a number you cannot defend is worse than no number.
If you want the honest picture of what people hire models for.
- How People Use ChatGPT — the largest usage taxonomy in public; the asking-versus-doing split is the finding most product teams get wrong.
- Marketing Myopia — the ancestor of this whole argument, sixty-six years early and still the sharpest version of it.
If you want the strongest objection to this essay, honestly named. Jobs-to-be-Done is not one thing, and the split matters for my argument specifically. Ulwick holds that jobs are stable over time and functional at the core, which is exactly the property I used to claim the job outlives the model. Klement’s January 2018 piece labels that reading “jobs-as-activities” and sets it against a “jobs-as-progress” reading, where a job is closer to an emergent story a person tells about their own change. Ulwick answered him the next day, at length and with heat, and the exchange is worth reading in both directions because the disagreement is substantive underneath the tone. If jobs are more emergent than stable, a job statement is a snapshot rather than a fixed point, and my portability claim gets much weaker. Christensen’s own camp never cited ODI in Competing Against Luck, which tells you how deep the split runs. Read Klement first, then Ulwick, then decide which one your product actually behaves like.