← AHMAD BILAL / WRITING ONE LOOP
PRODUCTJOBS-TO-BE-DONEAI FEATURES

Jobs-to-be-Done for LLM product features

A capability is not a job. Write the job as one solution-free line, break it into steps, and decide which single step a model may touch. That call survives the model being swapped.

Ahmad BilalAug 2026~8 minResearch to Product
FIG. 01 · THE SWAP A heavy-ruled box holding the job line, turn a complaint into an approvable estimate, wired to two ports labeled draft and approve. Beneath it three model boxes M1, M2, M3 sit on a dashed rail; only M2 is connected. The job holds still, the model does not.
The unit of analysis that survives model churn. The job line stays wired to the same two ports while the model boxes slide through underneath.
80%of features rarely or never used (Pendo, 2019)
8steps in the universal job map
49% vs 40%asking beats doing (NBER, 18B msgs/wk)
+20% / +15 / −23%ACV, NPS, churn across four AutoLeap initiatives

A capability is not a job. Most LLM features ship because the model can do the thing, and then they underperform, because “summarize this” names the tool and says nothing about the work a person was trying to finish. Write the job as one solution-free line, break it into steps, and decide which single step a model may touch. That call survives the model underneath it being swapped.

01Capability-led roadmaps were leaking before the models arrived

Pendo looked at feature usage across 615 of its subscriptions and reported in February 2019 that 80% of features in the average software product are rarely or never used. Their estimate of the cloud R&D behind those features: $29.5 billion. No language model anywhere in that data. Feature-led roadmaps were already burning most of their money before anyone shipped a chat box, and the models did not cause that. They made it cheap to build the next thing nobody hired.

You have probably seen the newer number. MIT’s Project NANDA put out “The GenAI Divide” in July 2025, and the press squeezed one line of it into a headline: 95% of enterprise GenAI pilots return nothing. Handle that one carefully. Kevin Werbach at Wharton read the report several times and could not find where 95% comes from, the nearest figure in it is much narrower, and “successful” there means a marked and sustained productivity or P&L impact, so “unsuccessful” does not mean zero. Scott Raynovich went further and called the whole thing unfounded. Take the trend, leave the decimal. Companies are buying a lot of model and converting a thin slice into anything a CFO will sign.

02Write the job so the technology can’t get into the line

Outcome-Driven Innovation gives you a grammar for this, and it is stricter than most teams expect. A job is a verb, an object of that verb, and a clarifier saying when. Nothing else. Tony Ulwick published the method in Harvard Business Review in January 2002, and the rule that makes it work is that the wording stays solution-free.

Here is the test. You can run it on your roadmap this afternoon. Read the job out loud, and if a technology appears in it, you wrote a capability wearing a job’s clothes. “Let advisors chat with the estimate” fails. So does anything with “AI-powered” in it.

Then break it into steps. Lance Bettencourt and Ulwick published the universal job map in 2008: define, locate, prepare, confirm, execute, monitor, modify, conclude. Eight steps, and every job you care about splits into them. Your AI call lives inside that split. You are never choosing whether to add AI to a product, you are choosing which step of a known job a model may touch and which step a person keeps.

At Strategyn, the firm behind Jobs-to-be-Done and ODI, I helped build the tooling for this. One product I shaped scores a customer need line and suggests fixes against ODI rules. Writing that scorer showed me how many needs people submit that are really feature requests with the verb moved to the front.

03One job at AutoLeap, from line to shipped feature

I was Principal Product Designer at AutoLeap, which makes shop-management software for auto-repair shops across North America. Service writing was the sore spot. Every new job started with a blank page, and the service advisor pieced together customer concerns, inspection findings, services and parts by hand. Slow. Error-prone. It pulled them away from the customer standing right there.

Written with no technology in it, the job goes like this: turn what a customer says is wrong with their car into a priced repair order they will approve.

Now walk the map. Locating the inputs and preparing a draft are two steps. Confirming that draft and closing the order are two more. The model got the first pair, the human kept the second, and that split was the design rather than an afterthought. Our AI Receptionist answers calls the shop would have missed and hands what the customer said into the estimate. Advisors can also type or speak plain commands into the estimate screen, “oil change, replace rear wiper, squeaking brakes,” and the model separates those into distinct jobs with parts and labor attached. The advisor reviews, edits, approves. Always. A shop’s name is on the invoice and the model does not get to sign it.

So an advisor opens a nearly complete order instead of an empty one. Across the four initiatives I led there, my résumé figures are +20% ACV, +15 NPS, and −23% churn. Those belong to the whole body of work rather than to service writing alone, and I am not going to split them after the fact to tell a cleaner story.

What the job line really bought was the shape of the argument in the room. Nobody had to debate whether a chat box was a good idea. We debated which map step was underserved, and evidence can settle that.

04What the job map could not tell me

Opportunity scoring in ODI has two axes: how important a step is, and how satisfied people are with it now. No slot for how sure the model is. Every line of a draft carries some chance of being wrong, which puts whoever reviews it into a job the framework never described.

A blank page is honestly empty. You know exactly how much work sits in front of you. A draft that is right most of the time is a different animal, because reading it costs something, and when the errors look plausible rather than obvious, reading it can cost more than starting clean. No outcome line I ever wrote at Strategyn had a term for that. Shipping surfaced it, and what we settled on was structural: review and approval stay with the advisor, and a wrong line has to be easy to spot and cheap to kill.

Take that away if you take nothing else. Jobs-to-be-Done tells you where to put the model. What a wrong answer costs at that spot is a number you only get by handing someone a draft they distrust and watching what they do.

05Why the job outlives the model

Models get replaced every few months now. Job lines do not, and that gap is the whole practical case for this framework in AI work. Service advisors have needed to turn a complaint into an approvable estimate for as long as repair shops have existed, and they will still need it when whatever you shipped this quarter is obsolete twice over.

Real usage points the same way. In September 2025 a team including Aaron Chatterji and David Deming published “How People Use ChatGPT” through NBER, drawing on roughly 700 million users and 18 billion messages a week. Practical guidance, seeking information, and writing cover close to 80% of chats. Asking beats doing, 49% to 40%. Read that list again. Those are jobs in the plainest language available, and not one of them names a capability. People showed up with work in hand and hired the model for whichever step of it hurt most.

Same loop I keep coming back to: research to product to code, run by whoever is holding the evidence. The job line is what makes the loop portable. Swap the model, keep the line, re-measure the same step. I wrote about the model-churn half of that in Building with LLMs.

06Where to start this week

Take the last AI feature your team shipped. Write down the job it serves, verb and object and context, with no technology in the wording. If you cannot finish it, you already have your answer. Then mark which map step the model touched and go look at your dashboard. Are you measuring that step, or counting prompts sent?

Prompts sent is theatre. The step is the thing.

One caveat before you commit. Jobs-to-be-Done is not a single framework, and the camp I trained in is the one that says jobs hold still. Alan Klement argues close to the opposite, that a job is a story a person tells about their own change, and if he is right my portability claim gets much weaker. I worked at Strategyn. Read that bias into everything above, and the branches below will point you at both sides.

What would you need to see before you let a model write the first draft of something your customer has to sign? For me it came down to how a wrong line gets caught, and I wrote up how the whole thing ran in the AutoLeap case study.

07References

08Further study

Branch by what you are trying to do next.

If you are writing your first job statements.

If you are explaining to a sceptical exec why the AI feature underperformed.

If you want the honest picture of what people hire models for.

If you want the strongest objection to this essay, honestly named. Jobs-to-be-Done is not one thing, and the split matters for my argument specifically. Ulwick holds that jobs are stable over time and functional at the core, which is exactly the property I used to claim the job outlives the model. Klement’s January 2018 piece labels that reading “jobs-as-activities” and sets it against a “jobs-as-progress” reading, where a job is closer to an emergent story a person tells about their own change. Ulwick answered him the next day, at length and with heat, and the exchange is worth reading in both directions because the disagreement is substantive underneath the tone. If jobs are more emergent than stable, a job statement is a snapshot rather than a fixed point, and my portability claim gets much weaker. Christensen’s own camp never cited ODI in Competing Against Luck, which tells you how deep the split runs. Read Klement first, then Ulwick, then decide which one your product actually behaves like.