← AHMAD BILAL / INDEX
WRITING
WRITINGAI · DESIGN · RESEARCH
Writing
Notes from Ahmad Bilal, a Product Design Engineer & Researcher, written at the seam of research, design, and engineering. Grouped by the ideas I keep returning to. Short pieces, mostly.
// follow RSS feed · no list, no tracking, nothing to unsubscribe from
AI as a design material
Behavior before screens.
Designing with the model's grain: guardrails and evals as part of the design surface.
-
The answer is the cheap partWhen drafting became cheap, the scarce thing became the reviewer's minute. Explanations move agreement, not accuracy. What an interface has to expose instead.
-
AI-native product design: what actually changesAI-native is not a chat box on an old product. It is the harness, evals, guardrails, approval gates, made part of the interface itself.
-
Agent experience design: when the user is sometimes an agentWhat changes in affordances, error states, and provenance when agents consume your product alongside humans.
-
AI as a Design MaterialDesigning with AI as a probabilistic raw material. You understand its grain, then build the guardrails, evaluation, and human-in-the-loop around it.
-
Designing AI Behavior, Not AI ScreensThe real work in AI agents UX. Confidence caps, a critique-repair loop, hypothesis-blind generation, and a harness that makes the behavior predictable.
-
Building With LLMs: Designing the HarnessTreat the model as a probabilistic material and wrap it in a harness: an eval harness, guardrails, model routing, and audit journals.
Auditable AI research
Defensible by construction.
Synthetic user research with evidence chains, eval harnesses, and audits a skeptic can run.
-
What my 92 was measuringA hobby TB screener scored 92 on the benchmarks and 78 at a real hospital. The audit that followed became a preprint: the benchmarks partly grade the acquisition pipeline.
-
Grounded personasA readable guide to the LLM persona-generation research: what a persona can hold, what it quietly flattens, and what grounding in evidence actually buys.
-
A simulated user research workflow, end to endGrounding sources to audit: how a simulated study runs when every finding has to trace back to evidence.
-
Harness engineering for research agentsThe agent-harness discourse is all about coding agents. Research agents fail differently, and need a different harness. Here is what that harness holds.
-
Grounded Simulation: faithful, not just fluentA first-principles architecture for keeping LLM “synthetic users” faithful and auditable, not just fluent.
-
Auditable AI ResearchWhy AI-generated research has to be inspectable and defensible. Research you can audit in thirty minutes, plus the mechanisms that get you there.
-
Less Expertise, More CoverageFraming an LLM as a narrow expert can shrink the coverage of an analytical task. A broader prompt often surfaces more of the real answer space.
Research to product to code
Fewer handoffs.
One loop from question to shipped code, with tighter evidence at every pass.
-
Jobs-to-be-Done for LLM product featuresA capability is not a job. Write the job as one solution-free line, then decide which single step a model may touch. That call survives the model being swapped.
-
Research → Product → CodeOne continuous loop with minimal handoff loss. What you gain when the person who frames the question also ships the code.
Public trust and AI for good
Design in the public interest.
Regulators people can verify, and AI that lowers barriers instead of raising them.
-
The grammar was the hard partThe build story of my peer-reviewed English/Urdu to Pakistan Sign Language translation system, and why a context-free grammar beat a neural model at 50,000 sentences.
-
Designing a regulator people can trustWhat trust means for a regulator's digital front door, and the design decisions behind the PVARA web experience work.
-
PVARA explained: Pakistan's virtual asset regulatorThe encyclopedia: history, licensing today, the Shariah layer, global peers, and what comes next. Sourced, dated, not legal advice.