← AHMAD BILAL / INDEX CASE 01 / 07 · ARTICOS · 2025–2026
AI · 0→1 · LEADRESEARCH PLATFORMPASSWORD-GATED

Articos: research you can audit in 30 minutes

Six months designing an AI research-report generator that survives the strict-reader test. This is what it looks like when a senior researcher has to defend the output. I led the research, the experience, and the production build.

Ahmad Bilal, Product Design Engineer & ResearcherLead on Articos
ROLE
Lead Researcher & Product Designer
COMPANY
Articos (Disrupt Labs)
TYPE
AI Research Platform
TIMELINE
2025–2026 · 6 months
4.4×more accurate than plain prompting · 15 studies
42%of the themes it reports are real · plain prompting: 6.7%
0.797research fidelity, scored 0 to 1 · 46 studies
18themes to read, not 141 · per study

What those numbers mean Point a plain AI at a pile of interviews and it hands back about 141 themes, of which roughly 7 in 100 are real. You then spend your week sorting them. This hands back 18, and about 4 in 10 are real. It finds slightly fewer of the true themes than plain prompting does, and I would rather say that than hide it. What you buy is that the ones it does put in front of you are worth acting on.

Overview

Traditional user research costs $5k–$30k per study and takes weeks. That blocks agencies on five-day timelines, SaaS teams shipping weekly, and founders without budget. Most synthetic user research tools collapsed into chat wrappers that produce fluent-but-unfaithful interviews. The hard problem is generating one a senior UX researcher would put their name on. I led the research, design, and production build of Articos around a single principle: the product refuses to ship a report it cannot defend.

The research figures above are published on SSRN and stay public. The full case study walks through the six-criteria rubric, the eval harness that made it executable, the critique-and-repair pipeline, and 13 production screens. That part is password-gated.

The method is public

Articos publishes the science behind the product. The five failure modes of AI research it was built against: no scientific grounding, no memory or life history, confirmation bias, persona homogeneity, and cost. What each system does about them: hypothesis-blind generation with context isolation, cognitive memory, stance-diverse personas grounded in NEO-PI-R and ACT-R, six-stage synthesis, and an adversarial quality review that scores every study before it ships. The validation evidence is downloadable, including the pre-print, the datasets behind the precision and F1 figures, and a worked persona example.

Read the full methodology on articos.com →

What stays gated here is my side of it: how the rubric was designed, how the eval harness gates a draft, and what the production screens look like.

Full case study · password required

The complete study covers the rubric built before the prompt, the hypothesis-blind persona architecture, the confidence caps, the eval harness that gated every draft, and the production screens. If we've spoken, use the password. If not, ask and I'll send it, or walk you through it on a call.

Wrong password. Email me and I'll send the current one.

The eval harness, confidence caps, and hypothesis-blind generation here are the applied form of Grounded Simulation, the architecture I coined and wrote up on SSRN for faithful synthetic UX research. For the formal paper and the rest of my auditable AI research, see the research page.