Where should AI live while you're designing a product, and after you've shipped it?

Threadwise tells a shopper if a piece of clothing is worth buying, based on its fabric. That's the backdrop for two decisions.
First: to design the prototype, three AI models proposed different directions, and I directed AI to combine the strongest parts of each into the final design.
Second: the product became research into cutting the cost of AI itself. The answer was that expensive AI reasoning should happen once, offline, never on every request while someone's using the app.
The final, human-curated Threadwise prototype. Fable's design identity, Sonnet's completeness, and Opus's discipline, rebuilt into one flow. Tap through it above.

ROLE

METHOD

TOOLS

HEADLINE OUTCOME

Product Designer, solo
AI-architecture design + 3 model controlled exploration
Zero runtime AI calls, the intelligence is frozen at build time
Claude Design: Sonnet, Opus, Fable

The product itself never calls AI.

01 THE REAL EXPERIMENT

Here's the second decision made concrete: not a claim, an architecture. Fabric data gets researched and scored once, checked by a person, then frozen into the app. A shopper opens Threadwise and gets an instant answer, with nothing running live.

A product that calls an AI model on every user request is slower, more expensive, and more fragile than it needs to be. And every answer is a fresh roll of the dice. Freezing the expensive reasoning at build time, and verifying it once by hand, makes the runtime experience instant, private, consistent, and cheap to run at any scale. Every verdict traces to a reviewed table, so the product is auditable in a way per-request inference never is.
A sustainability product shouldn't burn compute re-deriving the same answer millions of times. Threadwise applies its own ethic to its AI: use exactly as much as the job needs, exactly once.

why this matters

The label was honest. It just wasn't useful.

Garment labels give shoppers percentages, not judgment. Knowing an item is 80% cotton and 20% polyester doesn't tell you how that blend will actually behave: how it wears, how it washes, whether it sheds microplastics, or what happens to it at the end of its life. Sustainability information exists, but it's scattered, jargon-heavy, and never present at the one moment it would change a decision.

02 The problem

How might Threadwise translate fabric composition into a quick, honest, non-judgmental purchase decision?

Seven documents, before a single interface.

03 FOUNDATION

I directed these seven documents into existence rather than writing them solo. I set the thinking and made the judgment calls, then handed execution to whichever model suited the task: Sonnet for the core product goal and problem framing, Opus for structural and visual direction, Fable for the scoring logic's fiber math. Opus then compiled everything into the documentation site linked below and froze it there.

This is where the product's real judgment calls got made: what the score can and can't see, what stays out of the MVP, what tone the product refuses to take, so no model would have to invent them later.

The freeze had an economic job too: a settled brief means exploration happens once, against stable ground, instead of being regenerated every time the thinking moves.

One brief. Three models. No tow alike.

04 exploration

Sonnet
Resolved on a fast, minimal "Plan A"
Opus
Fable
Building the foundation meant matching each document to the model suited to it. The MVP exploration was the opposite move, on purpose: I gave the exact same seven documents and the exact same staged prompt to Sonnet, Opus, and Fable, so I could compare how three models handled one identical brief rather than three different tasks.

Each produced three MVP directions, resolved one into a hi-fi direction, and built it into an interactive prototype.

Resolved on "The Doorway"
Resolved on "The Signal"

Same brief, three different temperaments.

05 FINDINGS

I didn't pick a winner. I picked a base.

06 CURATION

I chose Fable's visual language and design judgment as the foundation, then pulled in Sonnet's completeness and Opus's discipline where they made the product better, and directed a final AI pass to bring the whole flow into one consistent design.

Where the curated version actually changed things.

07 before & after

Example 1 · The glance, without a nudge

The glance carried a nudge toward a different purchase: "Try: linen or Tencel →"
The shipped result screen: Verdict, score, one reason. Nothing steering the next purchase.

before

after

Remove the alternative nudge

Pulled straight from two real documents: the earliest lo-fi flow set against the refined screen library that shipped. The product didn't just get prettier. It got more honest about staying out of the way.

Example 2 · The detail, without a sales pitch

"Every result carried a "better path" ladder — repair, secondhand, or a different material to buy instead.
The shipped expanded view: Verdict, score, one reason. Nothing steering the next purchase.

before

after

Cut the better-path block

One Threadwise, built from three.

08 THE FINAL PRODUCT

What a controlled comparison actually looks like.

09 BY THE NUMBERS

7

3

0

foundation documents, written before any interface
runtime AI calls in the shipped architecture
Claude Design models run against the identical brief

12

evaluation criteria scored across all three outputs

What the experiment actually taught me.

10 takeaways

The tooling changed the speed of the work. It didn't change who had to make the decisions & which is why it's worth building a practice around a workflow that can move across models as the landscape shifts, rather than around any one tool.
Four takeaways from running one brief through three models, and from deciding where the AI belongs once the product ships.

The three-to-five-minute version is above. Here's the receipts.

11 explore in depth

© 2026 Vaidehi Yelkawar