A friend who runs a mid-sized apparel brand let me finish the Selrite pitch, then asked the question I should have seen coming. “So you ask four questions and calculate the answer. What are you going to ask me? How many dresses does she need per square metre?”
He was right, and it stung a bit. Our whole model rests on a property most catalogs don't have. When someone is buying epoxy for a 240 m² warehouse floor with forklift traffic, the requirement is a calculation. Area, coats, porosity, wastage, pack size. There is a correct answer, the customer usually can't reach it alone, and getting it wrong costs them a weekend and a floor. That is what makes the Finder motion work — and it is exactly what a dress does not have.
The temptation after that conversation is to conclude the idea doesn't travel. I think the opposite is true, but only if you separate the principle from the mechanism. The principle is that the store should do the expert's job rather than pushing it onto the shopper. That travels everywhere. The mechanism — a short structured interview that resolves to a quantity — travels almost nowhere outside computable categories.
So the real question isn't whether guided selling applies to your catalog. It's which kind of expert your customer is missing. Ask that, and you get a different answer for coatings than for cameras, and a different one again for dog food.
Every store is built like a menu. Nobody shops like a diner.
Start with what's actually broken, because it's the same everywhere. Almost every online store still runs the same twenty-five-year-old template: a search box, a category tree, a grid of SKUs. All three assume the customer already knows the catalog — which branch to open, which filter matters, which spec to compare. And every product page is a one-sided advertisement. It tells you what's good about the product and never tells you when not to buy it, what it's genuinely best at, or what else the job requires.
That's the gap. A good salesperson asks about your situation, rules things out, and right-sizes the purchase. None of that survived the trip onto the product page. I've written before about what a great rep knows that your website doesn't, and about what search-and-grid quietly costs you. The diagnosis holds across categories. The prescription doesn't.
Three questions that decide the motion
Before designing anything, it helps to know why categories differ. Three old ideas from consumer research cover most of it, and none of them require a literature review to use.
What can the buyer know before they buy? Economists separate search goods, where quality is assessable up front from attributes — a drill's torque, a laptop's specs; experience goods, where you only learn by using — how jeans fit, how a mattress sleeps; and credence goods, where you can't reliably tell even afterwards — whether the supplement did anything. Each has a different information gap. Search goods need translation and comparison. Experience goods need simulation and de-risking. Credence goods need evidence.
Head or heart, and how much is at stake? Cross involvement with think-versus-feel and you get four decision sequences. Expensive thinking purchases — cars, laptops, insurance — run learn, learn, feel, do, over days or weeks. Expensive feeling purchases — fashion, jewellery — run feel first and rationalise second. Cheap habitual purchases run do, learn, do. Cheap impulse purchases barely deliberate at all. This tells you when to intervene: before the shortlist for thinkers, after the inspiration for feelers, at the reorder moment for habits.
Is the requirement computable? This is our home turf and it is rarer than we'd like. Paint, flooring, fertiliser, cable, catering, insulation — the parameters resolve to a number. Most consumer purchases don't. The “requirement” is a body, a taste, a room, a skill level, or another person's preferences.
Run those three lenses over a catalog and you end up with six diagnostic questions. The first “yes” usually names your primary motion.
| Ask this about your catalog | If yes, the dominant motion is |
|---|---|
| Can the need be computed from job parameters? | Estimator — consumption maths, kit completion |
| Is physical or technical fit the main risk? | Fit guardian — try-on, size prediction, compatibility |
| Is taste or identity the main driver? | Curator — visual discovery, preference learning |
| Is the purchase repeated on a cycle? | Autopilot — consumption tracking, effortless reorder |
| Is quality unverifiable even after use? | Trusted advisor — evidence, education, expectations |
| Is the buyer shopping for someone else? | Concierge — proxy profiles, occasion memory |
Most real catalogs answer yes twice — a primary motion and a secondary one. A bike shop is spec-heavy and fit-risky. A skincare brand is credence-driven and replenishing. That's why the useful output is a playbook per archetype rather than one universal flow.
Where we started: computable jobs
Coatings, adhesives, flooring, tile, insulation, fertiliser, industrial consumables. The buyer has a job with measurable parameters, quantity is a formula, and a wrong pick has a concrete cost. Capture the job in three to five plain-language questions, compute the quantity with the working shown, pick the variant that suits the conditions, and complete the kit with the primer, sealer and applicator the job actually needs. How it works walks the whole flow if you want the mechanics.
Two things about this motion are worth stealing even if your catalog isn't computable. The first is showing the maths. Buyers don't trust a number they can't audit, and the moment you show the arithmetic the number stops being a sales figure and starts being an answer. The second is not-suitable-for enforcement: “this epoxy is wrong for your exterior deck — it yellows under UV, use the aliphatic urethane instead.” That sentence loses a sale roughly never and wins trust roughly always.
The metric here isn't conversion. It's right-order rate — the returns and support tickets that never happened. You can see the shape of that in the coatings, adhesives and industrial deployments.
When the risk is fit, not choice
Apparel, eyewear, furniture, PC parts, auto parts, printer cartridges, appliance spares. Here the customer usually knows roughly what they want. The risk is that it won't fit their body, their room, or their equipment. Two sub-patterns, and they need different machinery.
Body and space fit. The sequence is feel first, verify second, so don't put a quiz in front of the browsing — install a guardian at the decision point instead. A persistent fit profile — measurements, past purchases with their fit outcomes, or one photo — that every product page consults: “in this brand you're a medium, not your usual large; 34% of buyers your size returned this cut for tightness across the shoulders.” Virtual try-on is doing something more interesting than it looks: it converts an experience good partway into a search good. For furniture the equivalent is AR placement plus a doorway and stairwell clearance check, which is the one question customers reliably forget to ask. The signature honest behaviour is a return-risk warning before add-to-cart, mined from reviews and return reasons. Every prevented return is margin you keep and trust you earn.
Technical compatibility. Nobody has ever wanted to “browse RAM.” They want to know whether this stick works in their machine. The move is an equipment registry — my car, my printer, my build, my boiler — captured once and used forever, with every browse filtered through a compatibility check that flags the mismatch before the cart rather than after the delivery. Underneath, that's a graph problem: product fits product, product requires product, product conflicts with product. In these categories it's the highest-return thing you can build, because a wrong part isn't a disappointment, it's a guaranteed return.
When questions are the wrong instrument
Fashion styling, décor, jewellery, art, fragrance. Ask someone to describe their taste and they will describe someone else's. A ten-question aesthetics quiz reads as a tax, and worse, the answers aren't reliable. People know it when they see it.
So the motion inverts: show, observe, refine, instead of ask then recommend. Start from any seed the shopper offers — a saved board, a screenshot, a photo of their sofa, one liked item — and run a visual conversation. More like this, less like that. A minute of that converges better than any questionnaire, and vision models have made “find me something that goes with this photo” an ordinary capability rather than a research project.
Two behaviours here are worth naming. Ask about context rather than taste — “what's the occasion?” is one question a shopper can genuinely answer, and it prunes the catalog harder than ten questions about style. And sell completion, because the moment of highest receptivity is right after a choice, not before it: these two sandals you already own work with this dress, or this new pair finishes the look. That's where the basket grows, and it's why the strongest reported basket lifts come from categories where the assistant can assemble a set or a routine rather than sell an item. Honest behaviour here is unglamorous: occasion warnings (“this fabric creases badly — poor pick for a suitcase”) and duplicate alerts (“you bought something very close to this in March”).
When the buyer needs a translator
Laptops, cameras, TVs, appliances, mattresses, bicycles, power tools, strollers. High-involvement thinking purchases: infrequent, expensive, spec-dense, researched across many sources over days. The core problem is a language mismatch. The industry talks in nits and lumens and thread counts; the customer thinks in “will this be bright enough in my living room.”
Same-session conversion here is far lower than in consumables, and chasing it is a mistake. The sale lands later, often much later, so the design goal is to be the best research companion rather than to force a same-day close. Four things do that work. Elicit the requirement in life language and then show the translation — “video editing means 16 GB and a colour-accurate panel; it does not mean paying for 4K at fourteen inches.” Present three options, not forty: the safe one, the value one, the stretch one, each with reasons for and reasons against for this buyer. Bring the evidence inside instead of hiding it, because the shopper is going to verify you elsewhere anyway — summarise the review corpus honestly, hinge complaints and all. And remember the journey, so someone who returns four days later resumes where they stopped instead of starting over.
The honest behaviour that builds durable-goods brands is the down-sell. “For what you've described, the $700 model buys you nothing over the $480 one.” A menu store cannot say that sentence, because it doesn't know what you described.
When the best experience is no experience
Groceries, coffee, pet food, supplements, nappies, cleaning supplies, ink, blades. Low-involvement and habitual, and the customer-centric insight is uncomfortable for anyone who loves designing interfaces: the ideal number of shopping decisions is zero.
Guide once, at first purchase, then never ask again. After that the job is consumption modelling — from household size, pack size and order intervals, predict the run-out date and act just before it. “You're about five days from running out; reorder the usual?” The surface isn't a store, it's a basket the customer edits in ten seconds. Honest behaviour is pack-size and price transparency — telling someone the 2 kg bag is 18% cheaper per kilo costs you revenue today and buys the subscription — and, critically, never exploiting the inertia you've been given. One silent price creep, detected once, ends the relationship permanently. This is also where agentic commerce lands first, because a standing intent (“keep me stocked, cap it at $80 a month”) maps cleanly onto agent payment protocols with spend policies attached.
When nobody can verify the answer
Supplements, skincare efficacy, baby and child safety, health devices, sustainability claims. The customer can't judge quality even after using the product. Did the collagen work? Is this car seat actually safer? Anxiety runs high and the category is full of overclaiming, which makes it the hardest place to be believed and the best place to win by being believable.
The motion is evidence over adjectives, graded and sourced: “clinically studied at this dose — here's the study” sits in a different tier from “traditionally used,” and the assistant should say which tier it's in. Run a structured intake — skin type, concerns, current routine, sensitivities; or the child's age, weight and the car model — and then set expectations honestly, including the unwelcome ones. “Retinol takes eight to twelve weeks. Anyone promising two is lying to you.” Contraindication checking is this category's version of not-suitable-for enforcement, with a clear boundary about where a professional takes over.
The behaviour that wins outright is a willingness to recommend nothing. “Your routine already covers this — adding it would be redundant.” An assistant that visibly refuses to oversell is the one whose recommendations get acted on.
When the buyer isn't the user
Gifts, children's products, pet products, buying for a household or a team. The menu store handles this worst of all, because all it can do is show you what's popular.
The motion is proxy-profile building: elicit what's known about the recipient — “sister, thirties, into pottery and hiking, budget around $40, hates clutter” — and reason over it, including the inference that hates clutter points at consumables and experiences rather than objects. Occasion memory turns that into a relationship: the store that says “your father's birthday is in three weeks, last year you sent X, here are three that build on it” has quietly become the household's gifting service. Honest behaviours are gift-risk flags (“fragrance is intensely personal — here's a safer pick at this budget”) and a genuinely painless exchange path, because even good proxy guesses miss.
Trust is the binding constraint, and it scales with price
Everything above runs into the same ceiling. Product.ai's 2026 Trust in AI Commerce report, which surveyed 1,463 US online shoppers, found 43% had used an AI assistant to research a purchase in the previous 90 days — and that 86% of those still verified the recommendation somewhere else before buying. There's a price curve underneath it: 42% wouldn't act on an AI recommendation above $25 without a second source, more than 60% capped it at $50, and only 5% would trust AI on a purchase over $500. Ranked against seven research sources, AI came sixth, ahead of only YouTube creators.
That should change what you build, not just how you market it. If the shopper is going to verify you anyway, bring the verification inside the recommendation — cite the spec, quote the review, link the study. And design autonomy per price band rather than globally: at $20 the assistant can simply decide, at $2,000 it should coach, compare, cite, and leave the decision to the human. The single strongest trust mechanism available is still the willingness to say “this isn't for you, and here's why,” which is convenient, because it's also the honest thing to do.
One number I want to be careful with, since it gets quoted loosely. Guided quiz flows are widely reported at around 5.5% conversion against a typical 2% store average — RevenueHunt's benchmark across 45 million quiz responses. But that's a quiz-taker rate measured against a site-wide rate, and quiz takers are self-selected. RevenueHunt says so themselves. Part of that gap is the funnel, not the guidance. The effects I'd actually underwrite are the less quotable ones: fewer returns, bigger baskets when the assistant can build a set, and the compounding value of zero-party data you can only get by asking.
The timing argument is more solid. McKinsey puts agentic commerce at $3–5 trillion globally by 2030, and Deloitte projects roughly a quarter of global e-commerce sales will be agent-influenced by then. Whatever you build now should be legible to software as well as to people.
The eight, side by side
| Archetype | Example categories | First interaction | Signature honest behaviour | Metric that matters |
|---|---|---|---|---|
| Estimator | Paint, flooring, adhesives | “Describe the job” — 3 to 5 questions | Not-suitable-for enforcement | Right-order rate |
| Fit guardian (body/space) | Apparel, eyewear, furniture | One-time fit profile | Return-risk warning pre-cart | Return rate |
| Fit guardian (technical) | Parts, components, spares | “Add your equipment” | Incompatibility block pre-cart | Wrong-part returns |
| Curator | Fashion style, décor, jewellery | Seed image or 60-second swipe | Occasion honesty, duplicate alerts | Basket value on sets |
| Translator | Laptops, appliances, mattresses | “What's it for?” in life language | Down-selling; cons stated plainly | Long-window conversion |
| Autopilot | Grocery, pet food, supplements | One-time needs quiz | Pack-size and price transparency | Decisions eliminated |
| Trusted advisor | Supplements, skincare, baby safety | Diagnostic intake | Recommending nothing when apt | Repeat rate |
| Concierge | Gifts, kids, pets | “Tell me about them” | Gift-risk flags, easy exchange | Occasion repeat rate |
A marketplace or department store doesn't pick one of these. It routes. The first job becomes intent classification — work out which archetype this session belongs to, then hand off. That routing layer is itself the new front door, and it's the thing replacing the search box.
Eight front ends, one brain
Here's the part that makes this buildable rather than eight separate products. The motions differ at the surface and share almost everything underneath.
The non-negotiable layer is decision-grade product intelligence. Raw catalog data — title, price, benefit bullets — cannot power any of the eight. Every product needs a structured decision record: what it's suitable for, what it's explicitly not suitable for, what it requires or conflicts with, the sizing or quantity logic where there is any, the evidence behind each claim with its provenance and a confidence score, and where it sits against the alternatives. Language models made building this affordable — they can read spec sheets, manuals, transcripts and thousands of reviews and emit exactly that structure, with the low-confidence and safety-sensitive claims routed to a human before any customer sees them. Models commoditise. An evidence-linked decision graph of your catalog doesn't.
Above it sits customer memory — fit profiles, equipment registries, taste vectors, consumption models, constraints, journeys in progress, proxy profiles — built overwhelmingly from what customers deliberately tell you in the guided flows themselves. Two rules keep it healthy: the customer can see and edit all of it, and every fact they share has to improve their experience in the same session. Break the second rule and they stop sharing.
Then the reasoning layer, which is where most bolt-on advisors come apart. A router classifies intent and dispatches to the motion-specific agent, each grounding its answers in retrieval rather than recall — and the anti-sell rules live here as enforced policy, not model temperament. I've argued this at length in Why Most AI Product Advisors Fail: the problem is rarely the model, it's asking one model to do five different jobs. The retrieval half of that argument is in LLM Product Discovery Isn't SEO.
Finally, presentation stops being one fixed layout. The same intelligence renders as a conversation when exploration is open-ended, a four-question flow when the job computes, a swipeable deck for taste, a comparison table for a shortlist, a one-tap basket for replenishment. Classic browse never disappears — some people want the menu, and removing it costs you both trust and search traffic. The guided layer wraps it. And there's a fifth surface now: the agent interface, where structured feeds and machine-readable decision data determine whether a shopping agent can represent your catalog at all. The decision-grade data you built for layer one is exactly what those agents need. Stores that expose it get picked. Stores offering only marketing prose get skipped.
The loop that ties it together is the reason this compounds. Guided sessions produce zero-party data, which sharpens recommendations and follow-up, which produces more guided sessions and — more valuable — outcome data. Returns, satisfaction, repeat rates. That flows back into the product graph: buyers with narrow feet returned this shoe. The next customer gets better guidance because the last one did. A menu store has no such loop. Every session starts from zero.
Build it as a wedge, not a re-platform
The failure mode I'd most want to talk someone out of is trying to do all eight at once. A workable order:
- Pick the category where the payoff is measurable. Highest return rate if you're fit-risky, highest support-ticket volume if you're computable or spec-heavy, highest reorder share if you replenish.
- Enrich only that slice of catalog. With LLM pipelines this is weeks rather than months — but budget real human review time for anything safety-sensitive, because that's the part that bites.
- Launch the guided motion beside the existing browse, not instead of it. Instrument it end to end: flow completion, recommendation acceptance, conversion on a 30 to 90 day window, basket value, return rate, and which zero-party fields you actually captured.
- Wire the captured data into lifecycle marketing immediately. A meaningful share of the return shows up here rather than in the session.
- Then expand. Add the second motion, share customer memory once two exist, and add agent-facing endpoints when agentic traffic becomes measurable rather than theoretical.
Five pitfalls recur often enough to name. The bolt-on chatbot, which can only paraphrase marketing copy because nobody built the intelligence layer underneath it. Interrogation fatigue, where every question you ask should first have to prove it can't be inferred from behaviour, images or history. Fake honesty, where not-suitable-for is decorative rather than enforced — shoppers detect theatre, and most of them are verifying you anyway. Trust overreach, where the assistant acts with $20 confidence on a $2,000 decision. And privacy debt, where customer memory without transparency or control turns from an asset into a liability.
The same principle, a different expert
My friend with the apparel brand doesn't need a consumption calculator. He needs a fit guardian and a curator, and if he builds them properly the underlying machinery looks a lot like ours — a catalog enriched into decisions rather than descriptions, a memory of the customer, a reasoning layer that's allowed to say no, and a surface that adapts to how the decision is actually made.
A menu store sells what it has to whoever shows up.An AI-native store understands who showed up, what job they're doing, and sells them the right outcome.Sometimes that's less than they came to buy.
We proved the thesis where jobs are computable, because that's where the maths made it provable. The principle was never about the maths.
Sources. Shopper trust and verification data from Product.ai's 2026 Trust in AI Commerce Report (n=1,463 US shoppers, April 2026). Quiz benchmark from RevenueHunt's 2026 benchmark (45M+ responses). Market sizing from McKinsey and Deloitte.