Designing with LLMs: A UX Framework for E-commerce
A practical framework for integrating large language models into product experiences without losing the human touch.

LLMs in e-commerce feel smart. Until you actually use them
Last month I sat with a client who had just rolled out an AI chatbot on their webshop. Proud. 8,000 euro in implementation. "Customer service gets 60% cheaper," the COO said.
Two weeks later the bot had generated 340 tickets that had to be corrected by hand. Delivery times that were wrong. Return information that legally did not hold up. Product copy that read more like Wikipedia than conversion-focused text.
The model worked fine. The UX around it was a disaster.
Most AI features in e-commerce do not fail because of the model, but because nobody asked the UX question.
Two mistakes I see everywhere
The black box
At an outdoor brand I work for, their AI tool generated product titles for Amazon DE. Technically correct. But the titles missed the German search terms that actually drive traffic. The model had no access to current search data and filled the gaps with plausible-sounding but worthless keywords.
Nobody on the team knew. It looked professional. Only when the click-through rate dropped 23% in six weeks did someone check by hand.
That is the danger: the output looks confident, even when it makes no sense.
The replacement trap
A beauty brand I work with. 85 SKUs and 2 million euro in annual revenue. They wanted their entire product content generated by AI. "Saves us a copywriter."
It saved them a copywriter. It cost them 14% more returns in three months. The AI wrote that a serum "suits all skin types" while the product contains retinol. That was in the specs, but the model ignored the nuance.
AI can speed up work. But not every task should be automated end to end.
Five principles I apply everywhere
1. Set the confidence boundary
An AI-generated category text is something other than a technically validated fitment specification. A first product draft is something other than a listing ready for Bol.com.
But in practice most interfaces treat all output as if it were gold. The same styling. The same implicit message: "this is correct."
What I do: distinguish output visually based on certainty. High? Show it directly. Low? Frame it as a suggestion or a starting point. Sounds simple. Makes a world of difference.
2. Keep the human in the loop, but make it fast
The human has to stay involved. But if "staying involved" means you have to review every piece of output by hand, you lose the speed advantage you brought in AI for.
The trick: not "here is output, check everything." Instead: "here are three options, pick one and adjust the tone." AI writes product titles per platform. The specialist approves with one click or adjusts.
The strange thing: most AI tools are designed as if the human is no longer needed. While the human should actually be able to step in faster.
3. Show what the system did
Users do not need to know how a transformer works. They do need to understand where the output comes from.
At an automotive client we introduced this for AI-generated fitment data. Before, the team trusted the output blindly. Now they see on every record: "Source: TecDoc database, match confidence: 87%." That one label halved the number of errors that went live.
4. Design for correction, not for perfection
LLMs are never perfect in one go. Full stop.
The question is not "how do we make the AI error-free" but "how do we make correcting so easy that it is almost fun."
Rewrite with one click. Shift the tone from formal to casual. Put versions side by side. For teams working with Amazon, Bol.com, Channable and paid search that is not a luxury but a necessity.
5. Protect the moments where trust matters
At one of my clients an AI-generated answer about warranty terms went live through their chatbot. Subtle but factually wrong. One customer escalated. Legally recoverable, thankfully, but it cost three days and a lot of stress.
My rule: the higher the risk, the tighter the UX. More validation, stricter review, less room for interpretation.
Speed without control is not innovation. It is risk in a neat interface.
What I see in practice
Used well, LLMs help teams scale product content faster, structure search intent and support customer service. Not by replacing people, but by speeding up the repetitive work so people can focus where their expertise matters.
Used badly, you get generic content, factual errors and worst of all: extra review work that cancels out every bit of time saved.
At a marketplace team I saw AI product titles generated 3x faster. But without UX changes in the review process it led to 30% more correction time. Net, the team was slower than before.
The difference almost never lies in the model. It lies in the product thinking around it.
The teams that will win
The question is no longer whether AI can generate something useful. It can.
The question is whether you build an experience around it that also makes it reliable, understandable and commercially useful.
AI is not won by whoever builds fastest. It is won by whoever designs it best.
---
*Want to spar about how AI fits into your e-commerce operation? Get in touch for a no-obligation conversation.*
Want to spar about your marketplace strategy?
No hype. A sober look at where your growth is and where margin leaks away.
Get in touch