Field Notes: AI has a UX problem

Field Notes are thoughts and musings from building Moselle — real-time shower thoughts, building out loud, and saying the things others are already thinking. Very opinionated, but grounded in what I’m seeing in the field. That felt worth sharing out to the universe. No more, no less.

And a quick level set before we start: I think LLMs are one branch of AI, not the whole thing. Treating them as the whole thing is a disservice to other elements of AI, like linear programming, time-series forecasting, and constraint programming. But most of the “AI” being marketed today is really just LLMs. So when I say AI in this post, I really mean LLMs.

A few weeks ago I caught up with a founder friend over drinks. She runs a consumer brand, and lately, founders building AI products for commerce businesses keep reaching out to her — not to pitch her, but to ask how to sell to someone like her. Somewhere into the second glass she flipped the question on me: why is Moselle having success approaching brands while everyone else she talks to is struggling?

I’ve been chewing on my answer ever since. The honest version is that we stopped fighting user behavior and started designing around it. The longer version is this post, and it starts with two things I believe about this moment. The AI revolution is real: look at a benchmark like Finance Agent v2 and watch frontier models climb on core financial-analyst tasks. Top scores hover around 60%, though under the benchmark’s stricter all-pass scoring, where every detail must be right, none clears 51% (hold that thought). And the models are no longer the moat. Every serious player has access to roughly the same intelligence. What the technology can’t do, no matter how many parameters you throw at it, is change user behavior.

That’s why the headlines are so contradictory. Some companies report massive AI spend and consumption, while MIT’s “GenAI Divide” report found that 95% of enterprise AI pilots showed no measurable P&L impact. Both can be true. Building Moselle, a planning platform for fast-growing consumer brands, has given me a front-row seat to why. What I see on the ground is three problems, and none of them are about the models.

Problem 1: “Trust me bro” is not an adoption strategy

Users are ingrained in doing the work themselves. Their tools, their routines, their sense of professional competence — all of it is designed around them producing the output. Telling someone that AI will now do that work isn’t a feature announcement; it’s a massive UX shift, and we’re mostly asking people to make it on faith. On top of that, the industry’s default interface for all this power is an empty chat box. Type what you want the AI to do. That’s the product.

An empty chat box is a black hole from a UX perspective. It’s the blank Word document problem: you can write anything, which is exactly why most people write nothing. A blinking cursor doesn’t tell you what the system is good at, where its edges are, or what to try first. It hands the user all the ambiguity and none of the guidance.

The non-deterministic nature of AI makes this worse. When the same question can produce slightly different answers, users don’t experience “creative flexibility.” They experience uncertainty, and uncertainty reads as untrustworthy. Closing that gap takes heavy education on when to trust AI and when not to, and that can’t be solved overnight. Even platforms built AI-first still require users to develop judgment about when to hand work to the AI and when to step back. We see this friction at Moselle all the time, with customers who want to adopt.

Human nature being what it is, adoption will be gradual. But I’d argue the interfaces are a big part of why the rollout feels slow, and it could accelerate sharply the moment someone ships the killer app whose interface isn’t purely chat. Until then, gradual is still fast by any historical standard, but not fast enough from a VC perspective. Which means the name of the game isn’t speed of adoption. It’s retention, and consumption economics that prove the work is genuinely cheaper and better than what a human can do. That’s a proof you earn over quarters, not demos.

Problem 2: The commerce ICP isn’t the early adopter

Coding tools exploded first for an obvious reason: founders and engineers are the natural early adopters. They’re building the thing — of course they’re going to sell it to their friends. If your addressable market is startups and tech companies, the early-adopter pool is enormous.

The numbers back this up. In the Anthropic Economic Index, computer and mathematical work (largely software engineering) accounts for about 37% of all Claude usage, despite those roles being a sliver of the workforce. Meanwhile the Census Bureau’s Business Trends and Outlook Survey puts AI adoption in retail trade around 14%, near the bottom of the table, with professional services and finance at the top. The Fed’s write-up has a good chart of adoption by sector. Engineers are sprinting; commerce is walking.

Step outside the tech ecosystem and the math changes. Commerce brands barely understand the underlying tech (Python, SQL, programmatic tool calling), and no offense intended, because they shouldn’t have to. Talk to an average consumer brand and you’ll find a team highly reliant on SaaS, most of it bought precisely so they never have to think about what’s underneath. The market you can actually sell to shrinks to the small set of forward-thinking companies, because established brands are unlikely to be early AI adopters. Is that fair to say? Every founder I’ve compared notes with who sells AI outside of tech says the same thing.

If you want them anyway, the value prop has to be compelling and end-to-end: we can make your returns cheaper, we can get you better ROAS. A pitch like “we can help you with data and analytics using AI” is a setup for failure.

There’s a silver lining: abstraction can remove the need to understand the tech at all. But that just moves the problem, because validity of work becomes the new UX challenge. In a spreadsheet you can click into a formula and see how the numbers add up. AI does the same work inside a black box — and then asks for your trust. Customers can’t extend it unless you export the work in a format they already know, like a spreadsheet with the formulas built in, so they can check it the way they’ve always checked things.

And if the ICP isn’t the early adopter, that reinforces the long game — the game most VCs can’t stomach, but the one where the returns will be unparalleled.

Problem 3: Sabotage, misuse, and the modern Luddite

Not everyone is ready for this future of work, and some users, intentionally or not, work against it. There’s a pattern I’ve watched play out: a reluctance to lean on the LLM for tasks it’s designed for, paired with going out of the way to use it for tasks it isn’t, then declaring “see, it doesn’t work.”

A contrarian on your platform isn’t all bad: they’ll surface edge cases you should handle eventually. But these users create pressure to over-iterate on edge cases, and that’s a real drag on product velocity. You end up debating how much is a valid unsolved use case versus AI failing expectations simply because this was never an AI use case. It’s the Luddites of the industrial revolution breaking the machines because they were worried about automation — understandable in the moment and, in hindsight, on the wrong side of the arc.

Financial reporting is the sharpest example. AI is great at scaffolding a year-to-date summary by SKU. But the right architecture is to use AI to build the data asset, then answer the question deterministically: rerun the same query, get the same number. Because once you account for FX, variable selling prices, and a dozen other factors, passing the raw question to an LLM over and over will produce inconsistent results. And inconsistent financial results don’t read as “wrong tool for the job.” They read as “the technology doesn’t work.” That’s the moment a doubter points at, gives up, and churns. The problem gets misattributed to AI failure rather than misuse.

So where do we go from here?

If adoption is slow, you do the next best thing as a founder: triple down on the customers you already have. That means demonstrating that your AI product compounds. Build the flywheel where a brand finds value in the AI, uses it, finds more ways to use it, spends more, and the cycle repeats. Here’s the way I think about it: AI does a unit of work, and as anyone who’s run a business knows, there is an endless amount of work. An agent that reliably turns a dollar of spend into a unit of finished work is the path to venture scale, even in a market that adopts slowly. A consumer brand will stick with you if you can show quarter-over-quarter improvements to their actual business outcomes.

AI adoption will follow the same arc as every wave before it. The internet was before my time, but I’ve watched this play out with big data and mobile: some technologies just take time to mature. There’s always quick action inside the blast radius, but the compounding effects creep outward slowly, and the companies building platforms and defining the next industry best practices won’t be chasing the wave. They’ll be leading it.

The current tooling makes my case for me. Claude, Codex, MCP setups — for the average commerce user, it’s all too complex. Think about what we’re asking a customer to do: configure MCP servers, write project descriptions, set tool permissions correctly, write and test prompts, and understand that when the LLM does heavy work it’s writing little scripts under the hood. Influencer demos of ad creation or forecast building with Claude, some MCPs, and a spreadsheet look compelling. What actually plays out: customers start the journey, struggle, and quietly revert to familiar routines.

None of this is doom and gloom. It’s an argument for patience, and for continued investment in the right AI infrastructure. The platforms genuinely building toward the future of work will see adoption grow, the way it eventually grew for every prior wave. User behavior will conform, as it always has. It just won’t do it on an AI company’s clock. Sorry, Sam & Dario.