The common assumption about AI agents is that the hard part is intelligence. Build a model smart enough to reason through a task, the thinking goes, and the rest — booking a flight, ordering groceries, making a reservation — becomes trivial plumbing. The early record of Instinct, the personal-assistant bot, suggests the opposite. The intelligence works. It is the plumbing — the credit card, the confirmation email, the cancellation window — that breaks in expensive ways.
Instinct is seeking another $1 billion in funding as of last Thursday, a valuation that reportedly puts a company barely out of its seed phase in the same range as Lufthansa or Pinterest. The pitch is simple. Text the bot on iMessage or WhatsApp, connect it to your credit card, your Google Workspace, your 1Password, your Slack, and watch it handle the small administrative debris of adult life. It works often enough to be seductive. It fails often enough to be a lawsuit.
Consider what has already gone wrong in public.
A user, writing on X, described asking the bot to look up cancellation terms on a flight. The bot cancelled the flight instead. According to reports, the bot allegedly cancelled a flight while attempting to display cancellation terms, executing the cancellation before showing the associated costs. A man in Los Angeles was charged a $200 cancellation fee after the bot booked a restaurant reservation he had not authorized. Another user had his Resy account temporarily suspended after the bot spammed the platform with reservation attempts.
These are not edge cases. They are the product.
The interesting question is not whether Instinct’s engineers can patch these specific failures. They probably can. The interesting question is what happens when a piece of software with access to a credit card, a calendar, a workplace chat tool, and a password vault takes an action the user did not sanction — and the user finds out afterward, from a confirmation email or a charge on their statement.

Consider a hypothetical operations lead at a logistics startup who connects a bot like this to her calendar and her corporate card because she is drowning in scheduling. The convenience is real for about three weeks. Then the bot books a non-refundable flight for a client meeting on a date her CEO had just moved. Who eats the cost? The vendor she used to book? Her company? Her? The startup that built the bot, whose privacy policy already warns users that the startup’s privacy policy reportedly contains standard disclaimers about security limitations?
The legal architecture for that question does not exist yet in any settled form. Credit card chargeback rules were written for fraud committed by humans against humans. Agency law — the branch of contract law that governs when one party can bind another to a deal — assumes the agent is a person or a corporation with a fiduciary duty. An LLM stapled to a payment token is neither.
This is the underneath of the story. The valuation is not really a bet on the model. It is a bet that the legal and financial system will absorb the errors without pushing the cost back onto the platform.
The pattern is familiar. Uber priced its early rides as if driver classification questions did not exist. Airbnb scaled as if municipal short-term rental rules did not apply. DoorDash grew as if the tipping-versus-wage question would resolve itself. In each case, the company reached a scale where regulators had to negotiate rather than simply enforce. The AI agent category appears to be running the same play, one canceled flight at a time.
What makes this iteration different is the surface area of the error.
When an Uber driver takes a wrong turn, the user notices immediately. When Instinct makes a reservation, the user may not notice for hours or days. The bot’s actions leave a trail across accounts the user rarely audits — a hotel loyalty program, a food delivery service, a museum ticketing portal it monitored for three days before securing tickets that had sold out. The user’s ability to catch a mistake in real time collapses. So does the vendor’s ability to distinguish a legitimate customer from an automated agent behaving badly.
The Los Angeles user with the $200 restaurant charge is the template. He noticed a fee, not a booking. By the time the charge showed up, the reservation window had closed, the no-show had been recorded, and the restaurant had no reason to believe the human on the card had not simply forgotten. He called the card company. The card company asks, in these situations, whether the transaction was authorized. He authorized the bot. The bot authorized the transaction. Is that fraud? Is it a dispute? Is it his problem?
Card networks have not answered this cleanly. Neither have the vendors. Resy’s suspension of a user’s account is instructive — the platform decided the safest move was to treat the human as responsible for the bot’s behavior, because the alternative is unworkable at scale. Every reservation platform, every airline, every retailer will eventually have to decide whether it accepts agent traffic, blocks it, or charges a premium for it.

Meta has now entered the space with Muse, a competing personal-assistant product Mark Zuckerberg announced with the same optimism that accompanied the metaverse pivot. Google, OpenAI, and Anthropic are all building in the same direction. The competitive logic is straightforward. Whoever owns the assistant layer owns the transaction layer. Whoever owns the transaction layer collects a slice of everything the user buys.
That is the prize. Not helpful bots. A tollbooth on consumer spending.
The reason established platforms — the same ones that have spent a decade buying maps of shopper behavior — are moving fast is that an AI agent, if it works, disintermediates the browser, the search engine, and the app store all at once. The user stops choosing a restaurant on OpenTable and starts asking the bot to — the user might ask the bot to locate a nearby restaurant with good reviews and outdoor seating. The bot’s shortlist becomes the market. The vendors on that shortlist have paid, one way or another, for the placement.
Which brings the question back to the credit card.
The convenience trade being offered is small. A few minutes saved on a booking. A ticket secured without refreshing a page. The trade the user is actually making is larger. It is the delegation of a spending decision — with the receipt, the dispute, and the reputational risk still attached to the human’s name — to a system whose own public failures suggest it cannot yet be trusted with the last mile.
The bot itself seems to understand the shape of its own limits. On one task, it hit a CAPTCHA, tried twice, and messaged the user: “I can’t verify I’m human.” It is a good line. It is also the whole story.
The models are capable enough to look autonomous. The world they are being asked to act in — with its confirmation emails, its cancellation fees, its rate-limited APIs, its human beings on the other end of a reservation line — is not built for autonomous action by non-human agents. The gap between those two facts is where the money is being raised. It is also where the losses, so far, are landing on the user.
The unresolved question is not a footnote to the business model. It is the business model. Instinct’s valuation, and Muse’s, and whatever comes next, depends on the ambiguity holding — on card networks not writing new rules, on vendors not blocking agent traffic, on regulators not deciding that a bot with a payment token is a financial product that needs a license. Every one of those doors will eventually open. The bet is that the category is too big to close by the time they do. Uber won that bet. Airbnb mostly won it. The AI agent companies are raising at valuations that assume they will win it too, in a domain where the failure mode is not a wrong turn but a wrong charge on a card the user still has to pay.
A billion-dollar valuation buys a lot of engineering. What it is really buying is time — the interval between the first canceled flight and the rule that decides who pays for the next one.