Quick answer — Every AI tool looks extraordinary for about twenty minutes. The gap opens later: the model you're locked into, the agent that still needs a human reading every diff, and the bill that scales with usage you can't predict. When we test these tools past the honeymoon, the same three failures show up — and none of them are visible in the demo.
There's a genre of content you've probably hit fatigue with by now: "I replaced my entire team with one AI agent." It's a good story. It's also, almost always, told by someone with something to sell — a course, a template pack, an affiliate link.
We test this software for a living, and our experience is duller and more useful: AI tools are genuinely good, and the marketing is describing a product that doesn't exist yet. Both things are true. The interesting question isn't "is AI overhyped" — it's which specific claims break, and when.
Worth watching
Two Minute Papers — one of the better AI research channels — dug into where today's models still fall down. It's a sharp, technical counterweight to the founder-influencer genre, and it pairs well with the buyer's-eye view below.
Their angle is the research. Ours is what happens when you actually put these tools on a company card. Here's where the gap opens.
1. The agent still needs a babysitter
This is the claim that breaks first, and it's the one the demos lean on hardest.
When we ran the same task set across the major AI coding tools for our Cursor vs Claude Code vs Windsurf comparison, the honest finding was the same for all three: none of them ships an agent that reliably verifies its own output. They are excellent at generating and editing code. They are considerably weaker at knowing when they got it wrong.
That's not a small caveat — it decides your workflow. An agent that writes ten files in ninety seconds hasn't saved you an hour if a human has to read all ten. The tools that respect this are the ones that survive: Cursor's Composer shows an inline diff before anything is written; Claude Code is the most autonomous, which is exactly why it's best pointed at well-scoped tasks rather than "build my app."
The same applies outside coding. Every self-hosted agent in our directory — OpenClaw, Hermes Agent, OpenHands — carries the same trade-off in its cons list: real actions need real guardrails. That's not us being cautious. It's what testing them tells you.
The buying rule: ask what happens when the agent is wrong, not what happens when it's right. The demo only ever shows you the second one.
2. You're not buying a tool, you're buying a model
The pricing page says $20/month. It doesn't say which model you're marrying.
Claude Code runs on Claude, full stop. That's a defensible trade — the agent is deeply tuned to one strong model — but it is a trade, and it's invisible until the day a better model ships somewhere else. Cursor supports GPT, Claude, Gemini and your own keys. Windsurf sits in between. Same category, same price bracket, completely different lock-in profile.
This is the most under-discussed decision in AI tooling right now. Model quality is moving faster than any other part of the stack, which means the cost of being locked to one provider compounds faster than the cost of the subscription. A tool that's second-best today but lets you swap models is often the better 12-month bet than a tool that's best today and can't.
It's also why the open-source column in our guides keeps growing. Tools you can point at any model — or run locally via Ollama — aren't just a privacy play. They're an insurance policy against a market that reprices every few months.
3. The bill scales with usage you can't predict
Traditional SaaS bills per seat. You know your headcount, so you know your bill.
AI tools bill per usage — tokens, credits, requests, "fast premium requests" — and usage is the one number you cannot forecast before you've deployed. This is the same structural trap we mapped in how to choose email marketing software, where per-contact pricing quietly punishes list growth. AI pricing is that trap with a faster clock.
We hit a mild version of it ourselves: on monday.com, automation limits are tier-gated, and teams routinely blow through them running rules nobody remembers creating. The AI equivalent is worse, because a single agent left on a loop can spend real money while you're at lunch.
The buying rule: model your cost at 3× your current usage before you commit, and find the ceiling — the plan where you stop paying per unit — before you need it.
What actually holds up
Strip out the hype and there's a genuinely good product underneath. The claims that survive our testing:
- Autocomplete and inline editing. Boring, constant, real. Cursor's Tab is the strongest thing in this category, and it's the least exciting feature in any demo.
- Scoped, autonomous work. "Add Stripe checkout to this repo, open a PR" — a real ticket, a reviewable diff. This works today.
- Reading unfamiliar code. Whole-repo reasoning genuinely shortcuts the worst part of joining a codebase.
- The drudge work. Summaries, first drafts, tests, migrations. Unglamorous, and where most of the actual hours come back.
Notice what's missing: none of these is "it runs the business." They're all a competent assistant doing well-defined work under review. That's the product. It's worth paying for. It just doesn't make a good thumbnail.
How to buy AI tools without getting burned
- Test past the demo. Run your real task, not the vendor's. The first twenty minutes tell you nothing.
- Price the lock-in, not the subscription. Ask what happens when a better model ships next quarter.
- Find the ceiling. Know your cost at 3× usage before you're at 1×.
- Keep a human on the diff. Any tool asking you to skip that step is selling the story, not the software.
- Check for an exit. If your prompts, context and history can't leave, you're not a customer — you're a hostage.
The verdict
The people telling you AI changes everything are right. The people telling you it changes everything this quarter, with no supervision, for $20 are selling something.
Our directory has 400+ tools, and the AI ones are among the best software we've tested. They're also the ones where the gap between the marketing and month three is widest. Buy them — just buy them with your eyes open, and price the trade-offs the demo left out.
Want the honest version of this every week, before the hype cycle catches up? Join the ChooseMyStack Weekly Brief — one email, curated and filtered.