'AI-Powered' or Just an API Wrapper? How to Tell Before You Buy
Sit through three software demos this month and you will hear it three times: the product is "AI-powered". The phrase is doing a lot of work. It is in the price, for a start. And if you are the owner or the ops manager being asked to sign, you are entitled to a question the salesperson is not expecting: powered by whose AI, exactly?
Here is the open secret of the current software market. Many of the AI products being sold to small businesses contain no AI their maker built. They rent a model from one of a handful of vendors, the same models you or anyone else can rent, wrap it in an interface, add some prompts and some plumbing, and sell the result as a product. The industry calls these wrappers.
That word gets used as an insult, and this is where we part ways with most of what is written on the subject. A wrapper is not a scam by definition. A good one packages a general-purpose model into a specific job: it feeds the model the right context from your systems, catches its failures, fits the output into your workflow, and saves you building any of that yourself. That is legitimate work and it is worth money. The problem is that a bad wrapper looks identical in a demo, and the price tag rarely tells you which one you are holding.
The ingredient has a public price list
What makes this market unusual is that you can look up the cost of the main ingredient. The model vendors publish their prices per million tokens, a token being the unit models read text in. Anthropic, to take the vendor whose price list we checked while writing this, currently lists its mid-tier model at $2 per million input tokens and $10 per million output tokens, and the next tier up at $5 and $25. Its own rough estimate is that a token is about three quarters of an English word.
So price a concrete job. Say a tool drafts replies to customer emails: it reads about 500 words of email and instructions and writes about 150 words back. Call that roughly 700 tokens in and 200 tokens out. On the mid-tier model that is $0.0014 for the reading and $0.0020 for the writing, about a third of a cent per email. A thousand emails a month costs about $3.40 in raw model usage. Run the same numbers on the more expensive tier and a thousand emails still comes in around $8.50.
Now set that against the subscription you were quoted. If the tool costs $300 a month for that thousand-email workload, roughly $296 of your money is buying everything that is not the AI: the interface, the integration into your inbox, the prompt engineering, the error handling, the support desk, and the vendor's margin. That can be a perfectly fair trade. Software has always been priced above its ingredients, and nobody audits their accounting package by the cost of the electricity it runs on. But it does sharpen the question you are really answering in the demo, which is not "is the AI impressive". The model vendors made the AI impressive. The question is whether the layer on top is worth what the layer costs.
Five questions that expose the layer
You do not need a technical background to find out. You need five questions, asked in a normal voice, and some attention to how the answers land.
First: run it on our documents, today, in front of us. Not the demo data. A vendor whose product has real machinery around the model will treat this as routine, because the machinery exists precisely to handle unfamiliar input. A vendor whose product is a prompt in a trench coat will want to schedule a follow-up. Messy scans, the weird supplier invoice, the email written half in abbreviations: your ugliest real examples are the entire test.
Second: what does it do when it is wrong? Every model is sometimes wrong. A serious product has an answer that involves specifics: confidence thresholds, a review queue, validation against your other systems, a human sign-off step for anything above a money threshold. A thin product has an answer that involves the word "rarely".
Third: which model does this run on, and what happens when that model changes? The vendors behind these models update and retire them on their own schedule, not on yours. A product built as real infrastructure can swap one model for another and your workflow barely notices. A product that is one carefully tuned prompt against one specific model inherits every change that model vendor makes, and you find out at the same time the vendor does. You are not asking for the technical details. You are watching whether the question lands as normal engineering or as an accusation.
Fourth: where does our data go? The honest answers are specific: which third parties receive it, what they are contractually allowed to do with it, whether it is used for training, where it is stored. "It's all encrypted" is not an answer to the question you asked, and a vendor who cannot answer it about their own supply chain has told you how much of that supply chain they control.
Fifth: what number do you measure accuracy with, on data like ours? Real products carry real measurements, because their builders needed those measurements to build them. If nobody can tell you what percentage of invoices come through without human correction, the honest reading is that nobody has counted.
When the wrapper is exactly what you should buy
Keeping our own rule about honest answers: sometimes the thin product is the right purchase. If the task is generic, the volume is low, and the tool fits your workflow today, a $50-a-month wrapper that works is strictly better than a principled objection to wrappers. Drafting first-pass replies, summarising calls, tidying up text: buy the cheap thing, cancel it if it stops earning its keep, and spend your attention elsewhere. The markup over raw model cost is real, but at low volume the absolute numbers are noise.
The calculation flips when the workload is specific to you or the volume is not small. If the tool is reading your documents by the thousand, the per-unit markup compounds into real money. If the process it automates touches your scanners, your local server, your line-of-business database, then the vendor's generic layer was never going to fit anyway, and what you actually need is the model plus machinery built around your process rather than around the average customer's. That is the point where renting someone's wrapper stops making sense and building your own thin layer, which you then own outright, starts to. It is also worth saying that a good chunk of back-office automation needs no model at all: if the input is predictable, deterministic rules are cheaper, faster and easier to audit, which we wrote about in where AI actually pays for itself in an SMB back office.
That is where we sit in this market, so you know our angle: when a job genuinely needs a model, we build on the same public models the wrapper vendors rent, and we put the engineering into the part you are actually paying for, the machinery around it, wired into your systems and priced as a one-time build you own. When the job does not need a model, we say so and build the boring version.
If you are holding a quote for something "AI-powered" and you want a second opinion on what is under the lid, send us what the tool claims to do and we will tell you what we would ask, or whether the honest answer is that the subscription is fine. Details of how we scope and price custom work are on the custom software page.
Share this article
Juhász Ferenc
Founder & CEO