Your AI bill is growing. Is your ROI?

Robert Cizmas
October 2, 2026
Blog

Someone asks an AI assistant to analyse a spreadsheet.

The first answer looks polished but misses two columns. The second includes them but invents an explanation for a variance. The third is shorter, although it drops an important caveat. By the fifth attempt, the answer is usable.

Five prompts. Five responses. One completed task.

The individual charge may still be tiny. Across hundreds of employees doing this every week, the pattern becomes harder to dismiss.

What does an AI answer actually cost?

Most commercial AI models charge for some combination of input tokens and output tokens. Input includes your instructions, documents, conversation history and any additional context sent to the model. Output includes the answer and, depending on the model, reasoning tokens used to produce it.

Prices vary considerably.

Anthropic currently lists Claude Sonnet 5 at $2 per million input tokens and $10 per million output tokens. Claude Opus 5.5 costs $4 and $20 respectively. Google’s paid Gemini tiers also vary by model and service level, with some current rates starting at $0.375 per million input tokens and $1.875 per million output tokens. Microsoft converts token usage into Copilot Credits, with premium generative AI tools consuming 10 credits per 1,000 tokens.

Caching, batch processing and smaller models can reduce these costs. Longer context, reasoning models, web grounding and repeated generations can increase them.

One answer is rarely alarming. Scale changes the calculation.

Take an illustrative task using a model priced at $2 per million input tokens and $10 per million output tokens. An 8,000-token input followed by a 2,000-token answer costs roughly $0.036.

Run it five times and the generation cost becomes $0.18.

Now give that pattern to 1,000 employees completing 20 AI-assisted tasks each month. One successful generation per task would cost about $720. Five attempts would cost around $3,600.

Those figures are illustrative, not a prediction of what any organisation will spend. Real costs depend on the models, prompts, context and contracts involved. They do show why a cheap interaction can become expensive when poor results are repeated at scale.

The token invoice only catches part of the waste

The person refining the prompt is also working.

They read each answer, compare it with the source material, explain the request again and search for claims that look suspicious. Sometimes a manager or subject expert checks the final version because nobody is confident enough to use it.

If each failed attempt adds twelve minutes, five attempts have consumed nearly an hour. The API cost might be measured in pence. The salary cost is not.

There is also a third cost: a plausible answer that reaches a client, board or customer before anyone notices the problem.

A false statistic in an internal draft creates more work. The same statistic in a client recommendation creates a rather different meeting.

This is why the most useful AI cost metric is not price per prompt. It is cost per accepted result.

Better prompting helps, but it does not solve everything

Clear instructions reduce avoidable retries. Giving the model the audience, purpose, source material, constraints and expected format usually produces a better first answer than asking it to “analyse this” and hoping for the best.

People improve when they practise on real work and receive useful feedback. A generic prompting workshop delivered six months ago is unlikely to help someone decide what context a financial analysis needs this afternoon.

Etiq’s Prompt Coach works inside the task. It helps the user improve the request before sending it, while the document, deadline and desired outcome are still in front of them.

That should reduce wasted generations. It also builds a skill the employee can use with Copilot, ChatGPT, Claude, Gemini or another approved assistant.

A strong prompt can still produce a weak answer, however. The model may misread the source, mishandle a calculation or produce a claim that sounds far more certain than the evidence allows.

Prompting improves the request. Verification examines the result.

Verification has a cost too

Verification is not computationally free. Checking claims, comparing an answer with its sources and validating data may involve additional model calls and processing.

Any honest calculation needs to include that consumption.

The commercial question is whether verification costs less than the work it prevents: repeated generations, manual checking, senior review, corrections and mistakes that escape into a decision.

For a simple email draft, verification may add little value. For a financial report, client recommendation, policy summary or research document, the equation changes. The cost of an unchecked error is higher, and the person reviewing the answer is often more expensive.

Etiq’s Verification flags unsupported claims, questionable calculations and points that require human judgement. It does not tell employees to trust the machine. It shows them where trust needs to be earned.

Over time, users learn which tasks need careful checking, which instructions produce more dependable answers and where human judgement remains essential.

AI security cannot be separated from AI cost

Repeated prompting sometimes creates another problem. Employees paste more information into the conversation each time, trying to help the model understand.

The first prompt contains a summary. The third contains part of the spreadsheet. The fifth contains the customer’s name, internal numbers and the paragraph the model kept getting wrong.

Better results should not require people to improvise their own data-handling policy.

Etiq can work with an organisation’s existing LLM infrastructure and API keys. It can be deployed in public cloud, private cloud or on-premise environments, depending on the organisation’s requirements.

For applied learning, Etiq can create a digital twin of company data structures and metadata, then generate representative synthetic data. Employees can practise realistic tasks without moving the original dataset into a training exercise.

The aim is to make good AI use compatible with the environment the organisation has already approved.

When does the software pay for itself?

The calculation should include more than reduced token consumption.

An organisation can estimate the monthly value through:

  • generation costs avoided by reducing unnecessary retries;
  • employee time recovered;
  • senior review time reduced;
  • rework prevented when weak outputs are caught earlier;
  • incidents avoided through better verification and data handling.

Then subtract the cost of verification, implementation and the Etiq platform.

This gives a more credible business case than promising that every employee will become 30% more productive after attending an AI workshop.

Etiq is designed to reduce wasted consumption and improve the quality of accepted outputs. We should measure whether it succeeds rather than treating the claim as automatic.

A useful benchmark would compare employees completing the same realistic tasks with and without Etiq. Measure the number of attempts, tokens consumed, time taken, corrections required and unsupported claims that survive into the final answer.

That would reveal whether the platform pays for itself for a particular team, task and model.

Start counting completed work

Prompt volume tells you that people are using AI. It does not tell you whether the work is useful, accurate or safe.

The next time someone presents an AI usage report, ask how many outputs were accepted without being regenerated. Ask how much time employees spent correcting them. Ask whether the claims were checked and whether sensitive data remained inside the approved environment.

If an answer takes five attempts and one senior reviewer, price all six.

That is the real AI bill.

‍