shipwithjev

LLM Cost Calculator: Jev vs GPT vs Claude vs Gemini

Set how many calls you make and how many tokens each one uses, and see what that costs per month on Jev, GPT, Claude and Gemini, cheapest first. Then check which model actually fits the task.

Start from a preset

300,000 calls / month · 123M tokens

Monthly cost, cheapest first

  1. Jev (jev-latest)

    Cheapest · typed answers only

    $5.04/mo

    $0.017 / 1K calls

  2. GPT-6 Luna

    OpenAI

    $13.50/mo

    $0.045 / 1K calls

  3. Gemini 3.5 Flash-Lite

    Google

    $43.50/mo

    $0.14 / 1K calls

  4. Gemini 3.8 Flash

    Google

    $101.25/mo

    $0.34 / 1K calls

  5. Claude Haiku 4.5

    Anthropic

    $135.00/mo

    $0.45 / 1K calls

  6. GPT-6.1 Sol

    OpenAI

    $270.00/mo

    $0.90 / 1K calls

  7. Claude Sonnet 5.5

    Anthropic

    $270.00/mo

    $0.90 / 1K calls

Prices as of Oct 2026, per 1M tokens, list price, no batch or caching discounts. A month is 30 days. Jev bills input only and returns typed answers, so output tokens don't apply to it.

Price table and sources
ModelInputOutput
Jev (jev-latest)$0.042$0.00
GPT-6 Luna$0.10$0.50
Gemini 3.5 Flash-Lite$0.30$2.50
Gemini 3.8 Flash$0.75$3.75
Claude Haiku 4.5$1.00$5.00
GPT-6.1 Sol$2.00$10.00
Claude Sonnet 5.5$2.00$10.00

Which model fits?

Cheapest isn't the point if the model can't do the job. Two questions.

What's the task?
How fast must it answer?

Use

Jev (jev-latest)

The answer is a label or a probability, not prose. Jev returns typed answers you can threshold in code, runs many questions per request in parallel, and only bills input tokens.

Escalate the few low-confidence cases to a generative model if you need a written explanation.

At your volume: $5.04/mo · $0.017 per 1K calls

How it works

  1. 01Pick a preset close to your workload, or type your own calls per day and tokens per call.
  2. 02Read the monthly cost for each model, sorted cheapest first, with the cost per 1,000 calls beside it.
  3. 03Answer two questions about your task and latency to see which model fits, and what it costs at your volume.
  4. 04Copy the link: your numbers are saved in the URL, so a teammate sees exactly what you see.

Questions

What is the cheapest LLM API?

It depends on what you need back. For yes/no, label and score decisions, Jev bills $0.042 per million input tokens and nothing for output, far below any generative model. For generating text, the small models (GPT-6 Luna, Gemini 3.5 Flash-Lite) are the cheapest at list price. Compare them on your own call shape above, since output-heavy work changes the ranking.

How do I estimate my monthly LLM API cost?

Multiply calls per day by tokens per call, split into input and output, then by each price per million tokens and by 30 days. Input is everything you send (system prompt, context, the user message); output is what the model writes back. Most teams underestimate input, because retrieved context and chat history add up fast.

Why are output tokens more expensive than input tokens?

Input tokens are processed in parallel in one pass, while output tokens are generated one at a time, each needing its own pass through the model. That makes output several times more expensive to serve, typically 4 to 8 times the input price. Models that answer with a typed value instead of text, like Jev, avoid most of that cost.

How many tokens is a word?

In English, roughly 0.75 words per token, so 1,000 tokens is about 750 words. Code, non-English text and numbers use more tokens per word. A short classification prompt is a few hundred tokens; a RAG prompt with retrieved documents is often several thousand.

Is GPT-6 Luna cheaper than Claude Haiku?

Yes, at list price. GPT-6 Luna is $0.10 input and $0.50 output per million tokens; Claude Haiku 4.5 is $1 and $5. Price per token is only half the answer: run a few of your real prompts through both and keep the cheaper one if its output holds up.

When should I not use a generative model?

When the output is a decision: spam or not, which team, how urgent, does this match. Prompting a chat model to write a label and parsing it back costs output tokens and adds a failure mode. A model built for typed judgments returns the probability directly, which you can threshold in code.

Related reading

More free tools