All posts
2 min readby Romiel Inolino

Gemini 4 Argon is cheap per token but uses twice the tokens per task. Budget your automations by the job

GeminiAI modelsLLM costsAI agentsautomation

Google's strongest model in months looks cheap on the price sheet. The per task math tells a different story.

What happened

Google announced Gemini 4 Argon, its first frontier model in more than seven months. For now it is limited to a group of cyber defense partners in Google's Fairwind program. Google says paying API customers and AI Ultra subscribers come next, with no date beyond "as soon as possible."

The introductory price is $2 per million input tokens and $10 per million output tokens, with cached input 95 percent cheaper. Output limits jump to one million tokens, and a new Long Decode Continuation feature lets very long answers pause and resume across requests instead of timing out.

Independent testing from Artificial Analysis, reported by The Decoder, puts Argon at 53 on its Intelligence Index, tied with GPT-6 Astra and behind Claude Opus 5.5 (58) and Sonnet 5.5 (56). Argon took first place on AutomationBench-AA at 77.5 percent and showed a low hallucination rate of 15 percent. The catch: it averaged 62,000 output tokens per task versus 27,000 for Astra. At the promo price one index task cost $1.99. Once the discount ends, The Decoder estimates $3.98, about 20 percent above Astra.

My take

The headline number for anyone running automations is not the token price. It is cost per completed job. A model that is much cheaper per token but writes twice as much can end up costing more, and that only shows up on your bill a month later.

What I would do before moving any client workflow to Argon once it opens up:

  1. Pick three real jobs from production, like lead triage, invoice extraction and a reply draft.
  2. Run each through your current model and Argon, logging tokens out, cost and pass rate.
  3. Watch latency. Long reasoning runs slow down anything a customer is waiting on.
  4. Keep the model name in config so swapping back is one change, not a rebuild.

The low hallucination rate is the most interesting part for CRM work, where a confident wrong answer in a contact record does real damage. Test that claim on your own data.

If you are not sure what your AI steps cost per job today, that is the place to start. Tell me which workflow worries you most at romielwillautomate.dev and I will help you measure it.

More posts