Blog/Model Comparison/

Cheap LLM for AI Agents Benchmark: Pro vs Flash, 22 Models

Cheap LLM for AI agents benchmark: 22 models in 7 tier families, 17 tool tasks, 3 runs each. GPT-6 Luna scored 51/51 at about 1/20 of GPT-6.1 Sol's cost.

Cheap LLM for AI agents benchmark cover comparing Pro and Flash model tiers; editorial artwork, not a test result

GPT-6 Luna costs $0.10 per million input tokens. GPT-6 Astra, from the same vendor, costs $10. On our 17 agent tool tasks, run three times each, Luna got all 51 runs right and Astra got 50. Astra’s one miss was T2: it copied a like count of 2,780,636 as 2,782,636.

That result is the core of this cheap LLM for AI agents benchmark. We took the harness from our 12-model tool-calling benchmark and ran 10 more models through it on 2026-10-03 (UTC), so most vendors now have two to four tiers on the same test: Flash, standard, Pro, Prime, Max. The question is narrow. When the task is “call a data tool, read the result, answer in JSON”, does paying for the bigger tier buy you anything?

Key takeaway

  • GPT-6 Luna scored 51/51 at a median $0.00035 per task, about 1/20 of GPT-6.1 Sol’s $0.0069. Sol, Sol Pro and Luna all scored 51/51; GPT-6 Astra, the most expensive model tested, scored 50/51.
  • A higher tier clearly helped in two of six families. Claude Opus 5.5 scored 51/51 where Sonnet 5.5 scored 44/51, and Alibaba went from 38/51 (Qwen3.7 Flash) to 49 (Qwen3.8 Max 0902) to 51 (Max Prime). In the OpenAI, Xiaomi and Z.ai families, the top tier was no better than a cheaper one.
  • GPT-6.1 Sol Pro has the same list price as Sol ($2/$10) but used a median 9,026 input tokens per task against 2,887, so it cost about 3x more.
  • The cheapest model that passed everything is not the cheapest model. Qwen3.7 Flash had the lowest cost per correct answer ($0.00033) but failed 13 of 51 runs, 10 of them by truncation.
  • Scope: one task family (Chinese social data lookups plus aggregation), N=3, default settings, cost from usage × list price. It says nothing about long reasoning, coding or long documents.

The tier families and their prices

Prices are SandBase list prices on 2026-10-03, read from prompt_token_price / completion_token_price and price_formula in GET https://api.sandbase.ai/v1/models/<id> at about 11:16 UTC. Those fields are per-token USD values; the table converts them to dollars per million tokens. All 22 models showed enabled: true.

FamilyTiers tested (in / out)
OpenAIGPT-6 Luna $0.10 / $0.50 · GPT-6.1 Sol $2 / $10 · GPT-6.1 Sol Pro $2 / $10 · GPT-6 Astra $10 / $50
AnthropicClaude Sonnet 5.5 $2 / $10 · Claude Opus 5.5 $4 / $20
Z.aiGLM-5.3 Flash $0.15 / $0.50 · FlashX $0.37 / $1.25 · GLM-5.3 $1.40 / $4.40 · Prime $2.80 / $8.80
XiaomiMiMo v2.6 Flash $0.14 / $0.28 · MiMo v2.6 Pro $0.435 / $0.87
DeepSeekV4.1 Flash $0.30 / $1.20 · V4 Pro 0813 $1.32 / $3.96
AlibabaQwen3.7 Flash $0.03 / $0.13 · Qwen3.8 Max 0902 $2 / $6 · Qwen3.8 Max Prime $4 / $12
GoogleGemini 3.8 Flash $1.50 / $7.50 (no higher tier tested)
Single tierGrok 4.7 $2 / $6 · Kimi K3 $3 / $15 · MiniMax M3 $0.30 / $1.20 · Seed 2.1 Pro $0.89 / $4.45

Six models switch to a higher rate above 200K–272K prompt tokens; our largest run used 59,630 input tokens across all its turns, so every request stayed on the base rate.

SandBase language model catalog searched for gpt-6-luna, showing OpenAI GPT-6 Luna with 1.1M context, $0.1 input and $0.5 output per million tokens

Caption: The SandBase catalog lists openai/gpt-6-luna at $0.10 input and $0.50 output per million tokens, the rate behind its $0.00035 median cost per task (captured 2026-10-03).

“Flash” is a name, not a price band. Gemini 3.8 Flash lists at 15x GPT-6 Luna’s rates, and in this test it cost more per task than GLM-5.3 Flash, MiMo Flash and DeepSeek V4.1 Flash combined.

Test setup: same harness, same tasks, same cache

Only the model list changed from the 12-model run:

  • Six tools wrapping SandBase public data endpoints for Douyin, Weibo and Xiaohongshu. Tool responses are cached by tool name and arguments and replayed to every model, so all 22 saw identical data. Expected answers are computed from that cache in Python.
  • 17 scored tasks: T1 to T10 (lookups, an average, a date count) and H1 to H6 plus H8 (rankings, sums of 10 to 20 numbers, filtered lists, chained lookups). H7 stays excluded because our own task was ambiguous. Full task list: the 12-model article.
  • Three runs per task, 1,500 output tokens per turn, up to 10 turns, default temperature and reasoning, no prompt caching.
  • Claude Opus 5.5, like Sonnet 5.5, used SandBase POST /v1/messages; the other 20 used POST /v1/chat/completions with OpenAI function tools.

The system prompt, verbatim, for all 22 models:

You are a research agent with tools for public Douyin, Weibo and Xiaohongshu data. Use the tools to answer; never guess numbers. If something can't be found, say so. Finish with ONE line that starts with FINAL: followed by a compact JSON object with exactly the keys requested.

The grader parsed the JSON after FINAL: and compared each requested key: integers after stripping commas, booleans after lowercasing, IDs and names as exact strings, lists element by element in order. A missing key, unparseable JSON or no FINAL: line was a miss. Extra keys were ignored. We don’t publish the cached tool responses (third-party content about real accounts); the task list, prompt, grader and per-run aggregates are enough to rebuild the test.

One difference matters. The original 12 models ran from 07:36 to about 08:16 UTC; the 10 new ones ran from 10:22 to 10:42 UTC, all in parallel. The data is identical, but treat cross-set latency gaps of a few seconds as noise. DeepSeek V4.1 Flash and Qwen3.7 Flash also made 11 profile requests (10 distinct accounts) that weren’t in the cache; those went to the live API during the run. No expected answer depends on them.

Results: 22 models, grouped by family

These are our own benchmark results on our own task set, not vendor-certified performance claims. Median cost is per task. Cost per correct answer is a model’s total spend over its 51 scored runs divided by the number it got right, so failures count against it.

ModelT (30)H (21)Total (51)Median cost / taskCost per correctMedian timeTruncated
OpenAI GPT-6 Luna302151$0.00035$0.000387.6 s0
GPT-6.1 Sol302151$0.0069$0.00709.5 s0
GPT-6.1 Sol Pro302151$0.0213$0.022011.9 s0
GPT-6 Astra292150$0.0344$0.035213.6 s0
Anthropic Claude Sonnet 5.5271744$0.0130$0.016813.6 s0
Claude Opus 5.5302151$0.0285$0.030410.1 s0
Z.ai GLM-5.3 Flash301949$0.00068$0.0008021.4 s2
GLM-5.3 FlashX292150$0.0018$0.00199.8 s0
GLM-5.3292049$0.0066$0.007221.5 s1
GLM-5.3 Prime301949$0.0130$0.01579.8 s2
Xiaomi MiMo v2.6 Flash292150$0.00060$0.000635.8 s0
MiMo v2.6 Pro302050$0.0018$0.002012.7 s1
DeepSeek V4.1 Flash291948$0.0015$0.00186.3 s0
V4 Pro 0813282149$0.0069$0.007310.2 s0
Alibaba Qwen3.7 Flash261238$0.00025$0.000339.7 s10
Qwen3.8 Max 0902301949$0.0112$0.012512.3 s2
Qwen3.8 Max Prime302151$0.0213$0.02408.4 s0
Google Gemini 3.8 Flash271946$0.0058$0.00696.3 s0
Single tier Grok 4.7302151$0.0139$0.01488.7 s0
Kimi K3302151$0.0159$0.01678.5 s0
MiniMax M3282048$0.0019$0.002110.2 s1
Seed 2.1 Pro291948$0.0058$0.006717.1 s3

Across 1,122 scored runs, no model made a schema-invalid tool call, and all 22 said the non-existent account in T9 didn’t exist in all three runs. There were 48 misses in total; 26 of them came from the 10 new models, and 13 of those 26 were Qwen3.7 Flash.

What the bigger tier bought, family by family

OpenAI: nothing on this task set. Luna, Sol and Sol Pro all scored 51/51; Astra scored 50. Luna and Sol used the same median input (2,887 tokens) and similar output (139 vs 110), so the cost gap is almost exactly the price gap: Sol was 19.6x Luna’s median. Astra used the same tokens as Sol at 5x the rate and cost 5x as much.

Sol Pro: same price per token, three times the bill. Sol Pro lists at $2 / $10, identical to Sol. But its median input was 9,026 tokens per task against Sol’s 2,887. On T1 the gap was the same in every run: 10,374 input tokens against 2,788, for the same three model turns and two tool calls. Output was also higher (253 vs 110). Its median cost was $0.0213, 3.1x Sol, and it was slower (11.9 s vs 9.5 s). We didn’t isolate where the extra input tokens come from; usage reports them and the price formula bills them.

SandBase language model catalog searched for gpt-6.1-sol-pro, showing GPT-6.1 Sol Pro and GPT-6.1 Sol both at $2 input and $10 output per million tokens

Caption: GPT-6.1 Sol Pro and GPT-6.1 Sol share the same $2 / $10 list price in the SandBase catalog, so Sol Pro’s 3x cost per task comes from tokens, not rates (captured 2026-10-03).

Anthropic: the one family where the upgrade paid for itself on accuracy. Sonnet 5.5’s seven misses were all arithmetic over tool output: two T5 averages, one T7 date count, three H2 filtered lists that included videos below the average, and one H3 sum off by 10,000. Opus 5.5 got every run of those four tasks right. It used about the same tokens as Sonnet (median 5,041 in, 436 out, against 4,932 and 415), so its cost was roughly the 2x price ratio: $0.0285 against $0.0130. Its slowest run took 24 s; Sonnet’s took 259 s. If your agent has to use Claude, Opus is the safer pick here; otherwise Luna scored the same 51/51 at about 1/80 of Opus’s median cost.

Alibaba: the floor is real. Qwen3.7 Flash is the cheapest model per token here, and it shows. Ten of its 13 misses were truncations: it hit the 1,500-token cap on T5 (3 runs), H2 (3), H8 (2), H3 and H6, with a median output of 949 tokens per task, the highest of all 22. The rest were a sum off by 20,000 (H3), a wrong total (H4) and a follower count taken from a profile call instead of the search result T10 asked for. Qwen3.8 Max 0902 scored 49 and Max Prime 51, at 45x and 86x Flash’s median cost.

Z.ai: FlashX beat Prime. GLM-5.3 FlashX scored 50/51; its only miss was a T6 run that stopped after 208 tokens without a FINAL line. GLM-5.3 Flash scored 49, both misses H2 truncations, and it was the slowest new model: median 21.4 s, slowest run 77 s. GLM-5.3 and GLM-5.3 Prime each scored 49. Prime cost 7x FlashX per task and scored one fewer.

Xiaomi: Flash matched Pro. MiMo v2.6 Flash scored 50/51 at $0.00060 per task, a third of Pro’s $0.0018, and had the fastest median of all 22 models at 5.8 s. Its miss was a T5 average of 418,764 against 422,764, off by exactly 4,000. Pro’s miss was an H2 truncation.

DeepSeek: one extra point for 4.5x the cost. V4.1 Flash scored 48 against V4 Pro 0813’s 49. Two of Flash’s misses took numbers from the wrong source: on H6 and T10 it used profile follower counts (T10: 1,401,200) instead of the search results the tasks named (1,399,062). The third was an H2 list with three of the four IDs.

Gemini 3.8 Flash scored 46/51: T5 answered 422,739 in all three runs (correct: 422,764), plus extra videos in two H2 lists. Fast (median 6.3 s), but 16.5x Luna’s median cost for five more misses.

Cost per correct answer, and why it isn’t enough

At the medians, 1,000 tasks a day like these would cost about $0.35 on GPT-6 Luna, $0.60 on MiMo v2.6 Flash, $6.90 on GPT-6.1 Sol, $28.50 on Claude Opus 5.5 and $34.40 on GPT-6 Astra. That extrapolates from 51 short runs per model.

Cost per correct answer counts failures, but it can still mislead. Qwen3.7 Flash has the lowest figure ($0.00033) because it is so cheap that 13 failures barely move it. In production each failed run still has to be caught, retried or escalated.

Two models were on a catalog promotion when we checked the catalog and billing between 11:22 and 11:42 UTC, after the runs: Claude Opus 5.5 at 25% off ($3 / $15) and GPT-6 Astra at 35% off ($6.50 / $32.50). The GET /v1/models price fields still showed list price. A one-line Astra request in that window was billed $0.000215 for usage that costs $0.00033 at list price, which matches the 35%. We don’t know whether the promotion applied during the 10:22 to 10:42 runs, so every table here uses list price. At the promotional rates, Opus’s median would be about $0.0214 and Astra’s about $0.0224, still more than 60x Luna’s.

SandBase language model catalog searched for claude-opus-5.5, showing Claude Opus 5.5 with a 25% off badge at $3 input and $15 output, list price $4 and $20 struck through

Caption: On 2026-10-03 the SandBase catalog showed Claude Opus 5.5 at 25% off ($3 / $15 against a $4 / $20 list price); this article’s cost figures use the list price (captured 2026-10-03).

SandBase language model catalog searched for gpt-6, showing GPT-6 Astra with a 35% off badge at $6.5 input and $32.5 output, list price $10 and $50 struck through, above GPT-6.1 Sol at $2 and $10

Caption: The same catalog showed GPT-6 Astra at 35% off its $10 / $50 list price, while GPT-6.1 Sol stayed at $2 / $10 (captured 2026-10-03).

The 10 new models used $5.03 of tokens at list price across 510 runs. The data endpoints are Free in the catalog (base_price 0 on douyin/app-v3/user-post-videos, checked 2026-10-03).

How to choose a tier for your agent

The data suggests a routing rule rather than a single winner: run the cheapest tier that passes your own task set as the default, and keep a stronger model as the fallback.

If you needDefaultFallbackWatch out for
Lowest cost, OpenAI-compatibleGPT-6 Luna (51/51, $0.00035)GPT-6.1 Sol (51/51, $0.0069)One task family tested; check yours
Lowest cost, a second vendorMiMo v2.6 Flash (50/51, $0.00060)MiMo v2.6 Pro or GPT-6.1 SolOne T5 arithmetic miss; keep sums in code
Z.ai modelsGLM-5.3 FlashX (50/51, $0.0018)GLM-5.3 PrimeGLM-5.3 Flash was slow here (median 21.4 s)
Claude-specific featuresClaude Opus 5.5 (51/51, $0.0285)none needed on this setSonnet 5.5 missed 7 runs on hand arithmetic
Absolute minimum token priceNot Qwen3.7 Flash for multi-step tasksQwen3.8 Max Prime10 of 51 runs truncated at 1,500 tokens

GPT-6 Astra and GLM-5.3 Prime, the most expensive options in their families, scored no higher than a cheaper sibling. The upgrades that did help, Opus over Sonnet and Qwen3.8 over Qwen3.7 Flash, fixed concrete failure types: hand arithmetic and truncation. Move the arithmetic into code, as our Sonnet vs GPT-6.1 Sol test recommended, and part of that gap goes away. For coding agents the trade-off looks different: retries are only cheap when a test can reject a bad patch, which our Gemini 3.8 Flash retry-cost analysis works through.

A cheap-tier-first agent with a fallback

On SandBase, one API key reaches all 22 models, and 20 of them share the OpenAI-compatible Chat Completions contract, so a failover order is just a list of model ids. The program below tries GPT-6 Luna first and moves to GPT-6.1 Sol only when a run fails outright: HTTP error, finish_reason: length, no FINAL line, unparseable JSON or the turn limit. It doesn’t check whether the answer is right; that’s your task set’s job. The tool computes the statistics in Python and returns only aggregates. aweme_list and statistics.digg_count are field names observed in our 2026-10-03 responses, not documented guarantees, so they’re read with .get().

import os
import re
import json
import requests

BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
# Cheapest tier first, stronger tier as fallback. $/M input, output (SandBase list price, 2026-10-03).
TIERS = [("openai/gpt-6-luna", 0.1, 0.5), ("openai/gpt-6.1-sol", 2, 10)]
SYSTEM = "Use the tools to answer; never guess numbers. Finish with one line: FINAL: {json}."


def call_data_api(model: str, params: dict) -> dict:
    """Call a SandBase data endpoint and return outputs[0].data, failing loudly otherwise."""
    resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
    resp.raise_for_status()
    body = resp.json()
    outputs = body.get("outputs") or []
    if body.get("status") != "completed" or not outputs:
        raise RuntimeError(f"SandBase call did not complete: {body.get('status')} {body.get('error')}")
    return outputs[0].get("data") or {}


def douyin_video_stats(sec_user_id: str) -> dict:
    """First page of an account's videos, with the aggregates computed here, not by the model."""
    data = call_data_api("douyin/app-v3/user-post-videos",
                         {"sec_user_id": sec_user_id, "max_cursor": 0, "count": 20})
    likes = [(item.get("statistics") or {}).get("digg_count") or 0  # observed field names
             for item in data.get("aweme_list") or []]
    if not likes:
        return {"videos": 0, "found": False}
    return {"videos": len(likes), "likes_sum": sum(likes), "likes_mean_floor": sum(likes) // len(likes)}


TOOLS = [{"type": "function", "function": {
    "name": "douyin_video_stats",
    "description": "Like statistics for the first page (up to 20) of a Douyin account's videos.",
    "parameters": {"type": "object", "properties": {"sec_user_id": {"type": "string"}},
                   "required": ["sec_user_id"]}}}]


def run(model: str, task: str, max_tokens: int = 1500) -> tuple[dict, list[int]]:
    """One agent run. Returns (parsed FINAL json, [prompt_tokens, completion_tokens]); raises on failure."""
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": task}]
    usage = [0, 0]
    for _ in range(6):
        resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=240,
                             json={"model": model, "max_tokens": max_tokens,
                                   "tools": TOOLS, "messages": messages})
        resp.raise_for_status()
        body = resp.json()
        usage[0] += body["usage"]["prompt_tokens"]
        usage[1] += body["usage"]["completion_tokens"]
        choice = body["choices"][0]
        msg = choice["message"]
        messages.append({k: v for k, v in msg.items() if k in ("role", "content", "tool_calls")})
        if choice.get("finish_reason") == "length":
            raise RuntimeError(f"output cut at max_tokens={max_tokens}")
        calls = msg.get("tool_calls") or []
        if not calls:
            found = re.search(r"FINAL:\s*(\{.*\})", msg.get("content") or "", re.S)
            if not found:
                raise RuntimeError("no FINAL line")
            return json.loads(found.group(1)), usage  # JSONDecodeError also triggers the fallback
        for call in calls:
            args = json.loads(call["function"].get("arguments") or "{}")
            messages.append({"role": "tool", "tool_call_id": call["id"],
                             "content": json.dumps(douyin_video_stats(**args))})
    raise RuntimeError("turn limit reached")


def run_with_fallback(task: str) -> dict:
    """Try each tier in order; move to the next one only when a run fails outright."""
    for model, price_in, price_out in TIERS:
        try:
            answer, (tok_in, tok_out) = run(model, task)
        except (RuntimeError, ValueError, requests.RequestException) as err:
            print(f"{model} failed: {err}; trying next tier")
            continue
        cost = tok_in / 1e6 * price_in + tok_out / 1e6 * price_out
        return {"model": model, "answer": answer, "tokens": [tok_in, tok_out], "cost_usd": round(cost, 6)}
    raise RuntimeError("all tiers failed")


if __name__ == "__main__":
    task = ("For the Douyin account with sec_user_id "
            "MS4wLjABAAAA8U_l6rBzmy7bcy6xOJel4v0RzoR_wfAubGPeJimN__4, what is the average like count "
            "on the first page, rounded down? Keys: videos, average_likes.")
    print(json.dumps(run_with_fallback(task), ensure_ascii=False))

Tested on 2026-10-03 (UTC). We ran the program above verbatim at 13:29 UTC. Input: the public People’s Daily Douyin account used in T2, T5 and T7. Data request: POST https://api.sandbase.ai/v1/api/douyin/app-v3/user-post-videos with {"sec_user_id": "MS4wLjABAAAA8U_l6rBzmy7bcy6xOJel4v0RzoR_wfAubGPeJimN__4", "max_cursor": 0, "count": 20}. Trimmed response from that run:

{"id": "82db1811-8b65-4d7d-a5cb-3f0e323a7edf", "status": "completed",
 "model": "douyin/app-v3/user-post-videos",
 "outputs": [{"data": {"aweme_list": [
   {"aweme_id": "7692418769125199138", "statistics": {"digg_count": 6413}},
   {"aweme_id": "7692417622113111323", "statistics": {"digg_count": 2743}}],
  "has_more": 1}}]}

First 2 of 20 items in aweme_list shown; inside each item only aweme_id and statistics.digg_count are kept, and all other keys of the items, data and the envelope are omitted. The envelope (id, status, model, outputs[0].data) is documented; aweme_list, aweme_id, statistics.digg_count and has_more are field names we observed in this response, not documented ones.

GPT-6 Luna made one tool call and answered on its second turn with FINAL: {"videos":20,"average_likes":397485}. The program printed:

{"model": "openai/gpt-6-luna", "answer": {"videos": 20, "average_likes": 397485}, "tokens": [403, 102], "cost_usd": 9.1e-05}

The fallback never fired, so this run doesn’t exercise the Sol branch. GET /v1/tasks/<id>/cost billed the two Luna calls (task ids 3224d47a-8ffc-4c28-8b34-ac450818d8b9 and 216bfcea-3c0f-4875-8ed7-2027f5dc4e73) $0.000045 and $0.000047, matching usage × list price, and the data call $0.000000. That endpoint and GET /v1/models/<id> need the same Authorization: Bearer key as the calls themselves; without a key, the public model pages on sandbase.ai show the same prices. The average differs from the benchmark’s 422,764 because this was a live call hours after the cache was recorded, and the account had posted since: the first videos on the live page aren’t in the cached page.

To adapt it, the request fields for this endpoint are in the Douyin user-post-videos API reference, and the same key covers the 22 models above. Get a SandBase API key to run it against your own task set.

Scope of the data: these are public, read-only lookups that need a SandBase API key. SandBase isn’t an official Douyin, Weibo or Xiaohongshu partner, and nothing here touches private accounts, DMs, owner analytics or account actions.

FAQ

What is the cheapest LLM that works for AI agents?

On this test, GPT-6 Luna: 51/51 at a median $0.00035 per task. MiMo v2.6 Flash (50/51, $0.00060) and GLM-5.3 FlashX (50/51, $0.0018) were close behind. Qwen3.7 Flash is cheaper per token but failed 13 of 51 runs, mostly by running out of output budget.

GPT-6 Luna vs GPT-6.1 Sol: which should I use for agents?

Both scored 51/51 with the same median input (2,887 tokens); Sol cost $0.0069 per task, Luna $0.00035. Use Luna as the default and Sol as the fallback, after rerunning your own tasks.

Is GPT-6.1 Sol Pro worth it over Sol?

Not on short tool tasks. Both scored 51/51 at the same per-token price, but Sol Pro used about three times the input tokens (median 9,026 vs 2,887) and cost 3.1x more per task. We didn’t test the long, hard reasoning it may be meant for.

Flash vs Pro models for agents: does the Pro tier help?

In two of six families: Opus 5.5 over Sonnet 5.5 (51 vs 44) and Qwen3.8 Max Prime over Qwen3.7 Flash (51 vs 38). In the OpenAI, Xiaomi and Z.ai families, a cheaper tier matched or beat the top one.

Why did Qwen3.7 Flash score so low?

Ten of its 13 misses were truncations at the 1,500-token cap; its median output (949 tokens) was more than six times Luna’s 139. We didn’t rerun with a larger budget.

Limitations

One task family, 17 scored tasks, three runs each, default settings, a 1,500-token cap per turn, no prompt caching. Pro and Max tiers may earn their price on long reasoning, coding or long documents, which we didn’t test. The two batches ran at different times, so cross-batch latency is rough. Cost is usage × list price, not billing. With 48 misses in 1,122 runs, models a point or two apart aren’t meaningfully ranked. Treat this as a method: build 20 tasks that look like your traffic, run each tier three times, and pick the cheapest one that passes.