Blog/Model Comparison/

DeepSeek V4.1 Pro: Release Status and What to Use (2026)

Is DeepSeek V4.1 Pro out? Not as of Oct 4, 2026. The official timeline, why V4 Pro still runs, current prices, our agent test scores, and what to use now.

DeepSeek V4.1 Pro release status cover; editorial artwork, not a product screenshot

DeepSeek V4.1 Pro has not been released. We checked DeepSeek’s change log, its news page, its Models & Pricing page and the deepseek-ai Hugging Face organization on 2026-10-04 at about 03:15 UTC. None of them lists a V4.1 Pro model, model name, price or weights. The only official mention is a conditional one: the V4.1 Flash announcement of September 10 says a temporary routing plan “will continue until V4.1-Pro launches.”

So if you searched for DeepSeek V4.1 Pro, the practical answer is about the two models you can call today: V4.1 Flash and V4 Pro 0813. DeepSeek reversed its plan to switch V4 Pro off on September 14, and both still answer requests. If you’re deciding between those two on price and features, our DeepSeek V4.1 Flash vs V4 Pro comparison goes deeper. This post covers the status question, the dated timeline, and our own measured results for both models.

Key takeaway

  • As of 2026-10-04 03:15 UTC, DeepSeek has announced no V4.1 Pro: no model name on the pricing page, no change log entry, no Hugging Face repo.
  • DeepSeek’s September 10 news post planned to route deepseek-v4-pro to V4.1 Flash from 04:00 UTC on September 14 “until V4.1-Pro launches”. Its change log now says V4 Pro API service continues after September 14, with billing unchanged.
  • Official peak prices per million tokens: V4.1 Flash $0.30 input / $1.20 output; V4 Pro 0813 $1.32 / $3.96. SandBase lists the same two rates.
  • On our 51-run tool-calling test, V4 Pro 0813 scored 49/51 at a median $0.0069 per task and V4.1 Flash 48/51 at $0.0015. Start with Flash, keep Pro for the cases where it measurably helps.
  • On SandBase, deepseek/deepseek-v4.1-pro returned HTTP 404 “model not found” at 03:17 UTC.

Is DeepSeek V4.1 Pro out? What each source showed

A release usually leaves several traces at once: a change log entry, a model name on the pricing table, weights on Hugging Face, then routes on third-party platforms. We looked for each one. Here is what each source showed on 2026-10-04 (UTC).

Where we lookedWhat a V4.1 Pro release would addWhat we saw
DeepSeek change logA dated “V4.1-Pro” entryNewest entry is 2026-09-10, “DeepSeek-V4.1-Flash Release”
DeepSeek newsA release postNewest post is the V4.1 Flash release of 2026/09/10
Models & PricingA third model columnTwo columns: deepseek-flash (DeepSeek-V4.1-Flash) and deepseek-v4-pro (DeepSeek-V4-Pro-0813)
deepseek-ai on Hugging FaceA V4.1-Pro repoNewest repo is DeepSeek-V4.1-Flash, created 2026-09-10
SandBase catalogA deepseek-v4.1-pro idHTTP 404 for four Pro spellings (deepseek-v4.1-pro, deepseek-v4-1-pro, deepseek-v4.1-pro-preview, deepseek-pro); eight DeepSeek ids enabled

That’s a clean “not released”, not a “maybe”. Any V4.1 Pro price, context window or benchmark quoted today is a guess, and V4.1 Flash’s numbers don’t carry over to a larger sibling.

DeepSeek has said a bigger model is the direction. The change log describes V4.1 Flash as “the smallest model in our new architecture family” and says the architecture is designed for “scaling to larger models.” That is a design statement, not a release date. We found no official date or timeframe for V4.1 Pro.

The official timeline, and the plan that changed

DeepSeek’s own pages tell a story in three steps. The middle step is where most of the confusion comes from.

Date (UTC)Official sourceWhat it says
2026-08-13Change log, “DeepSeek-V4-Pro Update”V4 Pro GA rolls out; the model name stays deepseek-v4-pro
2026-09-10Change log, “DeepSeek-V4.1-Flash Release”V4.1 Flash released; call it as deepseek-flash; V4 Flash and V4 Flash Vision Exp retired, their old names “temporarily routed to V4.1 Flash”
2026-09-10News post, V4.1 Flash releaseFrom 04:00 UTC on September 14, all deepseek-v4-pro requests route to V4.1 Flash at Flash rates, “until V4.1-Pro launches”
Undated, under the 09-10 entryChange logDeepSeek will “continue providing API services for DeepSeek V4 Pro after September 14, 2026”, billing unchanged
2026-10-04Pricing pagedeepseek-v4-pro still listed as DeepSeek-V4-Pro-0813 at Pro prices; no routing footnote for it

The phrase people quote, “until V4.1 Pro arrives”, is a paraphrase. DeepSeek’s English news post says “This will continue until V4.1-Pro launches.” The Chinese version of the same post says the routing applies from noon Beijing time on September 14 until V4.1 Pro goes online. Both are still live on the news page as of October 4.

DeepSeek news post section saying deepseek-v4-pro requests route to V4.1-Flash from 04:00 UTC on Sept 14, 2026 until V4.1-Pro launches Caption: DeepSeek’s September 10 news post, still live, with the original plan to route V4 Pro to V4.1 Flash “until V4.1-Pro launches” (captured 2026-10-04).

The change log tells you the plan was dropped. Read the two pages together and you get the current state: the news post describes the original plan, and the change log describes the reversal.

DeepSeek change log entry dated 2026-09-10 with the API changes paragraph and the note that V4 Pro API service continues after September 14 with billing unchanged Caption: The 2026-09-10 change log entry. The second paragraph under “API changes” is the reversal: V4 Pro service continues after September 14, billing unchanged (captured 2026-10-04).

Two caveats. The reversal paragraph has no date of its own, so DeepSeek’s pages don’t tell us when it was added, or whether any V4 Pro traffic was routed before it. And “continue providing” is not a retirement date. DeepSeek says it will give further notice of any changes, so check the change log before you plan around V4 Pro long-term.

What to use today: V4.1 Flash or V4 Pro 0813

Without a V4.1 Pro, the choice is between the new small model and the older large one. The official table makes most of the spec side easy, because the two share more than you’d expect.

DeepSeek Models and Pricing table listing deepseek-flash as DeepSeek-V4.1-Flash and deepseek-v4-pro as DeepSeek-V4-Pro-0813 with peak and off-peak prices Caption: DeepSeek’s Models & Pricing table: two models, the same 1M context and 384K max output, vision only on Flash, and Pro at about 4.4x Flash’s input price. The feature check marks render as empty boxes in this capture (captured 2026-10-04).

We measured V4 Pro 0813 in two of our own harnesses and V4.1 Flash in one, on 2026-10-03, through SandBase. The table puts those results next to the official specs.

V4.1 FlashV4 Pro 0813
Official namedeepseek-flashdeepseek-v4-pro
SandBase iddeepseek/deepseek-v4.1-flashdeepseek/deepseek-v4-pro, deepseek/deepseek-v4-pro-0813
Peak price, input / output per 1M$0.30 / $1.20$1.32 / $3.96
Off-peak (official API only)$0.15 / $0.60$0.66 / $1.98
Context / max output1M / 384K1M / 384K
Image inputYesNot supported
Concurrency limit (official API)2,500500
Our tool-calling test, 51 runs48/51, median $0.0015 per task, 6.3 s49/51, median $0.0069 per task, 10.2 s
Our Claude Code coding test, 12 runsNot tested10/12, median $0.046 per task

The tool-calling numbers come from the same 17 tasks, three runs each: data lookups plus aggregation over cached tool responses, default settings. V4 Pro 0813’s two misses were a mistyped account ID and one run that skipped the profile calls the task asked for. V4.1 Flash’s three misses: two took follower counts from the wrong source, and one list had three of the four IDs. Details are in the 12-model tool-calling benchmark and the Pro vs Flash tier benchmark.

The coding number is from our Claude Code benchmark: six small Python tasks, two runs each. V4 Pro 0813 failed the refactor task both times by moving a rounding bug into the new module unchanged. We didn’t run V4.1 Flash on that harness, so we have no coding comparison between the two.

Here’s how we’d pick, given that data:

Your situationUseWhy
New agent, tool calls plus short answersV4.1 FlashOne point behind Pro on our test at about 1/4.5 of the median cost
Screenshots, charts or documents as imagesV4.1 FlashV4 Pro 0813 has no image input
High request concurrency on DeepSeek’s own APIV4.1 Flash2,500 vs 500 listed concurrency
A Pro-based app with accepted outputs you must keep matchingV4 Pro 0813, while it lastsIt’s still served at Pro billing; re-test those cases on Flash before switching
You want V4.1 Pro specificallyNeither yetIt doesn’t exist; build on Flash and swap the id later

The honest summary: our data doesn’t show V4 Pro earning its 4.4x input price on agent tool tasks. One point on 51 runs is inside run-to-run noise. Pro might pay off on long reasoning or larger codebases, but we haven’t measured that.

Switching between them with one key

On SandBase, both models sit behind the same OpenAI-compatible POST /v1/chat/completions endpoint, so switching is a one-string change. The program below sends the same tool-calling request to each id, reads the catalog price from GET /v1/models/<id>, and computes the cost in code. Both endpoints take the same Authorization: Bearer key; if you only want prices, the public model pages show them without a key.

"""Same tool-calling request to two DeepSeek ids on SandBase, one key.

Prints, per model: the id you sent, the `model` field the response came back with,
finish_reason, the tool call it made, token usage and the cost computed from the
catalog price (GET /v1/models/<id>). Arithmetic happens here, not in the model.
"""
import json
import os
import sys

import requests

BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
MODELS = sys.argv[1:] or ["deepseek/deepseek-v4.1-flash", "deepseek/deepseek-v4-pro"]

TOOLS = [{
    "type": "function",
    "function": {
        "name": "get_order_status",
        "description": "Look up the shipping status of one order by its ID.",
        "parameters": {
            "type": "object",
            "properties": {"order_id": {"type": "string", "description": "Order ID, e.g. A-1042"}},
            "required": ["order_id"],
        },
    },
}]
MESSAGES = [
    {"role": "system", "content": "You are a support agent. Use tools for order data; never guess."},
    {"role": "user", "content": "Where is my order A-1042?"},
]


def catalog_price(model):
    """Return (input $/M, output $/M) from the model card. Same Bearer key."""
    r = requests.get(f"{BASE}/models/{model}", headers=HEADERS, timeout=60)
    r.raise_for_status()
    card = r.json()["model_card"]
    return float(card["prompt_token_price"]), float(card["completion_token_price"])


def run(model):
    price_in, price_out = catalog_price(model)
    r = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=120, json={
        "model": model,
        "messages": MESSAGES,
        "tools": TOOLS,
        "max_tokens": 1500,
    })
    if r.status_code != 200:
        raise RuntimeError(f"{model}: HTTP {r.status_code} {r.text[:300]}")
    body = r.json()
    choice = body["choices"][0]
    calls = choice["message"].get("tool_calls") or []
    usage = body["usage"]
    cost = usage["prompt_tokens"] / 1e6 * price_in + usage["completion_tokens"] / 1e6 * price_out
    return {
        "sent": model,
        "returned_model": body.get("model"),
        "id": body.get("id"),
        "finish": choice.get("finish_reason"),
        "tool_call": [(c["function"]["name"], json.loads(c["function"]["arguments"])) for c in calls],
        "tokens_in_out": (usage["prompt_tokens"], usage["completion_tokens"]),
        "reasoning_tokens": (usage.get("completion_tokens_details") or {}).get("reasoning_tokens"),
        "cost_usd": round(cost, 6),
    }


if __name__ == "__main__":
    for m in MODELS:
        print(json.dumps(run(m)))

Tested on 2026-10-04 (UTC)

We ran the program unchanged at 03:17 UTC with three ids: python3 ds_check.py deepseek/deepseek-v4.1-flash deepseek/deepseek-v4-pro deepseek/deepseek-v4-pro-0813. The input is a made-up order number, not customer data. Full stdout, one line per model:

{"sent": "deepseek/deepseek-v4.1-flash", "returned_model": "deepseek/deepseek-v4.1-flash", "id": "fe138e9e-6933-41df-9a07-04fcc78bc9a7", "finish": "tool_calls", "tool_call": [["get_order_status", {"order_id": "A-1042"}]], "tokens_in_out": [335, 43], "reasoning_tokens": 0, "cost_usd": 0.000152}
{"sent": "deepseek/deepseek-v4-pro", "returned_model": "deepseek/deepseek-v4-pro", "id": "8729b374-ee50-4bd9-87bc-c84680076fa8", "finish": "tool_calls", "tool_call": [["get_order_status", {"order_id": "A-1042"}]], "tokens_in_out": [319, 85], "reasoning_tokens": 27, "cost_usd": 0.000758}
{"sent": "deepseek/deepseek-v4-pro-0813", "returned_model": "deepseek/deepseek-v4-pro-0813", "id": "47ba019f-4d06-4ef2-87c4-4963dfe36895", "finish": "tool_calls", "tool_call": [["get_order_status", {"order_id": "A-1042"}]], "tokens_in_out": [398, 64], "reasoning_tokens": 15, "cost_usd": 0.000779}

All three made the right tool call with valid JSON arguments. We then looked up each id with GET /v1/tasks/<id>/cost (same Bearer key): the settled costs were $0.000152, $0.000758 and $0.000779, matching the computed ones. So deepseek/deepseek-v4-pro on SandBase was billed at Pro rates, not Flash rates.

What this can’t tell you is which upstream version answered. The model field echoes the id you sent; it isn’t a version stamp (observed on this run, not a documented guarantee). One call per model shows the route works today, not how well.

Asking for the model that doesn’t exist fails fast. A minimal chat request (one “ping” message, no tools) with "model": "deepseek/deepseek-v4.1-pro" at 03:17 UTC returned HTTP 404 with this complete body; deepseek/deepseek-v4-1-pro got the same error:

{"error":{"code":null,"message":"model not found: deepseek/deepseek-v4.1-pro","param":null,"type":"not_found_error"}}

SandBase model page for DeepSeek V4 Pro showing the id deepseek/deepseek-v4-pro, $1.32 input and $3.96 output per million tokens, and 384K max output Caption: The SandBase model page for deepseek/deepseek-v4-pro lists $1.32 input and $3.96 output per million tokens, the same as DeepSeek’s peak rate. The page shows N/A for context window; DeepSeek’s own table lists 1M (captured 2026-10-04).

SandBase’s model cards list a single rate for each model, equal to DeepSeek’s peak rate; we didn’t see an off-peak rate in their price_formula. If your workload can wait for DeepSeek’s off-peak hours, its own API halves the price.

To try it, see the Chat Completions guide, the V4.1 Flash model page and the V4 Pro model page, then get a SandBase API key.

How we’ll know when V4.1 Pro lands

A rumor, an announcement and a usable route are three different events. These are the checks we’ll run, in order, and what each one lets you safely do.

StageEvidence to look forSafe action
1. AnnouncementA dated V4.1-Pro entry in the change logRead the stated limits and prices
2. Official model nameA new column on Models & Pricing, or deepseek-v4-pro remapped to a new MODEL VERSIONRun controlled tests on DeepSeek’s API
3. WeightsA V4.1-Pro repo under deepseek-ai on Hugging FaceSelf-hosting evaluation, if you do that
4. SandBase routeGET /v1/models/deepseek/<new id> returns 200 with enabled: trueSwap the id in a staging config
5. Our numbersThe same 51-run tool-calling harness and 12-run Claude Code harnessDecide with measured accuracy and cost per task

Stage 2 is the one to watch if you call deepseek-v4-pro by name. On September 10, DeepSeek was ready to point that name at a different model with four days’ notice. If V4.1 Pro takes over the Pro name, your requests could change model without any code change. Pin a version id where your provider offers one, and keep a small set of known-good prompts with their accepted outputs so you notice the switch.

When V4.1 Pro appears, we’ll run it through both harnesses with the same tasks and settings and update this post.

FAQ

Is DeepSeek V4.1 Pro released?

No. As of 2026-10-04 at about 03:15 UTC, DeepSeek’s change log, news page, pricing page and Hugging Face organization list no V4.1 Pro. The newest release is V4.1 Flash, dated 2026-09-10.

When will DeepSeek V4.1 Pro come out?

DeepSeek hasn’t given a date. Its only mention is the September 10 news post, which said V4 Pro routing to Flash would last “until V4.1-Pro launches”. That’s a condition, not a schedule.

Was DeepSeek V4 Pro shut down on September 14?

Not according to DeepSeek’s current pages. The original plan routed deepseek-v4-pro to V4.1 Flash from 04:00 UTC on September 14. The change log now says V4 Pro API service continues after that date with billing unchanged, and the pricing page still lists DeepSeek-V4-Pro-0813 at $1.32 / $3.96 peak.

Should I wait for V4.1 Pro or use V4.1 Flash now?

Use V4.1 Flash now unless you have a measured reason not to. On our 51-run tool-calling test it scored 48/51 against V4 Pro 0813’s 49/51 at a median $0.0015 per task versus $0.0069. Design so the model id is one config value, and swapping later is cheap.

Can I call DeepSeek V4.1 Pro on SandBase?

Not yet. deepseek/deepseek-v4.1-pro returns HTTP 404 “model not found”. The DeepSeek ids enabled on 2026-10-04 were deepseek-v4.1-flash, deepseek-v4-pro, deepseek-v4-pro-0813, deepseek-v4-flash, deepseek-v4-flash-0731, deepseek-v4-flash-vision-exp, deepseek-chat and deepseek-v3.2.

Limitations

This is a status check at one point in time, 2026-10-04 around 03:15 UTC; recheck the change log. Release facts come only from DeepSeek’s own pages, not news coverage.

The measured numbers come from two narrow harnesses run on 2026-10-03 at default settings. They don’t cover long reasoning, long documents, image input or large codebases. The live check is one single-turn call per model. We can see SandBase billing at Pro rates for deepseek/deepseek-v4-pro, but no response field told us which upstream model version served it.