DeepSeek V4 Pro 0813 is presented as an open-weight, MIT-licensed mixture-of-experts model from DeepSeek, aimed at reasoning, software engineering, and long-running agentic workloads, with a lighter V4 Flash counterpart. The supplied evaluations agree that its substantially lower inference pricing could make it attractive for production model routing, but conflict on capability: CAISI places it roughly eight months behind the frontier, while other reviewers describe it as near or equal to leading closed models on selected benchmarks. Reported latency and price-performance advantages are workload- and configuration-dependent, so independent production evaluations remain necessary; the snippets provide only indirect attribution to OpenRouter.
2026-08-20T11:30:21Z
No controlled Pro-versus-Flash evaluation or verified production economics emerged during the release window; repeated anecdotes and thin performance headlines never established a routing advantage. The episode has faded and can expire, with any future reproducible evaluation treated as a new delta.
2026-08-18T10:38:31Z
The second performance headline does not provide inspectable methodology or constitute independent corroboration of the runtime-control claim; it is repetitive amplification rather than a new evaluation result. V4 Pro remains a testable, harness-sensitive candidate with no demonstrated production-routing advantage.
2026-08-18T10:22:55Z
evidence attached: hn.story.49343330 — This independent-looking capability report bears directly on whether DeepSeek V4 Pro 0813 should affect frontier-model selection, though its evidence is currently thin.
2026-08-18T05:25:57Z
The runtime-control report identifies a plausible harness-level explanation for earlier weak results, making serving configuration a more explicit evaluation variable. With no inspectable settings, reproducible comparison, or independent corroboration, it does not establish an advantage or justify changing production routing.
2026-08-18T05:22:04Z
evidence attached: hn.story.49341271 — A reported runtime-control fix and comparative result directly bear on whether DeepSeek V4 Pro changes frontier-model selection, but require independent validation.
2026-08-17T20:31:28Z
No new evidence has arrived beyond repeated, provider-dependent anecdotes, so the release has settled into a parked evaluation candidate rather than an active routing shift. Keep watching for a controlled Pro-versus-Flash workload comparison, but slow the cadence.
2026-08-15T20:29:44Z
Refreshed discussion only amplifies the established, provider-dependent tradeoffs around Pro versus Flash; it adds no controlled workload comparison, reliability trace, or verified total economics. The case remains parked pending reproducible evidence capable of changing production routing.
2026-08-14T19:41:44Z
Refreshed comments add conflicting provider-dependent coding reports and one quantified self-hosted slowdown, but no controlled Pro-versus-Flash workload test or verified total economics. The case remains parked pending reproducible evaluation capable of changing production routing.
2026-08-14T13:37:18Z
Two tiny hands-on comparisons modestly reinforce that V4 Pro’s advantage over Flash is workload-sensitive and may be erased by failures or higher cost, despite occasional deeper coding behavior. The evidence is uncontrolled and non-reproducible, so it does not justify rerouting or advance the case beyond watching.
2026-08-14T13:23:11Z
evidence attached: reddit.post.1vo6ry6 — Additional hands-on comparison suggests V4 Pro 0813 may offer deeper coding reasoning without a clear enough capability jump to change model selection.
2026-08-14T13:23:11Z
evidence attached: reddit.post.1vo5x4i — Early hands-on evidence that V4 Pro 0813 may underperform Flash on practical tasks, directly informing model-selection judgment despite the tiny sample.
2026-08-14T05:27:59Z
The refreshed discussion is repetitive amplification of launch-day anecdotes and adds no controlled workload result, latency trace, reliability measurement, or verified serving economics. V4 Pro remains a testable routing candidate, but nothing new supports changing production selection.
2026-08-14T02:27:15Z
Refreshed comments and engagement only amplify established launch-day reactions and benchmark speculation; no controlled workload result, latency trace, or verified serving economics changes the production-routing judgment.
2026-08-14T00:37:24Z
The refreshed comments merely repeat launch-day artifact and benchmark reactions without a controlled workload result or verified serving economics. V4 Pro remains testable but provides no new basis for changing production routing.
2026-08-13T23:32:48Z
Refreshed discussion and engagement continue to amplify the already-known marginal, workload-sensitive tradeoff versus V4 Flash. No controlled workload result, latency measurement, reliability trace, or verified serving economics changes the production-routing judgment.
2026-08-13T22:33:26Z
Refreshed launch-thread comments and engagement remain repetitive amplification of the known pricing, availability, and mixed practitioner reactions. No controlled workload comparison or verified serving economics changes the production-routing judgment.
2026-08-13T20:32:34Z
Refreshed comments and minor engagement continue to amplify the established launch and mixed practitioner reactions without a controlled workload evaluation or verified serving economics. Nothing changes the production-routing judgment, so the case remains parked pending reproducible evidence.
2026-08-13T18:47:30Z
Refreshed launch-thread comments continue to repeat mixed capability, pricing, and artifact-access reactions without a controlled workload comparison or verified serving economics. The case remains testable but gains no evidence for changing production routing.
2026-08-13T17:44:52Z
Refreshed comments and engagement only amplify the already-known launch, pricing, and mixed practitioner reactions; they add no controlled workload result or verified serving economics. The case remains actionable for testing but can cool while awaiting reproducible evidence that could change routing.
2026-08-13T16:38:27Z
Restored Hugging Face availability removes the transient artifact-access concern and makes controlled self-hosted evaluation feasible. It adds no independent capability, latency, reliability, or total-cost result, so the production-routing judgment remains unchanged.
2026-08-13T16:23:51Z
evidence attached: reddit.post.1vnervw — shared external link with case evidence
2026-08-13T15:41:40Z
The first attached coding-agent comparison tilts against V4 Pro versus V4 Flash on practical quality and economics, but provider quantization and KV-cache compression materially confound attribution to the model itself. It sharpens the need for controlled, trace-backed tests without supporting a production-routing change.
2026-08-13T15:23:50Z
evidence attached: reddit.post.1vnci9h — Independent coding-agent use directly supports the open case by reporting weak results and unfavorable economics versus DeepSeek V4 Flash.
2026-08-13T14:43:52Z
Refreshed comments remain repetitive launch reaction and mixed practitioner anecdotes, with no reproducible workload, latency, reliability, or total-cost evaluation. A reported transient Hugging Face 404 does not yet overturn the established open-weight release or alter the production-routing judgment.
2026-08-13T13:30:42Z
The open-weight Hugging Face artifact expands V4 Pro from an API-only candidate into one that can be independently reproduced, quantized, and evaluated under controlled serving configurations. It enables stronger validation but supplies no independent workload result yet, so the production-routing hypothesis remains unsettled.
2026-08-13T13:23:02Z
evidence attached: reddit.post.1vn9it4 — The Hugging Face release is a confirmed first-party artifact that enables independent evaluation of DeepSeek V4 Pro 0813 for model selection.
2026-08-13T12:33:38Z
grounded: known/medium — Scott already holds the relevant position in Model Perishability and Capability Audit: models should remain swappable and be re-evaluated on representative prod
2026-08-13T12:31:12Z
DeepSeek’s official launch now resolves the designation and production-availability uncertainty, turning V4 Pro into an immediately testable routing option. The higher price tier and mixed early reports make trace-backed comparison against V4 Flash more consequential, but still do not establish a production advantage.
2026-08-13T12:22:55Z
evidence attached: reddit.post.1vn8m1x — DeepSeek’s launch announcement is first-party evidence that V4 Pro is available for the capability, latency, and price-performance evaluation in the open case.
2026-08-13T11:28:04Z
The refreshed discussion remains repetitive amplification of the known cheap-but-slower, reliability-sensitive tradeoff and adds no reproducible evaluation or verified economics. Production availability is established, but no new evidence changes the routing judgment.
2026-08-13T10:23:44Z
Refreshed comments add no reproducible capability, latency, reliability, or total-cost evidence beyond the already-known mixed anecdotes. Production availability was already surfaced, so the case can cool while awaiting an independent workload evaluation that could actually change routing.
2026-08-13T09:39:11Z
Official API documentation now establishes V4 Pro 0813 as a priced, directly testable production option, making immediate harness evaluation more actionable for Scott’s existing V4 Flash routing. Its selection advantage remains unproved because the new discussion supplies no reproducible capability, latency, reliability, or total-cost comparison.
2026-08-13T09:22:25Z
evidence attached: reddit.post.1vn5jbx — Announces the non-preview production release of DeepSeek V4 Pro 0813 on the API with pricing, directly updating the open case about its model-selection impact.
2026-08-13T08:32:27Z
The refreshed discussion remains repetitive amplification of a marginal, cheap-but-slower and reliability-sensitive tradeoff, with no inspectable evaluation, workload trace, latency result, or verified total-cost evidence. It does not change the production-routing judgment or advance the case beyond watching.
2026-08-13T07:43:50Z
The refreshed comments remain repetitive amplification of the same marginal, cheap-but-slower and reliability-sensitive tradeoff. No inspectable benchmark, workload trace, latency result, or verified total-cost change alters the production-routing judgment.
2026-08-13T06:32:00Z
The refreshed discussion continues to recycle the same cheap-but-slower, defect-sensitive comparison without reproducible latency, reliability, or total-cost results. It adds no basis for changing production routing, so the case remains on a slower cadence pending independent workload evaluation.
2026-08-13T05:26:25Z
Refreshed comments repeat the same cheap-but-slower, defect-sensitive tradeoff and unverified pricing concern without reproducible workload, latency, or total-cost evidence. The case still does not justify changing production routing and can move to a slower review cadence while awaiting independent evaluation.
2026-08-13T04:23:04Z
The refreshed comparison reinforces that Pro’s low token cost may be offset by slower execution and defect-driven rework, but it remains an isolated, non-reproducible workload anecdote. No independent latency, reliability, or total-cost evidence supports changing Scott’s production routing.
2026-08-13T03:36:12Z
Refreshed discussion still points to a marginal, workload-sensitive advantage over V4 Flash and raises the already-known concern about long-context coding reliability. It adds no inspectable evaluation, latency measurement, or total-cost evidence that would justify rerouting production workloads.
2026-08-13T02:31:59Z
The newly attached comparison points toward only a marginal Pro advantage over V4 Flash, sharpening the possibility that the release will not justify rerouting. Its methodology is not inspectable, however, so the case still awaits trace-backed workload tests and verified total economics.
2026-08-13T02:22:28Z
evidence attached: reddit.post.1vmxvud — This is an additional reported comparison relevant to whether DeepSeek V4 Pro 0813 changes frontier-model selection.
2026-08-13T01:23:03Z
The refreshed discussion remains anecdotal and reinforces the already-known cheap-but-slower, reliability-sensitive tradeoff. Neither the unverified pricing remark nor isolated coding comparisons provide reproducible economics or capability evidence that would justify changing production routing.
2026-08-13T00:23:34Z
The image-only comparison and refreshed practitioner anecdotes repeat the known cheap-but-slower, reliability-sensitive tradeoff without inspectable methodology or verified total economics. No new evidence supports changing production routing or promoting the case beyond watching.
2026-08-13T00:22:30Z
evidence attached: reddit.post.1vmv8ra — A direct DeepSeek 0813 comparison bears on whether the release changes frontier-model selection, although the image-only evidence is weak.
2026-08-12T23:30:43Z
Refreshed comments repeat the known cheap-but-slower, reliability-sensitive coding tradeoff without independent methodology, reproducible workload results, or verified total economics. The case still does not support changing production routing and should wait for trace-backed evaluation.
2026-08-12T22:33:39Z
The refreshed comments remain anecdotal comparisons reinforcing the known cheap-but-slower, reliability-sensitive tradeoff; they add no independent methodology, reproducible workload result, or confirmed pricing change. Repeated amplification without validation cools the case while it waits for trace-backed evaluations and verified total economics.
2026-08-12T21:33:56Z
The refreshed comments continue to amplify the same cheap-but-slower, reliability-sensitive tradeoff without independent methodology, trace-backed workload results, or verified total economics. The model remains worth testing for Scott’s routing stack, but this delta does not strengthen the case for changing production selection.
2026-08-12T20:34:25Z
The refreshed discussion remains repetitive practitioner testimony about the same cost, latency, and reliability tradeoff; it adds no reproducible evaluation or production evidence sufficient to alter routing decisions. The case still waits on trace-backed workload tests and confirmed economics.
2026-08-12T19:26:19Z
Refreshed practitioner comparisons reinforce a plausible cheap-but-slower coding tradeoff, including quality failures that can erase token-cost savings, but remain anecdotal and workload-specific. No independent evaluation yet justifies changing production routing from V4 Flash or frontier alternatives.
2026-08-12T18:34:41Z
The attached benchmark claims raise the possible upside substantially, but without independent methodology or reproduction they do not establish a production-routing advantage over V4 Flash or closed models. The release itself is established and already surfaced; the case now waits on trace-backed workload tests, latency data, and confirmed economics.
2026-08-12T18:22:38Z
evidence attached: hn.story.49276138 — The reported release and benchmark results materially bear on whether DeepSeek V4 Pro 0813 can change frontier-model selection, though the claims still need independent validation.
2026-08-12T17:42:55Z
Fresh practitioner anecdotes suggest a cheap-but-slower and potentially less reliable coding tradeoff, while some users still prefer V4 Flash; they sharpen the evaluation target but do not constitute a robust independent capability or price-performance result.
2026-08-12T17:36:48Z
grounded: known/high — The evaluation position is already carried by Scott’s Model-Plus-Harness Benchmark Unit and Capability Audit, while the radar already tracks the closely related
2026-08-12T17:34:07Z
origin walked (codex/luna, conf 0.94): anchor hn.story.49275114 -> echo.other.bb9241e978 by DeepSeek
2026-08-12T17:33:11Z
case created — First-party API documentation and substantial discussion jointly establish a newly available model release requiring rapid capability and economics validation.