DeepSeek-V4-Flash-0731: The $0.14 Model That's Rewriting AI Agent Economics
DeepSeek's V4-Flash-0731 update beats its own larger model on every agentic benchmark, at $0.14 per million tokens. Here's what it means for AI builders.
On July 31, 2026, DeepSeek quietly pushed an update that's forcing a lot of AI teams to redo their model-selection spreadsheets. DeepSeek-V4-Flash-0731 — the "cheap" tier of DeepSeek's lineup — now outperforms the company's own larger V4-Pro-Preview on every agentic and coding benchmark it has published. It's not a bigger model. It's the same architecture, retrained smarter, and priced at $0.14 per million input tokens. For agencies and builders stitching together AI agents, coding copilots, and automations all day, that combination is worth paying attention to.
DeepSeek-V4-Flash-0731 keeps the same 284-billion-parameter mixture-of-experts architecture as the earlier April preview, activating just 13 billion parameters per token and supporting a 1,048,576-token context window (with up to 65,536 tokens of output). Nothing about the underlying model got bigger. What changed is the post-training pipeline: DeepSeek rebuilt it specifically around tool use, multi-step reasoning, coding, and autonomous agent workflows — the exact skills that matter when a model is driving a workflow rather than just answering a chat message.
The weights are published on Hugging Face under an MIT license, and the official V4-Flash API moved into public beta the same day. That combination — open weights plus a hosted API — means teams can start on the API immediately and still have the option to self-host later without a licensing conversation.
The headline number isn't the price — it's that a "Flash" model is now beating DeepSeek's own bigger Pro model. On Terminal-Bench 2.1, a benchmark that measures how well a model can actually operate a command-line environment to complete tasks, the score jumped from 61.8 to 82.7 — a 20.9-point gain in one post-training pass. On DeepSWE, a software-engineering agent benchmark, the score went from 7.3 to 54.4, blowing past V4-Pro-Preview's 12.8.
Those are the kinds of benchmarks that map directly onto what agencies actually build: agents that read a codebase, make a change, run a test, and iterate; workflows that call tools, check their own output, and retry when something fails. A model that's dramatically better at that, without getting more expensive to run, changes the math on which tasks are worth automating.
DeepSeek-V4-Flash-0731 bills at $0.14 per million input tokens and $0.28 per million output tokens on the first-party API. The detail that matters most for anyone running agent loops is the cached-input price: $0.0028 per million tokens, a 98% discount off standard input pricing. Agent workflows tend to re-send the same system prompts, tool definitions, and conversation history on every turn — that's exactly the pattern cache discounts are built for. In a long-running automation, cache pricing alone can be the difference between a workflow that's economical to run continuously and one that only makes sense for occasional use.
For agencies pricing out AI-powered deliverables — a client-facing agent, a document pipeline, an internal automation — a frontier-competitive model at this price point widens the set of projects that actually pencil out. Tasks that were previously too expensive to run on a top-tier model at scale become viable on a budget line, without asking clients to accept a worse model to hit their number.
The release has been picked up on both ends of the AI conversation this week: Hacker News' front page carried a thread about running the model on a single AMD MI300X GPU, a detail that matters for teams who want the self-hosting option without needing a multi-GPU cluster, and it ranked in the top tier of Product Hunt's daily leaderboard for AI releases. That's a rare combination — a model release usually gets attention from developers or from the broader product crowd, not both at once. It's a signal that the update lands as genuinely useful rather than just a benchmark headline.
Our cutting-edge features simplify collaboration and creativity, making your workflow intuitive and efficient. Transform your vision into reality effortlessly with Hadidiz Flow.



