The 60-second briefing

DeepSeek released the general-availability version of V4 Pro on August 13 across its app, website, and API. The model keeps a 1-million-token context window, supports up to 384,000 output tokens, and now natively accepts OpenAI’s Responses API format—the interface used by several modern AI tools, including Codex. DeepSeek says the API call stays the same: set the model name to deepseek-v4-pro.

The useful numbers are on the price sheet. V4 Pro is listed at $0.435 per million input tokens and $0.87 per million output tokens today, with a concurrency limit of 500. From August 16 at 16:00 UTC, peak-hour pricing rises to $1.32 input and $3.96 output, while off-peak rates are half those peak prices. In plain English: the model is cheap enough for experimentation, but heavy users will need to schedule around the clock.

Anthropic made a different kind of AI change this week. Its updated help page says Claude models launched in the EU on or after August 2 carry invisible watermarks in generated text and signed provenance metadata on supported files such as PNG, JPG, and SVG. The marks apply worldwide across Claude, Claude Code, Claude Cowork, Claude Tag, and the API. Anthropic says a mark signals that Claude may have processed the work; it does not prove Claude wrote the underlying ideas.

What changed

A million-token context window sounds like a benchmark until you put a real project inside it. One million tokens is enough room for a large folder of notes, a long transcript, a product brief, or a codebase-sized collection of documents in one conversation. The benefit is fewer “here is part two” handoffs: you can ask the model to compare files, find contradictions, and keep a single set of instructions in view.

DeepSeek also added three thinking-effort levels—low, high, and max—for both V4 Pro and V4 Flash. That gives a builder a simple trade-off: use low for a quick rewrite, high for a normal multi-step task, and max for a harder investigation. The company’s own release reports 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified, but those are vendor-reported benchmarks. The more practical proof is whether the model can finish a small job without burning through your budget.

For an ordinary user, this is less about installing a new chatbot and more about the cost of trying. A long document review that would be expensive on a premium model can now be tested at a low per-token price. Use the John Savage AI tools page to find the current DeepSeek access point, then check the live terms before sending private material. Cheap does not mean private, and a 1M context window does not remove the need to verify important answers.

The money and trust signal

Axios reported on August 13, citing Bloomberg, that Cognition is in talks to raise funding at a $40 billion valuation. Cognition makes Devin, an AI coding agent. The report is not a completed financing, but the number shows where investors think software that can perform multi-step computer work might be headed.

The same report says Anthropic is discussing a roughly $6 billion purchase of Decart, a world-model and chip-optimization company that had raised more than $450 million and was last valued at $4 billion in May. That is a reminder that the AI race is moving beyond chatbots: companies are buying the tools and hardware knowledge needed to make agents cheaper, longer-running, and easier to deploy.

Anthropic’s watermark plan adds a trust layer to that expansion. If you use Claude to proofread your own draft, translate a customer email, or turn notes into a PDF, the output may still carry a Claude mark even when the original ideas came from you. The company says detection details are coming, and it warns that a missing mark does not prove content was never AI-processed.

Try this today

Run a long-context test with a project you already understand:

  1. Gather five to ten documents—notes, customer questions, a brief, or research links—and remove passwords and personal data.

  2. Ask DeepSeek to produce a one-page decision brief with three recommendations, three open questions, and a list of every source it used.

  3. Ask a second question that forces comparison: “Which recommendation changes if the budget is cut by 30%?”

  4. Check two facts against the original files and record the actual cost before you decide whether the workflow is worth repeating.

The point is not to admire a 1M-token number. It is to see whether one affordable model can hold the whole problem in view and leave you with fewer fragments to manage.

— John