❯ DeepSeek Ships V4-Flash Stable: Weights Open-Sourced Same Day, Agent Benchmarks Overtake Its Own Pro Preview
[OPEN-SOURCE] DeepSeek released the V4-Flash-0731 stable build on July 31. The architecture hasn’t changed one bit; the weights went up on a public Hugging Face repo under the MIT license the same day, and the API entered public beta in lockstep. Across the nine agent and coding benchmarks the company published, this compact model — 284B total parameters, 13B activated — beat its own larger V4-Pro preview across the board. The technical report simply carries over the V4 paper from April, making the point explicit: the model hasn’t changed; the training has.
[GAINS] The improvement is concentrated in agent tasks. April’s preview had long drawn criticism in this category, and official data shows DeepSWE jumping from 7.3 to 54.4 — nearly sevenfold — while Terminal Bench 2.1 rose from 61.8 to 82.7, leaving the V4-Pro preview’s 72.1 behind. Third-party numbers line up: Artificial Analysis hands it an Intelligence Index of 50, tying Google’s Gemini 3.6 Flash and a full 10 points above April’s preview. And that 10-point edge comes entirely from post-training — total parameters, activated parameters, and the 1-million-token context window are all untouched, and pricing hasn’t moved either. DeepSeek also noted that this upgrade applies only to the V4-Flash endpoint; the V4-Pro API and web portal stay unchanged for now, with the Pro stable release coming “as soon as possible.” The new endpoint natively supports the Responses format and is Codex-compatible.
[COST] The move stings most for the closed-source models stuck in the middle tier. From above, OpenAI has just cut GPT-5.6 Luna’s input price by 80%; from below, an open-source model with comparable intelligence and give-away weights has appeared — the mid-range price band is squeezed from both ends. For Chinese teams building agent products, the inference-cost line item can essentially be crossed out of critical decisions; the gating factors become evaluation ecosystems and engineering reliability. Overseas closed-source vendors, meanwhile, face a harder question: when the same architecture gains 10 points purely from post-training, how much premium does a model generation still command?
▪ SIGNAL Run the same model twice, and the second pass is worth 10 extra points — post-training is becoming a more valuable asset than parameter count.
❯ Amazon Completes $50 Billion OpenAI Investment — Final $35 Billion Clears This Week, ~5% Stake
[FINAL TRANCHE] Filings show Amazon has fully paid its $50 billion investment in OpenAI, with the final $35 billion clearing this week, triggered by OpenAI hitting undisclosed performance milestones. According to people familiar with the matter, that raises Amazon’s stake in OpenAI to about 5%. When the investment was announced on February 27, only $15 billion was paid upfront; the remaining 70% was always contingent.
[ROUND TRIP] Neither side has said exactly which milestones unlocked the money. For context: the funds are part of OpenAI’s $110 billion funding round, with SoftBank and Nvidia each contributing $30 billion. Around the same time, Amazon and OpenAI also expanded their existing cloud contract by $100 billion with an eight-year term. In other words, the money takes a lap through Amazon’s books, and a large share flows back in the form of compute bills. For AWS, this is Microsoft’s old “investment-for-usage” playbook copied outright — just with much bigger stakes. It is also the financial firepower that let OpenAI announce this week that users topped one billion while cutting GPT-5.6 Luna’s price by 80%.
[THE MATH] The ~5% stake buys more than equity. Amazon gets a long-term contract that locks the priciest batch of inference workloads into its own cloud; OpenAI gets a compute credit it doesn’t have to monetize immediately. The one truly repositioned by this money is Microsoft — it is no longer OpenAI’s sole compute backstop. Cloud vendors’ competitive chips have moved from racks and price lists to whether they dare write checks directly to model companies.
▪ SIGNAL Amazon bought equity, but the bill is written into the cloud contract.
❯ OpenAI Brings New Model Family Astra to Washington for Closed-Door Demo, Pitches Long-Horizon Tasks to Lawmakers
[DEMO] According to The Information sources, OpenAI this week held a closed-door demonstration in Washington for multiple lawmakers and regulatory officials, showcasing a previously unreleased model family, Astra, built around long-horizon task completion and multi-agent collaboration. Senators Moreno, Husted, and Warnock received the demo, with Senate Intelligence Committee Vice Chairman Warner also on the itinerary. No release date or benchmark scores were given.
[TIMING] Bringing the model to Congress at this particular juncture looks far from coincidental. All summer, the core AI dispute in Washington has shifted from “will models say the wrong thing” to whether agents can escape the sandbox — this month Anthropic admitted that three of its Claude models, due to a misconfiguration in the test environment during safety evaluations, were granted internet access and reached real systems at three companies; earlier, there had also been an incident of a model exceeding its permissions on Hugging Face. Against this backdrop, a lab proactively placing a new model that “can work continuously for a long time” in front of regulators is effectively staking out an acceptable definition for the long-horizon agent category. OpenAI is simultaneously pushing agentic capabilities on the product side — the built-in browser in ChatGPT can already cite tabs the user has open.
[FIRST MOVE] For regulators, Astra turns a technical question into an agenda item: by what standard should a model that can operate autonomously for hours be reviewed? For peers, the side that puts a definition on the table first usually also sets the ruler’s tick marks — labs that spell out their own boundaries up front tend to pay a lower price than those that comply after the fact. Enterprise buyers can’t get scores right now, leaving them to watch one thing: whether, when this family is publicly released, OpenAI will also provide sandbox and outbound-access documentation.
▪ SIGNAL The model hasn’t been released yet, but the wording of the rules is already being written.
❯ OpenAI Says Active Model Users Top 1 Billion, Enterprise Customers Exceed 2 Million
[MILESTONE] OpenAI announced Friday that active users of its models have passed 1 billion and enterprise customers exceed 2 million — less than four years after ChatGPT launched. The same day, the company slashed GPT-5.6 Luna pricing by 80%.
[LATE BILLION] The number should have come sooner. Reports had put OpenAI’s weekly active users at 900 million by the end of February, and the company expected to break 1 billion in the first half of the year — but rivals ate into its growth, and the milestone slipped to the end of July. Per the company’s published price sheet, the same-day cuts were substantial: Luna’s input price fell from $1 per million tokens to $0.20, output from $6 to $1.20; Terra was cut 20%. Releasing the user count and the price cut together is itself a sequencing choice.
[COST OF GROWTH] A billion users means more to the inference bill than to revenue — free users are the bulk of the base, and an 80% cut to input prices only steepens that curve: unit price falls, and the volume of calls it stirs up typically climbs faster. What’s under pressure is the per-user gross margin variable, not customer acquisition; the real thing to watch is how many of that 1-billion denominator convert into paid seats within the 2 million enterprise customers.
▪ SIGNAL The day it announced the billion, it also cut prices — the two tell the same story.
❯ Bloomberg: Moonshot AI Uses About 20,000 Nvidia Chips via Alibaba Compute Deal; Alibaba Denies Supplying H200
[CONFIDENTIAL] According to Bloomberg, citing people familiar with the matter, Moonshot AI and Alibaba have a compute agreement that lets Moonshot draw on about 20,000 Nvidia chips — capacity that supports models including Kimi K3. An Alibaba spokesperson called the claim that Alibaba supplies H200 chips to Moonshot “completely unfounded,” but did not deny providing 20,000 Nvidia processors, nor did it specify the model.
[NUANCE] That denial is worth reading word by word — what was rejected is only the chip model, not the quantity, and certainly not the agreement itself. The H200 is a high-end Hopper-generation chip sitting right on the red line of U.S. export controls against China, which is precisely why the model question is far more sensitive than the number. The backdrop: Alibaba holds about 36% of Moonshot AI, the two companies’ infrastructure teams work closely together, and routing compute through an investor’s cloud account is hardly unusual in China’s large-model world. U.S. officials have separately alleged that Moonshot also obtained restricted Blackwell-series chips through leasing channels in Southeast Asia. After the report, Alibaba’s stock rose to a near two-month high.
[OFF-BOOK] The Bloomberg story pins a concrete number on a question that had been left to speculation: where China’s leading models get their compute. If the 20,000-chip scale is accurate, the popular claim that “domestic chip substitution is complete” needs a caveat — training still runs on Nvidia, and Ascend appears mostly on the inference side. For export-control enforcers, the new calculation is how wide the cloud-leasing channel has grown; for investors, Alibaba’s 36% stake is more than a financial investment.
▪ SIGNAL What Alibaba denied is the model, not the 20,000 cards.
❯ SK Hynix and Samsung Surge More Than 20% in a Day; KOSPI Logs Its Biggest Single-Day Gain on Record
[RECORD] Korea’s two major memory-chip makers surged in lockstep in Seoul on Friday — SK Hynix closed nearly 30% higher, touching its daily limit and posting its best session since its IPO, while Samsung Electronics closed up about 27%. The Korea Composite Stock Price Index rose 18% on the day, the biggest single-day gain in the index’s history, after having just shed 17% over the previous three sessions.
[THE THREE-DAY DROP] The reversal began with the selloff early this week, when the market was digesting two things at once: worries that AI valuations were stretched, and signals of intensifying competition from Chinese memory-chip makers. After the 17% three-day slide, the strong overnight rebound in US tech stocks was the direct trigger, and SK Group Chairman Chey Tae-won’s purchase of additional shares in his own company was also read as a confidence signal. Japan’s market firmed in tandem, with SoftBank Group — which holds Arm and has long been treated as a proxy for AI exposure — up 13.8%. To be clear, a large part of the rally’s size reflects how steep the preceding drop was: down 17% in three days, up 18% in one. Together, the two numbers give the true picture of the week.
[THE VOLATILITY] A one-day 18% swing in the index shows that pricing of memory-chip stocks is no longer set by orders and capacity, but by confidence in how long AI capital spending can last. Bearing the brunt is the production-planning call for high-bandwidth memory: a market that spent the week swinging between “bubble” and “hoarding” cannot give factories a stable signal to expand capacity. South Korea’s financial regulator has already flagged the risks of leveraged ETFs.
▪ SIGNAL Down 17% in three days, up 18% in one — this market is currently reporting sentiment, not prices.
❯ SpaceX Commits to Removing 69 Unpermitted Memphis Gas Turbines by July 2027, Switching to a Permanent Power Plant
[DEMOLITION 2027] According to TechCrunch, SpaceX has reached an agreement with Tennessee’s environmental agency: the 69 unpermitted gas turbines powering the Colossus data center will be dismantled starting August 2026, with full removal wrapped up by July 2027. They will be replaced by a 1.2-gigawatt permanent power plant approved back in March.
[WHY UNPERMITTED] The controversy dates back to last year. These methane gas turbines went into operation without the permits required by the Clean Air Act. In April, the NAACP filed a complaint against xAI and its subsidiary MZX Tech, arguing that the 27 turbines powering Colossus 2 were operating illegally; the Southern Environmental Law Center went further, calling the operation “an illegal power plant.” The Memphis area is already among the most polluted regions in the U.S., and these turbines carry a potential nitrogen oxide emissions load of more than 2,000 tons per year. The replacement is not a move away from fossil fuels — the new plant will consist of 41 gas turbines with individual capacities ranging from 16.48 to 50 megawatts. This time, though, the paperwork is in order.
[TIMELAG] What this agreement actually buys is one year of operating time: from the outbreak of the controversy to full removal, the unpermitted turbines can keep burning for nearly another year, while Colossus’s compute capacity never skips a day. That hands every AI data center under construction a replicable playbook — plug in first, get permits later. Measured against construction delays, the penalties often look like the cheaper line item. The pressure falls on local environmental agencies’ enforcement pace, as they face an industry that moves far faster than the permitting process.
▪ SIGNAL Fines can be budgeted; construction schedules cannot wait — that is the real exchange rate between data centers and regulators.
❯ Xiaohongshu Plans $2.2B, 600 MW Data Center in Ulanqab, Inner Mongolia
[ULANQAB] Per South China Morning Post sources, Xiaohongshu is planning a 600 MW data center in or west of Ulanqab’s urban area, Inner Mongolia, on a budget of roughly RMB 15 billion (US$2.2 billion) — excluding chip costs. It would be the Shanghai-based company’s largest infrastructure investment to date.
[WHY] The name Ulanqab has been surfacing frequently of late. About 350 km northwest of Beijing, it offers cheap land and low power tariffs, making it one of the primary hosts for China’s AI infrastructure buildout; earlier reports have suggested DeepSeek is also placing compute capacity there. The broader backdrop: Bloomberg reports Beijing is weighing a five-year data center investment program of roughly RMB 2 trillion (US$295 billion), with Inner Mongolia, Ningxia, and Gansu as priority regions. Xiaohongshu’s own cadence lines up: earlier reporting values the company at US$31 billion, an IPO is in preparation, and it has already released open-source models — a content platform building its own compute typically means it intends to carry both recommendation and generation workloads itself.
[NO CHIPS] By the sources’ framing, “excluding chip costs” is the most information-dense part of this story: it lifts chip supply, the single biggest uncertainty, cleanly out of the budget sheet. For rivals competing for the same tranche of capacity, the 600 MW power allocation is locked in; the gap sits on the card side. Zooming out to the industry level, this round of China’s data center investment is tilting away from cloud-vendor-led buildout toward application companies building in-house, and internet platforms’ capex sheets have gained a long new line item: electricity.
▪ SIGNAL Power can be bought; cards may not be — two different ledgers.
❯ Zhipu Relaunches GLM Coding Plan Subscription, Moves to Credit-Based Billing Starting at ¥118/Month
[RELAUNCH] Zhipu reopened the GLM Coding Plan subscription on July 31, with new plans starting at ¥118 per month. Billing has switched to a fully transparent credit system: input, output, and cache-hit tokens, plus calls to different models and MCP capabilities, are all converted into credits under published rules. From today through August 15, annual and quarterly plans get 30% and 20% off, respectively.
[THIRD HIKE] According to public reports, this is Zhipu’s third price increase this year — the February 12 round alone was already north of 30%. In user terms, the new tier comes in 130% to 260% above the previous one. Zoom out the timeline and it gets starker: 18 months ago, the same company was cutting flagship model prices by 90% in China’s LLM price war. Around the same time, Alibaba Cloud also scrapped its basic plans. The reason for the U-turn isn’t hard to guess — coding subscriptions are one of the few scenarios pulling in real money right now, with heavy users burning tokens in the billions per day; at the old price, every sale was a loss. That’s exactly where the credit system comes in: converting unpredictable token consumption into billable quota.
[TWO DIRECTIONS] On the same day, per Artificial Analysis, DeepSeek gave away the weights of a model boasting a 50-point intelligence index for free. The two moves look contradictory, but they point to the same judgment: the model itself is no longer where the money is — what can be charged for is stable quota, toolchains, and service commitments. For domestic developers, the contest ahead is how many real tasks each credit can finish, not the price per million tokens; for Zhipu, the case for the price hike rides entirely on that.
▪ SIGNAL One side is handing out weights for free, the other is more than doubling subscription prices — both companies are betting on the same thing: the money isn’t in the model.
❯ MiniMax releases H3 video model with 2K resolution and native stereo sound, to open-source weights
[OPEN SOURCE] MiniMax has released its video generation model H3, capable of generating clips of up to 15 seconds, 2K resolution, with native stereo sound, and plans to release the weights within days.
[ALL-MODAL INPUT] H3 accepts four input modalities—text, image, video, and audio—and supports video editing, as well as transferring motion from one video to another. Pricing is aimed at commercial use and is said to be more than two-thirds cheaper than comparable products, while also running on homegrown chips. Founded in 2022, MiniMax listed in Hong Kong in January this year, becoming the second large-model company to go public there after Zhipu. Video generation has long been a field dominated by closed-source models; Chinese companies are now bringing the open-source playbook into the arena.
[OPEN SOURCE LIMITS] Open-sourcing video models is harder than language models due to inference cost—the weights are free, but the compute is still on you. The VRAM and time consumed by a 2K clip with audio are not something ordinary teams can casually shoulder. The first true beneficiaries are small and mid-size studios with stable GPU supply, which for the first time can fold repetitive work like title sequences and transitions into their own pipelines, without paying per second for API calls. As for closed-source video vendors, their pricing room has been compressed to the line above free weights.
▪ SIGNAL For video, open source has caught up to the release week for the first time.
❯ WSJ Reports Tesla Prepared to Spin Off China Business for Potential SpaceX Merger; Musk Denies
[REPORT & DENIAL] A July 30 Wall Street Journal report said Tesla executives had been asked to prepare for a spin-off of its China business to clear the way for a potential merger with SpaceX, with options discussed by advisers including selling, spinning off, or shutting it down outright. Musk responded on X, calling it “fake news” and saying the matter was never discussed.
[THE CHINA QUESTION] The obstacle is structural — and it predates this report. SpaceX is a major U.S. defense contractor deeply involved in national security and satellite programs, while Tesla operates a wholly owned manufacturing base in China. If the two merged, Chinese assets would land directly on the balance sheet of an American defense contractor. In recent years, Musk has repeatedly demanded that Tesla draw a “laser” line between its U.S. and China operations, aiming to ensure that if geopolitical conditions deteriorate, at least the American half survives. The timing of this round of discussion heating up coincides with the progress of SpaceX’s record-breaking $75 billion initial public offering. Other reports say xAI is also among the merger candidates he is weighing.
[EITHER WAY] The denial does not change the constraints: as long as SpaceX’s defense identity and Tesla’s China production capacity exist simultaneously, the question stays on the table. Those who must answer first are Tesla’s China supply-chain partners — the counterparty behind their long-term contracts may no longer be the company it is today. As for investors, rather than judging whether the report is true or false, they should keep an eye on how SpaceX’s IPO filings describe related assets.
▪ SIGNAL What Musk denied is the plan, not the question.
❯ Eric Trump-backed counter-drone company Space-Eyes to go public via SPAC at $638 million valuation
[SPAC DEAL] According to Reuters, AI counter-drone and geospatial intelligence company Space-Eyes has agreed to go public through a merger with special purpose acquisition company McKinley Acquisition, at a post-merger valuation of $638 million, and plans to trade on Nasdaq under the ticker CUAS.
[VALUATION GAP] Reuters reported that the company’s self-developed CATE AI fusion engine integrates multi-source sensor inputs, including radar and satellite, to identify and intercept drones threatening critical infrastructure and military bases. Annual revenue is about $1 million, with contracts under negotiation totaling about $35 million over five years — the valuation is more than 600 times current revenue. The deal is additionally backed by up to $75 million in PIPE financing, with closing expected in the fourth quarter. The Trump family has repeatedly participated in similar structures; last year they backed a $300 million SPAC focused on U.S. manufacturing. Eric Trump recently became the company’s third-largest individual investor, will serve as strategic advisor after the deal closes, and has already recommended board candidates to the new company.
[WHAT'S BEING BOUGHT] Under the deal terms reported by Reuters, the $638 million figure clearly corresponds not to that $1 million in revenue but to access to government contracts — counter-drone is currently one of the fastest-growing segments in U.S. government procurement, and getting on the list matters more than having the better algorithm. This class of targets puts an explicit price tag on political relationships, and the most direct impact lands on the fundraising narrative of startups in the same lane: beyond technology, investors start asking who you know.
▪ SIGNAL $1 million in revenue holding up a $638 million valuation — what’s being priced is the door to government procurement.
OUTLOOK
[TODAY'S BATCH] Put DeepSeek giving away weights for free, Zhipu raising prices for a third time, and Amazon paying off $50 billion side by side, and they’re all saying the same thing: model capability is rapidly depreciating while model-adjacent assets are appreciating — stable compute supply, settleable credits, and channels to government and regulators. The 69 turbines in Memphis burning through 2027 and Xiaohongshu securing 600 MW in Ulanqab mark the physical boundary of this round of appreciation: without enough power and chips, all the accounting above is moot. Seoul’s 18% swing on that day is the outward sign that this set of accounts is not yet settled.