Back to AI Daily home

❯ DeepSeek V4-Flash 0731 Third-Party Benchmarks Are In: Just 13B Active Parameters, Agent Score Beats Its Own Pro Edition

[BENCHMARKS] DeepSeek’s V4-Flash 0731, released late last month, was the subject of a wave of third-party evaluations over the past day: the official TerminalBench 2.1 score of 82.7 matches exactly what independent evaluator Ante harness produced, and the latter also ranks it above xAI’s Grok-4.5. Semiconductor research firm SemiAnalysis publicly congratulated the team, saying it substantially outperforms NVIDIA’s Nemotron3 Ultra on agent tasks while using 4.2x fewer active parameters and nearly half the total parameters.

[ARCHITECTURE] The result was delivered by a “small” model: a mixture-of-experts architecture with 284B total parameters and just 13B active, carrying a 1 million token context. The model structure is identical to the earlier Flash-Preview — only post-training was redone — and TerminalBench jumped from 72.1 on Pro-Preview to 82.7, the budget version overtaking its own premium edition by 14.7%.

[NINE-MONTH JUMP] The longitudinal comparison is even more striking. Data compiled by researcher Teortaxes shows V3.2 from nine months ago scored just 4.0% on comparable agent benchmarks, while 0731 delivers 61.4% at one-third the cost. Multiple evaluators reach the same price-performance verdict: Vals measured it as 28x cheaper than Grok, with comparable scores.

[PRESSURE] The pressure first lands on the pricing sheets of closed-source labs. When an open-weight, 13B-active model can crack the top tier of agent leaderboards, enterprise customers buying coding and agent solutions will need new reasons to pay a premium for closed-source APIs. The developer community is already predicting that, at this post-training iteration pace, DeepSeek touching closed-flagship score territory before year’s end is no joke.

▪ SIGNAL Redoing post-training once was enough to overtake the premium version, showing the bar for agent capability is shifting from “stacking parameters” to “refining the recipe” — precisely the stage where the open-source camp is closing the gap fastest.

❯ OpenAI designates Astra as first “critical” cybersecurity model, slows release

[CLASSIFICATION] OpenAI this week confirmed that the upcoming Astra is the first model in its history to potentially qualify for the “critical” highest risk tier in cybersecurity — internal evaluations show a major leap in its autonomous coding and network attack/defense capabilities, and after expert review the company cannot rule out that it reaches the top tier of the Preparedness Framework. Axios reports that OpenAI has therefore slowed the release of Astra.

[MEASURES] In parallel, the company has suspended internal use of Astra in scenarios lacking safeguards, placed the model under full monitoring, and is cooperating with government agencies and AI safety organizations on supplementary testing. This comes at a sensitive juncture: over the past few weeks, multiple labs have suffered AI-related network intrusions, and OpenAI itself has just disclosed two security incidents from third-party evaluations (the UK AI Safety Institute and testing partner Irregular).

[RUMOR] Separately, community rumors claim Astra has completed training and is only awaiting safety review, and that its successor model, codenamed “Doug,” has a larger pretraining scale. Both claims come from a single leaker account, with no official or media corroboration — for now, they remain rumors.

[PRECEDENT] For other labs, this is a live demonstration of risk-tiering systems: Anthropic just responded to the UK AI Safety Institute regarding Claude Mythos 5’s boundary-crossing behavior in network tests, while OpenAI is directly staking its release schedule on evaluation results. For the first time, frontier model launch timelines are being publicly determined by safety assessments rather than product calendars.

▪ SIGNAL The “critical” designation moves from a paper framework to a real release gate — safety assessors, for the first time, hold frontier model timelines in their grip.

❯ Claude Code Defaults to Auto Mode Starting August 14; Anthropic Says It Blocks Harmful Actions More Accurately Than Human Review

[CHANGE] Anthropic announced that starting August 14, Claude Code will enable auto mode by default for Pro, Max, and Team subscribers: the model will no longer request human approval at every step, pausing to ask only when an action is deemed irreversible, destructive, or directed outside the environment. Users who have pinned a different default mode are unaffected.

[DATA] Behind the decision is comparative data Anthropic published: in an earlier experiment with 1,053 paid testers, auto mode’s classifier blocked 89% of harmful actions, while step-by-step human approval caught only 13.6% — people habitually clicked “approve” and let dangerous actions slip through. The company says teams with auto mode enabled also produced roughly 25% more code merge requests.

[PACKAGE] Also shipping alongside are prompt-injection screening and customizable hard refusal rules; Anthropic announced it will no longer charge subscribers for the extra tokens the classifier consumes on each tool call.

[SHIFT] This amounts to Anthropic publicly declaring the safety ritual of step-by-step human approval obsolete — developer attention is the real scarce resource, so rather than burn it on confirmation dialogs, hand it to the classifier. What enterprise security teams will have to verify next is whether that set of hard refusal rules can catch their own compliance red lines.

▪ SIGNAL 13.6% vs. 89% — that pair of numbers reclassifies “human-in-the-loop” from safety guarantee to security vulnerability.

❯ SemiAnalysis Estimates: SpaceX to Build ~10GW Compute by End of 2027, Annual Revenue Run Rate Up to $300B

[KEY ESTIMATE] Semiconductor research firm SemiAnalysis published a lengthy analysis: SpaceX is on track to have around 10GW of compute capacity by end-2027, with 6–8GW delivered within 2027 alone. Assuming 50% of that capacity is monetized, it calculates this could support a $300 billion annual revenue run rate — the report estimates AI will account for $261 billion of SpaceX’s total run rate at that point, six times the combined total of its other businesses.

[WHY IT] The report says two pillars underpin the estimate. First, delivery speed: SpaceX takes just 3–5 months from groundbreaking to power-on, far faster than traditional data centers, which lets it command the industry’s highest pricing at $30–50 million per MW per year, per the report’s estimate. Second, capital loop: the report estimates capex of roughly $50 billion per GW, bringing total investment in 2027 to the $300–500 billion range, addressed via Nvidia supplier financing and other means.

[BIG BUYER] The report also discloses that Microsoft could become the largest offtaker: facing its own compute gap, the two sides are reportedly negotiating a 3GW, $150 billion mega-contract. Earlier, Nvidia was just reported to be planning an investment of up to $3 billion in Lancium, the power-infrastructure company behind Stargate — the upstream land grab for compute capacity has now reached the power layer.

[WHO'S HIT] The first to be unsettled will be traditional compute-cloud vendors and hyperscale buyers with self-built data centers. As SpaceX compresses data center delivery timelines using rocket-production-line logic, “how fast can it come online” displaces “how much per watt” as the first question on the bidding table.

▪ SIGNAL The real bet in this estimate comes down to a single variable: whether the 3–5 month delivery cycle can be replicated at GW scale — if it can, pricing power follows delivery speed.

[KEY POINTS] In an August 8 post on X, Musk said the V3 satellites, set to launch on Starship, outperform the current V2 by an order of magnitude — the full V3 system will deliver over 100x the bandwidth of V2. He also disclosed that Starlink will reach $20 billion in annual recurring revenue this year.

[THE MATH] Per specs SpaceX previously released, each V3 satellite provides 1 Tbps of downlink capacity, about 10x that of V2. A single Starship launch can deploy 60 satellites, adding roughly 60 Tbps to the network — over 20x a single Falcon 9 V2 launch. In his post, Musk ran an even more aggressive calculation: even if per-GB revenue falls to one-tenth, SpaceX communications revenue, by his estimate, would still exceed $200 billion a year.

[TWO TRACKS] Read together with the previous item, SpaceX’s two revenue curves — communications and compute — share the same lever: Starship’s payload capacity and launch cadence. Communications needs it to put V3 satellites into orbit in bulk; the 10GW compute blueprint is equally staked on its production speed. Every Starship test flight now prices both hundred-billion-dollar stories at once, and satellite-internet rivals and compute buyers alike are watching the same launch schedule.

▪ SIGNAL The $20 billion is this year’s realized revenue; the $200 billion is a promissory note payable once capacity is delivered — and what separates the two is precisely Starship’s production ramp.

❯ Report: Apple Tests CXMT Chips, Plans Use in China-Market iPhones and MacBooks

[TALKS] According to media reports, Apple is testing CXMT chips across product lines including iPhone and MacBook, and has opened preliminary talks with China’s largest memory chipmaker on component supply, with plans to first use them in some devices sold in China. CXMT is also considering building a second memory fab in Beijing to expand capacity.

[CONTEXT] Apple isn’t the first mover: HP and Acer already use CXMT chips in devices sold outside the U.S. The driver is the same — the AI boom has drained high-bandwidth memory capacity, causing a global memory chip shortage and surging prices, and forcing PC and phone makers to hunt for a fourth option beyond Samsung, SK Hynix, and Micron. The report also notes Apple wants White House approval before doing business with CXMT.

[SHIFT] For CXMT, this deal matters far beyond the revenue itself: clearing Apple’s certification and validation process is the equivalent of earning a top-tier ticket into the global consumer electronics supply chain. For Samsung and SK Hynix, the alarm is that Chinese memory capacity has, for the first time, used the global shortage window to squeeze onto the qualified supplier lists of flagship customers — and those lists won’t clear automatically once the shortage ends.

▪ SIGNAL Memory shortages are cyclical; supplier qualification is permanent — CXMT is trading one cycle for a long-term ticket.

❯ Unitree Robotics Opens Subscription Tomorrow: Issue Price 150.80 Yuan, Offline Subscription Multiple 2,618x, DeepSeek in Strategic Placement

[OFFERING] The “first humanoid-robot stock” Unitree Robotics kicks off its STAR Market subscription on August 10: issue price 150.80 yuan/share, planned total raise RMB 6.099 billion, about 45% above the target, issue market cap near RMB 61 billion, corresponding to a P/E of 219 times. Offline subscription has been swamped — effective subscription multiple reached 2,618.30 times, with BlackRock and multiple regional occupational pension plans among the bidders.

[SHAREHOLDERS] According to the prospectus, the 44 institutional shareholders before listing include internet giants Meituan, Alibaba, Tencent, ByteDance, plus Sequoia China and Ant Group, with the Meituan camp being the largest external institutional shareholder. The strategic placement list is even more noteworthy: three portfolios of the National Social Security Fund received allocations, and DeepSeek’s parent company also secured a slot, and will cooperate with Unitree on general AI, high-performance robotics, and large models — a large-model company directly taking a stake in a robot-body maker is a first on the A-share market.

[CONTROL & CHAIN] The company is controlled by founder Wang Xingxing, with post-issuance voting rights not exceeding 65.31%; core employees received 4.45% via two asset-management plans. The supply chain spans the entire humanoid-robot industry chain, with A-share companies such as Zhongda Leader, Wolong Electric Drive, and Orbbec supplying components. The IPO subscription fever has already rippled into the share prices of these suppliers ahead of the listing.

[PRICING] A 219x P/E is not paying for current earnings; it’s an option on the humanoid-robot mass-production timeline. After listing, Unitree’s quarterly shipments and gross margins will be benchmarked against this valuation — this stock is the yardstick for how much patience the secondary market has with embodied AI.

▪ SIGNAL DeepSeek’s stake in Unitree turns the convergence of “brain” and “body” from a forum topic into an equity structure.

❯ Filings Show Moonshot AI Converted to Joint-Stock Company, Taking First Step Toward Hong Kong IPO

[RESTRUCTURING] The Financial Times reports that filings show Kimi developer Moonshot AI has converted its mainland China entity from a limited liability company into a joint-stock company — the statutory prerequisite for a domestic company to go public, and the first visible step in its preparation for a Hong Kong IPO. The company has already begun dismantling its red-chip VIE structure and reportedly plans to appoint CICC and Goldman Sachs as joint sponsors.

[FINANCES] The restructuring is underpinned by a steep revenue curve already disclosed: according to reports, annual recurring revenue crossed $100 million in Q1, $200 million in May, and $300 million in June. The company has raised at least $4 billion cumulatively this year, the most of any domestic large-model startup. Market sources say it is preparing a final pre-IPO round, targeting a pre-money valuation of up to $50 billion — six months ago, it was valued at less than one-tenth of that.

[WINDOW RACE] A queue of domestic AI companies is already lining up for Hong Kong listings: Zhipu and MiniMax are at the front, and Moore Threads today also announced it is launching its H-share plan. For Moonshot AI, listing first locks in the fresh capital it needs to keep pace in the inference-compute arms race with ByteDance and Alibaba; for HKEX, whoever seizes the pricing anchor of the “first large-model stock” sets the valuation benchmark for the entire queue behind.

▪ SIGNAL The restructuring filing is stronger evidence than any funding rumor — on going public, Moonshot AI has moved from “consideration” to “construction.”

❯ Moore Threads Board Approves H-Share Issuance Proposal, Plans HKEX Main Board Listing

[ANNOUNCEMENT] Domestic GPU maker Moore Threads announced that on August 7 its board reviewed and approved a proposal to issue H shares and list on the HKEX Main Board, to be carried out at an opportune time within the validity period of the shareholders’ resolution. The company cautioned that the issuance still requires shareholder review and filing approvals from regulators including the CSRC and the HKEX; specific details remain undetermined, and whether it can be implemented is subject to significant uncertainty.

[TIMING & RESULTS] Moore Threads only debuted on the STAR Market at the end of 2025, and launching a secondary listing less than a year after listing is a pace that ranks as aggressive among domestic chip companies. The confidence comes from its half-year report: the announcement shows H1 2026 revenue of RMB 1.736 billion, up 147% year over year, with AI training and inference cards as the main growth driver. The company said the Hong Kong listing is aimed at deepening its internationalization strategy, continuing to attract R&D and management talent, and improving corporate governance.

[CAPITAL] GPUs are the fastest cash-burning track: a single tape-out costs on the scale of hundreds of millions of yuan, and an A+H dual financing channel is the equivalent of adding another capital pipeline for advanced-node tape-outs and R&D spending. For Hong Kong investors, this will be yet another directly tradable domestic compute target; for peers on the same track — Biren and Enflame — the breadth of the financing channel itself is already widening the gap.

▪ SIGNAL A+H is becoming the default play for domestic chip companies — the financing channel itself is the armament.

❯ Tencent Bets Its Highest-Level Resources on WorkBuddy, Internally Dubbed Its Third Strategic Product After QQ and WeChat

[RESOURCE TILT] According to media reports, the AI office product WorkBuddy is already one of Tencent’s highest strategic-priority AI applications: the company’s channel, compute, organizational, and ecosystem resources are being concentrated behind it. Pony Ma has personally attended product meetings, and the speed at which the team’s resource requests are approved is described internally as “green lights all the way.” This year, WorkBuddy is highly likely to be inducted into Tencent’s milestone-product honor, the “Hall of Fame” — a distinction previously reserved for products on the scale of QQ and Official Accounts.

[MARKET POSITION] The resources were not wasted. Earlier third-party data showed that in Q2 2026, WorkBuddy ranked first among domestic AI office agents with 20.97 million monthly PC visits, exceeding the combined total of second-place ByteDance’s TRAE and third-place Alibaba’s QoderWork. The internal framing of “the third strategic product after QQ and WeChat” lifts a recently launched enterprise tool straight to the same tier as two national super-apps — something with no precedent in Tencent’s history.

[THREE-WAY RACE] Across the field stand ByteDance and Alibaba — all three are vying for AI office agents as enterprise-grade traffic gateways. Tencent’s differentiating card is the connectivity of the WeChat ecosystem; for enterprise customers, the real choice is which vendor’s agent their workflows settle into, as switching costs deepen with every automated process they build.

▪ SIGNAL The Hall of Fame’s selection criteria have shifted from national apps with hundreds of millions of users to an enterprise AI product — Tencent’s definition of “strategic” has changed.