Back to home
This issue or one story

What would you like to share?

Share the full issue or choose a story worth reading.

❯ NVIDIA AI Server Systems Rise ~17%, Adding $5 Billion to a 1-Gigawatt Data Center’s Cost

PRICINGNVIDIA has notified major customers through server vendors that its mainline AI server systems will rise about 17%, with the increase applied to machines shipping early next year and spanning both the Vera Rubin and Grace Blackwell platform generations. According to The Information, this single item alone adds at least $5 billion to the system procurement cost of a 1-gigawatt data center — before power, cooling, networking, construction, and financing are even counted.

DRIVERThis round of increases isn’t NVIDIA marking up prices — it’s NVIDIA buckling under the pressure itself. Data from market research firm TrendForce shows server DRAM contract prices rose 53% to 58% quarter over quarter in Q2, with another 13% to 18% expected in Q3. Bloomberg earlier reported a more conservative figure: multiple top-tier customers were told of an increase of 15% or more, and the two outlets’ numbers point in the same direction. Memory price inflation has now rippled from consumer grade all the way to rack scale — the first time it has been explicitly written into NVIDIA’s price sheet.

FALLOUTFor the past two years, every hyperscale data center’s ROI calculation has rested on the default assumption that the unit price of compute declines every year. A 17% jump overturns that curve in one shot — and it lands only after projects have signed power agreements and locked in land. The first to be recalculated are cloud providers’ depreciation schedules and lease pricing — the same GPU fleet must absorb the added cost, either by stretching lease terms or raising the per-hour rate. And downstream application companies whose business models depend on inference costs continuing to fall may find next year’s gross margin model has to be torn up entirely.

▮ SIGNALThe memory shortage has finally gone from a supply-chain story to a deduction line in every AI project financial model.

❯ Nvidia in Talks to Invest in Perplexity at $30B+ Valuation, After Weighing Tech License and Staff Poaching

DEALAccording to The Information, Nvidia is in talks to join a new equity round for AI search company Perplexity at a post-money valuation above $30 billion — up more than 50% from roughly $20 billion in the year-ago round — with the round potentially reaching several billion dollars. The report also notes a detail worth pondering: Nvidia’s original plan was not to take a stake, but to spend billions licensing Perplexity’s technology and poaching some core employees, before shifting to a direct equity stake.

REVENUESupporting that valuation is a steep revenue curve. Perplexity’s annualized revenue has climbed from under $250 million at the start of the year to more than $750 million — tripling within the year. Bezos also stands behind the company. For Nvidia, Perplexity is just one in a string of moves this month — it has also invested in three data-center power and land developers in the same stretch, while open-sourcing its proprietary models for free.

DEMANDStrung together, these moves constitute the same playbook: using capital to lock in future chip demand. That the license-plus-poaching plan was dropped shows Nvidia wants more than the technology itself — it wants Perplexity to continue as an independent, steadily purchasing major customer. The founders who need to recalculate are those in the AI application layer: when the most upstream chip supplier is also your shareholder, valuations can be negotiated very high, but how much bargaining power and independence of technical direction get diluted in tandem is something that must be thought through well in advance.

▮ SIGNALJensen Huang is using equity to lock in customers’ future orders ahead of time — far more cost-effective than selling a batch of GPUs.

❯ SoftBank Plans ~$6.3 Billion Retail Bond Sale, a Japan Record, to Fund OpenAI Commitment

ISSUE SIZESoftBank Group plans to issue 1 trillion yen (about $6.3 billion) in retail bonds — the largest retail bond ever from any Japanese issuer and SoftBank’s third this year. According to Bloomberg, the 7-year bonds are expected to be priced on September 4, with a guided coupon range of 4.3% to 4.9%, raising funds for AI-related investments and repaying old debt. The size is nearly double SoftBank’s previous record of 600 billion yen from April 2025.

USE OF FUNDSSoftBank’s investment commitments to OpenAI already exceed $60 billion, and it is also accelerating data-center construction to expand computing capacity. The question is where the money will come from. Three consecutive rounds of retail bond issuance in itself signals that banks have little appetite for fully underwriting SoftBank’s AI exposure — the risk institutions won’t take is ultimately being placed on the counter of Japan’s individual investors. SoftBank shares fell after the announcement. The 600 billion yen retail bond in April was already a record at the time; four months later, the company has nearly doubled that figure itself.

RETAIL BUYERSA coupon of 4.3% to 4.9% is quite attractive in the Japanese market, especially for retail investors accustomed to near-zero interest rates. But buying this bond essentially ties household savings to OpenAI’s commercialization progress, with SoftBank’s own leverage structure in between. Japan’s individual investors are now standing on the exposed side of the AI capex cycle for the first time, so directly — and the information available to them is far less than what the banks that politely declined this deal had.

▮ SIGNALWhen institutional money starts to get picky, retail savings accounts become the last funding link in the AI infrastructure chain.

❯ Files Show Meta Agent Hatch Could Launch as Early as Late August; New Model Watermelon Set for October

TIMELINEAccording to internal documents seen by The Information reporter Jyoti Mann, Meta’s consumer-grade agent Hatch, built to rival OpenClaw, is planned to launch in late August or early September, with the company’s newest model Watermelon set for October. An explicit reason Meta built Hatch: OpenClaw is red-hot in the tech community but too complex for ordinary users. Hatch is designed to handle operations like shopping directly on Instagram; Meta has already tested it in simulated environments of DoorDash, Reddit, and Outlook.

CONVERGENCEHatch has been trained using Anthropic’s Claude Opus 4.6 and Sonnet 4.6, with a plan to switch to Meta’s in-house Muse Spark at official launch. Watermelon, though, is the heavier track — in July, Meta’s head of superintelligence, Alexandr Wang, told an internal all-hands that Watermelon has matched OpenAI’s GPT-5.5 on internal benchmarks, with training compute an order of magnitude higher than the previous-generation Muse Spark. The specific benchmarks remain undisclosed.

TIMINGPut the two dates together, and Meta has laid out a schedule in which model and product back each other up: Hatch seizes the entry point to consumer-agent use cases in September, then Watermelon swaps in as the underlying capability in October. The fallout hits Anthropic’s enterprise customer base — a product trained on your model drops you on launch day, and that script is likely to play out more than once. Model vendors, while supplying their biggest customers, are simultaneously cultivating their own replacements.

▮ SIGNALTrained on someone else’s model, launched on its own — Meta flips between API customer and competitor with zero transition.

❯ Musk’s First All-Hands Address to Cursor: Grok Lags, Anthropic Leads

ALL-HANDSAfter SpaceX completed its $60 billion acquisition of Cursor, Musk addressed all Cursor employees for the first time. According to The Information reporter Grace Kay, he bluntly said that Grok has fallen behind rivals, that he is “not used to losing,” and that Anthropic is currently leading the race, adding one judgment: AI will eventually become impossible for humans to control.

NUMBERSThis wasn’t false modesty. On Artificial Analysis’s Intelligence Index, Grok 4.6 is roughly on par with OpenAI’s GPT-5.6 Sol Max, trailing Anthropic’s Fable 5 Max. At the meeting, Musk previewed that Grok 4.7 will be released in three to four weeks, and after supplemental training on SpaceX’s vast data, it will “surpass all existing models.” xAI had already been folded into SpaceX; with Cursor added, compute, proprietary data, and a coding entry point now sit under one roof.

AFTERMATHA founder openly admitting defeat in front of a team he just bought for $60 billion is usually not an outburst but a tone-setter: resource allocation will now be organized around “catching up,” not “holding ground.” For Cursor’s engineers, with that statement, the product roadmap will likely give way to the pace of model catch-up. Those developers who treat Cursor as a neutral tool layer need to reassess — it is now a catch-up component of a model company, and how much neutrality remains will be answered by Grok 4.7’s performance in three to four weeks.

▮ SIGNALThe $60 billion bought an opening message of admitted lag — Musk wants not to reassure the team but to reprioritize.

❯ NVIDIA Groq 3 LPX Enters Mass Production: 100K-Token Long Context Benchmarked at 3,400 Tokens/Sec

PRODUCTION & BENCHPer NVIDIA’s announcement, the dedicated inference accelerator Groq 3 LPX has entered full mass production. In Artificial Analysis’s standard benchmark, running the open-source model Gemma 4 31B with a 100K-token active context, it delivered 3,400 tokens per second output — 4x the next-best platform under the same long-context conditions, and the fastest result ever recorded for this model. A single rack can hold up to 256 LPUs, with 128GB of high-bandwidth SRAM.

BACKSTORYThe chip is positioned as a complement to the Vera Rubin NVL72 platform, targeting agent workloads that need sustained token generation over long stretches. What’s more notable is the timeline: per The Register, eight months have passed since NVIDIA secured a non-exclusive technology license from Groq — the first time Groq technology has appeared in an NVIDIA rack-scale product. Cloud provider Nebius is the first to deploy it, making it available through its Token Factory service.

INFERENCE SPLITTraining and inference are being split into two independent hardware product lines, making tokens per second at long context an independently priced metric. The first constraint to loosen is the response-latency budget for agent products: an agent chaining dozens of sequential tool calls was previously limited by generation speed to asynchronous interaction. At this throughput level, synchronous real-time product forms become viable for the first time — and per-task inference costs are being recalculated accordingly.

▮ SIGNALPer The Register, that roughly $20 billion license has, for the first time, taken a shape that benchmarks can record.

❯ Xiaomi Unveils Three Xuanjie Chips; D100 Supports Local Deployment of 200-Billion-Parameter Model

TRIPLE LAUNCHXiaomi officially disclosed at a technical briefing that it is releasing three self-developed chips at once. The flagship SoC Xuanjie O3 uses a 3nm process, has a die area of 133 square millimeters, and integrates 24 billion transistors (up 26% from the previous generation). Its CPU is a 10-core architecture with six ultra-large cores and four large cores, running up to 4.35 GHz — a 60% performance improvement. The GPU is a 16-core design, 85% faster. The O3 has already entered volume production and will debut in the Xiaomi 18 Fold foldable next month.

AUTO & EDGEThe other two chips better illustrate the direction. The Xuanjie D100 is what Xiaomi officially calls China’s first 3nm high-compute AI chip for intelligent driving, with a 20-core CPU and a 16-core NPU, supporting up to 160GB of unified memory, which enables on-board local deployment of a 200-billion-parameter model. R&D is complete, and it will go into commercial use next year. The Xuanjie O100, meanwhile, is a high-bandwidth accelerator built specifically for on-device large models. It uses the industry’s first 6nm wafer-level vertical stacking and hybrid bonding process, with a bonding pitch of 1.4 micrometers and bandwidth of 1.22TB/s — 16 times that of traditional flagship phones — with on-device inference speeds of up to 330 TPS.

MEMORY WALLThe three chips point to the same conclusion: the bottleneck for on-device large models is not compute but memory bandwidth. The O100 uses advanced packaging to lift bandwidth by an order of magnitude, while the D100 uses 160GB of unified memory to raise the ceiling on “how large a model can run in a car.” If 200 billion parameters can stay local, cloud inference call volume, data-compliance paths, and the charging rationale for subscription models will all need to be reframed — and that line was drawn by Xiaomi itself, not handed down by a chip supplier.

▮ SIGNALXiaomi put its most aggressive chip in the car, not the phone.

❯ Xpeng’s Robotics Business Raises Over $900M at $6.3B Valuation; IDG Leads, Tencent and Alibaba Follow

RECORD DEBUTAccording to an Xpeng announcement, its robotics business has completed its first financing round, raising over $900 million at a post-money valuation exceeding $6.3 billion, setting a new record for a single private financing round in China’s embodied AI sector. The round was led by IDG Capital, with Tencent and Alibaba coming in as strategic investors; Gaorong Capital, an early backer of Momenta and Zhiyuan Robotics, also participated.

USE OF FUNDSThe use of proceeds is spelled out in specific terms: robotics hardware and software R&D, training and fine-tuning of physical AI models, high-quality data collection, construction of end-to-end mass-production lines, and overseas expansion. Xpeng’s previously set target is for its humanoid robot IRON to reach monthly production of 1,000 units by year-end, with initial deployment in its own stores and industrial parks; commercial sales and deliveries are set for 2027, launching simultaneously in China and overseas markets.

PRODUCTION WATERSHEDA robotics company that has yet to sell a single product has secured a $6.3 billion valuation — the basis is not revenue but the visibility of mass-production capability. The two figures — 1,000 units a month and 2027 delivery — happen to be among the few commitments in this industry that outsiders can verify. What now has to re-queue is the financing cadence of other domestic humanoid robotics companies — now that the leader has pushed the single-round record to the $900 million level, the first question mid-tier players face in valuation talks will be: where is the production line, and what is the monthly output?

▮ SIGNALThe valuation anchor for embodied AI is shifting from research papers and demo videos to production lines and delivery timelines.

❯ ByteDance Folds Trae and Coze Into Doubao, Launches Doubao Work This Week to Take On Tencent’s WorkBuddy

ORG-FIRSTBloomberg reports that ByteDance is folding the teams behind AI coding platform Trae and agent-building platform Coze into Doubao, and plans to launch Doubao Work as a standalone app as soon as this week, going head-to-head with Tencent’s WorkBuddy — currently China’s most widely used office AI tool.

CONSOLIDATIONThis is a textbook “all-in on one super-app” restructuring. Trae launched in China in March 2025, one of the first domestic AI coding tools to integrate the Doubao 1.5 Pro and DeepSeek models; Coze is ByteDance’s agent-building platform launched in 2024. Neither targets a developer audience that overlaps much with Doubao’s consumer base. Alibaba is doing the same, consolidating its AI tools into the office-focused Qwen Work.

OFFICE BATTLEChina’s consumer AI assistants have commoditized to the point of indistinguishability, while the office segment still shows clear willingness to pay and data moats. The first to feel the pressure is Tencent WorkBuddy’s usage lead — ByteDance’s playbook is to hit a rival’s enterprise share with consumer-scale traffic density. For enterprise buyers, lock-in depth is the real cost: once coding, agents, and office assistants sit inside a single app, switching is no longer just swapping software.

▮ SIGNALByteDance is moving its org structure, not product features — a sign it sees office AI as a battle that demands concentrated firepower.

❯ Alibaba Releases Video Generation Model Wan3.0, Which Turns Documents, Spreadsheets, and Slides Directly into 30-Second Videos

RELEASE & CAPABILITIESAlibaba has officially released video generation model Wan3.0, capable of generating up to 30 seconds at up to 1080p in a single run — twice the length of the previous generation. Its most distinctive feature is the input side: it directly accepts DOC, XLS, PPT, PDF, and Markdown files, and can read web pages, turning documents and spreadsheets into finished videos. The model entered public beta on August 6.

PRICING & TIMINGAccording to official pricing on Alibaba Cloud’s Bailian platform, the API is priced at $0.05 per second for 480p, $0.10 per second for 720p, and $0.20 per second for 1080p. Since the public beta opened on August 6, it has been applied to short-drama and film/TV production, advertising and marketing, cultural-tourism promotion, and music-video creation. The release timing is also worth noting: it came the day after Alibaba completed an approximately $10.2 billion share placement on the Hong Kong Stock Exchange, with the proceeds explicitly earmarked for AI infrastructure.

ENTRY SHIFTMoving from prompts to document input shifts the integration point for this class of tools — for the first time, video generation can hook directly onto the back of a company’s existing document workflow, with no need to write a description first. The real bill falls on marketing and content teams’ outsourcing budgets: whether a quarterly report or a product manual can go straight to video determines whether that money is saved — or simply becomes one more subscription.

▮ SIGNALOnly when the input format shifts from prompts to PPT does video generation truly become part of a company’s workflow.

❯ XPeng’s H1 Net Loss Widens 173% to RMB 3.121 Billion; R&D Up 39%, Poured Into Physical AI

LOSS WIDENSXPeng Group’s H1 2026 total operating revenue was RMB 32.777 billion, down 3.8% YoY, while the net loss attributable to shareholders was RMB 3.121 billion, widening 173.35% YoY. Over the same period, however, gross profit came in at RMB 6.766 billion, up 20.25% YoY, lifting overall gross margin to 20.6%, a 4.1-percentage-point improvement YoY — revenue is falling while gross margin is rising. That divergent pair of numbers is the entry point to this earnings report.

SALES & R&DThe report shows H1 deliveries of 165,977 vehicles, down 15.8% YoY, with vehicle sales revenue of RMB 28.05 billion, down 10.3% YoY. The real support for gross margin came from another segment: service and other business revenue of RMB 4.73 billion, surging 67.1% YoY. R&D spending, meanwhile, hit RMB 5.82 billion, up 39% YoY, explicitly directed at new models and at physical AI and humanoid robots. At period end, cash on hand totaled RMB 40.48 billion, with 20,632 employees — more than 8,700 of them in R&D.

LOSS NATURERead this report alongside the USD 900 million funding round for the robotics business announced the same day, and the math is clear: it sold fewer cars, but made more on each one — the savings and the new funding have both been poured into robotics. Q3 guidance calls for deliveries of 115,000 to 121,000 vehicles and revenue of RMB 21.7 billion to 23.4 billion. Valuing this company by delivery volume will become increasingly unreliable — the bulk of R&D is flowing into physical AI, a cost that won’t appear on any delivery statement in the near term, yet will materially shape next year’s cash-burn rate.

▮ SIGNALWhen a carmaker frames the source of its loss as R&D investment, readers must first determine whether this is bleeding or repositioning.

❯ Smart-ring maker Oura plans U.S. IPO to raise up to $3 billion at a valuation exceeding $16 billion

TIMING & SCALEAccording to Bloomberg, smart-ring maker Oura and some existing shareholders are seeking to raise up to $3 billion in a U.S. IPO at a valuation exceeding $16 billion, potentially as soon as September. That is nearly 50% above the $10.9 billion valuation set by last September’s $875 million Series E. The company confidentially filed for the listing in May this year, with underwriters including Goldman Sachs, Morgan Stanley, JPMorgan, Allen & Company, and Jefferies.

EXITSThe report specifically notes that existing investors are expected to sell a significant portion of their shares in the offering; terms are still under discussion and could change. That detail clarifies the nature of the deal: this is less a hardware company tapping the market for growth capital than a long-queued secondary exit. Oura’s $875 million Series E last September was led by existing investors, and the company followed with its confidential filing in May — eight months later.

VALUATION TESTA $16 billion market cap for a company selling rings rests not on hardware margins but on subscription renewal rates and the long-term value of health data. After listing, both metrics go public quarterly. The entire wearable-health sector gets repriced first — Oura is the first of this cohort to truly reach the public market, and its earnings multiple will directly become the valuation ceiling for every comparable company that follows.

▮ SIGNALExisting shareholders choosing to sell at this level is itself a verdict on the $16 billion figure.

❯ Pinduoduo Q2 revenue of 112.36 billion yuan misses estimates; net profit down 12% YoY but beats expectations

UP & DOWNPinduoduo posted second-quarter total revenue of 112.36 billion yuan (about $16.72 billion), up 8.1% year over year but below the 116.35 billion yuan analysts had expected. Net profit came in at 27.18 billion yuan (about $4.04 billion), down 12% from a year earlier, yet above the market’s 24.4 billion yuan expectation. As The Wall Street Journal reporter Tracy Qu reported, the company attributed the pressure to fierce competition in the Chinese market and a shifting overseas regulatory environment.

SQUEEZEDThe earnings release and analyst notes show weak consumer confidence at home — cautious employment expectations and a dragging property market kept spending intentions conservative, and even the 618 festival couldn’t turn that around — while price wars across the e-commerce industry directly compressed profit margins. Overseas, Temu faces the double blow of U.S. tariffs on China and the removal of the duty-free exemption for small parcels, pushing up freight and compliance costs and forcing some merchants to raise prices. Advertising revenue grew just 3.8% in the period.

SLOWING ENGINEPinduoduo has been one of the few Chinese internet companies able to sustain high growth in recent years; the 8.1% figure says that era is over. The first to revise their assumptions are rival e-commerce platforms on the same track: when even the industry’s fastest-growing player is down to single digits, the price war stops being a tool for taking share and becomes a cost everyone must carry.

▮ SIGNALAfter growth falls to single digits, the price war shifts from an offensive strategy into a war of attrition no one can exit.

❯ Geely 500Wh/kg Solid-State Battery: 2027 Multi-Brand Pilot, Volvo Seen as Candidate

TIMELINEGeely says it is accelerating solid-state battery commercialization, planning to launch pilot applications across multiple of its brands in 2027, with Volvo viewed as one of the first candidates to adopt it. The cell delivers an energy density of 500Wh/kg — nearly double that of current lithium iron phosphate batteries — putting equipped models on track for a range of over 1,000 km and a lifespan beyond 1 million km, while also outperforming liquid batteries in safety, lightweighting, and charge-discharge efficiency.

CONSTRAINTSMass production still faces hard constraints. The 2027 pilot is only small-batch testing — consumers won’t see production models any time soon, and the timeline depends on supply-chain maturity. Geely Holding’s brand portfolio spans seven or more tiers, covering Geely, Zeekr, Lynk & Co, Volvo, Polestar, Lotus, and smart. The company has yet to announce the first test brands and initial markets, and the industry broadly expects a phased rollout. Geely previously disclosed that the cell entered real-vehicle validation in 2026; mass production timing depends on the supply chain.

WEIGHTWhat 500Wh/kg truly rewrites isn’t the range figure — it’s the weight and volume budget of the battery pack. At the same range, the battery weighs half as much, freeing up design space across the chassis, suspension, and overall vehicle structure, and the whole-vehicle cost structure shifts along with it. The first thing to come under pressure is vehicle-platform development scheduling: a platform takes three to four years from definition to mass production, and given that 2027 is only small-batch testing, next-generation platforms must decide now whether to be designed around liquid or solid-state physical parameters.

▮ SIGNALBetween the pilot and mass production lies an entire supply chain — this time gap determines who is telling stories and who is scheduling production.

❯ Keep H1 Revenue RMB 825 Million, Adjusted Net Profit RMB 5.88 Million; Proprietary Sports Large Model Enters Core Scenarios

SLIM PROFITKeep released its 2026 interim report: first-half revenue RMB 825 million, adjusted net profit RMB 5.88 million. Two user-side metrics are improving: ARPU grew 21.3% year over year, and monthly active users’ average monthly exercise time grew 15.3% — more people are paying, and they are using it longer.

OWN BRAND LEADSThe report shows hardware is the mainstay of the revenue mix. Own-brand fitness product revenue grew 21.7% year over year to RMB 483 million, gross margin rose to 40.1%, and overseas revenue exceeded RMB 22 million. On the AI front, the proprietary sports-health large model Keepace.ai continues to be deployed into core App scenarios, while also exploring B2B capability output.

THE 5.88M WEIGHTAn adjusted net profit of RMB 5.88 million is essentially zero for a listed company — but it is the only figure in this report that proves this model can sustain itself at its current scale. What to watch is the causal chain between ARPU and AI deployment: if the 21.3% growth in per-user spend truly comes from willingness to pay driven by AI courses and personalized training, the story Keep is telling can keep going; if it was just a price adjustment, next year’s interim report will reveal it.

▮ SIGNALA company has just proven it can stay out of the red; next it must prove this profit can be replicated.

Pass along the stories worth reading.
Get each issue in your reader: Feedly Inoreader RSS feed