Back to AI Daily home

❯ NVIDIA Buys Poolside’s Technology and Team for $7 Billion, Builds Its Own Open-Source Model to Rival DeepSeek

[DEAL] NVIDIA has struck a deal worth roughly $7 billion with code-model company Poolside — $6 billion for the non-exclusive license to its Model Factory training system, plus a $1 billion investment at a $12 billion pre-money valuation. Per The Wall Street Journal, 109 Poolside engineers will fold into NVIDIA’s Nemotron open-weights model project; the founding team will not move over, and the company continues to do research independently.

[SUPPLY] The same week, NVIDIA reached into the data layer. Per The Information, existing investor General Catalyst is leading a round that would push AI data supplier Mercor’s valuation to $20 billion, with NVIDIA in talks to participate — double Mercor’s $10 billion valuation from last October.

[WHAT IT BUYS] Last quarter, NVIDIA already paid Mercor tens of millions of dollars for expert-annotated data in law, finance, and science — the very data feeding Nemotron. If the deal closes, Mercor will surpass Scale AI to become the highest-valued company in the data-labeling space.

[WHY NOW] NVIDIA’s books have long made the motivation clear: in the quarter ended April, it invested $18.58 billion in private equity, versus just $649 million a year earlier. No matter how many chips it sells, as long as Chinese models like DeepSeek and Kimi hold the top spot in open weights, where inference demand flows is not NVIDIA’s call. Buying Poolside fills out the training stack; investing in Mercor locks up the data source — the two deals together form an in-house pipeline from data to weights.

[STAKES] Open-source model vendors are first to recalculate their timelines. Nemotron used to be more of a giveaway to developers; now it has a dedicated 109-person team and an acquired training pipeline. Closed-source labs are feeling the squeeze too — a chip supplier that releases the strongest open weights for free will directly undercut the pricing of their mid- and low-end models. Chinese open-source players like Zhipu and Moonshot AI must reassess how much of their lead window remains.

▪ SIGNALA chipmaker building the strongest open-source model itself was never about selling the model — it’s about becoming the default option that routes everyone’s inference onto its own hardware.

❯ Bloomberg: Nvidia Tells Top Customers AI Server Prices Will Rise Over 15% Starting Early 2027

[PRICE HIKE] Nvidia has informed some of its top customers that AI systems — including Vera Rubin and Grace Blackwell — will see price increases of over 15% on shipments starting early 2027, with the exact increase varying by chip generation and memory configuration. Bloomberg first reported on Aug 22, and Reuters followed up the same day; Nvidia has not yet commented, so this remains a supply-chain account, not an official price announcement.

[COST DRIVER] The direct driver is memory. HBM and DRAM contract prices have been climbing, while the memory stacked into each AI server is doubling with every generation — memory has gone from a component to the biggest cost line item. Contract manufacturers building full systems for Microsoft, Google, and Oracle have already passed the expected increase on to their respective customers.

[PERF CLAIM CHECK] Running parallel to the price hike is another story: hardware analyst zephyr_z9 notes that Nvidia’s claimed 10x token throughput for Rubin versus Blackwell only holds at extremely high interaction rates. Labs typically serve users at 50 to 60 tokens per second; in that range, the improvement falls far short of 10x. The GB300 full system previously sold for around $4 million — value for money has to be calculated per token, not papered over with peak benchmark scores.

[STAKES] Cloud providers’ capex models need to move up a notch. Every compute budget for 2027 will have to be redone with at least a 15% increase, and that money will eventually flow downstream — first into hourly GPU rental rates, then into per-million-token API pricing. AI companies signing long-term contracts at current prices are locking in a cost dividend ahead of time; those still on the sidelines will pay more to board next year.

▪ SIGNALMemory prices have risen enough to rewrite system price tags — the bottleneck in AI infrastructure has shifted from compute to the supply chain itself.

❯ Ramp Data: Fable 5 Enterprise Spending Stalls at 11%, Overtaken by Cheaper Opus 5

[BILLING DATA] Payments company Ramp analyzed bills from 70,000 businesses and found that Anthropic’s flagship model Fable 5, released this June, accounted for only 11.4% of customer spending on Anthropic tools two months later — with its token-usage share even lower, at just 6%. Claude Opus 5, launched in late July at a lower price, has already overtaken Fable 5 in enterprise spending. The Financial Times reports the data was provided exclusively by Ramp.

[PRICE GAP] The gap is plain: Fable 5 runs roughly $10 per million tokens, about twice the price of GPT-5.6 Sol. Enterprise buyers would have to pay double for performance gains that are hard to quantify in day-to-day work, and after running the numbers, most companies chose not to pay. Cheaper alternatives are squeezing in from both sides at once — Chinese open-weight models and Anthropic’s own Opus 5.

[CORROBORATING EVIDENCE] Vercel’s AI gateway data points the same way: open-weight models now account for 29% of tokens on the gateway, up from 11% in April — nearly a threefold jump in two months — while consuming less than 4% of spending. Volume is on the open side, money on the closed side; that divergence is what “good enough” looks like on a ledger.

[STAKES] Frontier labs need to revisit the pricing assumption itself. The willingness to pay a premium for top performance is not unlimited — it has a ceiling, and Fable 5 just hit it. More awkward for Anthropic: half the pressure pushing it down comes from its own Opus 5 — the flagship’s premium headroom has been eaten by the sub-flagship. Enterprise tech leads, meanwhile, have been handed a ready-made bargaining chip: they can bring this data to the next renewal cycle and squeeze API pricing.

▪ SIGNALEnterprises accept the strongest model, but not its price tag. Between performance leadership and revenue sits a quantifiable price gap.

❯ Alibaba Placement Raises $10B, All Earmarked for Full-Stack AI From Chips to Models

[DEAL TERMS] Alibaba announced a placement of 710 million new shares at HK$112.70 each, a 3.6% discount to last Friday’s closing price, raising roughly $10 billion. Net proceeds will go 100% into “full-stack” AI capabilities spanning chips, infrastructure, and model R&D. It is the largest primary-market placement in Hong Kong stock-market history, and the third-largest globally this year, trailing only Alphabet’s $80 billion in June and Intel’s $15 billion in August.

[DEMAND & CONTEXT] The book was oversubscribed, prompting Alibaba to upsize the offering, with sovereign wealth funds among the buyers. The underwriting syndicate comprises Morgan Stanley, HSBC, UBS, and CICC. The real backdrop is the ledger: net profit fell 75% YoY in the April–June quarter, with nearly half of the three-year capex plan already spent. AI spending is eating into profit while the balance sheet needs more ammunition. Capex in the April–June quarter approached $10 billion, so this placement effectively sets aside a fresh war chest for the back half of the three-year plan.

[STAKES] The placement makes one thing unmistakable: Chinese cloud vendors’ AI arms spending has grown to a scale that operating cash flow can no longer sustain, forcing them back to the capital markets for funding. Alibaba shareholders must now judge how quickly the compute and model capabilities bought with this $10 billion convert into revenue rather than just depreciation. For Tencent and ByteDance on the same track, the financing window and pricing benchmark have just been recalibrated.

▪ SIGNALNet profit down 75% while raising $10B for AI — Alibaba is handing out a timetable with no way back.

❯ Hugging Face reportedly explores sale, valuation could top $13 billion

[RUMOR] The model-hosting platform is exploring a sale at a valuation that could land at over $13 billion, Business Insider reported, citing people familiar with the matter. The company has hired banks to gauge buyer interest. Talks are at an early stage; the report named no potential buyers, and the company has not responded. The news broke on August 23 and was picked up by multiple financial media outlets that day.

[VALUATION] Its previous public valuation was $4.5 billion in 2023, from a $235 million round led by Salesforce, with Google and Nvidia also investing. If the sale indeed closes at $13 billion, that is nearly triple in under three years. The platform currently hosts more than 2 million models, 1.5 million datasets, and 1.5 million AI applications, making it the de facto distribution gateway and default download destination for open-source models.

[STAKES] Buying it means buying the distribution layer of the open-source ecosystem. The same week Nvidia spent $7 billion building its own open-source models, the strategic value of this position is being repriced. Potential buyers could include chipmakers and cloud providers alike. Developers have another worry: once a neutral platform has an owner, will model ranking and hosting terms start tilting toward the shareholder side?

▪ SIGNALThe value of open-source models lies not just in the weights, but in who controls the address everyone defaults to for downloads.

❯ Trump Says Communities Resisting Data Centers Are “Making a Mistake” as Bipartisan Backlash Mounts

[STATEMENT] In an interview, Trump said communities resisting data centers are “making a mistake,” arguing the projects bring “lots of jobs and money.” Axios reported the interview on August 23, with the remarks landing as opposition intensifies across party lines.

[RESISTANCE] According to Data Center Watch, Q1 2026 saw the highest concentration of rejected and delayed data center projects on record, with signs the resistance will keep climbing through the year. The disputes now boil down to two questions: whether projects must pass local approval, and who pays the added electricity costs.

[BIPARTISAN PULLBACK] The backlash has reached the governor level, crossing party lines. Pennsylvania Democratic Governor Shapiro signed an executive order requiring new projects to first obtain local approval and specify who will bear electricity costs before they can move forward; Texas Republican Governor Abbott announced a pause on new data center construction pending the results of a statewide audit. Both had previously backed data center projects — but in this 2026 election year, multiple governors seeking reelection, including them, have shifted course at the same time.

[STAKES] At the core of the dispute is who pays for electricity. Data center tax revenue and jobs land at the county level, but higher power prices are spread across all ratepayers on the grid — an equation that is especially glaring in an election year. Hyperscalers’ site-selection models therefore must add one more variable: not just power supply and land prices, but the political cost of local approval. That compute capacity slated to come online around 2027 — a portion of it is still stuck in hearings.

▪ SIGNALCompute expansion has hit the ballot box for the first time — whether a project gets built no longer depends only on grid capacity, but also on that one vote in the county council.

❯ OpenAI President Brockman Takes Over Product and Expansion Teams, Consolidating Power After Executive Exodus

[TAKEOVER] OpenAI President Greg Brockman’s remit has expanded dramatically: he has taken over all duties of former product and business lead Fidji Simo, while also overseeing both the product and expansion lines. According to The Verge, he now directly runs the company’s most profitable and critical projects—from consumer subscriptions to enterprise deployments.

[DEPARTURES] The change comes after a dense wave of departures: within recent weeks, AI Workspace lead Kevin Weil, Sora lead Bill Peebles, and Enterprise Applications CTO Srinivas Narayanan all left in succession. Axios’s August 14 report described the moves as a pre-IPO executive overhaul. Brockman told employees the company is consolidating product strength to race toward an agentic future with maximum focus, aiming to surpass Anthropic in enterprise adoption.

[STAKES] The direct goal of the centralization is to merge scattered product lines into a unified agent platform. Enterprise buyers now face an OpenAI with more concentrated decision-making: previously, roadmaps were held by multiple leads, fragmenting interfaces and commitments; now they sit in one person’s hands, so iteration will be faster—but single-point risk is also higher. For Anthropic, its rival has just placed the enterprise market at the No. 1 spot.

▪ SIGNALA company consolidating product power into its co-founder’s hands before an IPO is betting that execution speed outweighs checks and balances.

❯ Stealth Model Ox Alpha Debuts on OpenRouter, Free 1M-Context Trial for a Week

[LAUNCH] A model named Ox Alpha quietly surfaced on OpenRouter on August 20 — no lab byline, no release notes. The spec sheet is generous: a 1,048,576-token context window, tri-modal input covering text, image, and video, maximum output around 130,000 tokens, and completely free for a one-week preview window.

[BENCH & ORIGIN] The long context stood up to scrutiny — the needle-in-a-haystack test still hit at about 934,000 tokens, and the model completed 113 DeepSWE tasks with a 58.4% solve rate. Its origin, however, is entirely unresolved: the model page claims a daily serving capacity above 100 trillion tokens, and the most widely circulated community guess is a variant of Zhipu’s GLM-5.3, with a secondary theory pointing at Xiaomi’s MiMo team. Researcher teortaxesTex has posted multiple tweets pushing back on the “it was Cursor” claim — every attribution so far is speculation; none has been confirmed.

[STAKES] Stealth launches are becoming a routine product-testing tactic: gather a round of real-workload feedback without brand baggage, then decide whose name goes on the listing. For developers, the near-term value is that one-week free million-context quota; for model vendors, OpenRouter’s blind-test reputation has become a distribution channel that bypasses launch events — a model’s first reviews now form anywhere but the vendor’s own blog.

▪ SIGNALAn anonymous listing, a free week, and the community running your evaluations — this playbook hands the power of the launch event to the users.

❯ Shein to List on Hong Kong Exchange September 1, IPO Aims to Raise Up to $1.8 Billion

[LISTING] Filing documents show Shein will list on the Hong Kong exchange on September 1, offering 280 million shares at an indicative price of HK$47.6 to HK$49.5 per share, with proceeds of up to HK$13.9 billion (about US$1.8 billion). Goldman Sachs, Morgan Stanley, and JPMorgan are joint sponsors. The offering has cleared CSRC filing, and ahead of listing the company also made about US$1.1 billion in allocation available to investors.

[VALUATION] This roadshow took three years and three continents — a U.S. filing first, then a London pivot, finally a Hong Kong landing. The price was a halved valuation: roughly US$40 billion to US$50 billion at the current range, versus a peak that once touched US$100 billion. In between came shifting tariff policies, the loss of duty-free treatment for small parcels, and regulatory scrutiny across multiple markets — each one cutting directly into its cross-border direct-shipping model.

[STAKES] For Hong Kong’s exchange, this is the year’s most significant new listing and a mettle test for US-listed Chinese names returning to Hong Kong. For fast-fashion peers, once Shein reports quarterly, the sector’s true profit margins will finally have a comparable basis. And for the private markets, the markdown from $100 billion to $50 billion is a yardstick for how much of the past three years’ high valuations has been digested.

▪ SIGNALThe listing journey took three years and three markets — the costliest part wasn’t the underwriting fees, but the half of the valuation that evaporated along the way.

❯ LandSpace Zhuque-3 Y2 completes China’s first land recovery of an orbital-class rocket

[RECOVERY] LandSpace’s Zhuque-3 Y2 launch vehicle lifted off from the Dongfeng Commercial Aerospace Innovation Pilot Zone on August 19. About 137 seconds into flight, the first and second stages separated, and the second stage precisely delivered the Honghu-03 satellite into orbit. At 07:41, the first stage successfully soft-landed at the landing site in Minqin, Gansu. This is China’s first land recovery of an orbital-class rocket first stage, and the first controlled land recovery using landing-leg technology, filling a gap in this technical field.

[TECH] The Y2 uses a “liquid-oxygen methane + stainless-steel airframe + landing legs” configuration. During the Y1’s maiden flight, the final landing-burn ignition failed. The Y2 took that lesson on board: it reduced the number of landing-burn engines to simplify the system and lower control difficulty, and added landing-point prediction safety control. The first stage is designed for reuse over 20 times; officially, that can cut the cost per launch by over 60%.

[REUSE] The airframe is currently at a critical stage of engineering reuse; inspections and assessments will follow, and the team is simultaneously optimizing the rapid post-recovery inspection process, with the goal of compressing the reuse cycle to within 30 days. LandSpace plans to fly the recovered airframe again within six months, completing the “launch-recovery-inspection-reuse” closed loop. All domestic satellite-constellation operators that have planned around expendable-rocket pricing will need to redo their math — the cost-per-kilogram-to-orbit baseline is moving.

▪ SIGNALLanding is just the passing line; flying again within 30 days is what truly earns a ticket to the cost curve.

❯ Rumor: Anthropic’s New Models Marshmallow and Melon Surface, Unconfirmed

[RUMOR] Developers claim to have spotted two unreleased model codenames, Marshmallow and Melon, reportedly corresponding to an Opus update and a new Haiku. So far, the report comes from a single account only, Anthropic has not confirmed anything, and no second independent source has followed up.

[CREDIBILITY] There is only so much to glean from the codenames themselves. Over the past year, Anthropic’s naming system has shifted from numeric sequences to word-based codenames like Fable and Mythos, so internal codenames do not necessarily map one-to-one to final product names. Previously, in March, the source code of the Claude Code command-line tool leaked a batch of unreleased features — code trails are indeed a common leak path at this company, but a trail is not a timeline.

[STAKES] The timing is worth examining: Fable 5’s enterprise spend was just surpassed by its own Opus 5, giving Anthropic a clear motive to quickly fill the cheaper and faster tier. If the Haiku line does get an update, the most direct impact will be on pay-as-you-go agent-style applications, whose cost structures are most sensitive to small-model pricing.

▪ SIGNALA codename leak is only worth as much as whether it lands at a moment when a vendor is being forced to accelerate.

❯ DeepSeek Adjusts API Billing — Entire Weekend Now Billed at Off-Peak Rate

[PRICING] As of 00:00 on August 23 Beijing time, DeepSeek has adjusted its API billing rules: Monday through Friday still uses the original peak/off-peak tiered pricing, while Saturday and Sunday no longer distinguish peak from off-peak, with all hours uniformly billed at the off-peak rate.

[PEAK/OFF-PEAK] The peak/off-peak pricing itself only launched on August 17, with peak windows set at 9:00–12:00 and 14:00–18:00 on weekdays, and the off-peak rate set at half the peak rate. In other words, a mechanism that has been running for less than a week has already loosened one of its seams. The official line is that this lets users schedule weekend tasks without worrying about time-of-day costs, while also balancing compute load across the network. Under the previous rules, the off-peak rate was just 50% of the peak rate — the entire weekend now effectively lands on that line.

[STAKES] Net result: weekend batch inference costs are cut in half. Teams running offline evals, batch data synthesis, and overnight agent jobs now have a solid cost-saving window. What’s even more worth pondering is the scale of compute sitting idle on weekends — the willingness to accept half price in exchange for filling that load suggests the gap between peak and off-peak is wider than the pricing sheet implies.

▪ SIGNALWhere the pricing sheet loosens is usually where overcapacity first surfaces.