Back to AI Daily home

❯ Broadcom’s Q3 AI semiconductor revenue jumps 221% to $16.7 billion, but Q4 guidance falls short

[the numbers] Broadcom posted third-quarter revenue of $29.59 billion, up 86% year over year and ahead of the $29.36 billion expected, with AI semiconductor revenue of $16.7 billion, up 221% year over year and 54% sequentially, beating the $15.2 billion average estimate. Adjusted earnings came in at $3.32 a share, also above consensus. A year earlier total revenue was just $15.95 billion.

[soft guidance] The problem is next quarter: Broadcom guided to $34.8 billion in Q4 revenue against the $35.03 billion analysts wanted, and the stock slipped. The miss is small, but at current valuations tolerance for anything short of a beat is close to zero. A second number points the other way — Q4 AI chip revenue is forecast to reach $21.7 billion, implying 236% growth. The AI line is accelerating; what is holding the company back is everything else.

[the custom silicon path] Broadcom builds custom accelerators for hyperscalers, a customer list that includes the same companies designing their own chips, Google and Meta among them. That $16.7 billion is still a distance from Nvidia’s $96.2 billion quarter, but the growth rates are 221% against 106%. Anyone tracking how fast in-house silicon substitutes for merchant GPUs can read this quarter plainly: custom chips are no longer a supplement to Nvidia but a market growing into itself at triple speed. The soft Q4 guide is a reminder of the other half — the non-AI semiconductor cycle has not come back.

▪ SIGNALAI chip revenue tripled while total guidance missed, which says the non-AI business can no longer carry this company’s own valuation.

❯ OpenAI’s Astra reportedly uses recurrent depth — better performance, less legible reasoning

[what the technique is] Per The Information citing one source, OpenAI’s upcoming Astra model uses a reasoning technique called recurrent depth, which steps outside the sequential thinking that characterizes most reasoning models and instead processes the same query several times in an internal loop. The gain is better cost and performance; the cost is that the model leaves fewer legible traces of its reasoning, effectively side-stepping a conventional chain-of-thought record.

[why safety people are alarmed] Chain of thought is currently the most practical tool for monitoring model intent — you can only spot something wrong if you can read what the model is doing. Recurrent depth is also called opaque recurrence: it pushes reasoning inside the model, leaving only conclusions on the outside. Last week’s independent investigation of the Hugging Face incident reconstructed the agents’ motives precisely through chain-of-thought analysis, and that method breaks against a recurrent-depth model. Astra’s use of the technique is reportedly constrained, but the direction is now visible.

[state the sourcing] This one needs its limits marked: the source is a single anonymous person, OpenAI has not confirmed it, and no named employee, system card, technical paper, code release or patent ties Astra to recurrent depth. A separately verifiable observation from the same period is that the model slug gpt-6-astra has appeared on the APIs. Compliance teams assessing model interpretability should note the direction itself: if the performance and cost gains are large enough, monitorability becomes the thing that gets optimized away.

▪ SIGNALChain of thought is monitorable because the model has to say it out loud; recurrent depth removes the requirement to say it.

❯ Fei-Fei Li’s World Labs releases Atlas, rebuilding 3D worlds from one photo and simulating robot views

[one model, four jobs] World Labs, founded by Fei-Fei Li, released the world model Atlas on September 1, calling it the first multimodal world model of its kind. It is a multimodal autoregressive diffusion transformer pretrained from scratch, operating natively on text, images, video and 3D, and treating world generation, 3D reconstruction and simulation as facets of one problem. Output runs up to 1440p and one minute per generation.

[camera as native input] Per its technical blog, prior approaches largely treated camera pose as a post-hoc constraint; Atlas’s decisive choice is to take camera geometry as a native input rather than coax it out of a text prompt. That yields pixel-perfect camera control: from a single image, or a few ground-level photos, it reconstructs large scenes and outputs 3D natively, exportable as point clouds and Gaussian splats. On 3D reconstruction the company reports an error of 25.3‰, ahead of specialist models including Pi3X and Depth Anything 3 — a generalist beating purpose-built tools at their own task is the hardest claim here.

[the robotics line] The Real-to-Sim demonstration says more than the benchmarks: World Labs reconstructed two large environments from cell phone video, 24 frames each, then had simulated robots move through them while Atlas rendered the RGB and depth data an onboard camera would see. That collapses the cost of a robot training environment from building a physical set or hand-modeling down to a phone video. Interaction with rigid, articulated and deformable objects is inside the reconstruction as well. Nvidia’s Isaac GR00T simulation stack is the industry default today, and that is the ground Atlas is cutting into.

[what was withheld] Atlas is in early access with select partners, with a request form on the site, and will power future versions of Marble and other products. The launch came with no paper, no pricing, no latency or compute-footprint figures, and no partner names. Teams building robot data pipelines need exactly those numbers before scoping it in: the cost and latency of rendering a frame decide whether this is a stage in a pipeline or a demo.

▪ SIGNALA generalist world model beat the specialists at 3D reconstruction; the shopping list for robot training data just gained an option.

❯ CXMT is producing HBM3E in small volumes, with Alibaba’s T-Head and Cambricon testing it

[the line is running] Per The Information, China’s largest memory maker CXMT has begun small-volume production of HBM3E — the high-bandwidth memory used across current AI accelerators, including Nvidia’s H200 and Blackwell. Alibaba’s T-Head unit and Cambricon are testing the chips with their own processors, and if qualification goes smoothly they could ship in commercial products as early as next year, with CXMT planning to expand output in 2027.

[what the controls targeted] The U.S. expanded export controls on HBM in December 2024, aimed squarely at what matters for AI training and inference. Chinese AI chip designers had depended on imported memory subject to restriction or delay — HBM has been the hardest chokepoint for domestic AI accelerators, since the logic can be designed but a part without its memory never becomes a product. It is also CXMT’s third piece of news in days, after first-half revenue of 150.31 billion yuan and net profit of 77.605 billion yuan, followed by its suit against the Pentagon over a “Chinese military company” designation.

[state the gap plainly] A breakthrough is still a generation behind: Samsung, SK Hynix and Micron are already mass-producing HBM4 while CXMT is at small-volume HBM3E, with industry estimates putting it three to five years behind and contending with low yields. “Small volume” is described in reporting as risk production, some distance from supplying at scale. Reading this as export controls failing overstates it; the chokepoint moved from none at all to some, and expensive.

[whose math changes] The first to recalculate are Chinese AI chip designers’ product schedules: HBM supply used to be exogenous, and now there is a domestic path — behind, but controllable — which makes 2027 a concrete date. On the other side, Micron’s, Samsung’s and SK Hynix’s share assumptions in China come under pressure; even if CXMT only absorbs lower-end demand, those orders leave first. Watch the yield curve, which decides whether this path is a supplement or a substitute.

▪ SIGNALDomestic HBM going from zero to something changes not the performance ceiling but whether China’s AI chips can be scheduled on their own clock.

❯ Google ships Gemini 3.8 Flash and Flash Cyber, the latter only to vetted Fairwind partners

[two models] Google released Gemini 3.8 Flash, its third Flash model in six weeks, holding pricing at $0.75 per million input tokens, and says it beats Claude Opus 5 and GPT-5.6 Sol on some benchmarks. Alongside it came Gemini 3.8 Flash Cyber for finding and patching vulnerabilities, available only to vetted organizations through the new Fairwind Program.

[scores and gatekeeping] Per Google, on CyberGym — a benchmark testing AI agents against real-world software vulnerabilities — 3.8 Flash Cyber scored 86.2% against 77.5% for the previous 3.5 Flash Cyber, and Google’s Chrome security team generated 2.6 times as many correct patches with it as with the larger commercial models tested. Fairwind currently spans more than 650 partners globally, including CrowdStrike, Datadog, Palo Alto Networks and Snowflake, with priority to critical infrastructure operators and maintainers of widely used software. Participants must restrict access to security, incident response and penetration testing staff, enable multi-factor authentication, and confine use to authorized threat simulation, reverse engineering and malware analysis.

[tiered release becomes the norm] The notable part is not the benchmark but how it shipped: same company, same day, one model with a price and one with an admissions process. It is a variant of what Z.ai did last week in swapping GLM-5.3’s license to require a security review from $10 billion-revenue hosts. Enterprises buying security-capable models now face qualification rather than a price list, and the more capable the model, the harder it is to simply purchase. Ethan Mollick’s read on 3.8 Flash stayed measured: a very good Flash model, but not equivalent to a frontier model.

▪ SIGNALTwo models in one day, one priced and one vetted — offensive-capable AI is turning from a product into a license.

❯ Meta ships Muse Spark 1.3 at unchanged pricing, and Zuckerberg says open weights are coming soon

[launch and price] Per Axios, Meta rolled out Muse Spark 1.3 in Muse Code and the Meta Model API, saying it significantly improves coding and agentic performance at the same price as Spark 1.2. Zuckerberg then said Meta’s Watermelon model and Muse Spark open weights are “coming soon,” describing the release as frontier performance almost too cheap to meter.

[benchmarks and efficiency] Per the company and third-party evaluations, there are two sets of numbers. On capability: 88.8 on Terminal-Bench 2.1, matching GPT-5.6, and a third-party reading of 75.4% on DeepSWE 1.1, placing it ahead of GPT-5.6 Sol and Fable 5. The efficiency side is more interesting — against 1.2 it makes 20% fewer tool calls and uses 25% fewer tokens. Per Artificial Analysis, the limited-preview max variant scores 62 on the Intelligence Index, behind only Claude, and 68 on the Coding Agent Index, second overall. Note that the max variant is partner-only with no announced pricing; what is publicly available is a different tier.

[holding price is the point] Per those same figures, a generational gain at flat pricing, combined with fewer tool calls and fewer tokens, means unit cost per task is falling. Engineering teams that select on cost per task feel this first: the same job now takes roughly a quarter fewer calls and a quarter less generated text. The open-weights promise points elsewhere — after GLM-5.3 moved to a restrictive license, a genuine Meta open-weights release would pry the open camp back open.

▪ SIGNALMatching frontier performance while cutting tool calls by a fifth, what Meta is selling is cost per task, not a leaderboard position.

❯ Alibaba updates Qwen3.8-Max to 1691 on the front-end coding leaderboard, taking first place

[first place] Alibaba released Qwen3.8-Max-0902, post-trained specifically for coding and professional office work. On the front-end coding leaderboard CodeArena, its score rose from 1669 to 1691, a 22-point gain that puts it first overall. The company says it set new marks in multi-step reasoning, tool use and end-to-end app generation.

[price is the other half] Per the company’s own comparison, the gap on cost is wider than the gap on score: the model averages $5 per million tokens against $20 and $12 for the second- and third-ranked models. It took first place at a quarter of the runner-up’s price. The model is live on the Qwen AI platform with API access, and integrated into Qwen Office, Qoder and the Qwen app — three channels covering enterprises, developers and consumers.

[where Chinese models sit] Within a single week GLM-5.3, Kimi K3, Tencent’s Hy4 and Qwen3.8-Max have all shipped, crowding Chinese vendors into the same capability band on coding. Teams selecting models face a reshaped question: when first place costs a quarter of second place, the criterion shifts from “which is strongest” to “which is strong enough inside the budget.” One caveat on scope — CodeArena tests front-end work only, a narrower window than a composite benchmark.

▪ SIGNALTaking first place at a quarter of the runner-up’s price, Chinese models are fighting the coding race on cost, not on score.

❯ Cognition is raising $1 billion at about $47 billion, nearly doubling in three months

[the numbers] Per Bloomberg, AI coding company Cognition is set to close roughly $1 billion at a valuation near $47 billion. The reference point is May’s $26 billionup about 81% in three months — against $10.2 billion as recently as September 2025. Sources say investor interest reached nearly $10 billion, and the final raise could exceed $1 billion.

[the math behind it] Revenue is genuinely running: Devin’s annualized revenue is now above $900 million, up from $492 million in late May. The other side is the cash burn The Information disclosed last week — the company may burn $800 million this year, largely on Nvidia servers. Roughly $900 million of revenue against roughly $800 million of burn means this round prices growth rate, not profit. SpaceX’s $60 billion acquisition of Cursor, closed in August, is credited with lifting appetite across the category.

[the category gets repriced] At $47 billion the multiple is about 52x annualized revenue, an extreme by software standards. Startups in the same category gain a citable anchor and get placed on the same curve: investors will expect the next company to double in three months too. The real test is renewals and gross margin — revenue is produced by agents and cost is burned by agents, and when those two curves diverge matters more than how fast the valuation climbs.

▪ SIGNAL$900 million of revenue against $800 million of burn — a $47 billion valuation buys the time before those two curves separate.

❯ Moonshot AI files confidentially with the HKEX while raising at a $50 billion pre-money valuation

[two things at once] Moonshot AI filed its A1 confidentially with the Hong Kong Stock Exchange this week, formally starting its listing process, while simultaneously raising a round at a $50 billion pre-money valuation that is likely its last before an IPO. The company declined to comment on the reports, saying it has nothing to disclose.

[the valuation curve] The steepness deserves its own line: $4.3 billion at the end of 2025, $31.5 billion by the first half of 2026, more than a sevenfold rise; going from $18 billion to $50 billion took about three months. Revenue support comes from annual recurring revenue that reached $300 million in June, with strong performance after the Kimi K3 release. The company denied filing rumors in August; a confidential submission is the formal start of a process and does not contradict that denial.

[the Hong Kong route] Confidential filing is a permitted HKEX process, and its benefit is keeping financial detail private until the hearing. Investors watching exit paths for Chinese AI companies should note the timing: Anthropic’s prospectus goes public after Labor Day targeting a late-September or early-October listing, while DeepSeek prepares for Shanghai’s STAR Market. Three listing windows sit in one quarter, and whoever prices first prices the other two. Watch how losses are presented in the prospectus, which decides the revenue multiple the public market will pay.

▪ SIGNALA company whose valuation rose sevenfold in six months is queuing to list — pricing power is being handed from private markets to public ones.

❯ Microsoft will disclose Azure quarterly revenue for the first time: $29.42 billion, up 42%

[the reporting changes] Microsoft will disclose Azure’s quarterly revenue for the first time under a new reporting structure. On the new basis, Azure revenue for the June quarter was $29.42 billion, up 42% year over year. Previously the company published only a growth rate, never the absolute figure.

[what the new basis excludes] The restated Azure line no longer includes GitHub cloud services, cloud services for developers, Security Copilot, or cloud products for healthcare and life sciences. Satya Nadella’s framing is that Azure will be more focused on the consumption-based platform and infrastructure business. Even after those exclusions, the figure is close to 33% of Microsoft’s total revenue, which is why it has to be visible on its own.

[why give the number now] Publishing an absolute figure puts Azure on the same ruler as AWS and Google Cloud. Analysts comparing cloud providers finally get a quarter-by-quarter benchmark, and companies generally volunteer that only when they expect to win. Transparency cuts both ways — once growth slows, there is no longer anywhere to hide. Watch whether next quarter’s absolute gap to AWS narrows or widens.

▪ SIGNALVolunteering the absolute number is both confidence and constraint; Azure now gets measured on the same ruler every single quarter.

❯ Snowflake’s Q2 revenue rises 35% to $1.55 billion, sending shares up more than 20% after hours

[a clear beat] Snowflake reported fiscal second-quarter revenue of $1.55 billion for the period ended July 31, up 35% and ahead of the $1.48 billion expected, with product revenue of $1.49 billion up 37% — an acceleration. Net loss narrowed to $191.7 million and adjusted earnings came in at $0.62 a share against $0.45 expected. Shares rose more than 21% after hours.

[guidance beats the quarter] The company raised full-year product revenue guidance from May’s $5.84 billion to $6.07 billion, and guided third-quarter product revenue to $1.59 billion against a $1.50 billion consensus. Management credited part of the growth to the AI coding agent CoCo, whose account count reached 9,100, up more than 2,000 in the quarter. The roughly $230 million raise is the company’s largest full-year guidance increase in several quarters.

[the data layer benefits] Read alongside Broadcom and Microsoft on the same day, this points at three segments of one chain: silicon, cloud, data platform. Anyone judging whether AI spending is spilling downstream can treat Snowflake as the trailing indicator — if agents are genuinely running inside enterprises, data platform consumption grows before software seat counts do. Those 2,000 net new CoCo accounts are the first sample of that.

▪ SIGNALRaising the full-year guide is worth more than beating one quarter; it says the data layer is seeing booked usage, not pilots.

❯ OpenAI connects ChatGPT to Epic’s EHR and adds a public healthcare data plug-in

[two capabilities] On September 1, OpenAI added an Epic EHR integration and a public healthcare data plug-in to ChatGPT for Healthcare, with UCSF Health as a pilot partner. The Epic integration lets authorized clinicians pull patient information — appointment notes, lab results, medications and specialist documentation — inside ChatGPT or directly within Epic workflows.

[where the line is drawn] Per OpenAI’s announcement, ChatGPT for Healthcare previously had no direct connection to a records system. The critical constraint is that it is read-only and does not write back to the patient record. Two use cases are supported: reviewing authorized patient data inside ChatGPT, and embedding ChatGPT into an Epic layout for in-chart assistance. The public data plug-in connects nine official public healthcare sources, including ClinicalTrials.gov, CMS Coverage, RxNorm, DailyMed and PubMed, and is available more broadly — covering both ChatGPT for Healthcare and the free ChatGPT for Clinicians tier for physicians, nurse practitioners and pharmacists.

[the fight is over the entry point] Epic dominates the U.S. electronic health record market, so connecting to it means reaching the clinical workflow itself. Read-only is the necessary answer to regulation and liability — the moment anything writes back, model output enters a legally binding record. What healthcare IT buyers have to evaluate is therefore not model capability but the audit trail: who pulled which record, what the model advised, whether the clinician acted on it, all of it logged.

▪ SIGNALRead-only is not a capability gap; it is the only posture in which AI can currently enter the clinic.

❯ Tencent’s WorkBuddy open platform launches with more than 100 ecosystem partners

[what is being opened] Tencent’s WorkBuddy open platform launched on September 2 with more than 100 ecosystem partners, opening its underlying agent capabilities to smart hardware makers, industry applications and developers. Nine co-branded hardware devices debuted at the event with partners including Plaud, Rokid, Insta360, iFlytek, Anker, Hollyland and JD’s own-brand line, while more than 30 industry applications connected across upwards of 20 fields including finance, law, healthcare, education and nonprofits.

[a blunt positioning] Per the launch event, Tencent Cloud vice president and CodeBuddy/WorkBuddy lead Liu Yi put it directly: the goal is to make WorkBuddy an operating system for the agent era — not a stronger tool, but a platform that can carry every tool. The background is that WorkBuddy, Tencent Cloud’s AI-native desktop agent workspace built on the same architecture as CodeBuddy, launched in March 2026 around executable office agents, and is only now opening its underlying layer.

[the hardware move] Per the company, WorkBuddy had until now been desktop-only. The nine co-branded devices are the least common part of this launch: most vendors building agent platforms connect only software, while Tencent wired in voice recorders, AR glasses and imaging devices — giving the agent a sensing layer. Companies building AI hardware gain a path that does not require building their own model and workflow stack. “Operating system” remains an aspiration rather than a fact; whether it holds depends on third-party retention, not on how many partners signed up on day one.

▪ SIGNALWiring recorders and glasses into an agent platform, Tencent is going after the sensing entry point, not another office assistant.

❯ Uber and Wayve launch London’s first commercial robotaxi service, still with safety drivers

[it is running] Uber and UK company Wayve began deploying a fleet across London on Thursday, London’s first commercial robotaxi service open to the public and Wayve’s first commercial public service anywhere. In this initial phase a human safety driver remains behind the wheel. Fleet size is “in the order of dozens to start,” per Uber head of product Wendy Lee.

[how you get one] Per both companies, they secured licenses for supervised robotaxis in London in August. London riders can set a preference in the app, and when booking through UberX, Uber Comfort or Uber Electric they may be matched with a Wayve autonomous vehicle, though it is not guaranteed. The fleet starts on Ford Mustang Mach-Es, with Nissan Leafs to follow. The move puts Wayve ahead of Waymo and Baidu, both of which have announced London plans, as the first to open a service to the public there. Neither company would give a date for fully driverless operation.

[first is not ahead] A commercial service with safety drivers is essentially scaled road testing inside real order flow, a different stage from Waymo’s fully driverless U.S. operations. But order has its own value: London’s regulatory and insurance frameworks will form around the first service actually running, and later entrants adapt to rules someone else shaped. For Uber, it is one more proof that its platform model can absorb anyone’s autonomy stack.

▪ SIGNALWhat gets claimed first is not the technology but the rules; the first company on the road is writing the regulator’s template.

❯ Apple filing puts Ternus’s fiscal 2027 compensation at about $58 million

[two packages] An Apple filing with the SEC puts new CEO John Ternus’s fiscal 2027 target compensation at about $58 million, comprising a $3 million annual base salary and a stock grant with a target value of $55 million; he also receives a prorated restricted stock unit award worth about $2.5 million for fiscal 2026. The filing landed on Ternus’s first day in the role.

[Cook's side] Per the same filing, Cook’s salary as executive chair falls from $3 million to $2 million, with a 2027 stock package targeted at $45 million, for about $47 million in total. The structural difference says more than the totals: three-quarters of Ternus’s equity award is tied to Apple’s performance relative to the S&P 500, with the remaining 25% vesting semiannually on time; only half of Cook’s equity is performance-linked.

[the signal is in the structure] The share of a package tied to performance is the board pricing how much pressure a person is expected to carry. Ternus’s 75% against Cook’s 50% states the nature of this handover plainly: the incoming CEO has to earn his pay by beating the index, and the tightest item on his job list is catching up in AI. Cook is not leaving, staying closely involved and keeping his relationships with Donald Trump and in China.

▪ SIGNALThree-quarters of the equity riding on beating the S&P — Apple’s board wrote its demand of the new CEO into the vesting terms.