Back to AI Daily home

❯ Anthropic says Claude now leads 26% of AI R&D work, up from 1% in six months

[Task share] Anthropic told Bloomberg that Claude can now lead end to end about 26% of model R&D tasks from a high-level instruction, up from 1% in March, while people largely set goals, review outputs and handle exceptions.

[Internal scale] The company also said Claude performs substantial work in more than 90% of R&D activity, while roughly 30,000 agents operate at any moment on its most-used internal platform. It analyzed more than 1 billion agent decisions in August. These are Anthropic’s own workflow categories, not a claim that 26% of jobs or hours have disappeared.

[Management controls] As models move from code completion to decomposing tasks, running experiments and submitting results, research teams need stronger evaluation, permission and rollback systems. The useful comparison for rival labs is how many outputs enter production R&D under close human supervision, and how often they require rework or reversal.

▪ SIGNALClaude’s internal R&D value is shifting from how much code it writes to how many tasks it can close.

[Vertical product] OpenAI launched Astra for Law, combining GPT-6 Astra with a legal-search index spanning U.S. case law, statutes, court rules and administrative decisions. Initial access is limited to select law firms for contract review, deal investigations and legal drafting.

[Tooling] OpenAI said the product includes 26 partner-built plugins and 47 community plugins, with data privacy treated as a core feature. This is more than a new prompt for a general-purpose model: retrieval sources, legal-analysis instructions and a firm’s existing tools sit in one workflow.

[Procurement test] Firms will look beyond accuracy to ask whether citations are traceable, client data is isolated and lawyers can review faulty conclusions. A limited initial rollout also gives OpenAI room to tune permissions and audit controls in a high-liability setting.

▪ SIGNALLegal AI competition is moving from raw answers to source coverage, traceable citations and firm workflows.

❯ OpenAI staff reportedly expect progress on the Hodge Conjecture, but no proof has been published or peer-reviewed

[Internal expectation] The Information, citing people familiar with the matter, reported that OpenAI employees expect the Hodge Conjecture could be solved relatively soon. It is one of the Clay Mathematics Institute’s Millennium Prize Problems, but OpenAI has not published a complete proof or announced a result.

[Earlier work] The September 17 report linked the effort to OpenAI’s earlier claimed work on the Navier-Stokes problem. Both sit among seven Millennium problems, each carrying a $1 million prize under the institute’s rules. OpenAI views stronger mathematics as a path toward automated AI research, but any proof still needs line-by-line inspection and independent peer review.

[Boundary] Staff expecting a solution is not the same as a conjecture being proved. Researchers and readers must wait for a reproducible proof, reviews by several specialists and the institute’s process. Until then, internal progress cannot be treated as a mathematical conclusion or a prize-winning result.

▪ SIGNALThe news value lies in AI potentially contributing original mathematics; authority rests with public proof and peer review.

❯ SEC offers a five-year exemption for tokenized securities, easing some exchange rules for platforms

[Regulatory window] Reuters reported that the U.S. Securities and Exchange Commission introduced a five-year Innovation Exemption allowing qualifying platforms for tokenized U.S. stocks to avoid some securities-exchange rules, creating a defined window for on-chain market experiments.

[Scope] The September 17 measure does not remove tokenized stocks from securities law. It eases selected registration and market-structure requirements for five years. Investor protection, anti-fraud duties and asset rights still depend on the final documents, platform eligibility and specific compliance conditions.

[Market effect] Brokers, crypto platforms and exchanges will compare compliance cost and settlement efficiency. Five years is enough to test products, but it does not guarantee the same structure after expiry. Platforms that cannot explain the shareholder rights behind an on-chain token will struggle to build lasting liquidity.

▪ SIGNALThe exemption lowers the cost of experimentation; rights mapping and liquidity will decide whether tokenized stocks become a market.

❯ Crusoe reportedly raises $3.9 billion at a $30.9 billion post-money valuation, betting on factory-built data centers

[Financing] AI data-center developer Crusoe raised $3.9 billion in a round co-led by Atreides, Valor and Mubadala, the Wall Street Journal reported. The deal values the company at about $30.9 billion post-money and supports further expansion of computing infrastructure.

[Construction model] Crusoe previously helped build a major OpenAI data center. In September 2026, it is also betting on factory-built data-center modules shipped to sites for assembly. Standardizing electrical, cooling and server-room units is intended to shorten schedules and reduce on-site engineering uncertainty.

[Capital test] The valuation raises the burden on delivery speed and capital turnover. If prefabrication converts land, power and equipment into billable capacity faster, it can narrow the lag in a capital-intensive business. Slower demand or grid connections would magnify idle-capacity risk.

▪ SIGNALCrusoe’s valuation ultimately depends on how much prefabrication can compress data-center construction time.

❯ Manus reportedly nears a $500 million round at a $4 billion valuation after the Meta deal collapsed

[New round] General-purpose agent company Manus is close to raising $500 million at a roughly $4 billion valuation, Bloomberg reported, citing people familiar with the matter. The transaction has not been announced, and its size, investors and final terms may change.

[Deal reversal] Bloomberg said this would be the company’s first major financing after Beijing unwound Meta’s short-lived acquisition of Manus. That deal was reported at about $2 billion. If the 2026 round closes as planned, Manus’s valuation benchmark would double after the takeover failed.

[Independent path] New capital can turn a popular agent into an independent platform only if paid retention, task completion and inference costs hold up. Regulation closed one acquisition exit, leaving Manus to show that the business can support a $4 billion valuation without a large platform’s distribution.

▪ SIGNALThe failed Meta deal did not end investor demand, but Manus must now support the higher valuation with standalone operating data.

❯ Kimi launches a finance suite with nine skills, more than 10 data plugins and managed enterprise agents

[Three surfaces] Moonshot AI launched a Kimi finance solution. Consumers can try nine finance skills and more than 10 data-source plugins at kimi.com, while professionals can invoke the same capabilities from the KimiWork desktop application and KimiCode.

[Enterprise offer] The company also introduced Kimi managed agents for enterprises, packaging data connections, skill configuration and runtime operations. Kimi previously reached users mainly through a general assistant and developer tools; the release now spans the web, desktop workspace, coding environment and enterprise deployment.

[Adoption hurdle] Financial institutions will examine data licenses, citations and audit logs, not plugin counts. When outputs enter research, risk management or client materials, data freshness, permission isolation and human sign-off become harder constraints than generation speed.

▪ SIGNALWith Kimi’s finance tools spread across several surfaces, competition shifts to licensed data and auditable delivery.

❯ Xiaomi’s MiMo-V2.6 reaches the middle of reinforcement learning, with the company showing training costs above $1.25 million

[Training status] Xiaomi said MiMo model lead Luo Fuli disclosed that MiMo-V2.6 is midway through reinforcement learning. A company training livestream showed cumulative spending above $1.25 million and described the model as approaching release.

[Scaling work] Luo said the team spent nearly six months exploring reinforcement-learning limits after open-sourcing MiMo-V2.5 in April. The current phase expands compute, environments and tools, and grader compute, with methods and engineering details due to be released over the coming weeks.

[Proof point] Public spending data makes the experiment easier to observe, but cost is not evidence of capability. Developers need post-release results on task success, tool reliability and unit inference cost before judging what the three scaling dimensions delivered.

▪ SIGNALThe livestream makes training more transparent; MiMo-V2.6 will still be judged by its released task performance.

GLM-5.3 helps optimize its own inference stack, with the company claiming a threefold throughput gain at 100,000-accelerator scale

[Engineering loop] Zhipu said GLM-5.3 helped analyze and modify the production infrastructure serving the model itself. With engineers setting goals and boundaries, it assisted with hypotheses, code changes and experiments that lifted GLM-5.3-Flash throughput threefold.

[Optimization details] The company said the system runs across more than 100,000 accelerators. Examples include tracing a transfer bottleneck to Python’s global interpreter lock, narrowing a performance gap from more than 20% to below 1%, and delivering a 1.71-times kernel speedup.

[Claim boundary] This is the team’s account of a model participating in infrastructure optimization, not autonomous self-upgrading. The immediate value for compute operators is better debugging speed and cluster utilization; replication across workloads requires code, test conditions and long-run operating data.

▪ SIGNALToday’s self-improvement looks like a supervised engineering loop whose output is measured first in throughput and debugging time.

❯ Chinese models reportedly reached 46% of U.S. OpenRouter weekly usage at one point, as open distribution expands reach

[Usage share] Public accounts of Rhodium Group data say Chinese-origin models exceeded 30% of weekly token volume among U.S. OpenRouter users and peaked at 46%. The measure covers OpenRouter only, not all U.S. model use or enterprise traffic through private interfaces.

[Distribution path] Open weights and low prices let developers self-host or call DeepSeek, Kimi and GLM through aggregators such as OpenRouter. That reach can dilute vendor revenue: usage may sit on third-party channels, fail to flow back as direct revenue, and resist comparison with official API data.

[Commercial split] U.S. model companies rely more heavily on closed APIs and subscriptions, while Chinese vendors use open distribution to reach global developers. Investors need to separate token share, paid revenue and model origin, and test whether the share persists across platforms before inferring market leadership.

▪ SIGNALChinese models’ overseas reach is visible in token volume, but usage share and revenue share remain different measures.

❯ Snap pitches its $2,195 Specs to enterprises, working with Salesforce, Amazon and Nvidia on applications

[Enterprise price] Snap is pitching its $2,195 Specs smart glasses to businesses and has struck partnerships with Salesforce, Amazon and Nvidia for visual overlays that help workers access information and perform tasks, Wired reported.

[Two-track product] Specs retains consumer features, but the price and partners show Snap looking first for paid demand in enterprise settings such as training, field work and visualization. Businesses can tolerate expensive hardware if deployments measurably reduce operating time or errors.

[Adoption conditions] Buyers will focus on comfort, battery life and integration cost, alongside privacy and workplace rules around cameras. If visual overlays cannot reliably connect with existing data, a $2,195 device risks remaining a demonstration project.

▪ SIGNALBy targeting enterprises first, Specs is looking for measurable workplace returns to justify premium hardware.

❯ Altman and Huang reportedly join the White House state-dinner guest list, putting AI leaders in a high-level U.S.-China setting

[Guest list] Bloomberg reported on September 16 that OpenAI CEO Sam Altman and Nvidia CEO Jensen Huang will attend next week’s White House state dinner for Chinese President Xi Jinping. A person familiar with the plans said Apple CEO Tim Cook will also attend the White House dinner.

[Industry context] The three companies occupy pivotal positions in models, AI chips and consumer-electronics supply chains. OpenAI depends on large-scale compute, Nvidia faces chip-export rules, and Apple is deeply tied to Chinese manufacturing and consumers. Attendance indicates an invitation, not a policy commitment.

[Observation boundary] Companies can watch for public developments in export rules, market access and supply-chain arrangements, along with any disclosures of business effects. Without a formal statement from the White House, regulators or the companies, the guest list cannot be extended into an agreement or policy shift.

▪ SIGNALThe presence of AI executives at the dinner places chips, models and supply chains near the center of high-level economic dialogue.