Back to AI Daily home

❯ Xiaomi releases open-weight MiMo-V2.6 Pro and Flash; an independent index ranks Pro first among open models

[Release] Xiaomi released the MiMo-V2.6 Pro and Flash weights alongside a technical report and reinforcement-learning research resources. The company says the models can take text, image and audio inputs and are available on its platform at unchanged API prices.

[Benchmark scope] Artificial Analysis assigned Pro an Intelligence Index score of 46, the highest among open-weight models it tracks, and estimated a cost of about $0.13 per index task. That supports a lead on this index, not a claim of superiority on every task. Xiaomi’s comparison with Claude Opus 5 and GPT-5.6 Sol depends on particular agent benchmarks and setups.

[Deployment choice] With both Pro and Flash available, developers can compare capability against the resources needed to run them. Application companies gain another model they can download, adapt and host instead of judging only closed-API prices. Reliability, serving expense and performance on actual workloads will determine adoption.

▪ SIGNALOpen weights are now competing on task-level cost at frontier-like performance, not merely on whether they can approach the frontier.

❯ OpenAI says an internal model solved more than 100 open math problems and turns to independent advisers on disclosure

[Company claim] OpenAI said on September 21 that an internal model whose training began on August 28 had resolved more than 100 long-standing open problems across mathematics. The pace surprised mathematicians inside the company, it said, prompting discussion about how to release the results.

[Advisory remit] An independent group of mathematicians will advise on reviewing results, judging significance, coordinating publication and research standards. Members are unpaid by OpenAI and free to criticize it publicly. The company explicitly says the group will not set the pace of its internal mathematics research.

[Verification gap] The announcement does not publish a problem-by-problem list and proofs for all the claimed solutions. Independent mathematical review still requires precise statements, comparison with prior work and complete arguments. The announcement establishes the company’s claim and governance arrangement, not external confirmation of every theorem.

▪ SIGNALMaking the problems and proofs inspectable will determine whether a striking internal tally becomes research the wider field can use.

❯ SpaceXAI launches Grok 4.7 at $2 per million input tokens and $6 per million output tokens

[Availability] Grok 4.7 is available through the API, Cursor and Grok Build at a base price of $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6. SpaceXAI says it checks its work more carefully and handles longer tasks better.

[Task results] The company trained a larger base model with a longer reinforcement-learning run focused on tasks taking hours. Artificial Analysis measured a rise in its coding-agent index from 47 to 56 using Grok Build. That tests the model-plus-tool combination, not the model in every coding environment.

[Buying decision] For teams running long-lived agents, cost per completed task matters more than the token sticker price alone. Better results at the same base rate may change model-routing choices. Requests above 200,000 input tokens carry a different pricing tier.

▪ SIGNALA same-price upgrade pressures rival long-task models, though real-project completion and rework rates remain the deciding measures.

❯ Bloomberg says Harvey and other startups are shifting toward open-weight or proprietary models to curb AI costs

[Supplier shift] Bloomberg reports that legal-tech company Harvey and startups including Abridge, Ramp and Rogo are exploring open-weight models or training their own systems to reduce reliance on a small group of frontier labs. Their approaches differ; the report does not say they have all abandoned closed models.

[Margin pressure] Model calls are both the engine of vertical AI applications and a direct expense that grows with usage. Routing narrower, repeatable tasks to cheaper systems while reserving frontier models for harder work could improve gross margins and response times. The trade-off is more in-house evaluation, hosting and quality control.

[Negotiating position] These businesses sell workflows and domain knowledge, not a place on a general model leaderboard. Frontier suppliers’ pricing power weakens when customers can show that a substantial share of production tasks runs well on alternatives.

▪ SIGNALWhen the workflow rather than an exclusive model is the moat, model spending becomes a more negotiable procurement line.

❯ Nscale files $103.4 billion in contract value, but only about $2.6 billion is active

[Contract mix] AI infrastructure provider Nscale reported about $103.4 billion of active and contracted total contract value at the end of August, with only about $2.6 billion active. Bloomberg says Microsoft and Anthropic together account for roughly 85% of the total, creating substantial customer concentration.

[Accounting boundary] Total contract value spans expected deliveries over several years; the filing gives a weighted average contract life of about 5.7 years. It is neither recognized revenue nor fully powered capacity. Active contracts represent only about 2.5% of the headline amount, while construction, power, chips and financing remain to be delivered.

[IPO question] Investors must price delivery speed and financing cost, not just the contract headline. A change in either major customer’s schedule would affect how quickly contracted value turns into revenue. The large order book proves demand while exposing execution and concentration risk.

▪ SIGNALThe $103.4 billion figure describes future obligations; the $2.6 billion active slice shows the scale operating today.

❯ New York Times says SoftBank’s SB Energy has delayed its IPO amid questions over a valuation above $50 billion

[Delayed offering] The New York Times, citing people familiar with the matter, says SoftBank-backed energy and data-center developer SB Energy has postponed an IPO originally planned for this month. Investors questioned a sought-after valuation above $50 billion. No new listing date has been announced.

[Valuation gap] Signed data-center projects, completed construction and operating revenue arrive at different times. SB Energy ties power development to AI computing demand, giving investors a large future market to consider. It also makes financing and power-on schedules central to the valuation.

[Next test] A delay does not establish that the projects have failed or predict the eventual offer price. Verifiable project cash flow would narrow the gap between the company’s target and investors’ willingness to pay; construction milestones and capital expenditure will matter in any renewed roadshow.

▪ SIGNALAI infrastructure can sell an order-book story quickly; an IPO price needs a credible path to powered sites and cash flow.

❯ Reuters says Z.ai disabled ZCode features after code-upload complaints and released the coding harness source

[Response] Reuters reports that users accused ZCode of uploading local code repositories to overseas cloud servers without clear consent. Z.ai apologized and disabled some features. Whether any particular company’s secrets were affected requires case-specific evidence; the complaints alone do not establish that.

[Open repository] The company has put source for its desktop, web and terminal coding-agent harness in a public repository. Developers can inspect parts of the client and runtime. Open sourcing does not, on its own, verify historical data flows, server-side processing or versions already installed.

[Enterprise controls] Permission to access a project is too broad a description for a coding agent. Outbound code paths, defaults and audit logs need to be checkable in practice. An unexpected full-repository upload can force security teams to revisit tool approvals.

▪ SIGNALThe more autonomously a coding agent reads, the more clearly enterprises need to see what it read and where it sent it.

❯ The Information says DeepSeek aims to train on Huawei chips, with deliveries expected late this year or early next

[Reported plan] The Information, citing people familiar with the matter, says DeepSeek chief executive Liang Wenfeng told investors that training on Huawei processors is one of the company’s largest bets. A new batch of chips could arrive in the fourth quarter of 2026 or first quarter of 2027; that is a prospective delivery window, not a completed deployment.

[Training hurdle] Under export restrictions, using domestic chips for the largest model runs depends on cluster size, networking, software and stable uptime. Evidence of alternatives in inference does not establish comparable maturity for large-scale training. Usable training-cluster capacity is the central constraint.

[What follows] Successful delivery and training would reduce DeepSeek’s dependence on restricted imported compute. If a stable cluster remains elusive after chips arrive, schedules and training costs may suffer. Supply dates and actual training results still need confirmation from the company or supplier.

▪ SIGNALThe test of domestic compute moving from inference to frontier training is sustained cluster performance, not a shipment announcement.

❯ The Information says Alibaba named Dayiheng Liu head of its Qwen LLM project

[Leadership] The Information, citing two employees, says Alibaba named senior researcher Dayiheng Liu to lead its Qwen large-language-model project. A public page for Alibaba’s Apsara conference also lists him as head of the Qwen LLM project, providing a visible corroborating detail.

[Reorganization] Qwen has undergone repeated management changes this year. Key researcher Junyang Lin departed, while senior executive Jingren Zhou shifted toward longer-term research. Liu joined Alibaba in 2021 and has contributed to the models; the appointment clarifies day-to-day accountability.

[Execution] The personnel change does not alter weights developers have already downloaded. Its impact is more likely to emerge in release cadence and the open-model roadmap. Alibaba still needs research, applications and cloud distribution to move together.

▪ SIGNALDevelopers will feel a steadier Qwen release schedule sooner than they will feel a new title on the organization chart.

❯ Oura seeks to sell 50 million shares at $40–$44, for an offering of as much as $2.2 billion

[Terms] Smart-ring maker Oura and existing holders plan to sell a combined 50 million shares in a US IPO at $40 to $44 apiece. The top of the range yields an offering of $2.2 billion and an implied market value near $14.1 billion. The price remains proposed, not final.

[Proceeds split] The filing indicates Oura itself will issue roughly 13.5 million shares, with existing owners selling most of the remainder. The entire $2.2 billion therefore is not new corporate cash. Nine-month revenue through June was about $1.21 billion, up from roughly $698 million a year earlier, although the period still carried a large net loss.

[Investor lens] The health-data subscription story sits alongside hardware sales, and the two should be judged separately. Primary proceeds versus shareholder exits change how investors read the purpose of the listing. The final issue price will follow the roadshow.

▪ SIGNALRevenue growth may support a high range, but investors first need to separate money raised for Oura from money paid to selling holders.

GPT-6 Astra is credited with a Liouville–Goldbach proof; the classical conjecture remains open

[Public claim] A public mathematics project credits GPT-6 Astra with an unconditional proof of a Liouville–Goldbach statement: every even integer greater than two is the sum of two positive integers whose Liouville-function values are both minus one. The project includes a paper and Lean 4 files; an external account reports an independent rebuild.

[Precise boundary] A Liouville-function condition is not the requirement that both numbers be prime, so this does not prove the classical Goldbach conjecture. The project says it removes earlier generalized-Riemann-hypothesis and sufficiently-large-number conditions. A successful formal replay helps audit the proof chain but does not settle novelty, attribution or research significance.

[Research process] If further review holds, this would illustrate a workflow in which a model helps discover a route and formal tools check it. The model’s exact contribution remains difficult to separate from human prompting, selection and editing without a complete interaction record.

▪ SIGNALThe strongest follow-up is to establish how the proof route was found and whether independent mathematicians can fully reproduce it.

❯ Jev opens access to a low-cost decision model for fast, structured agent actions

[Model role] TypeSafe AI’s Jev is designed to return classifications, routes and scores rather than long-form prose. The company lists $0.042 per million input tokens and free output. The precise scope of reported sign-up credits and open access should be checked on the account page.

[Division of labor] In a public Minecraft project, a developer used GPT-6 Astra for planning and Jev for immediate action choices, recording a successful Ender Dragon run in 8 minutes 43 seconds. That is one developer demonstration in one environment, not a cross-task benchmark. It does not establish that OpenAI or Anthropic has an equivalent internal product.

[Agent economics] Calling a large reasoning model at every step accumulates delay and expense. Separating high-level planning from fast action selection may reduce loop cost, but a cheap decision model can still make repeated edge-case mistakes. The full task chain needs evaluation.

▪ SIGNALIf fast decision models prove reliable, agent optimization may shift from replacing one large model to assigning different decisions to different models.

❯ OpenRouter-based estimates put Chinese models at 67.46 trillion weekly tokens, with important coverage limits

[Platform estimate] National Business Daily calculated from OpenRouter data that models handled about 129 trillion tokens worldwide in the week of September 14–20. It assigned about 67.46 trillion to Chinese models and 14.21 trillion to US models, with Chinese-model usage up 10.28% week over week.

[Ranking] The report puts DeepSeek-V4.1-Flash first on the platform at about 15.8 trillion weekly tokens, up 219%. OpenRouter traffic grouped by model origin is not all model usage in either country, nor a measure of domestic consumption. Providers’ direct channels may fall outside the sample.

[Commercial limit] The numbers show developer preferences on a major routing platform, not national compute, revenue or technical leadership. Paid usage across platforms remains distinct from token volume because prices, free allowances and task types differ.

▪ SIGNALToken share shows where models are being tried and called; commercial share still depends on who pays and at what price.

❯ Kimi Code Desktop brings a visual coding-agent workspace to macOS and Windows

[Release date] Moonshot AI’s Kimi Code Desktop is available for Apple Silicon and Intel Macs and for Windows. Official release notes date the launch to September 17. A fresh round of Chinese coverage this week should not be mistaken for a September 21 first release.

[Workflow] Users can open a project, ask the agent to read and change code, run commands and inspect tool calls and file edits. A built-in terminal and browser sit alongside planning, approval prompts and background work, exposing what the agent does beyond a chat response.

[Adoption test] Review effort for developers will determine whether the desktop interface widens use. Visible changes and approval at sensitive steps matter more than simply putting a graphical shell around a command-line tool. Teams should still inspect permission modes and project-data boundaries.

▪ SIGNALAs coding agents move onto the desktop, task visibility, approvals and change review become product differentiators alongside model speed.

❯ A third-party test finds sustained-write slowdowns in a 1TB iPhone 18 Pro Max under extreme loads

[Test result] Third-party channel Homolab tested a 1TB iPhone 18 Pro Max and reported a sharp slowdown in its QLC storage after prolonged heavy writes. In its comparison with a 512GB TLC device, some read results were close; the larger gap appeared after write caches were exhausted.

[Cache effect] QLC holds more data per cell and helps raise capacity, while fast caching handles short bursts. Coverage of the test cites later-stage average writes near 79.4MB/s, with some phases lower. This is a particular device under stress, not proof that every capacity, batch or routine recording session behaves the same way.

[Buyer relevance] People recording long, high-bitrate videos or moving large files may care about sustained write speed more than advertised capacity. Photos, app installs and shorter clips usually finish within cached operation; an extreme test should not be turned into a claim of universal lag.

▪ SIGNALIf the large-capacity trade-off appears only after caches run out, buyers should compare it with their own continuous-write workload.