Back to AI Daily home

❯ Apple unveils its first foldable iPhone Duo, joining A20 Pro, a new Siri and Watch audio intelligence in its fall lineup

[Foldable debut] According to Apple’s official announcements, the company introduced its first foldable phone, iPhone Duo, at its September 9 US event hosted by new CEO John Ternus, with a US starting price of $1,999. The iPhone 18 Pro and Pro Max, Apple Watch Series 12 and Ultra 4, and AirPods 5 also debuted. Larger screens, multitasking and personal AI connected the updates across the lineup.

[Two screens] Duo has a 5.4-inch outer and 7.6-inch inner display, an A20 Pro chip, a new display engine and Touch ID. An adapted iOS 27 supports side-by-side apps, different folding positions and transitions between screens. Two rear cameras have 48-megapixel sensors; the inner screen has an under-display calling camera. Apple lists up to 31 hours of video playback on the inner display and 44 hours on the outer display under its test conditions. Storage runs from 256GB to 2TB. Preorders start October 16, with sales October 23, subject to regional availability.

[Pro upgrades] The iPhone 18 Pro and Pro Max retain 6.3-inch and 6.9-inch displays, adding a variable-aperture main camera, a smaller Dynamic Island and a new vapor chamber for imaging and sustained performance. Both use A20 Pro. Apple says its six-core CPU is up to 20% faster than A19 Pro and memory bandwidth is up to 50% higher; those figures do not translate directly into equivalent gains in every application. US prices start at $1,199 and $1,299, with preorders September 12 and availability September 18.

[Wrist and audio] The watches add a Health Sensing System and readiness score; Series 12 can measure background heart rate every five seconds. Live Rewind on the two new watches displays text from the preceding 15 seconds of conversation, while Siri Recap provides summaries. Apple adds audible and visual indicators, and features have device and rollout requirements. Ultra 4 starts at $799 in the US. AirPods 5 bring active noise cancellation to the $129 tier. The $149 version adds a wireless charging case, volume swipes and longer battery life. Live Translation has language and regional restrictions.

[Staged delivery] Siri AI and Apple Intelligence connect personal context, image processing and everyday tasks across phones and wearables, with some capabilities arriving later. Hardware availability and feature availability need to be assessed separately, particularly by language, region and companion device. App developers need useful multitasking interfaces for the larger screen. Recording and summarization products face new Watch features in the same use cases. Sustained usage will provide a better test than launch demonstrations of whether these changes justify upgrading.

▪ SIGNALApple is adding large-screen collaboration, personal assistance and ambient audio processing to established devices. Upgrade demand will depend on which features work locally and whether people use them every day.

❯ OpenAI says 10,000 agents produced a Navier-Stokes proof in 88 hours using an internal model stronger than Astra

[Company release] OpenAI published a proof and Lean formalization for the Navier-Stokes existence and smoothness problem on September 8, saying the successful group used about 10,000 concurrent agents over 88 hours. They ran an internal model still in training that the company describes as substantially more capable than GPT-6 Astra. This is not a newly available model for ordinary users.

[Scope of result] The company describes a finite-time singularity under smooth external forcing, addressing a specific formulation of the Millennium Prize problem. Astra then performed formalization and verification over an additional 17 hours. The project began September 1. Researchers reallocated resources, supplied problem variants to different groups and consolidated intermediate results. The effort therefore cannot be reduced to a model independently performing every part of the research without organization.

[Review and credit] Discussion also concerns unpublished work that mathematicians conducted using Codex. OpenAI denies accessing their specific user data while leaving open the possibility that de-identified data helped improve its models. Proof validity still requires independent examination, and research-data boundaries require a separate explanation. For research teams, adopting AI involves reproducible results, confidentiality and priority. A company publishing a proof does not substitute for formal acceptance by the mathematical community.

▪ SIGNALLarge agent groups turn computing resources into research attempts, but their results still have to withstand proof review, reproduction and examination of data provenance.

❯ Cognition raises $2 billion in Series E at a $48 billion valuation as annualized revenue approaches $900 million

[Four-month return] AI coding company Cognition announced a $2 billion Series E, Reuters reported, at a valuation of $48 billion, up about 85% from $26 billion in May. New investors a16z and Accel co-led alongside existing backers Founders Fund, General Catalyst and Avenir, roughly four months after the previous round.

[Revenue growth] The company says revenue run rate rose from $492 million to nearly $900 million over the same period, approximately 83%, close to the valuation increase. May’s $1 billion round was half the size of this one; both large raises support expansion within a short period. Annualized run rate extrapolates recent business activity. It is not recognized full-year revenue, profit or the total value of renewable contracts.

[Delivery demands] Flagship product Devin handles software engineering tasks such as planning, coding, testing and deployment, with a proposition centered on completing whole assignments. Using nearly $900 million as the denominator, the valuation is still around 53 times annualized revenue, with the exact multiple dependent on the final figure. Enterprise retention and delivery costs will determine how much growth translates into gross profit. Financing supplies resources for expansion while raising expectations for reliable execution on complex engineering work.

▪ SIGNALValuation and annualized revenue have grown at nearly the same pace over the same period. Task delivery, customer retention and operating costs will underpin the next valuation.

[Official figure] Legal AI company Harvey announced $550 million in funding on September 9, co-led by Diffusion and Lightspeed. Its announcement gives a $15.5 billion valuation, compared with $15.6 billion in Bloomberg’s report. The official figure is about 41% above the $11 billion valuation in March, with a small difference between the two latest accounts.

[Clients and tools] Harvey says 80% of Am Law 100 firms use its product, alongside five of the Fortune 10. The funding follows its first post-trained open-weight model and the Harvey LAB legal agent benchmark. The company is bringing industry workflows, model adaptation and testing into one product direction, with proceeds supporting hiring and capability development.

[Buying criteria] Adoption by a law firm or legal department does not mean every practice is deeply using the product, and customer coverage does not substitute for revenue. Accuracy on legal tasks has to show up in research, documents and multistep delivery, while confidentiality and review costs influence deployment depth. Whether specialized models and evaluations reduce rework will matter more to expanded purchasing than adding another generic chat interface.

▪ SIGNALInstitutional purchasing must support legal AI valuations. Model and evaluation spending is more likely to win sustained expansion when it reduces review and rework costs.

❯ DeepSeek reportedly engages CITIC Securities for a STAR Market IPO, potentially starting this year with deal size undecided

[IPO preparation] Reuters reported on September 9, citing two people familiar with the matter, that DeepSeek has engaged CITIC Securities to prepare a listing on Shanghai’s STAR Market, aiming to begin the IPO process this year. Deal size and timing remain unclear. This is a source-based report on preparations, not confirmation of a completed filing or listing approval.

[Separate tracks] The Hangzhou-based AI company continues to develop models and API services, including arrangements for V4.1 Flash. Product launches and IPO preparation follow different timelines: notices and live endpoints can establish the former, while formal documents are needed for the latter. Current reporting does not provide enough information to establish an offering valuation, use of proceeds or share structure.

[Disclosure ahead] Public documents would allow investors to assess model competitiveness alongside revenue quality, research spending and inference costs. Low-priced APIs can attract traffic without demonstrating steady profit. The depth of financial disclosure will influence how the market interprets growth and investment. Until then, deriving a valuation conclusion directly from listing discussions lacks support.

▪ SIGNALIPO preparations provide a potential financing route. Formal revenue, cost and capital-demand disclosures would reveal more about the economics of the model business.

❯ DeepSeek plans to route V4 Pro requests to V4.1 Flash after launch and charge Flash rates

[Migration plan] A DeepSeek platform notice says V4.1 Flash is scheduled for release around September 10 Beijing time. After launch and before V4.1 Pro arrives, V4 Pro requests will be routed to V4.1 Flash and billed at its price. The notice attributes the change to technical optimization, rather than asking users to manually move to another paid plan.

[Name versus execution] The company says the new Flash outperforms the old Pro on performance and cost metrics, with migration tied to the new release. That is a company comparison, not evidence of superiority on every task. Backend routing means the executed model may change while callers retain the same model name. The activation condition is formal launch; the notice alone does not establish that every request has already switched.

[Regression testing] Developers using V4 Pro in production should verify lower bills alongside task regression tests. Structured output, tool use, long context and response style can all affect an application. An unchanged name does not ensure unchanged behavior. Version, test-set and billing records help distinguish a model migration from changes in business data and measure the actual cost per successful task.

▪ SIGNALBackend substitution reduces migration work but leaves developers needing version identification and regression tests. Lower bills and stable task performance need to be accepted together.

❯ OpenAI product lead says Astra demand is unprecedented and new Pro subscriptions may be paused if pressure persists

[Capacity warning] OpenAI product lead Tibo said on X that Astra demand is unprecedented, with the team using available resources to sustain service. If the situation continues, it may pause new Pro subscriptions temporarily. He prioritized quality service for existing users, a position Sam Altman subsequently echoed.

[No suspension announced] For Astra users, this is a conditional warning, not a formal announcement that Pro sales have stopped. Tibo did not provide the size of the shortfall, a recovery date or changes to individual plans, nor details separating compute from scheduling or other bottlenecks. His comparison with previous steep growth describes pressure; it does not establish a user count or per-customer resource consumption.

[Service commitments] For people relying on Astra for work, subscription value depends on availability and stability. Completing a task at peak times is as immediate a concern as model capability. Restricting new subscriptions would put existing commitments ahead of short-term revenue expansion. Enterprise and development teams also need to check actual limits and terms when scheduling critical work, rather than interpreting a management statement as a permanent capacity guarantee.

▪ SIGNALNew subscriptions bring revenue and consume service capacity. When demand exceeds supply, the experience of existing customers becomes a constraint on expansion.

❯ Anthropic reportedly withheld Mythos 5.1 prerelease access from UK AISI, exposing a divide in cross-border safety evaluation

[Testing access] Anthropic did not provide the UK’s AI Security Institute with prerelease testing access to Mythos 5.1, the Financial Times reported, prompting British concerns about US AI protectionism. The question is whether an independent body can assess a new model early, not whether the model underwent any safety testing at all.

[Restricted availability] Anthropic’s official page says Mythos 5.1 launched September 1 for cybersecurity and biology research and is currently available only to some vetted US organizations. The company previously worked with UK AISI, and earlier Mythos access was affected by US government measures. Whether the government directly requested this particular denial of UK prerelease access has not been publicly confirmed.

[Independent evidence] Cross-border testing requires the actual model, sufficient permissions and preparation time. Reducing those conditions undermines the timeliness of independent evaluation and makes comparison harder for overseas customers. Company testing and third-party assessment serve different functions. Governments and buyers deciding whether restricted models can enter critical operations need to examine the verifiable scope of testing.

▪ SIGNALIf model restrictions also block independent testing, overseas users face a wider timing gap between access to capability and access to evidence about its risks.

❯ Anthropic alignment lead Evan Hubinger puts AI extinction risk above 10% over the next decade, citing recursive self-improvement

[Personal estimate] In public posts, Anthropic alignment science lead Evan Hubinger said he believes the chance of AI causing human extinction within the next decade is above 10%, and that superintelligence alignment remains unsolved. The figure is a researcher’s subjective risk estimate, not an incident statistic or a scientifically validated forecast probability.

[What concerns him] Hubinger subsequently clarified that he considers current-model risk low. His concern is superintelligence emerging through recursive self-improvement: systems helping improve the next generation, potentially accelerating capability gains further. The comments respond to debate about laboratory conduct and concern safety conditions for future development and release. They should not be edited into a claim that today’s Claude carries the same extinction risk.

[Testable decisions] Disagreement over progress and loss-of-control pathways does not disappear behind a striking percentage. Safety research and regulatory debate benefit more from specifying capabilities that trigger restrictions, evidence for checking improvement and responses to dangerous behavior. Separating personal probabilities from executable measures preserves the warning without presenting an uncalibrated judgment as a settled conclusion.

▪ SIGNALA subjective probability expresses the strength of a concern. What can be examined is which stopping conditions laboratories set and what they do when those conditions are met.

❯ Anthropic launches an interactive model of the US economy in 2030, exploring AI effects on growth, wages and jobs

[Scenario explorer] Anthropic introduced an interactive tool for the US economy in 2030, allowing users to change assumptions about capability, adoption, autonomy, productivity and job-transition speed. It shows resulting growth and labor-market outcomes and compares expectations with those of 10,980 US respondents. It illustrates the transmission of assumptions rather than offering a single certain forecast.

[Growth scenarios] In the tool’s modest, substantial and extreme scenarios, GDP in 2030 is respectively 1.6%, 8.3% and 32.4% above a no-AI baseline. These are differences in economic size by that year, not three annual growth rates. The most aggressive case assumes extensive automation of knowledge work, bringing job-transition pressure. Changing adoption and reemployment conditions changes the results.

[Distribution] A larger economy does not ensure equal gains for every class of worker. By representing jobs as tasks, the model helps separate capability improvements, business adoption and worker transitions. For corporate staffing plans, the pace of job restructuring is closer to operational reality than a single capability score. Public debate also needs to examine income distribution. These are scenario mechanisms, not labor-market changes that have already occurred.

▪ SIGNALAI capability, business adoption and worker transitions move at different speeds. Whether growth becomes broadly shared income gains depends on how those speeds interact.

❯ Jeffrey Katzenberg reportedly teams with former Sora head on a video-model startup for professional filmmakers

[Planned team] Film executive Jeffrey Katzenberg, former OpenAI Sora head Bill Peebles and former Dropbox CFO Sujay Jaswa plan to start an AI video company that would train models for filmmakers, The Information reported. The lead describes a venture being planned, not commercial progress from an already released product.

[Professional needs] The proposal brings video-model development and filmmaking requirements into one team, aiming to serve professional production workflows. Beyond producing a watchable clip, delivery requires continuity between shots, consistent characters and repeated revisions. Those are requirements of the target market, not capabilities already demonstrated by the new company. Verifiable details on financing, training data and launch timing remain limited.

[Adoption hurdles] Filmmakers’ adoption decisions depend on controllable delivery and rights arrangements. Technical and industry experience may help identify real needs, but paying customers and workflow validation will test viability. Whether the company secures project budgets will depend on its actual product, delivery results and partnerships. It is too early to determine that outcome.

▪ SIGNALProfessional video models need to produce material that is editable, deliverable and covered by clear rights arrangements before they can earn a place in film budgets.

❯ China reportedly raises humanoid IPO bar, seeking recurring revenue, a path to lower losses or genuine innovation

[Window guidance] China’s securities regulator has issued informal window guidance to some banks and companies raising the listing bar for humanoid robot businesses, The Information reported. Reuters relayed the report while saying it could not verify it independently and that regulators did not immediately respond. This remains a reported review trend, not a published blanket prohibition.

[Operating evidence] The report says companies must demonstrate recurring revenue and progress toward lower losses or actual innovation. It follows volatility after Unitree’s listing, with a private-market funding boom and more listing candidates also cited as background. Attention extends from robot sales to revenue durability. One-off purchases, recurring business and trial projects should not be treated as the same evidence of steady demand.

[Funding timelines] If implemented, these requirements would make operating evidence more consequential for financing plans, with order quality and technical differentiation affecting progress. Listing uncertainty also influences cash use for teams without stable revenue. Formal documents and individual project milestones are still needed; anonymous reporting does not establish that the industry’s entire listing route has closed.

▪ SIGNALRobot companies need to demonstrate sustained operations beyond demonstrations. Tighter review expectations bring revenue quality and cash consumption into financing discussions earlier.

❯ Physical AI simulation company Antioch raises $32 million Series A to reduce hardware-validation work through cloud testing

[Simulation funding] Physical AI simulation company Antioch raised a $32 million Series A led by Greylock, with A* and other investors participating, Forbes reported. It builds high-fidelity simulations for customers’ hardware and disclosed Amazon’s Ring as a customer, aiming to speed development that otherwise requires physical testing after each change.

[Faster experiments] The product expands customer experiments into parallel cloud evaluations for robotics, drones and smart security. Ring says simulated and physical results were close in some scenarios withheld from calibration. The approach retains real-data calibration and uses simulation to expand testing. It aims to reduce physical trials, not eliminate hardware validation; previous real-world results remain an important reference for judging simulation quality.

[Validation value] For physical AI teams, inexpensive simulation is useful only when it predicts hardware performance. The simulation-to-reality gap determines which tests can move to the cloud, while iteration time determines engineering savings. Benefits shrink if rare scenarios and hardware changes still require extensive manual rebuilding. That is part of what buyers need to validate.

▪ SIGNALThe commercial value of physical AI simulation depends on how much repetitive testing it replaces and whether retained real-world calibration controls the error.

❯ Xiaomi opens MiMo Desktop invitation testing with two MiMo-X preview models for desktop work and multi-agent collaboration

[Limited trial] Xiaomi opened invitation testing for MiMo Desktop on September 8. Approved users receive time- and quota-limited free access to MiMo-X-Pro-Preview and MiMo-X-Flash-Preview. Existing MiMo platform users receive priority, according to the company. This is a trial rather than a generally available commercial service.

[Task coverage] Official targets include complex reasoning, coding, office work, web design, video editing, 3D and music generation, and computer control. The client can show interactive previews within conversations and supports browser operation and multi-agent collaboration. These define testing directions. The models are still changing, and actual completion quality must be assessed task by task; a coverage list does not establish production readiness in every category.

[Desktop competition] Xiaomi is moving from model APIs toward delivering desktop work, entering the agent market alongside Tencent, Alibaba and ByteDance. Users benefit from fewer interventions and editable deliverables more than a single demonstration. Invitation testing gives Xiaomi real-task feedback. Free quotas and preview capabilities should not be treated as lasting commitments, with formal pricing and availability still to come.

▪ SIGNALDesktop agents need to reduce intervention and rework on users’ own tasks. During invitation testing, the most valuable feedback identifies the exact step where execution fails.