Back to AI Daily home

❯ OpenAI Switches Free-Tier Default Model to GPT-5.6 Luna and Opens Unlimited Text Chat

[FREE TIER SHAKEUP] OpenAI is swapping the default model on ChatGPT’s free tier to GPT-5.6 Luna and, starting next week, opening unlimited text conversations to free users — the first time in ChatGPT’s three-year history that the free tier has dropped its message cap. Also launching alongside it is a Think button that lets users manually switch into reasoning mode for hard questions. Per Axios reporter Herb Scribner, limits on file uploads and image generation remain in effect.

[PAID TIER UPDATE] The paid side is updating in tandem: both the Instant and Thinking paths for Plus and Pro users unify on the improved GPT-5.6 Sol, plus a new reasoning-intensity slider that lets users decide how much reasoning compute each answer burns. Sam Altman confirmed on X the same day that “5.6 Sol is much better in chat scenarios,” while Greg Brockman placed the update on the path to “making the product simpler over time.” The reasoning here isn’t complicated: for the past year, the free tier’s message wall has been ChatGPT’s most obvious shortfall against Gemini’s free tier and several free apps in China — and Luna, this generation of small models, has driven unit reasoning costs down to a level where dropping the cap is viable.

[COMPETITION] What this shift really rewrites is the customer-acquisition cost baseline for free AI assistants. Teams building consumer AI apps have typically won users on “more generous free quotas” — a moat OpenAI has now filled in one move. The reasoning-intensity slider also hands part of the reasoning-budget decision to users, effectively admitting that a one-size-fits-all default reasoning tier is suboptimal on both cost and experience. Expect this design to be followed.

▪ SIGNAL Removing the free tier’s message cap is OpenAI trading falling inference costs for distribution — and taking away the best card its competitors had to play.

❯ Bloomberg: OpenAI’s First Hardware Is a Hockey-Puck-Sized Smart Speaker, Priced Above $300, Set for 2027

[FORM] According to Bloomberg reporter Mark Gurman, OpenAI’s first hardware device is a screen-free, donut-shaped smart speaker roughly the size of a hockey puck, compact enough to carry around the house in one hand. It’s priced above $300 (another report cites a $300–$400 range), with release set for 2027.

[DETAIL] The device was designed in collaboration with Jony Ive’s LoveFrom team and uses high-cost materials such as metal. It packs a battery, camera sensors, and dynamic lighting — plus movable mechanical parts meant to give it “personality” rather than making it a passively responsive speaker. OpenAI hopes to preview the product publicly this year, then start selling it next year. Measured against the hype of its 2024 acquisition of io, the form factor is actually quite conservative: neither glasses nor a pendant, it lands squarely in the desktop-speaker category Amazon and Google have fought over for a decade.

[ASSESSMENT] The speaker poses a question on behalf of smart-speaker makers: how much hardware premium can a smart-enough model support? Amazon Echo and Google Nest have long kept prices at the $50 tier; OpenAI is starting at above $300, betting that conversational quality alone can justify a sixfold price gap. That bet won’t be settled until 2027 — with an entire round of hardware supply-chain price increases in between.

▪ SIGNAL A price six times that of the Echo puts “is the model any good” directly on the price tag of consumer electronics.

❯ DeepSeek Backend Notice Flags Major API Price Hike, Hitting an Inflection Point in the Low-Price Era of Chinese Large Models

[ANNOUNCEMENT] DeepSeek has posted a notice in its user console saying it will significantly raise API pricing across the board in the near future, without disclosing the exact scale or effective date — the company said only that an official notice will follow. The company that drove domestic large-model prices to rock bottom is, for the first time, voluntarily pulling prices back — its V4-Flash output price was previously as low as 2 yuan per million tokens, and cache-hit input cost as little as 0.02 yuan per million tokens.

[CAUSE] The price hike is not a sudden whim. What the low prices bought was call volume too heavy to sustain: weekly call volume for a single model once exceeded tens of trillions of tokens and repeatedly ranked No. 1 globally — at the cost of frequent timeouts and lag on API endpoints during weekday peak hours. In mid-July, DeepSeek had already tested the waters, rolling out time-of-day pricing — prices doubled during the two weekday peak windows of 9:00–12:00 and 14:00–18:00, while off-peak hours and weekends stayed unchanged. This time, it is swapping the time-based patch for an overall price adjustment. Researcher teortaxesTex argued on X that if DeepSeek also unlocks vision capabilities on the API at the same time, the hike would come with something worth paying for — its image understanding is already good enough, and the cache mechanism makes that part extremely cost-efficient.

[FIRST HIT] The first to be hit are small and mid-sized developers who use DeepSeek as their default backend: over the past year, countless applications built their cost models on the premise that “tokens are so cheap they don’t need to be counted” — and that premise is now falling apart. The pricing logic of large-model APIs is thus switching from “burning cash for scale” back to “billing by compute,” and it comes precisely at the moment when rivals like Kimi K3 and Qwen3.8-Max are pushing per-task costs even lower. The price war among Chinese large models may not be over, but the phase propped up by losses has run its course.

▪ SIGNAL The price slasher is raising prices itself — the ledger behind rock-bottom pricing no longer balances.

❯ AMD Acquires Toronto Chip Startup Taalas, Etching Model Weights Directly into Silicon

[DEAL CLOSED] AMD announced on August 6 that it has reached a definitive agreement to acquire Toronto chip startup Taalas; the deal value was not disclosed. Taalas’s approach is extreme: model weights are etched directly into the silicon, bypassing high-bandwidth memory loading, with the company claiming more than an order of magnitude improvement in inference performance. This is AMD’s second acquisition bet on the inference side, following its Cerebras deal a few weeks ago.

[TECH & TRADEOFF] Taalas’s chip is currently divided into two parts — weights are fixed in a mask read-only region, while KV cache and fine-tuning adapters sit in an SRAM region. Of the hundred-plus layers, only two need to change with model design; the company says its in-house tools can complete a tape-out in about two months. The trade-off is equally blunt: a finished chip can only run the model it was built for — switch models and you switch silicon. Its second-generation HC2 chip is planned to raise parameter capacity to 20 billion. AMD’s integration plan is to have the Taalas chip work alongside Instinct GPUs, plug into Helios racks and the Epyc platform, and run on the unified ROCm software stack.

[INDUSTRY IMPACT] What this acquisition is truly betting on is the structure of inference costs: as long as model iteration slows down, welding weights into silicon eliminates the most expensive part — high-bandwidth memory. And right now, HBM is the industry’s most constrained material. Cloud providers running inference services must start weighing a new question: which models are stable enough to justify a dedicated tape-out. The more answers there are, the more inference share gets carved away from general-purpose GPUs.

▪ SIGNAL Welding weights into silicon is a bet that model iteration will eventually slow — AMD is willing to place it first.

❯ Nvidia Reportedly Evaluating Lower VRAM Configurations for Rubin Ultra to Address High-End HBM Shortage

[SPEC CUT] According to multiple supply-chain media reports, Nvidia is evaluating reducing the VRAM configuration of its next-generation flagship GPU, Rubin Ultra, and has been running parallel tests on at least three versions, with some versions falling below previously disclosed specs. The original plan was to use HBM4e 12hi across the entire lineup; now lower-spec options such as HBM4e 8hi, HBM4 12hi, and HBM4 8hi have entered the evaluation sheet.

[SUPPLY GAP] The impact is quantifiable: maintaining the top-tier configuration can lift I/O speeds from the previous generation’s 8 to 11.7Gbps up to 14 to 16Gbps; if forced to fall back to lower HBM4 specs, the gain may only reach 11 to 12Gbps. The root of the cut lies in supply — HBM shipment capacity in 2027 is expected to grow 50% to 60% year-on-year, yet still won’t close the gap in AI compute demand, while the most advanced HBM4e 12hi remains uncertain in both validation progress and mass-production yield. This shortage chain has already spilled over to the consumer side: memory-chip price increases are pushing up system costs, and companies like Apple have already raised hardware prices.

[CASCADE] VRAM cuts do not mean performance is halved — they can be compensated through other means, but the math needs to be redone: AI companies running large models, with smaller per-card VRAM, will need more chips to fit the same model, raising procurement costs per unit of compute. Data center procurement budgets must therefore be rebalanced between “number of cards” and “card specs” — the linear extrapolation based on card count that has been the habit over the past two years no longer holds.

▪ SIGNAL Even Nvidia has to make concessions on its own flagship; memory is the true limiting factor in this round of compute expansion.

❯ The Information Says Stripe in Exclusive Talks to Acquire OpenRouter at Nearly $10 Billion Valuation

[DEAL] According to The Information, payments company Stripe has entered exclusive negotiations to acquire model-routing platform OpenRouter in a cash-and-stock deal valued at close to $10 billion. The reference point is striking: OpenRouter just closed a $113 million Series B in May 2026 at a valuation of only $1.3 billion — in three months, the valuation has risen nearly sevenfold.

[PARTIES] OpenRouter was founded in 2023 and acts as the intermediary layer between model developers and enterprise users: developers connect once via API and can compare, call, and switch among hundreds of closed-source and open-source models at any time. It already has a business relationship with Stripe — OpenRouter’s own payments run on Stripe. The report says several other large tech companies also evaluated a bid. For Stripe, this is a step from its core payments business toward AI infrastructure: model calls are becoming high-frequency, usage-based consumption, and metering and settlement happen to be Stripe’s bread and butter.

[LANDSCAPE] This deal puts a price on the model-routing layer for the first time. Startups that aggregate models and provide gateways were once questioned for having “no moat and eventually being eaten by model makers themselves.” The $10 billion bid offers a different answer: when callers need to compare and switch among dozens of models, the middle layer becomes a toll booth. Model makers’ pricing power will be diluted a notch — the more users get used to entering through the routing layer, the weaker single-model brand loyalty becomes, and pricing power shifts to the middle layer along with call volume.

▪ SIGNAL The sevenfold valuation jump in three months isn’t about the technology — it’s about the spot in front of every model toll booth.

❯ Qwen3.8-Max Takes Fifth in Intelligence Index, First in Agentic Index, but Costs Still Higher than Kimi K3

[SCORES] The latest results from independent evaluation body Artificial Analysis show Alibaba’s Qwen3.8-Max scoring 56 in the Intelligence Index, ranking fifth, and taking first in the Agentic Index; average cost per completed task is USD 1.14. Alibaba’s Tongyi official account confirmed both rankings on X that day, citing the agency’s public leaderboard.

[CONTRAST] But the leaderboard has another half: the open-weights camp’s leader Kimi K3 scores 1 point higher, with a per-task cost of just USD 0.86 — roughly 25% lower than Qwen3.8-Max. In other words, at the same level of intelligence, this closed-source version has not bought a cost advantage. Looking at the domestic lineup, Qwen3.8-Max has 2.4 trillion parameters and K3 has 2.8 trillion; the two are so close that they’re separated by a single decimal place.

[TAKEAWAY] Enterprise buyers choosing models are now really comparing cost per task, not leaderboard rankings. When the score gap is only 1 point but the price gap is 25%, ranking position is no longer a procurement rationale. Taking first in the Agentic Index is a more tangible win for Alibaba — agent scenarios demand high stability in multi-turn tool calls, and being first in this category converts to orders better than a higher overall score.

▪ SIGNAL A 1-point score gap and a 25% price gap — no need to hesitate over which way procurement decisions will tilt.

[LEAD] On August 3, MiniMax officially open-sourced its next-generation general-purpose multimodal generation model MiniMax H3, which took the No. 1 spot globally in video editing capability on Artificial Analysis and topped the Hugging Face trending chart, surpassing DeepSeek V4-Flash, the previous leader. According to MiniMax’s official announcement, the model supports video generation at up to 2K resolution, up to 15 seconds in length, and native stereo audio.

[ECOSYSTEM] Even more notable is the scale of adaptation on launch day: chipmakers including Huawei Ascend, Moore Threads, MetaX, Hygon, Kunlunxin, Tianshu, and Biren, along with AMD and Intel, as well as Hugging Face and multiple cloud inference platforms, completed Day-0 adaptation in sync; more than 100 domestic and international partners had integrated within 24 hours of the open-source release. The capital markets responded just as directly — Jefferies reiterated its Buy rating with a target price of HK$1118, while Citi and Goldman Sachs also maintained Buy ratings, citing H3’s reinforcement of MiniMax’s position in AI video generation and a cost advantage that supports customer acquisition and commercialization. As of the close on August 5, MiniMax shares were up more than 10%, and the stock was also added to the Stock Connect eligible list during the same period.

[COMPETITION] What truly widens the gap in this round is the breadth of chip adaptation. Teams looking to run video models on domestic compute have long been dogged by the “model is ready, but the cards can’t run it” problem; Day-0 coverage of eight domestic chips effectively compresses the deployment cycle from months to a single day. The competitive standard for open-source models is thus shifting from pure benchmark scores to “how many hardware and platform players are willing to show up on day one.”

▪ SIGNAL Eight domestic chips adapted on the same day says more about what MiniMax has a grip on than a No. 1 spot on the leaderboard.

❯ LatePost Reports ByteDance Discussing a Model with Over 5 Trillion Parameters, the Largest Known in China

[SCALE JUMP] LatePost reports ByteDance is discussing training a model with more than 5 trillion parameters, surpassing Alibaba’s Qwen3.8-Max at 2.4 trillion and Moonshot AI’s K3 at 2.8 trillion — the largest known in China to date. The plan is still early-stage and may ultimately not be released.

[PEOPLE & PATH] The project is led by Seed Foundation head Xiang Liang, in collaboration with LLM pre-training data lead Shen Ke. Xiang Liang joined ByteDance in 2016 and worked on the AML machine-learning middle-platform team before becoming head of the Doubao LLM Foundation team; Shen Ke joined right after graduating from Tsinghua in 2018 and now focuses mainly on pre-training data. The report also cites two directional principles from Zhang Yiming: no distillation, and don’t be swayed by short-term hotspots like coding. In the first half of the year, multiple Seed teams repeatedly reassessed their work; they are now reworking organization and resource allocation — rather than continuing to chase at existing model sizes, they want to push parameters to several times peers’ scale in one move and go straight for the lead.

[VERDICT] This is ByteDance’s classic “brute force creates miracles” play, but the cost is out in the open: training and inference costs for a 5-trillion-parameter model will both climb, and HBM and advanced packaging capacity are both in a tight cycle right now. Other domestic model makers need to decide in advance whether to join this round of the parameter arms race — sit out and risk falling behind on capability, or join and bet an entire budget cycle at the most expensive moment for compute. The project isn’t finalized yet, but the pressure has already spread.

▪ SIGNAL Skip the catch-up and go straight for scale — ByteDance is betting that compute can buy time.

❯ Alibaba Cloud Video Generation Model Wan3.0 Enters Public Beta, Generates 30 Seconds per Run and Supports Document Input

[BETA] Alibaba Cloud’s next-generation video generation model Wan3.0 has opened public beta, generating up to 30 seconds of video in a single run. On top of the four base modalities of text, image, audio, and video, it supports document-format input for the first time, including doc, xls, ppt, pdf, and md. The beta spans Alibaba Cloud Bailian, Wanjing Yike, the Wanxiang official website, and Qwen Creation on PC, with the Qwen App in staged gray release.

[PRICING] API pricing is split into three tiers by resolution: 480P, 720P, and 1080P at 0.3 yuan, 0.6 yuan, and 1.2 yuan per second, respectively, and the API will be fully opened in the near term. A caveat worth flagging: these are beta-period reference prices only; official pricing has yet to be announced. The 0.6-yuan-per-second rate at 720P is a meaningful threshold for teams mass-producing content such as short dramas and ad clips—the model cost for a finished 30-second video lands at around 18 yuan.

[USE CASES] Document input deserves more attention than the duration number. Teams producing enterprise content can now drop a PPT or a financial-statement spreadsheet straight into the model and get a finished video out, removing the manual middle step of “first rewriting the document into storyboard prompts.” The input side of video models thus expands from prompts to structured files, pulling users a big stride from the creative side toward the enterprise-office side—and directly reshaping the labor-cost structure of content teams.

▪ SIGNAL A video model that can chew through PPTs and spreadsheets isn’t just taking the editor’s job anymore.

❯ Unitree Robotics Reveals Strategic-Placement Roster; DeepSeek Allocated About 141 Million Yuan

[ROSTER] Unitree Robotics has published the strategic-placement roster for its STAR Market IPO, with 9 investors participating and 8.0893 million shares allocated in total. According to the prospectus, Hangzhou DeepSeek (DeepSeek) was allocated about 933,400 shares — 2.31% of the offering — worth roughly 141 million yuan. It is the first time the model company has appeared as a strategic investor in a hardware company’s IPO lineup.

[COHORT] Alongside DeepSeek in the same batch: Tencent’s Shanghai Qishan Investment, CNPC’s Kunlun Capital, China Southern Power Grid’s Industry-Finance Holdings Group, and Tianyi Capital Holdings — each allocated 2.23%, all carrying the status of “large enterprises with strategic cooperation ties or long-term cooperation vision with the issuer, and their affiliates.” The offer price is set at 150.80 yuan per share, implying a total market capitalization of about 60.993 billion yuan based on post-offering share capital. Unitree’s rationale in the filing is blunt: bringing in the large-model leader is to focus on joint R&D of large models and embodied AI, improving robots’ understanding of complex scenarios and their generalization capabilities.

[SIGNAL] A company whose core business is models — and which just announced API price hikes — is shelling out 141 million yuan for a strategic placement in a robot-body maker. That suggests the boundaries of the division of labor in embodied AI are loosening. What humanoid-robot startups must reassess: whether their brain-side suppliers could become shareholders, or even rivals. When model makers start holding equity in body makers, procurement and competition get bound together. With CNPC, China Southern Power Grid, and Tianyi — three state-owned players — entering at the same time, Unitree’s downstream scenarios shift from consumer showcases to industrial inspection and energy operations.

▪ SIGNAL A model company is paying for robot-body equity — upstream and downstream in embodied AI are crossing into each other’s turf.

❯ US Sets Minimum Import Prices for Solar Supply Chain; Modules No Less Than $0.38 per Watt

[PRICES SET] The US has set minimum import prices across the solar supply chain: $21 per kilogram for polysilicon, $100 per kilogram for polysilicon ingots and wafers, $0.22 per watt for solar cells, and $0.38 per watt for solar modules, plus a 15% ad valorem tariff on specified polysilicon ingots and related derivative products — reportedly taking effect in early December.

[CONTEXT] Washington’s stated rationale is not trade but national security: polysilicon is also a critical semiconductor material, tied to radar, communications, and the control systems of missiles and drones, so domestic capacity is deemed critical. Previous solar interventions relied mainly on tools like anti-dumping duties and the 15% ad valorem tariff — but a minimum price is a harder instrument: a tariff only adds cost, a price floor draws a line no one can undercut, shutting down the entire playbook of winning on scale and cost. The package reportedly takes effect in early December, covering all four segments: polysilicon, wafers, cells, and modules.

[SPILLOVER] It looks like a solar story, but the landing point is compute. Polysilicon is a shared upstream for both semiconductors and solar, and AI data centers are the most voracious buyers of incremental electricity right now. The first to feel the pressure are operators building data centers in the US: with module prices held above $0.38 per watt, the construction cost of supporting solar plants rises along with it, and the math on self-built green power turns ugly immediately. Where the power comes from and what it costs per kilowatt-hour is becoming a line item in the compute buildout every bit as important as buying GPUs.

▪ SIGNAL The floor is drawn under solar; the pressure lands on the data center’s power bill.

OUTLOOK

[TODAY'S BATCH] Today’s five items are all really saying the same thing: the era of cheap compute is winding down. DeepSeek posts a price-increase notice, Nvidia trims VRAM on its flagship card, AMD buys a solution that welds weights into silicon, the US draws a price floor under solar modules, and ByteDance is discussing 5 trillion parameters — memory and power, the two hard constraints, are tightening simultaneously while model scale keeps charging upward. Caught in the middle are all the application teams whose cost models are built on the premise that tokens will keep getting cheaper.