Back to AI Daily home

❯ Alibaba prices Qwen3.8-Max at $2 per million tokens, open-sources weights next week

[PRICE MOVE] Alibaba has priced its 2.4-trillion-parameter Qwen3.8-Max at $2 per million input tokens and $6 per million output tokens. According to The Information, that undercuts Kimi K3 — which shipped days earlier at $3 / $15 — by one-third on input and 60% on output. Bloomberg reports Alibaba says the model beats Kimi K3 on some benchmarks and plans to release the weights of two models together next week.

[VALIDATION] This time Alibaba didn’t just toss out benchmark scores. It produced evidence outsiders can audit line by line: the model started from an empty folder, worked autonomously for 16 straight days, produced a command-line tool, and left behind 265 commits and 127 pull requests — the entire process public on GitHub. The move targets one of the hardest capabilities for today’s large models to credibly demonstrate: staying on track through long-horizon tasks. The release cadence is tight, too. Kimi K3 landed just days ago; MiniMax’s H3 went live the same day as Qwen3.8-Max, hours apart. Independent evaluator Artificial Analysis posted it to its leaderboard that same day, with an AA-Briefcase total score of 1430 Elo and a 48% rule-pass rate — above Claude Sonnet 5 and GPT-5.6 Sol at 42%, but still trailing Kimi K3.

[REALITY] The internal gap on this report card is worth unpacking: Qwen3.8-Max scores 1595 Elo on analysis quality but only 1340 on presentation quality — a 255-point spread. It reads more like a model that can think its way through a problem but stumbles on the final delivery. Researchers who’ve already gone hands-on offer a more mixed read; some find it unusually sensitive to the calling framework, with the same task able to burn through over a million tokens in different environments. The teams that actually get the bargain are those willing to write their own orchestration layer; teams expecting to save money by simply swapping API endpoints will likely burn their per-token savings right back into token volume.

▪ SIGNAL Open-sourcing the weights cedes pricing power; the only layer left to monetize is the engineering that makes the model run smoothly.

❯ Artificial Analysis Estimates: DeepSeek V4-Flash Costs 3 Cents per Task

[COST CLIFF] Independent benchmarking outfit Artificial Analysis estimates that a single test run of DeepSeek’s newly released V4-Flash costs about 3 cents on average. On the same basis, Kimi K3 comes to 86 cents, GPT-5.6 Sol to $1.86, and Claude Fable 5 to $3.15 — the most expensive is more than 100x the cheapest, a set of figures Reuters picked up. V4-Flash’s list price is $0.14 per million input tokens and $0.28 per million output tokens.

[METRIC MATTERS] Per Artificial Analysis, the key to this calculation isn’t the list price but the fact that it measures “what it costs to actually finish a task”: every token the model consumes to reach an answer is folded into the number. For the past year, enterprises picking models have generally looked only at list prices — a brutal basis for reasoning models. Many models look cheap per token, but once the chain-of-thought runs long, the real bill multiplies several times over. V4-Flash is a 284-billion-parameter mixture-of-experts model whose official weights store routed experts directly in MXFP4, with the rest in FP8 or BF16 — effectively baking quantization into training. What it saves at inference isn’t just VRAM. Last week, third parties already released 14 quantized versions, one of them specifically flagged for machines with 128GB of memory.

[PRESSURE POINTS] This estimate lands first on how procurement compares prices: for the past year, enterprise selection has revolved around per-million-token quotes; now that number in isolation means little. The first to feel the squeeze are inference providers selling a premium on token price — when the cost to complete the same task can differ by two orders of magnitude, “we’re expensive for a reason” needs something harder than model quality to back it up. DeepSeek’s own gross margin, per its earlier disclosures, is not low; the cost edge comes from engineering, not subsidies, which makes matching its pricing even harder.

▪ SIGNAL Per-million-token pricing is losing its reference value; what counts is the total bill for getting a task done.

❯ Palantir Q2 Revenue Up 93%, US Commercial Business Grows 1.5x

[BEAT] Palantir posted second-quarter revenue of $1.94 billion, up 93% year over year, above the $1.81 billion the market expected; US commercial revenue reached $764 million, up 149% year over year, the main driver of overall growth. The company also raised its full-year 2026 revenue guidance to $8.15 billion to $8.16 billion, implying 82% annual growth, and shares rose more than 14% in after-hours trading.

[PROFIT] Less discussed than the growth rate is profitability: GAAP operating profit was $912 million this quarter with a 47% operating margin, adjusted operating profit was $1.19 billion with a 62% margin, and EPS was $0.41, above the $0.35 expected. This is Palantir’s eighth consecutive quarter of beats. The company also raised its full-year US commercial guidance to above $3.424 billion, implying at least 134% growth — a figure that reflects management’s judgment that demand will not cool in the second half. Palantir has long been viewed as a company that lives on government contracts, but US commercial’s share has climbed steadily over the past two years, and this quarter’s growth rate was already several times that of the government business.

[VALUATION] This report pushes the debate back to valuation itself: there is almost nothing to fault in the results, and with the stock already down about 40% from its highs before the earnings, the market’s disagreement is not about fundamentals but about the multiple. The investors who need to redo their homework are those using software-company valuation frameworks — a company growing 93% with a 62% operating margin is hard to place on any existing comps table. The signal from the enterprise customer side is also worth singling out: US commercial has posted triple-digit growth for several straight quarters, showing that companies actually spent their AI deployment budgets rather than stopping at the pilot stage.

▪ SIGNAL Enterprise AI spending has finally left its mark on an income statement, not just an intention announced at launch events.

❯ White House Says Advanced AI Model Evaluation Framework Completed on Schedule, But Details Withheld

[ON TIME, NOT PUBLIC] A White House official said the voluntary evaluation framework for advanced AI models required by the June 2 executive order has been completed on schedule — but the White House has not disclosed the framework’s contents, who has reviewed it, or when companies will begin using it. According to Axios, a working-level meeting with companies is scheduled for Tuesday to introduce the finished product to vendors.

[SECRECY CLAUSE] This opacity is not an ad hoc decision; it was written into the original text itself. The executive order explicitly states that the benchmarking process for assessing models’ advanced cyberattack capabilities is classified, and that the threshold determining which models fall under oversight is likewise classified. The White House says the industry partners it has engaged go well beyond OpenAI, Anthropic, and Google. Many had expected a discussable evaluation standard to be made public, since a “voluntary framework” typically derives its binding force from transparency — once companies sign on, outsiders can hold them to their commitments. What has emerged instead: a framework whose contents no one outside can see, paired with a pledge of voluntary participation.

[THE DILEMMA] This design creates a dilemma for policy researchers and corporate compliance teams alike: the former cannot determine where the threshold is set or which models are covered, making it impossible to judge whether the framework is strict or lenient; the latter must decide how many resources to dedicate to aligning with a standard without seeing the full text. The most directly affected area is how model capabilities are publicly disclosed — if cyber capability evaluation results are themselves classified, the boundary of what vendors can and cannot write in release notes shifts accordingly. After Tuesday’s meeting, most likely only attendees will know what the framework looks like.

▪ SIGNAL An invisible framework — its binding force ultimately rests on whether signatories are willing to voluntarily disclose what they’ve done.

❯ Sources Say Dario Amodei Worries New Hires Join Anthropic for Money, Not Mission

[INTERNAL CONCERN] Per Axios, citing people familiar with the matter, Anthropic CEO Dario Amodei has voiced internal concerns: some new hires are here for the money, not the mission. The same report also cites a retention rate of roughly 80% over the past two years, and finds OpenAI engineers are 8 times more likely to jump to Anthropic than the reverse.

[COMP WAR SCALE] The backdrop is compensation that has inflated to the point of distortion. Sam Altman has previously acknowledged signing bonuses as high as $100 million used to poach top researchers; Meta has poached Joel Pobar, who led reasoning work at Anthropic. Amodei’s response runs counter to most peers — he has made clear Anthropic will not broadly raise salaries to counter poaching, a rare choice in a market where talent is fought over with a trifecta of cash, equity, and compute quotas. Mission alignment has always been Anthropic’s chief differentiator in hiring; as compensation gaps stretch to dozens of times, how many people that pitch filters out — and how many it holds onto — is now being tested by reality.

[TWO-SIDED PRESSURE] The remarks push hiring-stage screening standards to the fore: a company that won’t match offers can only trade off along the line of “willing to earn less for the mission,” and the tighter that line is drawn, the smaller the candidate pool. For candidates, the center of judgment is also shifting — in the past it was about comparing numbers; now they must weigh whether the compute resources and research freedom gained by declining a high-premium package are worth it. An 80% retention rate is high for this industry, but it reflects the past two years, while compensation packages have only jumped in magnitude over the last six months.

▪ SIGNAL The cost of filtering for mission rises with every increase in rivals’ offers.

❯ Alibaba’s QwenWork Opens Public Beta, Three Products Unified Under Chen Yusen

[LAUNCH] Alibaba’s enterprise-grade AI agent QwenWork (千问办公) opened public beta on August 3. Both individual and enterprise users can experience it through the official website; the web version and standalone PC client are already live, with the DingTalk entry point to follow. The product was forged from the merger and restructuring of QoderWork, Wukong, and MuleRun — the three original products no longer exist as standalone offerings, and their talent has been consolidated as well. Under the hood it runs the just-released Qwen3.8.

[CONTEXT] This integration has moved fast: Chen Yusen took over DingTalk in June, pushed Wukong and MuleRun to merge within a week of taking office, folded QoderWork into the same lineup in early July, and went straight to public beta under the new brand in early August. Organizationally, the QwenWork business unit was upgraded from the former Wukong unit to a first-tier business unit under the ATH business group, with Chen Yusen at the helm. The three products had each held their own stretch of the battlefield: QoderWork was strong on desktop-level agents, able to invoke local applications for file organization, data processing, and document generation; Wukong was an enterprise-grade work platform deeply embedded in DingTalk; MuleRun was a cross-platform agent execution engine aimed at overseas markets. Alibaba has made AI office one of its strategic directions in AI, positioning it as an enterprise-facing productivity platform.

[STAKES] The merger itself is just tightening formation — what truly determines the outcome is whether it can plug into enterprises’ real data flows and workflows, which is also the officially stated next step. Pulling together the three capabilities — desktop agent, cloud agent, and collaboration agent — is not the hard part; the hard part is making them run inside a company’s actual processes, which requires permissions, approval chains, and historical data, not model scores. DingTalk’s installed base is Alibaba’s biggest bargaining chip — and what sets it apart from pure model vendors. The international version and standalone app are still in the works; whether MuleRun’s existing users in overseas markets can be carried over is the first open question left by this restructuring.

▪ SIGNAL Merging three products into one cuts internal friction, but hasn’t yet secured a place in enterprise workflows.

❯ MiniMax Open-Sources 33B Video Model H3, Runs on a Single RTX 5090

[CONSUMER BAR] MiniMax has put the weights of its 33-billion-parameter omni-modal video model H3 on Hugging Face, with the official description stating the model can generate videos up to 15 seconds, 2K resolution, 24 fps, and natively output 32kHz stereo audio — all runnable on a single RTX 5090. It was released a few hours apart from Alibaba’s Qwen3.8-Max on the same day, a collision Simon Willison specifically called out.

[FINE PRINT] This version went live on July 31, and consumer-grade runnability comes with real trade-offs. Per official and community notes, ComfyUI offers day-one support, but the optimized stack totals about 40GB, fitting into a consumer GPU only through dynamic offloading between RAM and SSD; early benchmarks on a 5090 put generating 5 seconds of 768p-level footage at about 5.5 minutes. MiniMax’s video models were previously API-only, and this open-sourcing still comes with an asterisk: core pieces like context orchestration, 2K regeneration, and sparse attention remain server-side, so what arrives locally is not the full capability. The community license also restricts public use in the United States, the European Union, the United Kingdom, and South Korea, citing copyright litigation — an unusual clause for an open-weight model.

[RECIPIENTS] The most concrete impact of the weight release lands on the cost structure of independent creators and small teams: short videos with sound used to require the API; now a single GPU plus one night of compute turns out a finished clip, paid for in waiting time. What gets squeezed are the per-second-billed video generation services — once the base capability runs locally, all the cloud has left to sell is speed, resolution, and those few un-open-sourced modules. For enterprise users, that regional restriction clause deserves a look from legal before the parameter count does.

▪ SIGNAL The weights are out, but half the capability stays server-side — “open source” is being taken apart piece by piece.

❯ ByteDance Releases Seedance 2.5: 30-Second Single Generations, Mixing Up to 50 Reference Assets

[DURATION] ByteDance has released video generation model Seedance 2.5, with single-session output doubling from 15 seconds to 30 seconds, native 4K support, and up to 30 images, 10 video clips, and 10 audio tracks — 50 reference assets in total per run. The official starting price is $0.097 per second.

[DISTRIBUTION] The model is rolling out across ByteDance’s Jimeng AI and Doubao Pro, with API access to land on Volcano Engine’s Ark platform. Compared with the previous generation, the key optimizations here are shot continuity and scene transitions, plus multi-round extension support — all aimed at the same problem: AI video was previously stuck in chunks of a few seconds that could not be cut into a coherent narrative. The 50-asset mixing capability effectively packs storyboards, character consistency, and music references into a single generation.

[PRICING] That official per-second price, applied to 30-second one-take capability, pushes batch production costs for short video below outsourced editing. The first to redraw their budget sheets will be content agencies and e-commerce asset teams — short clips once outsourced per piece can now be drafted at a cheaper tier, with a chosen take then rendered at a higher-quality tier. ByteDance’s own distribution stack (Jimeng, Doubao, Volcano Ark) means this workflow does not have to leave one ecosystem.

▪ SIGNAL Competition in video generation has shifted from image quality to how long a story can be told in a single continuous run.

❯ Snap Q2 Revenue Up 19%, DAU Hits 493M, Shares Jump More Than 10% After Hours

[DOUBLE BEAT] Snap’s Q2 revenue came in at $1.60 billion, up 18.9% year over year, above the $1.53 billion expected; daily active users reached 493 million, up 5.1%, also beating the 487 million consensus. The company guided Q3 revenue above expectations as well, sending shares up more than 10% after hours.

[GROWTH DRIVERS] Breaking it down: advertising revenue was $1.28 billion, up 9% year over year, while other revenue surged 85% to $316 million — the growth engine is no longer ads themselves but subscriptions and other non-advertising businesses. Adjusted profit came in at $250 million, above the expected $192 million. Monthly active users were 971 million, with global ARPU at $3.25 — and North America ARPU rose 23% to $10.26. North American user numbers have been stagnant for years, but per-user monetization is moving up.

[THE STRUCTURE] The piece of this report most worth isolating is the shift in revenue structure: a social company, with ad growth down to single digits, pulled overall growth to nearly 19% on subscriptions. Advertisers need to adjust their read accordingly — if Snap’s revenue depends less and less on ad inventory, it can afford to take a harder line at the negotiating table. North America DAU has barely moved for years; what held up the business this time was that 23% ARPU gain. Monetization efficiency, not user growth, is now the company’s main storyline.

▪ SIGNAL Ad growth dropped to single digits and Snap still delivered these numbers — that 85% is what held the quarter up.

❯ Ant’s Embodied Intelligence Subsidiary Lingbo Launches First External Funding Round

[EXTERNAL FUNDING] Ant Group’s embodied intelligence subsidiary Lingbo Technology (Robbyant) has launched its first external funding round — per LatePost, it plans to raise 1.5 billion yuan and aims to close a second round before year-end. The company confirmed the funding rumors on August 3, saying it will stay focused on the general-purpose robot brain and increase investment in the embodied-native technology route.

[FINANCING SHIFT] Lingbo was incorporated on December 17, 2024, wholly owned by Ant Intelligent (Hangzhou) Technology, and has already rolled out its first humanoid robot, Robbyant R1, built on a self-developed embodied intelligence foundation model, with pilots in scenarios such as guided tours, pharmacy sorting, and health consultations. One reason public reports cite for the move from wholly-owned group incubation to external funding is resource squeeze inside the group — Ant is spending heavily across AI overall, and its compute budget faces allocation pressure. Independent fundraising effectively moves this business’s funding source off the group’s books.

[VALUATION HURDLE] Independent fundraising also puts one thing squarely on the table — the company is facing external pricing for the first time. Inside the group, embodied intelligence was a strategic bet; outside, a 1.5-billion-yuan raise requires a valuation and a commercialization path the market will accept. What gets re-examined is the delivery cadence of the “general-purpose robot brain” route — the gap between pilot scenarios and scaled orders is what investors will actually ask about. The next checkpoint is whether the year-end round lands as planned.

▪ SIGNAL The moment a strategic business leaves the group’s books, it has to start speaking the language of project returns.

OUTLOOK

[TODAY'S BATCH] Qwen3.8-Max’s pricing, V4-Flash’s 3 cents, and the weights MiniMax released are all the same story: the price of model capability itself is being pinned to the floor. Palantir’s 93% growth offers the counterpoint — money is flowing from the model layer to the layer that can embed models into enterprise workflows, and Alibaba’s Qwen office three-in-one is fighting for position in that layer too.