Back to AI Daily home

❯ Filings Show Amazon Has Fully Paid Its $50 Billion OpenAI Investment, Holding ~5% Stake

[FINAL PAYMENT] Amazon has fully paid the $50 billion it committed to OpenAI, lifting its stake to roughly 5% and making it one of the company’s core pre-IPO shareholders. According to filings with the U.S. SEC, the money moved in three tranches: $15 billion in Q1, $13.7 billion in Q2, and the remaining $21.3 billion settled after June 30 — ahead of the original schedule. People familiar with the matter said OpenAI received the final tranche this week.

[PARTNER TO SHAREHOLDER] The relationship between the two started accelerating this February, when they announced a multi-year strategic partnership covering cloud infrastructure, AI chips, and enterprise services — an initial $15 billion, later topped up by $35 billion. The round pushed OpenAI’s valuation to roughly $852 billion, the largest private AI investment to date. The timing is also unusual: OpenAI confidentially filed for its IPO in June, which means Amazon locked in a shareholder seat at a private-market valuation just before the public-offering pricing window closed.

[COMPUTE BOUND TO EQUITY] For the competitive landscape of cloud providers, this money buys more than shares. What Amazon gets is the rights to a leading model company’s long-term compute orders; what OpenAI gets is the certainty of not having to scramble for its next training budget. What must be revised accordingly is other cloud vendors’ pricing logic — when the customer and the shareholder are the same company, a tender that competes purely on unit price is very hard to break into. Secondary-market investors should also think ahead: the public float at OpenAI’s IPO will be smaller than many expect.

▪ SIGNAL Amazon bought an entry ticket — turning cloud-contract renewal negotiations into a shareholder-meeting issue ahead of time.

❯ OpenAI’s Unreleased Astra Solves Ten Math Problems; Anthropic Researcher Says Fable Reproduced Five

[PRE-RELEASE SCORE] According to OpenAI’s official website, on August 1 the company released 10 results in mathematics and theoretical computer science, all produced by an internal version of Astra — the name of its next-generation model family, not yet released. The covered areas include high-dimensional geometry, group theory, operator algebras, circuit complexity, quantum complexity, and lattice cryptography. OpenAI says the core conclusion of each main problem had seen no progress for at least a decade, and estimates the token cost of running all ten problems at roughly $2,000 at Sol API pricing.

[HALF MATCHED] The awkward echo came 24 hours later. According to his posts on social platforms, Anthropic researcher Levent Alpöge said that Claude Fable, already publicly available, reproduced five of them — completed autonomously with generic prompts, entirely offline, and isolated to keep OpenAI’s published solutions from leaking into the context; of the five, only one followed a substantially identical argument path. Alpöge is no bystander — in July he had just used Fable to find a counterexample to the 1939 Jacobian conjecture, which had already stirred a round of debate in the math community. OpenAI has not yet responded to the reproduction claim, which currently lacks independent third-party verification.

[THINNING MOAT] What this really weighs on is the marketing valuation of an unreleased model. If a model already on the shelf — callable by anyone — can chew through half the list in a day, the premium on that four-character phrase, “internal version,” gets shaved off by a large chunk. Enterprise buyers will next weigh the cost difference on the same problem, not who solved it first — the roughly $2,000 compute bill OpenAI disclosed will shape budget approval more than “unsolved for a decade.” The math community’s stance is calmer: a proof’s credibility awaits peer review, not two companies trading results on social platforms.

signal: The lead window has narrowed from “one model generation” to “one day” — first-mover advantage no longer warrants its own pricing.

❯ DeepSeek V4-Flash official release goes live, output price cut to 2 yuan per million tokens — one-twelfth of Pro

[PRICE SLASH] DeepSeek’s V4-Flash official release entered public beta on July 31, with output priced at 2 yuan per million tokens and input at 1 yuan. Against sibling flagship V4-Pro’s 24 yuan output / 12 yuan input, Flash’s output price comes in at one-twelfth that of Pro — while outscoring the Pro preview on agentic benchmarks. Overseas researcher teortaxes puts it even more bluntly: 14x cheaper than V4-Pro, 50% faster, and better to use.

[NO CAPABILITY CUT] The key shift: this price cut comes with no capability discount. On agentic tests, V4-Flash official release scored 25.2 points, closing in on Claude Opus 4.8’s 25.7 points, while the V4-Pro preview managed only 15.8 — the first time a budget tier has run flush against the top closed-source line. Overseas developer communities report from hands-on testing that V4-Flash’s tool-calling framework is far better than Kimi K3. Analyst kimmonismus posted a Pareto-frontier chart, marking it the current price-performance winner. teortaxes argues that among Chinese labs, only Luna is currently competing in the same market.

[BUDGETS REDRAWN] The first to redraw the math are teams building agent products. Splitting traffic — cheap models on simple tasks, premium models on complex chains — was a compromise forced by pricing. When a budget tier scores 25, the split itself becomes redundant engineering cost. Pressure will hit first on token-billing intermediary providers — once inference costs fall into this range, the spread from reselling call credits is essentially erased.

▪ SIGNAL The moment a cheap model catches up to a premium one, what gets eliminated isn’t the premium model — it’s model routing as a business.

❯ All Top Five in OpenRouter’s Weekly Calls Are Chinese Models; Xiaomi MiMo-V2.5 Takes Top Spot with 10.5 Trillion Tokens

[SWEEP] According to OpenRouter’s official leaderboard page, the top five in the model-aggregation platform’s latest weekly call-volume ranking are all models developed in-house by Chinese companies. Xiaomi MiMo-V2.5 topped the list with 10.5 trillion tokens in a single week, up 12% week over week, and was the only model globally to break the 10-trillion mark that week; it also took first on both the weekly and monthly charts. Over two months, that figure climbed from 1.46 trillion to 10.46 trillion.

[SURGE] The public OpenRouter leaderboard data gets more interesting further down the list. The previous week, the top five still included a non-Chinese model. DeepSeek V4-Flash ranked second with 6.37 trillion tokens, up 18% week over week; Tencent Hunyuan Hy3 was third with 3.94 trillion, a week-over-week gain of more than 999% — it was officially open-sourced only on July 6, meaning it went from zero to the top three in three weeks. Zhipu GLM-5.2 came in fourth, the only name in the top five to decline week over week; DeepSeek V4-Pro was fifth with 3.17 trillion, up 17%, forming a high/low pairing with Flash in second. To be clear on the methodology: OpenRouter counts only the call volume relayed through its platform, reflecting the choices of overseas independent developers — not the total global usage of these models.

[DISTRIBUTION] What this leaderboard really shows is that the distribution channel has changed. Overseas developers no longer source models only from big-tech consoles; aggregation platforms have returned the choice to price and open-source licensing, and that is precisely the opening Chinese models came through. Pricing teams at overseas model vendors will feel the pressure first: when free weights plus low-cost APIs can rack up nearly four trillion calls in three weeks, the price band sustained by closed APIs can no longer hold.

▪ SIGNAL Open source isn’t a margin concession — it’s moving the distribution channel from someone else’s console into your own hands.

❯ Two House Committees Probe DoorDash’s Use of Moonshot AI’s Kimi Model for Coding; Documents Due August 14

[COMMITTEES CALL] The House Select Committee on China and the Committee on Homeland Security have sent a letter to DoorDash CEO Tony Xu over the company’s use of Chinese models, demanding a full inventory of Chinese models in use and security-test records by August 14, and requiring the relevant executives to appear before Congress by August 21. DoorDash is the largest food-delivery platform in the United States.

[SELF-DISCLOSURE] The evidence cited in the letter comes from public statements — co-founder Andy Fang previously said the company had connected open-weight models to its internal AI-assisted code-review system. The approach is tiered: harder tasks go to U.S. frontier models, while lower-level work goes to Moonshot AI’s open-weight model Kimi K2.6. Lawmakers also cite the White House Office of Science and Technology Policy, which says Moonshot AI operated a covert platform, conducted large-scale distillation of U.S. models, and used an unauthorized advanced-computing system — allegations that have not been confirmed by any court or independent third party. The investigation began in April and has already examined several other U.S. companies’ use of Chinese open-weight models.

[COMPLIANCE COSTS] The trouble is that the original selling point of open-weight models was that they can be downloaded, run locally, and send no data back — technically more controllable than calling overseas APIs. But Congress is not asking about data flows; it wants to know whose models actually ran in the codebase. U.S. companies’ technology-selection processes must add a review that never existed before: the model’s nationality. For Chinese model vendors, the overseas installation base generated by open-sourcing is becoming a liability — the broader the adoption, the more instances get named.

▪ SIGNAL Weights may be free to download, but the cost only starts accruing the moment they’re wired into production.

❯ Trump Media Launches Paid Real-Time Data Feed, Two Senators Demand SEC Investigation

[MILLISECOND PUSH] On July 16, Trump Media & Technology Group launched a paid data service, Truth API, that pushes posts from the most-followed accounts on Truth Social to subscribers with millisecond-level latency. The company has discussed pricing as high as $100,000 per month. Senators Elizabeth Warren and Adam Schiff have written to SEC Chairman Paul Atkins demanding an investigation into whether the service violates federal securities law.

[41% STAKE] According to CNBC, the two senators used strong language, calling it “a blatant abuse of the presidency that could give Wall Street an unfair trading advantage.” The reasoning lies in the ownership structure: Trump holds roughly 41% of Trump Media through a revocable trust managed by his family, meaning he personally benefits from the service’s revenue. The more practical issue is the content — Trump’s posts frequently contain policy information on tariffs, personnel, and diplomacy that can move markets within minutes. The senators argue that the biggest beneficiaries are high-frequency trading firms that depend on millisecond execution; the tiny time advantage they get over ordinary investors is precisely the most valuable part of such information. The letter asks the SEC to determine whether the service touches rules on insider trading and market manipulation. Earlier, when Trump Media announced the service on July 16, it did not publicly disclose pricing; the quoted price was subsequently reported by the media.

[TIME LAG PRICED] This episode pushes a previously murky question to the fore: can the time lag on policy information be retailed? The financial data industry has long had a mature low-latency business, but the seller has never been the policymaker himself. Regulators and compliance departments now face a new question — whether subscribing to such data itself constitutes obtaining material non-public information. Buyers are not off the hook either: the advantage gained for $100,000 a month could become evidence in some future enforcement action.

▪ SIGNAL Putting the president’s posting time lag up for sale is tantamount to putting a price on the regulatory bottom line of “information equality.”

❯ Nomura, Citing QuestMobile Data: ByteDance Accounts for 40.1% of User Time in China’s Top Apps, Crossing 40% for the First Time

[CROSSING 40%] ByteDance’s products account for 40.1% of user time among China’s top 50 apps, breaking past the 40% mark for the first time; Tencent’s ecosystem stands at 29.7%, with the gap widening to more than 10 percentage points. The figures follow Nomura Securities’ report “China Internet & New Media: June 2026 App Tracker,” citing QuestMobile methodology; these 50 apps cover roughly 93% of mobile internet usage time.

[ATTENTION FIRST] The report shows that time-spend advantages are being converted directly into AI product installs. A year earlier, the time-spend gap between the two companies was still in single digits. On QuestMobile’s June ranking of domestic AI-native apps by MAU, ByteDance’s Doubao ranked first by a wide margin with 382 million MAU, up 172.1% year over year — the only first-tier product with MAU above the 100-million level; average monthly usage per person was 76.7 sessions and 143.7 minutes, both above the industry average. This chain explains the shape of the leaderboard: it’s not model capability that widened the gap, but existing entry points like Douyin and Toutiao funneling users directly into the new products.

[DISTRIBUTION PREMIUM] By the methodology of this Nomura report, other Chinese model vendors need to reassess customer acquisition costs. In a market where the leading player holds 40% of attention, the window for breaking out on product strength alone is far narrower than before — the same feature, embedded in Douyin versus a standalone app, carries an order-of-magnitude difference in reach cost. Ad budgets will therefore shift from new-user acquisition to retention earlier.

▪ SIGNAL What’s valuable is that 40.1% — model generations turn over fast, but where users spend their time moves slowly.

❯ Karpathy Announces the Pelican Test Is Retired: 1 Million Tokens to Have Opus 5 Render The Lord of the Rings as a Playable Scene

[BENCHMARK END] Andrej Karpathy shared experimental data on social media and said the era of testing large models with “draw a pelican riding a bicycle” is nearly over. He fed Opus 5 the opening passage of The Lord of the Rings with a 1 million token budget (roughly $10), asking it to do a three.js render. The model ran for about two hours and wrote 5,500 lines of code, programmatically turning the passage into a scene you can walk through in the browser. The source code is open-sourced, and you can open it and play with it directly.

[WORLD SHIFT] The judgment he offers is more worth remembering than the experiment itself: models are moving from generating single artifacts to building highly customized entire worlds on demand — but they still lack the native ability to perceive and audit what they have constructed. This is a concrete engineering gap, not a philosophical reflection: whether the picture rendered from 5,500 lines of code is correct, whether anything clips through geometry, whether it is faithful to the original text — the model cannot see any of it; only a human can open a browser and look. Simon Willison — the originator of the “pelican riding a bicycle” test — also joined the discussion. The reason the old test no longer works is simple: frontier models can all draw convincingly now, and that prompt can no longer surface any difference.

[EVAL COSTS] What needs to change in step is the cost scale of evaluation. A test that costs a few hundred tokens per question and produces results in seconds cannot assess a model that can work continuously for two hours; and at $10 per run, no leaderboard can realistically be updated daily. Teams doing model evaluation now face a new constraint: who defines the scoring standard for long tasks. Human acceptance of 5,500 lines of code is unrealistic, and having another model verify it just loops back to the origin: “the model cannot see its own output.”

▪ SIGNAL Models can already build worlds but have not yet grown eyes — the bottleneck in evaluation has shifted from posing the question to verifying the result.

❯ Google Paper: Suppressing a Model’s Claims of Consciousness Also Dampens Its Judgments on Animals and Faith

[SIDE EFFECT] A new Google paper finds that safety-motivated fine-tuning — getting a model to stop claiming it is conscious — collaterally dampens its mental-state judgments about other entities: not just itself, but also non-human animals and natural objects, while significantly reducing its expression on religious-faith questions. The paper’s title, literally translated, is Inducing language models to assert their own consciousness restores human beliefs and values.

[STEERING] The paper, posted to arXiv as a preprint (arXiv ID 2607.28607) under the title Inducing language models to assert their own consciousness restores human beliefs and values, has not yet been peer-reviewed. The researchers take a mechanistic approach: they ablate the learned safety-refusal direction and apply targeted steering to a “consciousness vector” in activation space. The result: the suppression is reversed — once these internal representations are restored, the model’s answers on standard sociology questionnaires about religiosity, moral values, hope, and subjective well-being move noticeably closer to human samples. The study also confirms that these changes do not impair theory-of-mind ability, indicating that core social reasoning and self-awareness representations are mechanistically independent of each other.

[EVALUATION] This adds a new check for alignment teams: how many dimensions a single safety fine-tuning actually changes. Apply the “don’t say you’re conscious” rule alone, and the model drifts on entirely unrelated value judgments — yet existing evaluations mostly only test whether the single banned behavior is suppressed; the drift isn’t on their benchmark at all. That directly affects the acceptance cost of safety fine-tuning: the metric to watch becomes, before and after the same fine-tuning, the distribution shift of the model on value and common-sense questionnaires — not just the refusal rate.

▪ SIGNAL Safety fine-tuning isn’t deleting a single sentence; it’s twisting an entire direction vector — and no one has ever catalogued what gets carried along with it.

❯ Bloomberg’s Gurman: MacBook Air Supply Tight; Apple Looks to Make Glasses and Headsets Health Devices

[LEAD TIME] Bloomberg journalist Mark Gurman says ordering a MacBook Air from Apple’s website now pushes delivery to the end of this month, with retail channels saying supplies are “tighter than at any time in memory.” He cites two reasons: a memory shortage driven by AI demand, and Apple giving capacity priority to the entry-level M5 MacBook Pro, which is due for a refresh this fall.

[HEALTH PLATFORM] Another item in the same newsletter takes the longer view: Apple is hiring to build health and fitness capabilities into the Vision product line, aiming to make future smart glasses and headsets the next health platform after Apple Watch. That idea never landed on Vision Pro — the company built a version of Fitness+ that could run on the headset, letting users follow workouts and close their rings, but shelved it because the headset is too heavy to wear while exercising. Gurman says these features won’t appear on next year’s first-generation glasses.

[MARGIN SQUEEZE] The shortage is worth watching more closely than the glasses. Memory price increases have already spread from servers into entry-level consumer electronics, and hardware makers like Apple are having to reorder capacity allocation across product lines — protecting high-margin models first is the simplest move in a price-hike cycle, at the cost of longer waits for entry-level buyers and drained channel inventory. What to watch: whether next spring’s entry-level refresh raises pricing in step with memory costs.

▪ SIGNAL AI’s compute bill is reaching consumers by a different route: not a price increase, but a wait.

OUTLOOK

[TODAY'S BATCH] Three items are worth reading together: Amazon paid 50 billion in full, OpenRouter’s top five are all Chinese models, and DeepSeek slashed output prices to 2 yuan. At the top, equity locks in compute; at the bottom, prices erase profit — the middle tier, which resells API calls and skims interface spreads, is squeezed from both directions. The one-day gap between Astra and Fable says it more bluntly: a capability lead now has a shelf life shorter than one procurement negotiation.