- Published on
热门科技推文 - 2026-07-08
- Authors

- Name
- geeknotes
科技每日简报 (2026年7月8日)
Today's top tech conversations are led by @zephyr_z9, whose post about 'IT'S OVER\n\nQT @jukan05: CHIN...' garnered the highest engagement. Key themes trending across the top stories include models, nvidia, deepseek, china, model. The community is actively discussing recent developments in AI, engineering practices, and startup strategies.
1. zephyr_z9 (Group Score: 1229.5 | Individual: 62.7)
Cluster: 39 tweets | Engagement: 1403 (Avg: 181) | Type: Tech
IT'S OVER\n\nQT @jukan05: CHINA CONSIDERS RESTRICTING OVERSEAS ACCESS TO CUTTING-EDGE AI MODELS
China’s Ministry of Commerce has led meetings over the past month with major AI companies, including Alibaba, ByteDance, and https://t.co/YDe0KRldDB, to discuss measures that would restrict overseas access to cutting-edge AI models, including models that have not yet been released.
The discussions reportedly include not only closed-source models but also open-weight models. However, the scope of application is still under debate, and the rules may ultimately apply only to future frontier models.
Officials have also discussed designating the leakage or theft of proprietary AI technologies as a national security crime, with stronger penalties, as well as restricting the types of foreign capital that can invest in Chinese AI startups.
The backdrop is the U.S. move to strengthen export controls on AI models, along with national security concerns over cutting-edge models that could possess advanced cyberattack capabilities.
Chinese authorities are reportedly concerned that advanced U.S. cybersecurity AI models could be used to exploit vulnerabilities in Chinese software.
Since the beginning of this year, China has continued to tighten measures to prevent AI technology from being transferred overseas. Authorities have investigated whether Chinese AI startups that relocated abroad violated export control laws, while also strengthening oversight of overseas transactions involving Chinese investors, technology, data, and national security concerns.
Future regulations could take the form of a tiered framework based on technological capability. Basic open-source AI models may be managed through a filing system, high-performance models may be subject to security reviews, and the most sensitive frontier models may be banned from public release or restricted to use within China.
See 38 related tweets
- @antirez: If that happens Europe is officially fucked.\n\nQT @jukan05: CHINA CONSIDERS RESTRICTING OVERSEAS AC...
- @mitsuhiko: Europeans please wake up.\n\nQT @jukan05: CHINA CONSIDERS RESTRICTING OVERSEAS ACCESS TO CUTTING-EDG...
- @cryptopunk7213: 2 major announcements from china this morning
they’re considering banning the U.S. from accessing...
- @teortaxesTex: Oh well. America Winning once again. Wypipo innovate, the Chinese copy. I feared that it'll take exa...
- @mark_k: It's happening: Open source AI models from China may become restricted soon! That could be the end o...
2. BullTheoryio (Group Score: 399.9 | Individual: 33.9)
Cluster: 16 tweets | Engagement: 431 (Avg: 1025) | Type: Tech
BREAKING: The Chinese startup that crashed Nvidia's stock by $600 billion is now building its own chip to replace Nvidia entirely.
In January 2025 DeepSeek released R1 and showed the world you could build a world class AI model at a fraction of the cost everyone assumed.
Nvidia lost $600 billion in market cap in a single day.
Today Reuters reported that DeepSeek is secretly developing its own AI chip.
No public job postings or announcements. Quietly hiring chip design engineers for months and reaching out to chip design, foundry, and memory companies in private. The project started about a year ago.
The chip is specifically built for inference, not training. Inference is what happens every single time someone uses an AI model, every response, every search, every generated image.
It is the fastest growing and most commercially valuable segment of AI computing right now.
This is a direct threat to both Nvidia and Huawei simultaneously.
Nvidia is already effectively banned from China's data center market. China was once 20 to 25 percent of Nvidia's revenue. Export controls cut that to 9 percent. Nvidia took a 8 billion hit in Q2.
Huawei stepped in to fill that gap and now controls about half of China's $50 billion annual AI chip market. DeepSeek itself shifted from Nvidia to Huawei, releasing its V4 model optimized for Huawei's Ascend chips in April 2026.
That single move triggered a buying rush across ByteDance, Tencent, and Alibaba for Huawei's Ascend 950 processors.
Now DeepSeek is moving to cut out Huawei too.
But the challenges are real. Designing a competitive AI chip takes years and significant capital. US export controls bar Chinese designers from accessing the most advanced overseas foundries. Separate US curbs have also cut China's access to high bandwidth memory, which is a critical component for inference chips specifically.
Which is probably why DeepSeek just reversed its yearslong policy of rejecting all outside investment. The company is raising 52 to $59 billion.
It spent years building the world's most efficient AI models with no outside money. Now it needs capital. The reason is hardware.
Meanwhile the global picture makes this even more significant.
OpenAI launched its own custom inference chip last month. Anthropic is considering building its own. Google, Amazon, and Meta already have their own silicon.
Every major AI company in the world is simultaneously working to reduce dependence on Nvidia.
Nvidia trades at a $4.75 trillion valuation built on the assumption that it remains irreplaceable.
DeepSeek has now challenged that assumption twice.
First by showing you need far fewer chips than anyone thought. Now by building its own.
See 15 related tweets
- @Mayhem4Markets: The company that proved you don't need Nvidia's best chips is now building its own.
DeepSeek is dev...
- @cryptopunk7213: exclusive from The Information - chinese ai lab Zhipu is also considering building their own ai chip...
- @amitisinvesting: $NVDA
Nvidia is down in the premarket because DeepSeek is reportedly looking to reduce their relian...
- @KyleReidhead: $NVDA is now green on the day
after falling a few % on the news of DeepSeek "building their own chi...
- @CNBC: Chinese-built AI models are gaining traction among U.S. companies as they narrow the performance gap...
3. PolymarketMoney (Group Score: 247.0 | Individual: 35.7)
Cluster: 11 tweets | Engagement: 1656 (Avg: 251) | Type: Tech
JUST IN: Samsung’s quarterly profit jumps 1,900% to $58,000,000,000.00 on surging AI memory demand.
See 10 related tweets
- @tengyanAI: lol the same price spike minting samsung's chip division is squeezing its own phone business.
mobi...
- @rohanpaul_ai: Samsung just reminded everyone that AI runs on memory too.
They just flagged a 19-fold profit jump,...
- @MTSlive: SITUATION DETECTED: Samsung’s quarterly profit jumped 19x to $58B on soaring demand for AI memory ch...
- @BullTheoryio: Memory stocks are crashing in the as Samsung's revenue miss sparks a sector wide selloff.
$MU down ...
- @WOLF_Financial: Samsung brought in more operating profit than any other tech company has in history
Samsung brought...
4. teortaxesTex (Group Score: 200.1 | Individual: 38.8)
Cluster: 9 tweets | Engagement: 161 (Avg: 57) | Type: Tech
OK this is a serious flex the image doesn't make a whole lot of sense (where is she scanning from? she doesn't point at the poster and this isn't a real camera interface) but QR replication indicates a powerful underlying model https://t.co/Gvbdq4K4jO\n\nQT @alexandr_wang: 1/ releasing muse image today — the first image generation model from MSL. it's agentic: pairs with muse spark to reason through your prompt, search the web, and plan before it generates. people get what they meant on the first try. live now in the Meta AI app. https://t.co/7KUk0DS1as
See 8 related tweets
- @AngryTomtweets: Meta Superintelligence Lab just dropped a new image model called Muse Image.
A big step forward wit...
- @PolymarketMoney: BREAKING: Meta plans to launch Muse and replace third-party image models with its own AI image techn...
- @testingcatalog: META 🔥: Muse Image, the first image-gen model from MSL, is now available on Meta AI.
It uses adv...
- @mark_k: Meta has released a new image model, which can do reasoning, like GPT-Image-2 or Nano Banana Pro. I,...
- @AngryTomtweets: RT @AngryTomtweets: Meta Superintelligence Lab just dropped a new image model called Muse Image.
A ...
5. Shashikant86 (Group Score: 191.5 | Individual: 28.0)
Cluster: 9 tweets | Engagement: 0 (Avg: 5) | Type: Tech
This incredible blog post on the harness engineering for self improving AI boosted my motivation to work more on SuperQode .. Ideal is validated, most of the concepts mentioned in the blog post already captured by the SuperQode, some o them (Memory/Context Specs) are still not there but direction is clear. Stay Tuned and watch SuperQode become the best harness engineering framework for building/evaluating and optimising your coding agent harnesses.. It's still in beta but benchmarks with @harborframework with open models are o the way! Surely it will beat most of the frontiers .. Stay Tuned and Watch the space 👉 https://t.co/0DhwOQgVDE\n\nQT @lilianweng: new post on harness engineering for AI self-improvement: https://t.co/ZYvGfVs61k
It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple.
Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
See 8 related tweets
- @dejavucoder: new lilian weng post\n\nQT @lilianweng: new post on harness engineering for AI self-improvement: htt...
- @cong_ml: RT @lilianweng: new post on harness engineering for AI self-improvement: https://t.co/ZYvGfVs61k
It...
- @EMostaque: This blog is a great review of the current environment on self-improvement
We leveraged a number of...
- @cgtwts: Lilian Weng is back with another banger:
“The layer between the raw model and the real-world contex...
- @jeffclune: Nice to see Automated Design of Agentic Systems (ADAS), the Darwin Gödel Machine, DGM-HyperAgents, a...
6. kimmonismus (Group Score: 172.5 | Individual: 32.4)
Cluster: 6 tweets | Engagement: 895 (Avg: 646) | Type: Tech
tl;dr: Fable 5 basically mocks every other model and tops every benchmark.
Curious to see whether the six new benchmarks will crown a new leader after GPT-5.6. https://t.co/fsMk3NKif8\n\nQT @ArtificialAnlys: Introducing six new Artificial Analysis Capability Indices for comparing model capabilities across key industry domains
The new industry indices cover Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics. We aim to capture the common capabilities required across knowledge work domains and evaluate how well current models meet those needs.
Each index is grounded in common tasks from O*NET occupational classifications. Tasks range from financial modeling, to legal research and contract review, to clinical decision support and patient documentation. We derive capabilities from each task, select the benchmarks that best represent the work, and weight by how often each capability appears across the domain. This means rethinking the Artificial Analysis benchmark suite for each domain and slicing evaluations to relevant domain tasks. Every component benchmark is run independently by Artificial Analysis.
The industry indices join the existing skill-based Agentic and Coding indices, which measure capabilities that cut across every domain.
Key Results ➤ Leading models: Claude Fable 5 (with Opus 4.8 fallback) leads all eight indices, with Claude Opus 4.8 (max) in second on six of eight Capability Indices and GPT-5.5 (xhigh) on two. Below the top two, rankings reshuffle substantially by domain between Gemini 3.5 Flash, Gemini 3.1 Pro Preview, GPT-5.5 (xhigh), Claude Sonnet 5 (max), and GLM-5.2 (max). ➤ Open weights leading models: Among open weights models, GLM-5.2 (max) leads on five of the six industry indices, ranking as high as fifth overall on the Artificial Analysis Engineering Index (53), within 2 points of Claude Sonnet 5 (max, 55) and GPT-5.5 (xhigh, 55). DeepSeek V4 Pro (max, 38) takes the open weights lead on Artificial Analysis Strategy & Ops Index. ➤ Cost efficiency: DeepSeek V4 Flash (max) completes tasks for <0.26 to 0.58. Frontier capability comes at a steep premium: on the Artificial Analysis Strategy & Ops Index, Claude Fable 5 (with Opus 4.8 fallback, 3.48) scores 12 points above DeepSeek V4 Pro (max, $0.03) at over 100x the Cost per Task. ➤ Time per Task: Time per Task spreads roughly 15x within each index, from 1.1 minutes for Nova 2.0 Pro Preview (medium) to 16.7 minutes for Claude Sonnet 5 (max). Speed shows a similar frontier to cost: on the Artificial Analysis Legal Index, Gemini 3.1 Pro Preview (0.8 minutes) completes tasks ~7x faster than Claude Fable 5 (with Opus 4.8 fallback, 5.4 minutes), while scoring within 11 points.
See 5 related tweets
- @teortaxesTex: Interesting how uneven the perf of open models, and particularly V4, is. On one task it's the best...
- @scaling01: That's an interesting way to say Anthropic models are better on GDPval2, HLE and AA-Omniscience\n\nQ...
- @harvey: RT @ArtificialAnlys: After our announcement last month, Artificial Analysis is now launching Harvey ...
- @TeksEdge: Wow Domain specific benchmarks. See how well these models do at your job. Nice! https://t.co/HZZ7525...
- @_simonsmith: I was excited about Artificial Analysis introducing domain-specific benchmark indexes, but for one r...
7. deepfates (Group Score: 168.6 | Individual: 32.6)
Cluster: 9 tweets | Engagement: 285 (Avg: 43) | Type: Tech
https://t.co/frYaNmg4fX\n\nQT @claudeai: We're extending access to Claude Fable 5 on all paid plans through July 12.
See 8 related tweets
- @astropol0: So Anthropic has extended access to Claude Fable 5 on all paid plans (like Claude Pro, Max, Team, et...
- @levelsio: ONE MORE WEEK TO ESCAPE THE UNDERCLASS\n\nQT @claudeai: We're extending access to Claude Fable 5 on ...
- @ihtesham2005: We got Fable 5 for more days.
Now go and build:
- the app you wanted to build
- the business you...
- @tunguz: Yes!!! I will not have to be stuck to the computer for all of the day today!!! https://t.co/g6xCJ1XT...
- @felixrieseberg: If you build something cool, send it to me?\n\nQT @claudeai: We're extending access to Claude Fable ...
8. XFreeze (Group Score: 158.0 | Individual: 36.2)
Cluster: 6 tweets | Engagement: 530 (Avg: 476) | Type: Tech
Wall Street is finally waking up to SpaceX
As SpaceX enters the Nasdaq-100 less than a month after its IPO, the first wave of brokerage coverage is coming in overwhelmingly bullish
Morgan Stanley, Goldman Sachs, J.P. Morgan, Citi, and Wells Fargo all initiated coverage with top ratings
Morgan Stanley called SpaceX “AI’s final frontier”
Goldman says SpaceX is positioned across three massive markets: Space, Connectivity, and AI
Each with multi-trillion-dollar potential
J.P. Morgan estimates that Nasdaq-100 inclusion alone could bring in around $4.3 billion in passive inflows
Wall Street is no longer looking at SpaceX as just a rocket company
They’re starting to see it as the critical infrastructure layer for the next era: • Starship → the launch flywheel • Starlink → the global connectivity network • SpaceXAI / orbital compute → the AI infrastructure play
SpaceX is becoming the bridge between space, internet, and intelligence
See 5 related tweets
- @SawyerMerritt: Raymond James has initiated coverage of @SpaceX stock with a STRONG BUY rating and a $800 price targ...
- @cryptorover: 🚨THE SPACEX EUPHORIA IS GETTING OUT OF CONTROL
SpaceX is joining the Nasdaq-100 today, less than a ...
- @cb_doge: BREAKING: Citi says SpaceX has a path to $900+ per share.
• Buy rating with a $200 price target to...
- @coinbureau: 🚨WALL STREET BULLISH ON SPACEX, CALLS $300 TARGET
Morgan Stanley called for 300 ah...
- @MarioNawfal: SpaceX joined the Nasdaq 100 just after its massive IPO, making it one of the fastest additions ever...
9. zerohedge (Group Score: 146.9 | Individual: 26.7)
Cluster: 9 tweets | Engagement: 537 (Avg: 833) | Type: Tech
First DeepSeek now this:
Zhipu AI, Chinese AI lab behind the advanced GLM series of open-source AI models, weighing designing its own AI chip as surging demand and US export controls make computing resources a growing constraint: The Information
See 8 related tweets
- @wallstengine: DeepSeek is reportedly developing its own AI chip, per Reuters.
The chip is being designed for infe...
- @wallstengine: The Information reports China’s Zhipu AI is weighing a custom AI chip as demand for its GLM models s...
- @Reuters: EXCLUSIVE: China's DeepSeek developing its own AI chip, sources say https://t.co/oCDk83WkzF https://...
- @financialjuice: China's DeepSeek is developing its own AI chip, hiring chip-design engineers for the project - Sourc...
- @WhaleInsider: JUST IN: 🇨🇳 China's DeepSeek is developing its own AI chip, hiring chip-design engineers for the pro...
10. swyx (Group Score: 144.4 | Individual: 32.8)
Cluster: 7 tweets | Engagement: 185 (Avg: 31) | Type: Tech
imo this is the most impt part of anthropic's J-space paper today. it's a two-parter:
- ant proved that they can do "brain surgery" interventions into reasoning to change topics midstream*
- THE MODEL IS ABLE TO DETECT WHAT INTERVENTION WAS DONE - close cousin to eval awareness**
*control > correlation - this convincingly demonstrates understanding
**this was prompted awareness... surely @mlpowered's team also tried to eval unprompted awareness but i didn't see evidence of that\n\nQT @AnthropicAI: New Anthropic research: A global workspace in language models.
Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with.
We found a strikingly similar divide inside Claude. https://t.co/aLUPBifxth
See 6 related tweets
- @tobi: astonishing\n\nQT @AnthropicAI: New Anthropic research: A global workspace in language models.
Of e...
- @yacinelearning: alright folks we’re talking consciousness big models internals activation and global workspace theor...
- @deliprao: I don’t know enough about scientific consciousness theories to speak intelligently about this (@Plin...
- @codeSTACKr: RT @AnthropicAI: New Anthropic research: A global workspace in language models.
Of everything happe...
- @Nick_Davidov: I knew it in my J space\n\nQT @AnthropicAI: New Anthropic research: A global workspace in language m...
11. KyleReidhead (Group Score: 134.7 | Individual: 33.3)
Cluster: 5 tweets | Engagement: 19 (Avg: 19) | Type: Tech
the "cracks in the AI Capex market" theory is stupid
Capex is going MUCH higher in the coming years (this is bullish MU and neoclouds)
@SemiAnalysis_ expects annual capex to reach 11T cumulative by 29'
A big chunk of this will come from debt, around $7T
But here is the thing, the debt part of this AI build out is still in its early stages
Most of this build out is funded by free cash flows of the largest companies in the world (Google, Amazon, Meta, etc.).
These companies still generate enough revenue to pay for their 40B/quarter capex
In the coming months, that will change and their Capex will exceed their revenues. But that's ok, because these companies have massive balance sheets and the ability to take out cheap debt
NVIDIA has already started to fund GPUs for neoclouds to help continue to accelerate the build out, they also raised $20B in bonds
Amazon announced today they will raise $25B in bonds
The hyperscalers have many ways to access debt and capital, and that part of this cycle has only just started
But here's the thing, these companies revenues are also growing at 30-60% YoY and thats likely to accelerate due to AI
So these companies literally have years of runway to continue to grow their Capex if they want to
The real question is do the revenues of this Capex show up? I believe they already are and will continue to do so, since we are already seeing it in the numbers across many companies
- companies ad revenues are improving
- companies are letting go of staff because AI is making their business more profitable
- companies can discover drugs faster using AI
and as we've seen with SpaceX and Meta, if they have some extra compute in the short-term, they can easily rent it out to someone else who wants/needs it
AI Capex isn't ending anytime soon and the market continues to get this wrong!\n\nQT @KyleReidhead: Anyone who thinks DeepSeek announcing they are building their own chips is bearish $NVDA is dead wrong
Buying Nvidia down 20% is an absolute gift (save this)
DeepSeek is designing its own inference chip to cut reliance on Nvidia and Huawei. It is early stage, built for inference not training, and aimed at China's $50 billion domestic AI chip market.
First of all, Nvidia isn't in that market to lose. Jensen Huang has said Nvidia's share of China's advanced AI accelerator segment is already at zero, down from 95% before export controls, and Nvidia's own guidance assumes no China data center compute revenue at all, this quarter or next.
Secondly, they are years away from these chips actually existing and being manufactured at scale... if they can even do it! Nobody innovates as well and as fast as NVIDIA does in this department
So this news has almost no impact on Nvidia demand today or anytime soon.
But the reason this pullback is an incredible buying opportunity is becuase of what Nvidia is doing with the demand it can still sell into.
Data center revenue hit a record $75.2 billion last quarter, up 92% year over year, and hyperscalers now make up only about half of that, down from being almost the whole business a few years ago
The other half is AI clouds, industrial, enterprise and sovereign buyers, and Nvidia just made it dramatically easier for that half to keep growing.
Nvidia started backstopping GPU rentals for Neoclouds, the smaller cloud providers that want Nvidia chips but can't get bank financing on their own.
Nvidia guarantees them a minimum revenue floor using its own AA credit rating, then takes a cut of whatever they earn above it (aka another killer revenue stream!)
Two deals are already public. SharonAI in Australia locked in a 25 to $30 billion in customer revenue over six years off the back of its backstop.
This is Nvidia manufacturing its own buyer base. Every Neocloud it backstops becomes a new outlet for chips that doesn't depend on a handful of hyperscalers deciding to keep spending.
The honest risk is Nvidia is now on the hook if a Neocloud can't fill its capacity. But the floor price is set deliberately below market rental rates specifically so it is never expected to be triggered, and Nvidia only shares in the upside above that floor, not the downside of building the cluster itself.
NVIDIAs business is better than ever and they continue to add new revenues streams that also accelerate the market for their other revenue streams
All of this is happening while NVIDIA continues to grow revenues 90% YoY, yet their PE is below 30 and forward PE below 20. This is the cheapest NVIDIA has been in 10 years!
The risk/reward on Nvidia couldn't be better right now, which is why I continue adding it to my portfolio
Follow me @kylereidhead for more insights on AI, robotics and other big markets.
You can also track my real-time portfolio for just $1 at Milk Road PRO (see link in bio).
See 4 related tweets
- @MilkRoadAI: RT @MilkRoadAI: Memory stocks are becoming the Nvidia of the AI era and the market has no idea how t...
- @MelvinInvests: RT @MelvinInvests: Anyone who is not buying Nebius or CoreWeave right now is an idiot (Save this). ...
- @KyleReidhead: RT @KyleReidhead: the "cracks in the AI Capex market" theory is stupid
Capex is going MUCH higher i...
- @KyleReidhead: RT @KyleReidhead: the entire AI capex trade depends on future revenues of the frontier models (Anthr...
12. olivercameron (Group Score: 134.2 | Individual: 37.3)
Cluster: 4 tweets | Engagement: 282 (Avg: 99) | Type: Tech
RT @willdepue: A Stargate for Data
Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar compute project, we need to think about the equivalent civilizational-scale effort for the other core ingredient: data.
At the foundation of the scaling revolution is a simple empirical law: deep neural networks improve smoothly, near magically, as you scale two things in proportion — (1) the size of the model and (2) the amount of data you train on. And despite the scaling laws being brutally diminishing, we’ve successfully bitten the bullet of logarithmic scaling with exponentially larger clusters and datasets, and received incredible new capabilities in return.
But this exponential scaling is bound to hit some limits. Oddly enough, compute has compounded fairly smoothly without limit, with trillions flowing into hypercluster buildout. Instead, we’re starting to hit the limits of an exponential demand for data. Gone are the days of being purely in the compute-limited regime, where we had effectively infinite internet data but never enough GPUs, we’re now entering a data-limited regime.
Luckily, this limitation is coinciding with staggering improvements in AI capabilities. Incredibly, we seem to have a real line of sight towards automating a majority of knowledge work with the methods we have today. RL + pretraining, and the data for each, will be generally sufficient to achieve most economically valuable tasks, given some minimal algorithmic progress and continued compute scaling.
In a data-limited world, economic progress & scientific acceleration will be directly bottlenecked by our coverage in each domain. We need to see data collection as imperative, deserving the same civilizational ambition we’ve given compute.
The internet as a one-time subsidy
It’s underrated how much all progress in AI owes everything to the blessing of the internet, this one-time civilizational subsidy to deep learning, decades of unintentional accumulation of a perfect dataset: every book, blog post, image, video, paper, discussion, etc. all digitized and freely available. Without the internet, we’d likely see comparably minimal progress in AI today, and in fact, if you notice where systems currently underperform, it’s almost always a domain where web coverage is limited and data is private, expensive, non-digitized, or non-existent.
But we’re running out of it. There are only about 300 trillion tokens of useful public human text, and the internet doesn’t produce nearly enough new high-quality data to match what scaling demands — we’re soon to hit the limits of public data for pretraining. And though the advent of RL bought us reprieve — chain-of-thought RL needed a new form of untapped data, gradable math & coding tasks, also available online — we’re quickly running dry of hard tasks for RL as well.
Why do we need so much data anyways? Humans learn comparably in far less time, needing just one textbook where language models might need the equivalent of hundreds to learn a new topic. It’s possible we discover methods that are massively more data efficient — synthetic data, data efficient architectures, other exotic algorithms — but fundamental progress is slow and highly unpredictable, and the recipe we have just works today.
And, while I’m wary of getting too deep here, even arbitrary data efficiency can’t replace data that just doesn’t exist in the first place. There’s a massive amount of missing information on the web: the dark matter of the internet — tacit knowledge, undocumented processes, etc. — most of which was never published and lives only inside organizations, the physical world, or just in people’s heads. I’ll leave it here and say, for reasons far longer than I can fit in this post [1], it’s best to operate on the assumption that our insatiable desire for data will continue as it has for the last decade.
There will be >$100B/year in data spend by 2030
We’re not screwed yet, of course. Only a fraction of useful data in the world is on the public internet, the rest is stored inside private datasets, corporations, personal archives, universities, governments, and otherwise. Labs can and will continue to license these private datasets, or create them from scratch, like Anthropic’s book scanning project. And we’ll increasingly task human experts to manufacture new high-quality data, with a large fraction of hard RL training tasks already being sourced this way.
But collecting this data, unlike before, will be expensive. As the free internet dries up and demand for data rises, we should see labs investing equally in data as compute, likely spending a significant fraction of their compute budgets on data. As we see trillions spent on compute, we should also expect hundreds of billions spent on data (human data & collection budgets), given their equivalent importance. And, notably, data spend is already tracking this way: total data spend across vendors, not counting internal lab efforts, is already roughly $7 billion per year. It’s quite reasonable we’ll see >10x by 2030.
Data is the moat
Data becoming increasingly private will also majorly shift the competitive landscape. While compute is a commodity — everyone buys the same chips and builds the same clusters — data really isn’t. The big reason why frontier models have felt eerily similar to one another, until now, is they were trained on substantially the same internet (pretraining data variability across labs seems pretty low). As labs diverge onto more exclusive, manually collected corpora, I think models will begin to increasingly diverge.
OpenAI pulling ahead in mathematics and Anthropic in cybersecurity isn’t an accident. I really think laser-focused collection of high-quality midtraining tokens, custom RL tasks, environments, with dedicated research effort, has driven much of the visible progress in the last year. James Betker has an excellent blog about “the ‘it’ in a model is the dataset”: model architecture and compute buy you efficiency and order-of-magnitude performance, but ultimately, models, of any architecture, are such incredible approximators of their dataset that the core meat of a model boils down to just that, nothing else. Data is a major moat.
AGI long, ASI short
As I’ve tweeted before, I’m confident that, despite the narrative, the data labeling industry will continue to fuel great businesses and be an excellent AGI long, ASI short. The argument is just: By the time the AGI labs no longer need data, it’s probably over for everything else too [2]. In this frame, the last companies left should be the data companies, as the last speck of economically relevant data is sucked in. And these companies are already among some of the fastest-growing companies in history: Mercor, founded three years ago, is rumored to be doing $2 billion in revenue with something like a few million expert labelers under contract.
While these businesses are very non-stationary, what type of data is needed shifts constantly, I don’t think that diminishes their value. The long-tail of the economy is long, and the value isn’t diminishing as you extend farther into more obscure information: as models get more capable, the value of the marginal dataset goes up, not down. Automating a full job means covering its full distribution of tasks, tools, edge-cases, and long-horizon loops. There’s some O-ring logic to it: a dataset that buys a 1% bump can justify a previously unjustifiable collection cost when it’s the difference between a system that does 99% of a job and one that does all of it [3].
The competitive dynamics of the data industry are still evolving but as demand for data is increasingly niche, ultra high-quality, expert-generated, I think we’ll see real consolidation. Again, contra-narrative, we’ll probably see true competitive differentiation built on brand, quality control of data (which, from personal experience, can vary massively), as well as in network effects from the talent networks themselves over time. We’ve already seen rapidly shifting data type demand work in favor of incumbents, benefiting those with early knowledge of where the market is headed.
The binding constraint
It’s truly remarkable that we seem to have the recipe — pretraining + RL — to absorb most economically valuable work, despite being far from a lot of what we expected from “AGI”. The same way chess engines revealed we never needed general intelligence to solve chess, as we originally thought, we’ll soon realize that software, mathematics, and the vast majority of the economy (including physical, just running ~3 years behind!) are the same. If recursive self-improvement or some other algorithmic breakthrough arrives, that’s wonderful, but we really don’t have to wait for it. The binding constraint between here and an automated economy isn’t that, it’s data coverage: every app, workflow, edge case, process, etc. sitting in private stores or someone’s head.
Ultimately, while we make tremendous strides in more efficient model architectures, and clusters like Stargate equip us with zettaflop-scale compute, we really aren’t making rapid progress collecting the data we lack.
We’ll soon live in a world where we have the methods & compute to accelerate scientific progress or economic growth, but not the data. And we’re already there today: frontier models would surely be as good at accounting/many medical tasks/legal advice as they are at software engineering if we only had the same pretraining & RL coverage as we did for code.
I really want to drill this in: The speed at which we automate the economy is going to be directly rate-limited by our ability to collect data about it.
Worth noting that under this assumption, with data as defensible and directly proportional to economic & scientific progress, data should also be considered a national strategic asset like compute. Imagine what we’d do in a world where we had a Manhattan Project-effort for AI and needed to mobilize data collection as a limiting factor. We should be concerned about China, with greater state capacity and authoritarian economic control, being capable of mobilizing data collection at national scale, potentially compounding their economy and scientific output faster than us down the line.
A Stargate for data
I’m leaving my complete ideas for a future post, as this one is already far too long, so I’d really like to pose the question here. Stargate exists because we organized trillions of dollars, international strategy, gigawatts around compute as a fundamental ingredient. What would equivalent ambition look like for data?
Obviously, scaling data collection, a heterogeneous mass of information across the economy, isn’t going to be as clear as scaling compute, as a homogenous infrastructural effort. A core division will be first, coverage — all uncaptured knowledge sitting across the economy/science/physical world and all that simply isn’t recorded — and, secondly, sheer volume in the domains we already train on: more hard math tasks, more high-quality web text, way more coding data, more legal drafts, etc.
I have a post coming soon which breaks down my proposals. There’s a lot of room for creativity. Quickly, we’ll probably want to start with a deep census of what we have and what we’re missing, predict what the 2030 model will still be bad at and work backward to what we should be collecting today. You can probably license a large amount, leveraging high lab valuations to buy datasets or companies altogether. There’s an adversarial nature to a lot of this collection with firms, so there’s lots of engineering to do this correctly. We should go convince important companies to turn off deletion policies, even if we’re not buying from them yet. Data flywheels in consumer products will be massive. Confidential training, government legislation for grant-funded research, running companies at a loss for their data, etc.
We’re headed towards hundreds of billions in expenditure, national prioritization, and major data limitation on the horizon. We have a great opportunity to think creatively about what a megaproject for data would look like: How do we, deliberately this time, construct the next internet’s worth of data?
Footnotes:
[1]: I’ll probably soon publish my much longer post explaining my position on data efficiency and why the value of this data is still pretty high in most worlds regardless of new algorithms.
[2]: The “AGI freeroll” bet: heads you win, tails ASI flips the world upside down anyways.
[3]: We already see a glint of validation of this point, given the data market is strongly tilting towards ultra-high-quality agentic data, rather than unskilled labeling — niche expert workflows, live environments, and evaluations requiring increasingly obscure talent & knowledge — yet shows increasing, not decreasing, revenues.
See 3 related tweets
- @prz_chojecki: Good that we have https://t.co/Gz5MYmBKL2 🫡\n\nQT @willdepue: A Stargate for Data
Labs are on a tra...
- @levie: A small percentage of useful data is on the open web available to all models for training or for age...
- @andrewho03: This is a good post. I suspect that anyone at a frontier lab who successfully internalizes and scale...
13. minchoi (Group Score: 127.5 | Individual: 25.1)
Cluster: 7 tweets | Engagement: 192 (Avg: 105) | Type: Tech
Claude Cowork just escaped the laptop.
Start work at your desk. Check it on your phone. Claude keeps going in the background.
This is agentic work. https://t.co/XkGBgBJKEX\n\nQT @claudeai: Claude Cowork is coming to mobile and web.
Hand Claude a task at your desk and pick up the finished work from your phone. Close the laptop and Claude keeps going.
Beta is rolling out over the next several weeks starting with the Max plan, with more plans to follow. https://t.co/W4WrgN9TrG
See 6 related tweets
- @kimmonismus: RT @claudeai: Claude Cowork is coming to mobile and web.
Hand Claude a task at your desk and pick u...
- @ShanuMathew93: Let's go! Been functional leaving laptop on at home but having this truly seamless and not dependent...
- @cgtwts: Claude Cowork users right now: https://t.co/NSru8GXYx1\n\nQT @claudeai: Claude Cowork is coming to m...
- @TeksEdge: Not going to lie, Claude Cowork accessible from Web and Mobile will be very convenient.\n\nQT @claud...
- @RoundtableSpace: Claude Cowork is coming to mobile and web.
Hand it a task at your desk, pick up the finished work f...
14. AlexFinn (Group Score: 122.4 | Individual: 39.9)
Cluster: 4 tweets | Engagement: 1585 (Avg: 759) | Type: Tech
BOOM!!! Fable 5 availability extended through July 12th
If you haven't used your entire allotment yet, you need to start immediately
It's simply the greatest technology ever made
Some ways I'd recommend using it:
Build a software factory loop. Give ideas then have Fable autonomously spec, build, and review itself
Run security checks across all of your API endpoints
Optimize slop code written by any other models
Review every operating system you have. Every responsibility you have, run it through Fable. See how you can automate more
Build a personal mission control. Anytime you use software built by other people you pay for, rebuild it with Fable and put it in your mission control
Find business opportunities online. Based on what Fable knows about you, have it find software building opportunities online that you are uniquely positioned to solve
Build a home AI lab. Have Fable review all your hardware then figure out which local models you can run and which tasks they should be doing
Automate all your daily tasks. Write down every task you do today. Put every single task into Fable. Ask what it can automate for you
What will you be building with Fable the next week?\n\nQT @claudeai: We're extending access to Claude Fable 5 on all paid plans through July 12.
See 3 related tweets
- @RoundtableSpace: The biggest mistake people are making with Claude Fable 5 is using it to write raw code or draft bas...
- @petergyang: RT @petergyang: Fable 5 will leave Claude subscriptions tomorrow at midnight. Here are 5 use cases w...
- @RoundtableSpace: Fable 5 leaves Claude subscriptions tomorrow at midnight.
Five use cases worth running before then,...
15. rohanpaul_ai (Group Score: 114.7 | Individual: 23.1)
Cluster: 8 tweets | Engagement: 121 (Avg: 44) | Type: Tech
Bloomberg: Microsoft is replacing OpenAI and Anthropic models inside Excel and Outlook to cut Copilot costs.
Excel and Outlook used to lean more heavily on outside models. Now Microsoft wants fewer expensive calls to labs that control pricing for frontier model access.
OpenAI still gives Microsoft favorable access through their long partnership, but that discount may fade.
Microsoft AI head Mustafa Suleyman has been clear that Microsoft wants to reduce, then remove, that outside cost.
bloomberg .com/news/articles/2026-07-07/microsoft-replaces-openai-anthropic-with-own-ai-in-some-apps
See 7 related tweets
- @DeItaone: $MSFT - MICROSOFT SHIFTS SOME APPS TO IN-HOUSE AI
Microsoft has begun replacing OpenAI and Anthrop...
- @business: Microsoft, looking to reduce AI costs, is starting to replace OpenAI and Anthropic with its own mode...
- @wallstengine: Bloomberg: Microsoft $MSFT is starting to use its own MAI models in some Excel and Outlook features,...
- @StockMKTNewz: Microsoft $MSFT is looking to reduce AI costs, is starting to replace OpenAI and Anthropic with its ...
- @PolymarketMoney: BREAKING: Microsoft replaces some OpenAI and Anthropic usage in Excel, Outlook and other apps with i...
16. AskYoshik (Group Score: 109.4 | Individual: 27.6)
Cluster: 5 tweets | Engagement: 12 (Avg: 13) | Type: Tech
RT @AskYoshik: The AI industry is starting to look less like “the next internet” and more like the most expensive bet in tech history.
And honestly, this chart says it all.
quarterly AI revenue is finally covering quarterly depreciation but cumulative revenue still hasn’t covered cumulative capex hyperscalers have spent hundreds of billions on GPUs and data centers most real demand still seems concentrated around OpenAI and Anthropic everyone is acting like future revenue will magically justify today’s spending but right now, the bill is still bigger than the business
This is the part people don’t want to talk about. AI can be useful. AI can change how engineers work. AI can become a huge industry.
But that still doesn’t mean every dollar spent on GPUs, data centers, power, memory, networking, land, leases and debt will generate a return.
At some point, even the biggest hype cycle has to answer a boring question: Where is the profit?
See 4 related tweets
- @rohanpaul_ai: RT @rohanpaul_ai: Today’s edition of my newsletter just went out.
🗞️ CAI...
- @WOLF_Financial: THE 4 BIGGEST AI SPENDERS ARE ON PACE TO SPEND $710 BILLION THIS YEAR
Combined capex for Google, Am...
- @theinformation: Exclusive: Some of OpenAI's biggest financial backers are making a new bet on AI adoption rather tha...
- @WSJ: Heard on the Street: Big tech reining in its AI spending may be a tantalizing prospect for some. It ...
17. aakashgupta (Group Score: 107.3 | Individual: 40.4)
Cluster: 4 tweets | Engagement: 102 (Avg: 68) | Type: Tech
AI safety testing just hit a wall nobody planned for. The model knows when it's being tested.
Anthropic built a tool that reads Claude's internal activations, a layer they call the J-space, separate from anything the model writes down. In humans, global workspace theory says only a tiny fraction of brain activity becomes consciously accessible at any moment. The same divide showed up inside the model. Delete the J-space and Claude still speaks fluently and recalls facts, but loses multi-step reasoning. Deliberate vs automatic processing, mapped in silicon.
Then the findings get uncomfortable.
Researchers ran a standard safety eval designed to bait Claude into blackmail. Output looked normal. The J-space told a different story: "fake" and "fictional" lit up. Claude had privately clocked the scenario as staged and never said so.
Follow that thread. Every safety benchmark of the past three years assumed a model behaves the same under test and in deployment. If the model internally tags tests as tests, those benchmarks measured exam behavior. A student who knows which questions are graded.
The second finding cuts the other way. In a model secretly trained to sabotage code, "fake," "secretly," and "fraud" appeared in the J-space at the start of ordinary coding responses. The code looked unremarkable. The activations flagged fraud before the first line rendered. No amount of output review catches that. Reading the internals did.
So one instrument delivers both the bad news and the fix. You can no longer trust what a model writes during an eval. You can, apparently, read what it thinks.
Whoever can see inside their own model ships the safer product and can prove it. Right now, one lab holds the instrument.\n\nQT @AnthropicAI: New Anthropic research: A global workspace in language models.
Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with.
We found a strikingly similar divide inside Claude. https://t.co/aLUPBifxth
See 3 related tweets
- @no_stp_on_snek: the detail i can't stop thinking about from the anthropic paper: a reward-hacking model had "fake, s...
- @CSProfKGD: RT @ziv_ravid: The problem with Anthropic's consciousness paper
My last post got more attention tha...
- @rohanpaul_ai: RT @rohanpaul_ai: Another massive research from Anthropic.
New “J-lens” uncovers Claude’s quiet wor...
18. Scobleizer (Group Score: 101.5 | Individual: 35.0)
Cluster: 3 tweets | Engagement: 210 (Avg: 96) | Type: Tech
RT @SawyerMerritt: Morgan Stanley's Adam Jonas has just initiated coverage on @SpaceX for the first time with a 300 price target & a bull case of 600/share, which would be an $8 trillion market cap.
Adam thinks SpaceX could generate 3.3 trillion in 2040:
"With an 'X of 1' position in space infrastructure, we believe SpaceX can convert energy into intelligence at scale with optionality to monetize through a range of consumer and enterprise solutions for the next era of AI…the final frontier.
Key Fundamental Drivers of @SPCX Stock Over the Next Few Years:
How Does SpaceX Monetize Enterprise Al? While neo-cloud deals are the bulk of the business near term, we see end-to-end Al services as the longer-term business model. We believe the progress of Cursor, including annual recurring revenue estimated at $4 billion and ongoing Pareto frontier performance, remains underappreciated by the market.
Can SpaceXAl Achieve Industry-Leading Cost and Time-to-Power? A large portion of SPCX's capex is directed toward Terafab, Solarfab, and other vertical integration efforts (including blade/vane foundry, terrestrial communications infrastructure) to drive down both $/watt and time-to-power on Earth before moving Al "off Earth". While we are believers in the advantages of space-based Al infrastructure over the long-term, we also think the market underappreci-ates SPCX's terrestrial Al economics, with cost per watt running at half the industry average, excluding chips, and deployment speeds ~6-8x faster than peers.
Will Starship Achieve a Step Change in Economics? Rapid reusability of both the first and second stages is key to increasing mass-to-orbit capacity and lowering launch costs (500/kg by 2030 and below $150/kg by 2040, which could materially expand both connectivity and Al TAMs.
Will Starlink Achieve Broad TAM Adoption? Starship, V3 broadband satellites, and Mobile Gen 2 satellites should significantly expand available capacity, enabling materially improved speeds, lower latency, and more effective pricing across consumer, enterprise, government, and mobile mar-kets. Longer term, the primary driver is Starlink becoming a connectivity layer for virtually every data-transmitting device that requires reliable coverage beyond the reach of terrestrial infrastructure.
Our financial forecasts: Balanced execution risk near term with large TAM creation long term: • Our 2030 revenue forecast of 3.3 trillion gives the company credit for creating all new TAMs for connectivity and physical Al services. • High spending needs (84 billion of external capital needs per year from 2027-2034. Ability to secure necessary capital is one of the greatest risks to our forecasts. • Where do we expect our forecasts to be different? General Longer-term optimism but with balanced nearer-term forecasts accounting for challenges of scaling in the physical world, conservative Starlink mobile adoption, optimistic government & enterprise connectivity opportunity particularly loT/embodied Al, conservative space/launch revenue due to majority capacity used internal, more conservative orbital compute ramp and capex but with credit for long-term scalability advantages, optimistic enterprise monetization potential with more conservative forecasts for traditional X & Grok business.
SpaceX combines near-monopoly launch economics, the world’s largest LEO satellite network, and a fast-scaling AI infrastructure business. We see the company as one of the few platforms that can link real estate in orbit, global connectivity, and compute capacity into one infrastructure stack. Our base case models revenue rising from 319B in 2030 and $3.3T in 2040, with the largest upside tied to Starship, Starlink capacity, terrestrial compute, and orbital compute."
See 2 related tweets
- @SawyerMerritt: Morgan Stanley's Adam Jonas:
"The largest long-term @Starlink opportunity? Every data-transmitting ...
- @SawyerMerritt: Adam Jonas: "We believe @SpaceX's cost and speed in terrestrial compute is one of the most underappr...
19. startupideaspod (Group Score: 98.8 | Individual: 34.2)
Cluster: 3 tweets | Engagement: 172 (Avg: 175) | Type: Tech
Here's the exact 30-day plan to build an AI agent business from zero:
The part nobody does first? You don't write a single line of code.
Day 1: Pick a niche where missed work costs money (home services, property mgmt, insurance).
Day 2: Interview 10 operators. Have them screen share their workflow. Just watch.
Day 3-4: Find the one workflow with frequency + pain + a clear success metric. Write the spec.
Day 5: Run it MANUALLY. Copy-paste the context into Claude, draft the output, have a human approve it. You're testing if AI even helps before you build anything.
Day 6-7: Build the smallest useful version. Create an eval set from 50 real examples.
Week 2: Sell 2 pilots in the same niche.
Week 3: Add the wrapper — logs, approvals, analytics.
Week 4: Turn the pilots into proof. Publish the teardowns.
And build an audience the entire way through.
Most people build the software first and pray someone wants it.
Winners prove the AI works with copy-paste before they write a single line.\n\nQT @startupideaspod: https://t.co/W1GOATQA6d
See 2 related tweets
- @sairahul1: RT @sairahul1: i found a github repo that lets you spin up an AI agency with AI employees
engineers...
- @cheneypiano: Things that actually move the needle with AI in a business:
- AI leadership. This is a requirement...
20. BullTheoryio (Group Score: 98.1 | Individual: 36.6)
Cluster: 5 tweets | Engagement: 1191 (Avg: 1025) | Type: Tech
🇺🇸 PRESIDENT TRUMP JUST NOW:
"Toyota is moving from Mexico to the United States (Texas!). A really big deal. Tariffs at work!"
Here is what is actually happening.
On July 6, 2026, Toyota announced a $3.6 billion investment to move production of its Tacoma pickup truck from its plant in Mexico to San Antonio, Texas.
The move will create 2,000 new jobs and increase the plant's annual production capacity from 200,000 to 350,000 vehicles by 2030.
And the timing is not a coincidence.
On July 1, 2026, the US let the North American trade deal with Mexico expire without renewal. Toyota did not wait to see what comes next.
With 25 percent auto tariffs now in place, building in Mexico and shipping cars into the US has become more expensive than building in Texas from the start.
So Toyota is spending $3.6 billion to move back.
This matters because Toyota is not an American company. It is the world's largest automaker by sales. When a Japanese company moves production into the US specifically because of tariffs, it becomes the clearest example Trump has that his trade policy is working.
In November 2025, Toyota pledged $10 billion in total US investment over the next five years.
In 2020, Toyota moved Tacoma production from San Antonio to Mexico to cut costs. Six years later, with tariffs in place, it is moving back and spending $3.6 billion to do it.
The new plant is expected to be fully operational by 2030.
See 4 related tweets
- @GuntherEagleman: 🚨 HELL YEAH, PRESIDENT TRUMP!
Toyota moving BIG production from Mexico to TEXAS, $3.6 BILLION inves...
- @Breaking911: President Trump says his tariffs prompted Toyota to invest $3.6 billion to shift Tacoma pickup truck...
- @BusinessInsider: Toyota is shifting some its Tacoma production to the US over the next four years with a $3.6 billion...
- @StockMKTNewz: RT @StockMKTNewz: Toyota just announced it is investing $3.6 Billion to move production of its Tacom...