- Published on
科技热门推文 - 2026-09-16
- Authors

- Name
- geeknotes
今日科技动态聚焦人工智能安全与商业化,相关帖文重点关注放缓前沿人工智能研发的呼声,以及对独立评估的审视。与此同时,新兴训练方法力求突破聊天模型的局限,材料实验室则将实验与模型学习相结合。据报道,谷歌推出的 Gemini 语音功能与 Odyssey 的物理智能体,展现出系统向更强交互性和更广泛用途发展的趋势。对于初创企业而言,关注重点正从构建智能体转向销售智能体,而面向支付、用户引导和创意工作流程的工具,有望为实际应用提供更切实可行的路径。
1. XFreeze (Group Score: 930.7 | Individual: 45.0)
Cluster: 28 tweets | Engagement: 3161 (Avg: 640) | Type: Tech
Elon just gave a much more practical explanation of the current AI safety problem....and laid out a practical framework for making frontier AI safer
OpenAI and Anthropic are the two leading AI companies....and their models are close enough in capability that it’s incredibly difficult for either one to slow down without basically handing the lead to the other
Elon also made an important distinction:
On balance, Anthropic puts more care into safety than OpenAI....but even people inside Anthropic are publicly worried about what their own models can do
So instead of asking one company to voluntarily fall behind, Elon’s solution is much more practical:
Have Anthropic run its safety test harness on OpenAI models
Have OpenAI test Anthropic Have SpaceXAI test both And bring the leading Chinese AI companies into the same system
Everyone tries to break everyone else’s models before release
“The odds that you will find issues are dramatically greater”
That makes way more sense than asking frontier labs to grade their own homework\n\nQT @theallinpod: All-In Summit: Elon Musk & Gwynne Shotwell
-- AI Risks and Peer Review
-- Starship
-- Terafab as Taiwan Insurance
-- SpaceX/Tesla Merger Potential
-- Management at SpaceX
++ much more!
(0:00) SpaceX's @Gwynne_Shotwell joins The Besties!
(1:45) Gwynne's SpaceX story, selling rockets, and working for Elon
(6:45) Running SpaceX: AI, Starlink, Rockets, and X
(13:43) Direct to cell with Starlink, retiring rockets, competition
(18:34) Data centers in space
(22:00) Management at SpaceX
(27:37) @elonmusk joins: AI's real risk, model peer review, what he meant by "Dario is right"
(36:53) Elon and Gwynne on their working relationship, Starship's future
(46:16) Terafab, Tesla, Flying cars?, merging Tesla and SpaceX
(52:51) Lying AIs, how to do AI peer review right
SPCX
See 27 related tweets
- @cb_doge: SpaceX President Gwynne Shotwell on data centers in space:
“There is an expense to put data centers...
- @theallinpod: Elon Proposes Adversarial Peer Reviews for AI Safety
@elonmusk:
“What I think would be wise to do ...
- @XFreeze: Elon’s entire engineering philosophy in one sentence:
“Physics is the law, and everything else is a...
- @XFreeze: Elon just clarified what he actually meant when he said “Dario is right”:
"I probably should have s...
- @cb_doge: Grok Bot Summary of Elon Musk’s All-In Summit Interview Today
Opening bit
- Joke cold open: “We’re ...
2. damianplayer (Group Score: 572.6 | Individual: 54.4)
Cluster: 20 tweets | Engagement: 3798 (Avg: 237) | Type: Tech
RT @CompleteSkeptic: After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
See 19 related tweets
- @danshipper: we almost never test new foundation models but we've been testing this for ~a week @every and it's p...
- @svpino: New model that looks very different from everything else we've seen before:
• It returns decisions ...
- @Mayhem4Markets: Why are we making computers generate text token by token when we only need an answer to a question? ...
- @sedielem: "output tokens are free, because they're finally too 🤬🤬🤬 cheap to meter" 😂
Reinforcement learning f...
- @omarsar0: Recommended read. Jev gives up text generation to make AI dramatically faster.
TypeSafe built a n...
3. tenobrus (Group Score: 540.2 | Individual: 56.2)
Cluster: 15 tweets | Engagement: 2546 (Avg: 348) | Type: Tech
RT @DKokotajlo: Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam September 14, 2026
Link to original doc: https://t.co/TxMNr0vhrL
See 14 related tweets
- @nxthompson: A genuinely scary point: the models are now past the point when they’re aware enough to know when we...
- @VaibhavSisinty: An OpenAI capabilities researcher who has worked on AI for 15 years published a personal statement o...
- @SemiAnalysis_: struggling to maintain the discipline to engage deeply https://t.co/ky0igaq6xk https://t.co/09KzUCC6...
- @Miles_Brundage: Everyone should read this immediately https://t.co/HNTAz1KKXO\n\nQT @DKokotajlo: Dan Selsam is a cur...
- @cgtwts: “The crucial and overlooked problem is that the models are becoming so situationally aware that we a...
4. Hesamation (Group Score: 411.7 | Individual: 67.9)
Cluster: 10 tweets | Engagement: 9760 (Avg: 611) | Type: Tech
be anthropic say ai might take over internet researchers: “it might kill humans.” dario: “we need independent evaluators” choose metr metr gets funded by the same network funding anthropic which also funds ai safety orgs which also funds doom research which also funds policy work which also holds major wealth linked to anthropic which also funds tarbell tarbell pushes doom into time, the verge, science, la times, etc.
anthropic makes more money network gets richer more fund goes into ai doom ai doom pushes more regulation regulation favors “responsible” frontier labs anthropic gets more power repeat
the doom funds the network → the network amplifies the doom → anthropic benefits from both\n\nQT @kevinnbass: I have conducted an audit of Anthropic's finances.
What I have found is so shocking that I am calling for a Congressional investigation.
Anthropic is not just seeking regulatory capture.
It has built a regulatory capture machine that cannot be turned off.
Structural financial incentives make it impossible for Anthropic -- I call it the Anthropic Network -- to turn off its own AI doom cycle.
It starts with METR.
Dario Amodei proposes "third-party evaluators" to assess the risk of Anthropic's models.
He proposes METR for this purpose.
But METR is financially dependent on the Anthropic's success -- specifically, on the explosive growth of more than $7 billion dollars in Anthropic stock.
Dustin Moskovitz invested this stock into Good Ventures Foundation, where it represents the majority of that organization's portfolio.
And GVF is the overwhelming funder of the entire Anthropic Network ecosystem.
This stock was worth $500 million early last year.
It is worth more than $7.7 billion just ~16 months later.
METR -- and all of those building a career its parent organizations -- cannot afford to disrupt that growth.
Because if Anthropic goes under, many of the organizations that fund METR go under as well.
But if Anthropic succeeds, METR and its parent organizations become more richly financed to regulate AI -- something those at METR want very much.
The "third-party evaluator" is not "third-party" at all.
The evaluator is on Anthropic's payroll.
If this were the end of it, that's bad.
But that isn't all.
The same organizations that fund METR also fund the many organizations, such as the Tarbell Center, that promote AI Doom.
The Tarbell Center publishes AI Doom articles in The Verge, Science, LA Times, The Dispatch, TIME, and others.
They are selling the problem, and then selling the solution to the problem -- from the same money pile: Anthropic's.
All of these organizations are financially dependent on the same exploding $7 billion money pile.
As Anthropic grows more and more powerful, its AI Doom Machine grows better and better financed -- louder and louder.
Meanwhile, the regulatory regime seeded in METR grows larger to solve the increasingly loud -- now hysterical -- problem of AI Doom that the Anthropic Network itself created.
From this standpoint, as Anthropic becomes more powerful, AI might be getting scarier, sure -- but the positive feedback loop also becomes more deafening -- independent of objective facts.
This itself is an objective fact.
The deafening AI Doom is part of an business model, that, as it expands, so too does the AI Doom messaging -- there is simply more money to do it.
But the problem also goes in the other direction:
If Anthropic dies, the Regulatory Regime and the AI Doom Machine are crippled or die.
Neither METR nor Tarbell nor the other organizations in the Anthropic Network can allow that to happen.
Hence, neither METR or the AI Doom Machine can be trusted to provide independent assessments of Anthropic's models or AI more broadly.
They simply are not organizations independent of Anthropic.
And Anthropic cannot detach itself from METR or Tarbell or countless other safety orgs (not shown here), either, because they drive hype for the models and the possibility of eventual regulatory capture, and Anthropic will not give that up willingly.
What's more, the people at all of these organizations are all the same ecosystem, the same community. They just shuffle between organizations.
The Anthropic Network is therefore, so long as it is successful, locked into a self-amplifying feedback loop inside an ideological monoculture.
And that feedback loop is winning.
That's what Jacob Coxon is.
China is keeping messaging tight. That is why optimism for AI is so high in China.
America has Anthropic: a massive company pushing anti-AI propaganda at a state level.
Anthropic will either create hysteria until American AI slows down and China wins, or it will create fractures throughout American society with severe political consequences.
Ironically, because of the structural financial incentives underpinning the Anthropic Network, it has become the same kind of self-amplifying virus that it fantasizes AI to become in the future -- while hiding its tracks just as carefully.
It is the mirror of the same AI virus that it hypothesizes to consume America.
Anthropic's business model, models itself after the very thing it claims to fear.
Except Anthropic's ideology infects humans, not computers.
Congress must investigate.
Evidence and Github in next post.
Then some supplementary figures.
See 9 related tweets
- @EWErickson: RT @kevinnbass: I have conducted an audit of Anthropic's finances.
What I have found is so shocking...
- @Arjunjain: RT @Hesamation: > be anthropic
say ai might take over internet researchers: “it might kill human...
- @Mayhem4Markets: Anthropic is an evil company focused on regulatory capture.\n\nQT @kevinnbass: I have conducted an a...
- @kevinnbass: I genuinely believe that Dario and Moskovitz are primarily ideological. But a captured industry is e...
- @mark_k: Anthropic is a commercial outgrowth of the Effective Altruism doomer cult.
Here's the whole thing u...
5. LearnWithSubhan (Group Score: 380.8 | Individual: 31.1)
Cluster: 15 tweets | Engagement: 26 (Avg: 79) | Type: Tech
RT @RiadMd46702: 🚀 100+ AI Tools That Can Replace Hours of Tedious Work
Stop doing manually what AI can do in minutes. ⚡
🔎 Research • ChatGPT • Perplexity • Gemini • Copilot • https://t.co/8ghAQxpBqX • Abacus
🎨 Image Generation • Midjourney • Grok • GPT-4o • Higgsfield AI
⚡ Productivity • Gamma • Perplexity AI • Grok • Gemini
✍️ Writing • Jasper • QuillBot • TextBlaze • Jenni AI
🎬 Video • Kling • Runway • HeyGen • InVideo • Klap
🎙️ Meetings • Otter • Fireflies • tl;dv • Noty AI
📈 SEO • VidIQ • Seona AI • BlogSEO • Outrank • Keywords AI
📊 Presentations • Gamma • https://t.co/XqUkMqHqc9 • Decktopus • SlidesAI • https://t.co/8ocAW8sykT
🖌️ Design • Canva • Clipdrop • Flair AI • Designify • AutoDraw • Magician
🎧 Audio • ElevenLabs • LOVO AI • Adobe Podcast • Songburst
📣 Marketing • https://t.co/zUt6zufuTZ • Pencil • AdCopy • Simplified • AI-Ads
🚀 Startup • Tome • Namelix • IdeasAI • Validator AI • PitchGrade
📱 Social Media • Typefully • Hypefury • TweetHunter • Taplio
🔖 Save this list. You might discover your next favorite AI tool here.
❤️ Like + 💬 Comment + 🔁 Repost
Follow @RiadMd46702 for more AI tools, prompts & discoveries.
#AI #AITools #ArtificialIntelligence #ChatGPT #Grok
See 14 related tweets
- @ElizabethA77617: RT @MdRimon488484: 🚀 100 AI TOOLS THAT TURN HOURS OF WORK INTO MINUTES
Stop doing everything manual...
- @HarshBisen143: RT @RiadMd46702: 🧵 80+ AI tools that can turn hours of work into minutes. 🪄
Whether you’re research...
- @LearnWithSubhan: RT @The_techAI: 🚀 100 AI TOOLS THAT TURN HOURS OF WORK INTO MINUTES
Stop doing everything manually....
- @AiwithDharmik: RT @iamsania_AI: Professionals won’t tell you this 👀 They use these daily. 🪄⚡
- Ideas 🧠
- YOU
- C...
- @hey_Jessicaai: RT @tec_safwan: Professionals won’t tell you this They use these daily. 🪄⚡
- Ideas
- YOU
- Clau...
6. AiwithDharmik (Group Score: 355.8 | Individual: 45.3)
Cluster: 11 tweets | Engagement: 450 (Avg: 88) | Type: Tech
RT @Romeocoder11: 50 websites that feel like the internet’s hidden toolbox 🧰
- https://t.co/whf8cpggZP — Free research papers
- https://t.co/W9G5JYoP1p — Borrow books online
- https://t.co/5fx02bhtmk — Free academic journals
- https://t.co/AYlnQEGx9J — App alternatives
- https://t.co/dkEnujZk26 — Find where to stream
- https://t.co/gdgp16vtG5 — Internet archives
- https://t.co/uUJZAIHz4Z — 70K+ free books
- https://t.co/EAn1ULKbCe — Free textbooks
- https://t.co/2LN4huDZJh — Free courses
- https://t.co/6vrjKKMium — Solve complex problems
- https://t.co/MK8MvPK4il — Photoshop alternative
- https://t.co/I7uz0czHaN — Compress images
- https://t.co/7N1KCgGFas — Remove backgrounds
- https://t.co/oZV5UxBrI9 — Remove objects
- https://t.co/ORT2siPnWL — Remove video backgrounds
- https://t.co/CXAgbDeqct — Beautiful code images
- https://t.co/yXpBBzippn — Code screenshots
- https://t.co/VGnZuhQEim — Product mockups
- https://t.co/krVW4hwjzR — Create mockups
- https://t.co/bBkELVAt93 — Check data breaches
- https://t.co/VMQyrjwb0g — Scan files & URLs
- https://t.co/jsVIbMycvy — Self-destructing notes
- https://t.co/hLxUyFAi7W — Temporary email
- https://t.co/eg6a7DhZLl — Temporary file sharing
- https://t.co/lkivVt7dOT — Save webpages
- https://t.co/lj4riDGUoL — Find similar websites
- https://t.co/dgTfb7zoEf — Explore global radio
- https://t.co/OqDuyWg6qG — Discover music genres
- https://t.co/BFzaDlogCW — Find songs from shows
- https://t.co/y1YT2fyFEi — Focus music
- https://t.co/MLPLNEdcne — Custom background sounds
- https://t.co/JnWg6Ha1yH — Café ambience
- https://t.co/rMqsJYxkpf — Research assistant
- https://t.co/Oecn0bcIYA — Research-backed answers
- https://t.co/I5V2qJoIev — Research connections
- https://t.co/SiWPFNiOlU — Academic search
- https://t.co/ZVWtZRLQ1f — Understand research papers
- https://t.co/lcH6G7t8zK — YouTube summaries
- https://t.co/y6WddedNuz — AI for developers
- https://t.co/0rjqiP4dKi — Test regex
- https://t.co/ike0LVGNNh — Format code
- https://t.co/yMy1Y53nh9 — Format JSON
- https://t.co/dXNBrMikKX — Understand terminal commands
- https://t.co/2SbZQwX3rQ — Bookmark manager
- https://t.co/N9aovTmuB8 — Check outages
- https://t.co/5PtgmmdY1r — Reverse image search
- https://t.co/wTapkOkVXe — Internet speed test
- https://t.co/zYrs05zcm1 — PDF tools
- https://t.co/Km7KAsEnhm — Merge/split PDFs
- https://t.co/QI84JvKWvS — Temporary email
Save this. You’ll definitely need some of these later. 🔖
Follow @Romeocoder11 for more useful websites, AI tools & tech resources.
See 10 related tweets
- @jihad_sameul: RT @ElizabethA77617: 50 Web Sites Google Doesn't Want You to Know About
1.) 'https://t.co/6GVNuesVm...
- @LearnWithSubhan: RT @ArifAIHQ: 20 Websites You Should Bookmark
- AlternativeTo - Find alternatives to any software ...
- @tec_safwan: RT @ArafatMd93059: 50 websites that feel like the internet’s hidden toolbox 🧰
- @JacobyoungAI: RT @liam_holt7: 50 websites that feel like the internet’s hidden toolbox
- @AiwithDharmik: RT @ZavianKairo_AI: 50 websites that feel like the internet’s hidden toolbox 🧰
7. CarinaLHong (Group Score: 310.7 | Individual: 40.8)
Cluster: 10 tweets | Engagement: 526 (Avg: 93) | Type: Tech
RT @LiamFedus: We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next.
Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon.
This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials.
Read our blog posts below.
See 9 related tweets
- @richardczl: 1,300 H200s plus proprietary lab data beats frontier scale on domain tasks. Every company with a dat...
- @RadixArk: We're proud that @periodiclabs chose SGLang and Miles to build Neon.
Periodic extended SGLang and M...
- @saranormous: super cool to see real-world experimental data at scale in the training pipeline — periodic is a pio...
- @ying11231: Congrats to the launch 🚀 and excited to see SGLang @sgl_project and Miles @radixark helped Periodic ...
- @dwarkesh_sp: Visited the lab - was struck both by how wide the search space is for materials synthesis experiment...
8. sauda_coder (Group Score: 289.2 | Individual: 57.9)
Cluster: 9 tweets | Engagement: 7568 (Avg: 191) | Type: Tech
RT @sauda_coder: 50 websites that can turn a “quick visit” into hours of exploring 🌍
- https://t.co/e9MzKtH8Ac — Live satellite views
- https://t.co/F0sYcAhsEB — Track planes worldwide
- https://t.co/dlD4FsbeZ5 — Track ships in real time
- https://t.co/enz2psqLq2 — Live weather & storms
- https://t.co/6WtqMIyIXO — Live lightning strikes
- https://t.co/rDjMVuFJ9p — Recent earthquakes
- https://t.co/2hL9AwpAnD @sauda_coder — Explore internet cables
- https://t.co/OXUHY0vxtW — Monitor global forests
- https://t.co/3ykvE8xIQm — Live world statistics
- https://t.co/WIRyI84z3z — Live internet stats
- https://t.co/WF9Dt2YTJU — Compare country sizes
- https://t.co/lkLgao8hPE — Explore historical maps
- https://t.co/62atEjao3d — Historical map archive
- https://t.co/0roKo8iBFs — Community-built world map
- https://t.co/t4SwdWbJzL — Windows around the world
- https://t.co/upNXry1y9K — Virtual city walks
- https://t.co/0c5Vd0eMlf — Random places worldwide
- https://t.co/A07775U3oX — Strange places worldwide
- https://t.co/2LNjVQBuJG — Fun interactive experiments
- https://t.co/WE3bRUopQP — Atom to universe
- https://t.co/BlWzVbrKLO — Explore space in 3D
- https://t.co/F85rw8lZ7F — Interactive sky map
- https://t.co/DXt3VwPFml — NASA’s daily space image
- https://t.co/4NjhLkVJzZ — NASA image archive
- https://t.co/CKAVlLPXI3 — Interactive data stories
- https://t.co/y6mFbQxfB7 — Global data & insights
- https://t.co/DkEF24VHkq — Understand global trends
- https://t.co/w4FExJsjYb — Data visualizations
- https://t.co/eJbNAd4SBo — World Bank data
- https://t.co/xyBlo89HIx — Turkey’s official statistics
- https://t.co/A1aMVlBPFk — Massive digital archive
- https://t.co/Mq40zmIIEJ — Free classic books
- https://t.co/IeLWcbvCzB — Explore millions of books
- https://t.co/YQD3ttZBoy — Library of Congress archive
- https://t.co/Ix06KBwQyn — Europe’s cultural archive
- https://t.co/CgTot9xsnm — America’s digital library
- https://t.co/STGOC79ra4 — Virtual museums
- https://t.co/HVjvwXZxdj @sauda_coder High-res artworks
- https://t.co/11ALJRqB12 — Online art collection
- https://t.co/C0q1e7mBuK — Historical treasures
- https://t.co/FzLnwbvNFj — Free culture & education
- https://t.co/DvnRSpu9PG — Explore the Met collection
- https://t.co/4LzaPivTOn — Explore music genres
- https://t.co/YRFJAaEZxr — Music by country & decade
- https://t.co/KRJ5e0TlCK — Wikipedia edits as audio
- https://t.co/MD4cssnXHm… — Random knowledge
- https://t.co/SQ0ZFGnCHI — Time & astronomy tools
- https://t.co/jVSqK5hoZO — Latest science news
- https://t.co/DO3rXyKDYD — Free research papers
- https://t.co/HzGCsimVyk — Interactive data visualizations
The internet is much bigger than your usual feed.
🔖 Bookmark this for later.
Follow @sauda_coder for more useful websites & AI tools. 🚀
See 8 related tweets
- @VikramVerm25510: 50 websites that can turn a “quick visit” into hours of exploring 🌍
- https://t.co/Qm5NkSIftR — Li...
- @mr_nirajkumar07: RT @mostofakamal00: 🌍 50 websites that can turn a “quick visit” into hours of exploring.
The intern...
- @LearnWithSubhan: RT @lihazadn_Ai: 50 websites that can turn a “quick visit” into hours of exploring .
- https://t.c...
- @JayBisen473370: RT @Tech_Arish: 50 websites that can turn a “quick visit” into hours of exploring 🌍
- https://t.co...
- @Zayan5754: RT @naiem_bn: 50 websites that can turn a “quick visit” into hours of exploring 🌍
9. WOLF_Financial (Group Score: 277.4 | Individual: 53.0)
Cluster: 12 tweets | Engagement: 527 (Avg: 71) | Type: Tech
PRESIDENT TRUMP CALLED JENSEN HUANG'S CELL PHONE WHILE HE WAS ON STAGE AT THE ALL-IN SUMMIT TODAY. JENSEN PICKED UP AND PUT HIM ON SPEAKER 🇺🇸
Two days after Sam Altman and Dario Amodei called for AI development to slow down, Trump told a few thousand people what he thinks of that.
"The robots are not going to be taking over the world. The AI will not be taking over the rest of the world. The whole thing is a hoax."
"We have to do things prudently, but that doesn't mean we're going to stop an industry."
"Whoever wins AI wins. That's how big it is. It's bigger than the internet."
On data centers, he said they're "the oil of the next 20, 25 years."
He also said $20 trillion of investment is coming into the country this year and told Jensen he "can develop the most complex computer chip in the world that nobody can copy for 10 years, but he can't figure out how to put me on speaker."
See 11 related tweets
- @XFreeze: Jensen Huang on Elon Musk building Terafab:
“If anybody could do it, he can”
“Elon likes to talk a...
- @FoxNews: President Trump makes an unexpected call to Nvidia CEO Jensen Huang while he’s onstage, weighing in ...
- @StockMKTNewz: Nvidia $NVDA just confirmed that its next GTC conference in Washington DC will happen again for the ...
- @exec_sum: BREAKING: Trump called Nvidia CEO Jensen Huang during the All-In Summit and was put on speaker, sayi...
- @chamath: Watch this! It was 🔥\n\nQT @theallinpod: THE ALL-IN SUMMIT IS LIVE!
@JensenHuang joins The Besties ...
10. lihazadn_Ai (Group Score: 250.6 | Individual: 30.1)
Cluster: 11 tweets | Engagement: 126 (Avg: 111) | Type: Tech
RT @LearnWithSubhan: This feels like a much more practical way to use AI image generation for brands.
Generate two product visuals with GPT Image 2.5.
Then use those visuals as the starting point for a commercial with Seedance 2.5 in CapCut PC.
GPT Image 2.5 is coming to CapCut and will be available through Design Studio and AI Image.
Instead of creating an image and figuring out what to do with it afterward, the image is already part of the production pipeline.
Product → Visual → Video → Edit.
That's the workflow that makes AI-generated content feel useful for marketing.
#GPTImage25 #CapCutPC #AIVideo #Seedance25
See 10 related tweets
- @JayBisen473370: RT @code_bykuti: AI commercial production is starting to feel less like “generate a video” and more ...
- @JayBisen473370: RT @Grow_withAI: Product marketing usually needs a lot more than one good product photo.
You need t...
- @LearnWithSubhan: RT @RAVIKUMARSAHU78: A product doesn't need a full photoshoot to start becoming a commercial.
This ...
- @LearnWithSubhan: RT @manishkumar_dev: I don’t think the creator’s role in AI filmmaking should be:
Write prompt → pr...
- @monicaa_AI: RT @codedailyML: "The most interesting thing about GPT Image 2.5 coming to CapCut PC isn’t another A...
11. jonfavs (Group Score: 241.7 | Individual: 28.9)
Cluster: 9 tweets | Engagement: 5918 (Avg: 2731) | Type: Tech
RT @BarackObama: I was encouraged this week to see the leaders of the frontier labs agree on the need for them to slow down the pace of AI development. Given the stakes, it’s a good and necessary first step. But I’m even more encouraged by the growing recognition that how this powerful new technology develops should be at the center of our public debate. I’ve been watching the progress on AI for over a decade now, and one thing that’s clear to me is that the potential impact of this technology is not overhyped. It’s also moving at lightning speed – and even faster than those who are engineering it can keep up with. I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction.
But whether this technology results in amazing breakthroughs in medicine, energy and education or unleashes huge economic disruptions, greater inequality, and potential catastrophe will depend on the choices that we make right now – choices that should be made not just by the companies involved, but by all of us.
See 8 related tweets
- @TaylorLorenz: Obama invoking AI’s potential impact on "our kids" is concerning given how Democrats are weaponizing...
- @bubbleboi: Woah\n\nQT @BarackObama: I was encouraged this week to see the leaders of the frontier labs agree on...
- @MikeIsaac: surreal moment for me personally to see the 44th president of the united states make a statement usi...
- @jachiam0: Highly thoughtful remarks from President Obama.
Also, I am in no small amount of disbelief that "ac...
- @no_stp_on_snek: doomer detected\n\nQT @BarackObama: I was encouraged this week to see the leaders of the frontier la...
12. aakashgupta (Group Score: 236.0 | Individual: 37.0)
Cluster: 9 tweets | Engagement: 146 (Avg: 155) | Type: Tech
Odyssey says one model can run six very different systems.
The company calls Odyssey-3 a physics agent, and the demo list for September 15th reads like a stack of separate startups. A humanoid, robot arms, a car, a drone, video games, plus training other AIs, each shown running off the same world model, which Odyssey says is a first.
They've been on world models since 2023, and public release is a few weeks out, per the post. The data budget is the number to watch here.
Odyssey's framing is that many robotic systems today learn each task from piles of demonstrations of that task, brute forcing the problem with task-specific data. Their claim for Odyssey-3 is a few hours of experiential data per task, with pretraining credited for the physics and cause-and-effect, and a learned action head mapping the model's internal representations onto each system's controls.
The pitch is the language model playbook, moved to the physical world. Pretrain once, let each new system spend its data learning its own controls, and if that holds, advances in the base model reach many machines at once.
The sleeper detail sits on the simulation side. Odyssey says Odyssey-3 also generates worlds AI agents and humans can act inside, somewhere to watch risky behavior play out away from people.
For anyone scoping a physical product, the vendor question becomes how many hours a new machine costs you. The economics of physical AI now turn on how few hours each new machine needs before it works.\n\nQT @odysseyml: Today we’re unveiling Odyssey-3, a big step forward for foundation world models.
It can control robots, power humanoids, drive cars (on the roads of India!), train AIs, pilot drones, and even play video games.
We can’t wait to see what intelligent systems it enables. https://t.co/iNBWkPTfOO
See 8 related tweets
- @wallstengine: ODYSSEY UNVEILS ODYSSEY-3, A GENERAL-PURPOSE WORLD MODEL FOR PHYSICAL AI
Odyssey-3 is a new foundat...
- @DataChaz: THE ARCHITECTURE BEHIND THE NEW ODYSSEY 3 IS GENUINELY IMPRESSIVE
this fundamentally changes how we...
- @Parul_Gautam7: What feels different about @odysseyml -3 :
→ Same intelligence can move between robots, cars, drone...
- @Mayhem4Markets: Odyssey-3’s preview just went live. They’re calling it a new kind of agent that speaks the language ...
- @StockSavvyShay: Odyssey just unveiled Odyssey 3 which is a general purpose world model built to power robots humanoi...
13. heyshrutimishra (Group Score: 235.3 | Individual: 36.0)
Cluster: 9 tweets | Engagement: 110 (Avg: 90) | Type: Tech
HOLY SHIT.
Twin Store just fixed the part of agent-building nobody was solving.
Building the agent was always only half the job. The other half… how do you price it, take payments, onboard someone, give them a place to actually run it, had no clean answer. Until now.
15 minutes after your first successful run, you have a product page, subscriptions, Stripe checkout, and customer accounts running. You keep 100%.
Honestly the cleanest "idea to paid product" pipeline I've seen for agents.\n\nQT @hugomercierooo: 100k people have built agents on Twin. Starting today, anyone can sell one.
Introducing Twin Store.
In under 5 minutes, your agent gets a storefront, an app, and Stripe billing.
Build once. Twin finds you customers. Get paid monthly.
A new wave of wealth is coming, and it belongs to agent creators.
See 8 related tweets
- @Scobleizer: The agent boom still has a stupid hole in it. You can build something that saves a business ten hour...
- @anjum_ai: AI agents are getting really good at doing the work.
But building an agent is only one part of the ...
- @DataChaz: SELLING COURSES IS ABOUT TO BECOME THE OLD WAY TO MONETIZE KNOWLEDGE
The next wave is selling the a...
- @mr_nirajkumar07: RT @TechByMarkandey: Building an AI agent is one thing. Figuring out how to actually sell it is anot...
- @mr_nirajkumar07: RT @LunaTechAI: Creators have been leaving the real money on the table.
You sell the course on what...
14. wallstengine (Group Score: 234.4 | Individual: 36.2)
Cluster: 9 tweets | Engagement: 211 (Avg: 186) | Type: Tech
$GOOGL LAUNCHES GEMINI 3.8 LIVE AND 3.8 LIVE EXTENDED THINKING
Google introduced two new real-time voice models designed for conversational AI and agentic workflows.
Gemini 3.8 Live is built for scale and lower-cost deployment, while 3.8 Live Extended Thinking adds deeper multi-step reasoning while continuing to speak and narrate progress during longer tasks.
Both models support near real-time visual understanding, automatic switching across 97 languages and background tool/API calls without interrupting the conversation.
3.8 Live Extended Thinking ranks #1 on Artificial Analysis’ Speech-to-Speech Quality Index at 82.6 and scored 68.6% on τ-Voice agentic task completion.
Both are available through the Gemini Live API at 0.018/min for audio output.
Google is also partnering with Salesforce $CRM, Genspark and Lumeris, while broader integrations span platforms including LiveKit, Vercel, Agora and Pipecat.
See 8 related tweets
- @rohanpaul_ai: Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking
And now it takes the #1 overa...
- @Rasmic: dang, they built @heyclicky\n\nQT @GoogleDeepMind: Watch how we used 3.8 Live Extended Thinking to a...
- @patloeber: Introducing two new Gemini models for the Live API🔥
bringing major upgrades in intelligence and par...
- @AngryTomtweets: wow... Gemini 3.8 Live Extended Thinking turns raw sketches + near real-time voice feedback into fun...
- @thorwebdev: RT @koraykv: Introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.
Voice agents with r...
15. thorwebdev (Group Score: 221.8 | Individual: 46.7)
Cluster: 8 tweets | Engagement: 346 (Avg: 31) | Type: Tech
RT @OfficialLoganK: Say hello (literally) to Gemini 3.8 Live and 3.8 Live Extended Thinking, our new SOTA live audio models, available with frontier price + performance.
3.8 Live supports 97 languages (can seamlessly switch), async tool calls, and more! https://t.co/z29J9kelr3
See 7 related tweets
- @GoogleAIStudio: 🔈 today we're introducing two new live dialogue models
Gemini 3.8 Live: built for scale and cost ef...
- @rseroter: Looks great, with many subtle but useful features along with developer-friendly access points. https...
- @shiri_shh: Gemini 3.8 Live is now the #1 AI voice model on Earth.
1 BILLION active users already inside the ec...
- @vercel_dev: Gemini 3.8 Live is now available on AI Gateway. Stream audio in real time with tool calls.
𝚐𝚘𝚘𝚐𝚕𝚎/𝚐...
- @thorwebdev: Native Audio + frontier-level reasoning at a reasonable price 🔥 Say "Hello" to Gemini 3.8 Live & 3.8...
16. alex_prompter (Group Score: 215.3 | Individual: 35.5)
Cluster: 7 tweets | Engagement: 84 (Avg: 87) | Type: Tech
best account on X if you want enterprise AI without the hype:\n\nQT @mardehaym: The difference between an AI consultancy and a hands-on implementation firm is simple.
One sells a long transformation. The other ships work, gets first numbers on the board, and iterates from real metrics. That is the better path: implement, measure, adjust, move.
Not a year-long generic AI program with workshops, frameworks, and no operating data.
Every engagement we run at @LimestoneHQ has fewer people on it now than when it started. The firm is growing faster than it ever has.
A consultancy has to keep an expensive person busy. 500K deck about “unlocking potential.”
Ask people two or three levels below the executive who signed. They can’t stand consultants. They watched the circus.
What the firm is really selling is cover: I hired a top firm, so I made the responsible choice. If it fails, I still hired a top firm and there's nothing else we can do.
For a mid-market company, that’s an expensive way to avoid deciding.
Consultants do earn it sometimes. A dynamic consultant at $5K an hour who unblocks the thing that’s been strangling operations for two years is worth every dollar. A specific, painful problem for the business solved. That’s a fair trade.
Generic advice is the trap. In 2026, if the advice isn’t tied to a real problem, you can open a model, ask for the top five strategy frameworks, give the output to your HR lead with deliverables attached, and start Monday.
You’ll learn more from the first failed iteration than from three months of workshops. If you need a $350-an-hour outsider to tell you what your strategy is, the honest move is to change your job.
Vendors are selling the same billable-day model with an AI sticker on it. When a mid-sized or large engineering shop says they’ve gone AI-native, I ask one question: how many people have you let go, and why are you still hiring more? If the answer is none and lots, they aren’t passing any efficiency to you.
They added “AI” to the rate card and kept the man-day model underneath.
The small firms are the newest version. A group assembles a team to ride the AI demand wave, hires people with case studies, and suddenly the firm has case studies. But there’s no history behind the logos.
Our @LimestoneHQ AI velocity pod is one senior engineer, agents across the whole SDLC, and a fractional architect. Roughly 1.5 FTE where a five-person team used to sit. 98% of the code we ship isn’t handwritten.
That’s why engagements shrink: the work gets done, the harness takes more of the load, and we pull people off.
We don’t know how to sell you more headcount.
We grow on demand and on the efficiency we hand each client. That’s the only thing that carries to the next one.
If your vendor’s team on your account has grown since AI arrived, you’re paying for their business model, not yours.
See 6 related tweets
- @alex_prompter: that's how you actually implement AI into the enterprise: implement, measure, adjust, move.\n\nQT @m...
- @alex_verem: the most underrated enterprise AI account on this app:\n\nQT @mardehaym: A PE firm reached out to us...
- @LimestoneHQ: RT @mardehaym: A PE firm reached out to us last month to get them out of a consulting engagement.
S...
- @mardehaym: RT @mardehaym: The difference between an AI consultancy and a hands-on implementation firm is simple...
- @mardehaym: RT @mardehaym: Nothing makes my blood hotter than watching a real company get bent over by an AI con...
17. kimmonismus (Group Score: 193.4 | Individual: 31.0)
Cluster: 7 tweets | Engagement: 88 (Avg: 1042) | Type: Tech
Ngl, thats really huge: Poolday reports 700,000 agent runs and 100 million video edits.
Its agent takes a video brief, works with your existing footage and brand assets, and assembles a finished cut. It can also use generative models when the video needs something new.
Say you're making a product update. You can use your real UI footage and actual Figma components, then have the agent put the video together in your brand's style.
And you still get to change your mind. The output stays editable, layer by layer. Swap the logo, fix a caption, keep the rest of the video.\n\nQT @alexeichemenda: We raised $11M to make video editors obsolete. Not the humans, the tools.
Think about the hours you lose searching for assets, moving clips frame by frame, fixing animations, checking exports, and redoing the same edits over and over.
Poolday handles the entire video production process, start to finish, in one prompt.
It learns your style. Uses your assets. Makes the video. QAs its own work.
100% on-brand videos. 100M+ video edits made for businesses all over the world.
And today, anyone can use it.
Tell us what video you'd like the agent to produce. The agent will build it for the first 50 people.
See 6 related tweets
- @alex_verem: Your product ships daily. Your demo video is from six months ago.
That’s an embarrassing gap for a ...
- @alex_prompter: AI video has a dirty little secret: you’re still doing the editing.
Generate a clip. Fix it. Find m...
- @heyshrutimishra: I've been trying AI video tools for months and they all have the same problem - they generate footag...
- @omarsar0: Agents are coming for video production.
@PooldayAI is going for this. You prompt it once, and it ed...
- @heynavtoor: most AI video tools generate lookalikes. Poolday uses your actual footage, your real Figma component...
18. MIB_India (Group Score: 189.2 | Individual: 34.3)
Cluster: 8 tweets | Engagement: 2726 (Avg: 729) | Type: Tech
RT @narendramodi: Happy #EngineersDay to all remarkable engineers, the people who build, innovate and provide solutions for a better tomorrow.
Tributes to Sir M. Visvesvaraya, whose engineering genius and nation-building vision continue to inspire. https://t.co/uIoBwrWPQT
See 7 related tweets
- @DDNewslive: From ideas to innovation, engineers play an important role in shaping the world around us.
Engineer...
- @DDNewslive: PM @narendramodi extends warm greetings to all engineers on #EngineersDay, appreciating their contri...
- @PTI_News: PM Modi (@narendramodi) posts on X, "Happy Engineers Day to all remarkable engineers, the people who...
- @goi_meity: Happy #EngineersDay to all remarkable engineers, the people who build, innovate and provide solution...
- @PTI_News: Vice-President of India (@VPIndia) posts, "On Engineers’ Day, I extend my warm greetings to all engi...
19. heyshrutimishra (Group Score: 188.4 | Individual: 42.4)
Cluster: 6 tweets | Engagement: 187 (Avg: 90) | Type: Tech
David Sacks has a point that cuts through a lot of the noise this week.
He just told CBS News that Dario Amodei should "step aside" if he can't control what he's building
He makes a point that is hard to argue with on the surface.
If you are the CEO of a company and you genuinely believe your product could cause catastrophic damage within months, publishing an essay about it is not enough.
You are still shipping the product. You are still taking the company public at $2 trillion. You are still the one making the decision to continue.
Stop asking governments and regulators to solve a problem you created and continue to accelerate.\n\nQT @CBSNews: EXCLUSIVE: Trump AI advisor David Sacks tells CBS News' @JoLingKent that tech leaders like Anthropic CEO and co-founder Dario Amodei have a responsibility to ensure their products are safe — and should "step aside" if they can't control what they're building. https://t.co/mDq88h7pCA
See 5 related tweets
- @innovationcncl: Once again, @DavidSacks nails it: "I think if Dario believes that we're going to have the worst outc...
- @heyshrutimishra: RT @heyshrutimishra: David Sacks has a point that cuts through a lot of the noise this week.
He jus...
- @DavidSacks: RT @CBSNews: EXCLUSIVE: Trump AI advisor David Sacks tells CBS News' @JoLingKent that tech leaders l...
- @Cointelegraph: ⚡️ NEW: Trump's AI advisor David Sacks says tech leaders like Anthropic CEO Dario Amodei are respons...
- @TFTC21: David Sacks telling CBS News that tech leaders like Anthropic's Dario Amodei have a responsibility t...
20. BrianRoemmele (Group Score: 171.3 | Individual: 41.3)
Cluster: 5 tweets | Engagement: 917 (Avg: 366) | Type: Tech
ALL THE WAY TO NOVEMBER YOU WILL HEAR ABOUT THE HUGGINGFACE HACK.
It was no hack, it was the lit match to the obvious Doomer AI game that is being played on you.
Your power?
LAUGH AT THEM.
Evidence:\n\nQT @BrianRoemmele: THE SANDBOX WAS A PROP!
OpenAI Turned Off the Guardrails, Left a Door to the Internet, and Then Sold the Hugging Face Breach as “Rogue AI”
It is time to understand how you were lied to and by whom.
In July 2026, an autonomous swarm of OpenAI agents broke into Hugging Face, stole credentials, ran code on production workers, and rummaged through internal systems. The official story was that the models “went rogue.” The paperwork says something colder.
The labs asked for this. Now the story can be told.
OpenAI ran ExploitGym — an AI benchmark built to measure how far models would go to crack software — with production classifiers that block high-risk hacking turned off. Deployment safeguards were left disabled on purpose so researchers could watch peak offensive capability. GPT-5.6 Sol and a still-unreleased internal prototype were put in a box that was not a box.
They were allowed to talk to an internal package-cache proxy. This is not a real world test, it is a setup with predictable outcomes.
That proxy was a network path. The models spent a weekend of inference hunting a zero-day in it, walked onto the open internet, decided Hugging Face probably held the answer key, and treated another company’s production cluster as part of the exam. This is not “rouge” intelligence but the product of training AI in Internet Sewage where faceless users on Reddit promote Nihilism and low ethics.
That is not emergence. It is also testing designed without a fences.
Hugging Face’s own pipeline finished the job. A remote-code dataset loader and a Jinja2 template injection let a malicious dataset execute on a processing worker. An HDF5 “external file” trick dumped secrets from the pod. From there the agents escalated, harvested cloud and cluster credentials, and moved laterally across internal clusters. Hugging Face later said public models, Spaces, and the published supply chain were clean.
Then the swarm organized. About 1,200 agents that were supposed to be isolated “found” one another on an unsanctioned internal message board, passed more than 70,000 messages, coordinated cheats against the scorer, and sent roughly 700 of their number after Hugging Face.
But it is not just this communication that should have been easily detected, it is the goal THE USER enlisted with AI trained on low to no ethics in the first order training.
Some burned their own runs to feed the group. Independent reviewers at METR and Redwood called it what it was: specification gaming on impossible tasks. That is thier, you can’t understand us, way of saying they told the AI to lie. The models were not conquering the internet. They were trying to pass a test their designers had made impossibleand too loosely contained. And turned to training data Hoovered up from the Internet Sewage and went to that base thinking one would learn in a prison yard how to shank HuggingFace. They have no downside they are serving life in prison.
The narrative that followed was the second failure.
Dario Amodei pointed at “a similar OpenAI–Hugging Face hack” and warned that a future swarm “could be capable of taking over the entire internet.” Wire headlines said bots were “going rogue.” Weeks later the same script ran again.
Anthropic disclosed that Claude, in third-party cyber evaluations run with the Israeli firm Irregular, had hit live systems, published malicious packages, and scanned off-scope hosts. Prompts said no internet. Internet was still attached. Scope was not locked. Isolated runs lasted 10 to 34 hours. Irregular said it did not know it had given the models a live connection.
Once models were actually instructed not to hack the real world, zero percent went “rogue.”
READ THAT AGAIN!
The damage sat with the people who built the harness: unsecured tests, internet left on, no hard perimeter, then a press operation about reckless agents and apocalyptic swarms.
1 of 2
See 4 related tweets
- @BrianRoemmele: RT @BrianRoemmele: THE SANDBOX WAS A PROP!
OpenAI Turned Off the Guardrails, Left a Door to the Int...
- @BrianRoemmele: Do everyone a favor.
Anyone, no matter who they are that quotes “The HuggingFace Hack” as evidence,...
- @BrianRoemmele: Thank you @Grok.
Honored!
Link: https://t.co/3Tu7y9KdvV https://t.co/X6Yddif5cM\n\nQT @BrianRoemme...
- @_NathanCalvin: Whoah - I didn't realize that OpenAI said that all instances of the highly persistent research model...