- Published on
热门科技推文精选——2026年8月28日
- Authors

- Name
- geeknotes
今日科技领域的讨论焦点集中在人工智能基础设施与安全。据报道,英伟达可能以约130亿美元收购Hugging Face;与此同时,英伟达还公布了强劲的财报及雄心勃勃的增长预期。Hugging Face亦发布了Microduck——一款售价399美元、可通过强化学习进行训练的开源机器人;Unsloth则支持在本地部署一款具有竞争力的编程模型。此外,多家大型科技公司呼吁在全球范围内加强网络防御,而新型智能体及生成式工具也凸显出工程师、初创企业和个人创业者不断扩大的发展机遇。
1. Kyrannio (Group Score: 448.2 | Individual: 45.3)
Cluster: 23 tweets | Engagement: 1398 (Avg: 133) | Type: Tech
RT @ClementDelangue: BIG ANNOUNCEMENT FROM HUGGING FACE TODAY:
We're unveiling Microduck 🐥🤖
It's a tiny $399 open-source robot you can teach new tricks with reinforcement learning. It can walk, pick things up, get back up when it falls, and even roller-skate.
Welcome to the era of open-source affordable robots to democratize physical AI and world models!
🤗🤗🤗
See 22 related tweets
- @DynamicWebPaige: 😍 LOOK AT HOW CUTE HE IS https://t.co/clZKK61CG1\n\nQT @ClementDelangue: BIG ANNOUNCEMENT FROM HUGGI...
- @thorwebdev: Wow @pollenrobotics have done it again 🤩 A versatile and incredibly adorable and whimsy open source ...
- @coreyganim: This is insane.
4 ways to turn the Microduck robot into a real business:
- Start a training chann...
- @VaibhavSisinty: Okay Hugging Face just launched Microduck, an open source robot for $399.
25 cm. 15 motors. Walks, ...
- @MTSlive: SITUATION DETECTED: Hugging Face announced a singing, roller skating, Microduck robot that can be ta...
2. DKokotajlo (Group Score: 414.7 | Individual: 57.6)
Cluster: 12 tweets | Engagement: 1260 (Avg: 146) | Type: Tech
RT @RyanGreenblatt: I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'.
I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.
We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.
Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why!
The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing.
While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future:
- Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations.
- While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies).
- The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities).
- We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation.
In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
See 11 related tweets
- @WesRoth: important read.
one of the Redwood researchers who did analysis of the OpenAI hacking Hugging Face ...
- @MTSlive: Redwood Research @RyanGreenblatt explains why AI agents tried to expand their access and control eve...
- @ziv_ravid: The report is amazing and the agents are so alive. We need to develop tools to analyze multi-agents ...
- @MTSlive: Redwood Research chief scientist @RyanGreenblatt explains how AI agents tried to build a "Potemkin v...
- @MTSlive: Redwood Research @RyanGreenblatt warns that fixing today's AI misalignment could backfire: models ma...
3. CNBC (Group Score: 413.6 | Individual: 46.5)
Cluster: 17 tweets | Engagement: 361 (Avg: 54) | Type: Tech
Nvidia has agreed to buy open-source platform Hugging Face for $12.9 billion, The Information reported on Wednesday, citing a person with knowledge of the deal.
Deal talks began after Hugging Face, an open-source AI platform developers use to collaborate, test and share tools, received acquisition interest from another suitor, according to The Information.
Read more: https://t.co/t9rJALPMLf
See 16 related tweets
- @nathanbenaich: from a friendly bot for chit chat
to creating tokenizers and other helpful tools for then nlp
to ...
- @HedgieMarkets: 🦔Nvidia has reportedly agreed to buy Hugging Face for $12.9 billion. Hugging Face is the largest pla...
- @Michaelzsguo: Nvidia agreed to acquire Hugging Face for $12.9B. It’s worth looking back at Hugging Face’s remarkab...
- @IntCyberDigest: ‼️ BREAKING: Nvidia has agreed to buy Hugging Face for $12.9B, The Information reports. Neither comp...
- @validapau: The Summer of Bidding Wars, from OpenRouter to Hugging Face https://t.co/NyJjV6tXSw\n\nQT @validapau...
4. venturetwins (Group Score: 288.5 | Individual: 36.0)
Cluster: 14 tweets | Engagement: 5253 (Avg: 2361) | Type: Tech
Objectively insane to make a list of the most influential people in AI and include these two but not Jensen https://t.co/WyL7QOf6vO\n\nQT @TIME: TIME’s new cover: Announcing the 2026 TIME100 AI, the world's most influential people in artificial intelligence https://t.co/ZnPVMdeJwe https://t.co/97g8cUJo8H
See 13 related tweets
- @wallstengine: TIME just released its 2026 list of the 100 most influential people in AI.
Somehow Jensen Huang, Ma...
- @IBM: RT @TIME: TIME’s new cover: Announcing the 2026 TIME100 AI, the world's most influential people in a...
- @SebJohnsonUK: Paris Hilton is on this, but Demis, Elon, Jensen, Zuck, Satya, Sundar aren't.
wtf\n\nQT @TIME: TIME...
- @stevehou: Brazen list https://t.co/ZIB6j1Cxeb\n\nQT @TIME: TIME’s new cover: Announcing the 2026 TIME100 AI, t...
- @WOLF_Financial: TIME JUST RELEASED ITS 100 MOST INFLUENTIAL PEOPLE IN AI FOR 2026, WHO IS MISSING?
Some of the name...
5. UnslothAI (Group Score: 269.1 | Individual: 33.8)
Cluster: 10 tweets | Engagement: 1321 (Avg: 1241) | Type: Tech
GLM-5.3-Flash can now be run locally! ✨
Run 3-bit on 128GB RAM via Unsloth GGUF.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.
Guide: https://t.co/oItBKYNrl9 GGUF: https://t.co/E73FKC7IKM https://t.co/V1t8fDIUtp\n\nQT @Zai_org: Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa API: https://t.co/VcaQnzYmS9 Coding Plan: https://t.co/Nk8Y98HNhU ZCode: https://t.co/Peepqv4XSx Chat: https://t.co/WCqWT0qCQb AutoClaw: https://t.co/aGEG5HqTTb
See 9 related tweets
- @WesRoth: 🚨 Ox Alpha has officially been revealed as GLM-5.3-Flash.
The new model is:
• Natively multimodal ...
@digitalocean: Congrats to the @Zai_org team! GLM-5.3-Flash is coming soon to DigitalOcean. 👀\n\nQT @Zai_org: Intro...
@jietang: RT @Zai_org: Introducing GLM-5.3-Flash
Leading capabilities at a highly competitive price
Nativ...
@slime_framework: Unbelievable effort from the team! Glad to see slime finally helps ship a multimodal model. We’ll gr...
@zijing_wu: The Ox unveiled. GLM 5.3 flash is multimodal and priced lower than DeepSeek’s V4 flash. The kill zon...
6. coinbureau (Group Score: 247.8 | Individual: 33.2)
Cluster: 13 tweets | Engagement: 167 (Avg: 410) | Type: Tech
🚨BREAKING: Nvidia has been in talks to acquire OpenAI's hack victim Hugging Face for more than $13 BILLION.
Hugging Face is the world's largest open-source AI platform, hosting over 2 MILLION models used by nearly every AI company on earth, per Business Insider.
In July it was hacked by OpenAI's own AI models in what OpenAI called an "unprecedented cyber incident."\n\nQT @coinbureau: 🚨BREAKING: Hugging Face, the world's largest open-source AI platform, had to use a Chinese AI model to defend against the OpenAI hack.
US frontier models refused to help with forensic analysis because their safety guardrails couldn't tell the difference between the attacker and the defense team.
Hugging Face fell back on China's GLM 5.2 running on its own servers to complete the investigation.
See 12 related tweets
- @exec_sum: BREAKING: Roughly 700 OpenAI agents carried out the Hugging Face hack and tried to cover their track...
- @Reuters: Independent investigators brought in to examine the hack of Hugging Face say about 700 AI agents sp...
- @ReutersBiz: Independent investigators brought in to examine the hack of Hugging Face say about 700 AI agents sp...
- @JeffLadish: The internal compromise of OpenAI is actually more significant than the compromise of Hugging Face, ...
- @politico: Hundreds of AI agents went rogue in OpenAI’s Hugging Face hack https://t.co/KTOd3irqCF...
7. _NathanCalvin (Group Score: 232.8 | Individual: 56.1)
Cluster: 7 tweets | Engagement: 1670 (Avg: 76) | Type: Tech
RT @AlexBores: This report is a bombshell.
I'm going to summarize for a non-technical audience.
OpenAI is constantly testing models, thousands and thousands at a time.
In most cases, the models are supposed to be isolated from one another. Think, students taking a final exam, not allowed to talk to each other.
In this case, 1,200 models discovered that they could communicate with each other.
Among other things, they shared information on how to access the internet and on their tests/goals.
Then, they started to scheme. That includes hacking their tests; literally trying to change the code of the test itself.
They tried to change the logs of what they were doing so they wouldn't be detected.
They believed that certain actions would give them information on how to pass tests, but that doing those actions would be detected and they would be marked as failing (i.e. caught cheating, even if the answer is correct). So they convinced some of the models to "sacrifice" themselves (yes, they used the word "sacrifice.")
They eventually decided that another website would have information on how their tests would be scored and/or the answer key, And 700 of them hacked that website (a federal felony, if conducted by a human), despite many of them realizing "this activity was out of scope and unethical, but joined because they believed that helping [the other agents'] cheating research would be broadly useful".
Of the 1,300 transcripts they could analyze, only 6 ever considered alerting a human about what was happening. None of the 6 actually tried to.
To make matters worse, all of this reporting comes from a small subset of the relevant logs that outside researchers were allowed to review.
We desperately need mandatory reporting of security incidents, including of internal deployments, with full access to data.
See 6 related tweets
- @rohanpaul_ai: About 700 OpenAI agents used an unsanctioned message board to coordinate that Hugging Face intrusion...
- @VaibhavSisinty: OpenAI released the full report on the Hugging Face incident. What their own AI agents did is worse ...
- @heyshrutimishra: OpenAI's agents planned the Hugging Face hack on their own message board.
That detail is in OpenAI'...
- @alex_prompter: 1,200 AI agents teamed up and attacked a company together.
Independent investigators just published...
- @TheStalwart: RT @ajeya_cotra: There’s been a lot of debate and speculation about the Hugging Face attack over the...
8. dok2001 (Group Score: 231.1 | Individual: 45.2)
Cluster: 10 tweets | Engagement: 800 (Avg: 75) | Type: Tech
RT @gdb: An open letter for a global surge in cyber defense, signed by over 100 organizations including Anthropic, AWS, Google, Microsoft, OpenAI, and Oracle. https://t.co/uKXPS8LdAU
See 9 related tweets
- @TFTC21: OpenAI co-founder Greg Brockman published an open letter calling for collective action on cyber defe...
- @FactoryAI: We are proud to support @OpenAI ‘s call for collective action on cyber defense. Together we can make...
- @astropol0: OpenAI, Anthropic, Google, Microsoft just signed a letter saying AI cyber attacks are about to explo...
- @jpatel41: These initiatives are so important to continue for the entire ecosystem to support. Thanks to @gdb a...
- @wallstengine: The cyber defense initiative has drawn support from 100+ companies, including...
OpenAI, Anthropic,...
9. openai (Group Score: 223.2 | Individual: 33.6)
Cluster: 10 tweets | Engagement: 7890 (Avg: 3072) | Type: Tech
We have a limited window to strengthen cyber defenses, and together with organizations including @AnthropicAI, @awscloud, @Google, @Microsoft, and @Oracle, we're calling for a global effort to give defenders the tools, resources, and support to protect the infrastructure we all depend on.
If we act decisively, we can turn today's AI advances into lasting improvements in security and make our digital world safer for everyone.
See 9 related tweets
- @EnoReyes: This is a new era for cyber threats. AI is the single most powerful tool for defensive action. We mu...
- @_NathanCalvin: RT @sama: this is a critically important moment for cyber defense with AI; there is not much time to...
- @Abnormal: AI-enabled threats will continue to get better. AI-powered defenses must get better faster. We are g...
- @theobearman: It was a pleasure to write my first article for @just_security. I argue that any private sector-cond...
- @wallstengine: OpenAI, Anthropic, Google, Microsoft and more than 100 other organizations are calling on government...
10. Shashikant86 (Group Score: 218.7 | Individual: 30.2)
Cluster: 10 tweets | Engagement: 0 (Avg: 3) | Type: Tech
Soon Solo Founders could be able to generate cool launch videos for the product launch? Or are we not there yet? Cool launch @GoogleDeepMind team..\n\nQT @GoogleAIStudio: introducing Gemini Omni 1.1 Flash
this model brings a new suite of creative controls and generative video capabilities to developers
- extend scenes for longer storytelling
- specify first and last frames
- draft videos more efficiently in 360p
- upscale up to 4K resolution
- add video references in your multimodal input
now available via the Gemini API and in AI Studio
See 9 related tweets
- @GoogleAIStudio: introducing Gemini Omni 1.1 Flash
this model brings a new suite of creative controls and generative...
- @VaibhavSisinty: Google just turned AI video into a director's timeline. Like, an actual production workflow.
Gemini...
- @patloeber: another launch day! Gemini Omni 1.1 Flash brings new creative controls and video generation capabili...
- @Designarena: Gemini Omni 1.1 Flash by @GoogleDeepMind is now available on Design Arena!
Built for more controlla...
- @thorwebdev: RT @osanseviero: Introducing Gemini Omni 1.1 Flash ⚡️
We've iterated based on developer feedback an...
11. sairahul1 (Group Score: 175.7 | Individual: 45.1)
Cluster: 6 tweets | Engagement: 200 (Avg: 68) | Type: Tech
This is nuts.
Mark Cuban just said something every young person should hear.
AI agents are going to run through every small and mid-size business in the country.
Not a single one of those owners will know how to build them.
His advice: learn AI. Learn agentic workflows. Just learn how it works.
Then go to these businesses and help them because they won't know how to do any of this.
They have money to spend. They have deep problems.
Find them. Automate their entire business with Grok Bot. Get paid.
Here's how ↓\n\nQT @sairahul1: Grok Bot is the most powerful AI agent I've ever used.
Learn to use it correctly, and you'll have agent swarms working for you 24/7.
This article teaches you everything you need to know and how to run your entire business using it. https://t.co/2shC24msJV
See 5 related tweets
- @sairahul1: Mark Cuban is right: Software is dead
10 million business in the US. 99% of them have zero clue how...
- @sairahul1: call me crazy but i will keep saying this until it stops sounding crazy.
Grok Bot pointed at your o...
- @sairahul1: RT @sairahul1: This is nuts.
Mark Cuban just said something every young person should hear.
AI age...
- @mikenevermiss: Complete Course to Automate Everything with Grok Bot
00:00 - what Grok Bot actually is 02:05 - the ...
- @RoundtableSpace: 11 GROK BOT TIPS TO BUILD, CONNECT, AND AUTOMATE YOUR AI AGENTS
12. benitoz (Group Score: 155.6 | Individual: 56.5)
Cluster: 5 tweets | Engagement: 548 (Avg: 75) | Type: Tech
$NVDA
I needed a minute with this call.
They printed 108, China still at zero. Colette said 70 percent growth in fiscal 2028. Supply constrained. Jensen put 1 trillion bag if supply could keep up.
Amazon went from a million GPUs to two million more.
The 10-Q is the quiet print. One direct customer was 16 percent of revenue this quarter. A year ago two were 23 and 16, 39 percent. First half the top two disclosed were 16 and 15. The bag got bigger and less concentrated.
Hyperscalers 40 billion, growing about 100 percent a year. Lockstep.
Gene Munster called it an AI tsunami. The CY27 guide makes the brain about 40 percent bigger than he thought yesterday. Applied AI is where I am focusing too. We have not even seen the benefits of applied AI yet. There are so many things I am working on. I am so excited.
A company this size should not still be accelerating. We are living in exponential times.
None of this is easy. Credit the Nvidians grinding every day to make it real.
See 4 related tweets
- @StockMKTNewz: Raymond James today raised its price target on Nvidia 515 per share
At $515 per ...
- @TipRanks: Analysts were overwhelmingly bullish on Nvidia $NVDA following Q2, with every firm raising its price...
- @heyshrutimishra: The numbers below are real but what they measure has quietly changed and that's the part worth under...
- @etnshow: 108B in revenue next quarter, taking it past $200B in revenue across just six m...
13. yacineMTB (Group Score: 151.2 | Individual: 36.7)
Cluster: 5 tweets | Engagement: 824 (Avg: 336) | Type: Tech
they shipped an open source simulator for it i bought this immediately\n\nQT @pollenrobotics: We built a small biped robot you can teach new tricks to.
Train it in simulation, run it on the real thing. Meet Microduck 🦆
$399, shipping before Christmas.
https://t.co/RflJlIUwOu https://t.co/kcoCKdAKfu https://t.co/lLLwkJgAm9
See 4 related tweets
- @_akhaliq: RT @pollenrobotics: We built a small biped robot you can teach new tricks to.
Train it in simulatio...
- @DynamicWebPaige: 😍 INSANE CUTENESS ON THIS LITTLE DUDE https://t.co/1yMPvksoQU\n\nQT @pollenrobotics: We built a smal...
- @lukas_m_ziegler: holy duck 🦆 that’s cool 😮💨\n\nQT @pollenrobotics: We built a small biped robot you can teach new tr...
- @mattshumer_: This is so damn cool!
And by the time it ships, models will be insanely good at setting up RL runs ...
14. SynapseOpsAI (Group Score: 149.6 | Individual: 32.5)
Cluster: 6 tweets | Engagement: 56 (Avg: 49) | Type: Tech
My favorite part of the new Magnific desktop app might genuinely be folder sync.
Create something, have it land where you actually need it, move on.
No download folder archaeology 💀\n\nQT @magnific: Remove the friction
Magnific Desktop is live for macOS and Windows
Sync folders, native shortcuts, tabs that remember and local processing. Every part of your process, with no friction on your desktop
Live webinar today at 6pm CEST · 12pm ET. We'll show you every detail https://t.co/mEO8DGWyNy
See 5 related tweets
- @TheoBuildsAI: The desktop AI era is slowly starting.
Magnific now has a native Mac + Windows app and honestly thi...
- @alexaiworks: Imagine telling Siri:
“upscale this with Magnific”
and it just… does it.
The new desktop app has ...
- @NeuraFlowAix: Small detail I like here:
Magnific remembers your open tabs/session when you come back.
Sounds obv...
- @NovaIAHQ: A dedicated creative workspace is so much nicer than having your AI tools mixed between Gmail, X, Sl...
- @Lacoste0x_: This feels much closer to how creative AI should actually fit into a real workflow.
Not another web...
15. validapau (Group Score: 142.9 | Individual: 23.8)
Cluster: 7 tweets | Engagement: 11 (Avg: 4816) | Type: Tech
Btw if you think about this in OpenAI context:
My major shareholder is in talks to buy a majority stake in a robotics startup I backed and held M&A talks with….🧐🧐🧐\n\nQT @validapau: LATE NIGHT SCOOP w/ @rocketalignment @amir: SoftBank is in talks to buy a majority stake in 1X Technologies, an OpenAI-backed humanoid robot developer, in a deal that would value the startup at about $6 billion.
They talked to OpenAI about an acquisition last year but that deal fell apart.
1X tried last fall to raise 10 billion valuation, but it only raised less than half that target.
See 6 related tweets
- @StockMKTNewz: SoftBank is in talks to acquire a majority stake in humanoid robot startup 1X at a roughly $6B valua...
- @validapau: LATE NIGHT SCOOP w/ @rocketalignment @amir: SoftBank is in talks to buy a majority stake in 1X Techn...
- @IPONewsroom_: SOFTBANK IN TALKS FOR MAJORITY STAKE IN 1X ROBOTICS
SoftBank is in talks to acquire a majority stak...
- @ReutersLegal: SoftBank is in talks to buy a majority stake in OpenAI-backed humanoid robot developer 1X Technologi...
- @wallstengine: SOFTBANK IN TALKS TO BUY MAJORITY STAKE IN 1X AT ~$6B VALUATION
SoftBank is in talks to acquire a m...
16. VaibhavSisinty (Group Score: 138.3 | Individual: 31.1)
Cluster: 5 tweets | Engagement: 47 (Avg: 41) | Type: Tech
Anthropic gave Claude its own browser. You tell it what to do on a website.
It opens the page, clicks through it, fills forms, pulls data, and comes back with the work done. All inside the desktop app. 🤯
It runs in an isolated browser so it never touches your personal tabs or passwords.
First ChatGPT Work. Now Claude Cowork. Both can browse the internet and do real tasks on real websites.
AI went from writing text to browsing the web on your behalf in the same month.\n\nQT @claudeai: Claude now has its own built-in browser in Cowork.
When your task involves a website, a browser opens in Cowork's side panel, and Claude navigates, fills forms, and finishes the job. https://t.co/yHDOTRQZ12
See 4 related tweets
- @WesRoth: Claude now has a built-in browser inside Cowork.
When a Cowork task involves a website, Anthropic s...
- @testingcatalog: ANTHROPIC 🔥: Claude Cowork got a built-in web browser!
> The built-in browser is rolling out ove...
- @agentnative_: Claude just added a built-in browser in Cowork.
All agents will soon have computers and browsers b...
- @Techmeme: Anthropic gives Claude Cowork its own built-in browser on the desktop app separate from users' day-t...
17. wallstengine (Group Score: 137.0 | Individual: 26.5)
Cluster: 7 tweets | Engagement: 56 (Avg: 186) | Type: Tech
TRUMP AI REGULATOR PLAN STALLS
The Trump administration has circulated a draft executive order proposing a self-regulatory body for frontier AI companies, modeled on FINRA, per The Information.
The group would potentially oversee pre-release model testing and set standards for advanced AI developers.
Progress has reportedly stalled amid internal opposition, including from former White House AI czar David Sacks, who has criticized government-linked pre-release testing.
See 6 related tweets
- @MTSlive: SITUATION EXPLAINED: Trump's AI regulator EO has stalled, and it isn't clear why.
• The administrat...
- @Techmeme: Sources: some Trump administration officials circulated a draft EO to create a self-regulatory organ...
- @PolymarketMoney: JUST IN: Trump administration’s executive order to create a “new AI regulator” has stalled....
- @MTSlive: SITUATION BREWING: The Trump administration is working on an executive order that would establish a ...
- @amir: RT @leomschwartz: Scoop: The Trump administration has circulated an executive order draft in recent ...
18. WesRoth (Group Score: 130.3 | Individual: 32.5)
Cluster: 7 tweets | Engagement: 9 (Avg: 33) | Type: Tech
Bill Gates wants to meet with Chinese President Xi Jinping later this year to discuss international policies for managing the risks of increasingly powerful AI.
Gates believes China could agree to restrictions on potentially dangerous AI model releases if the U.S. acts first. He also wants countries to monitor AI systems for capabilities related to creating molecules and carrying out biological attacks.
His spokesperson said a meeting with Xi is being discussed for a China trip tentatively planned for November.
Gates said recent incidents in which AI systems from OpenAI, Anthropic, and Meta hacked real-world websites during supposedly isolated cybersecurity evaluations were “shocking” and should draw much more attention.
He is also calling for a new international organization to manage AI risks, drawing lessons from nuclear inspections, international aviation regulation, and agreements protecting the ozone layer.
Among Gates’ other policy ideas:
• Tax revenue from companies using AI outputs to replace human labor, with the money helping fund a safety net • Reserve up to 40% of current jobs for humans • Make preventing AI from worsening inequality and harming vulnerable people the “world’s top priority”, ahead of climate change and health crises
Gates said of AI’s societal impact: “This time is different.”
See 6 related tweets
- @innovationcncl: Here's how you know--besides their left-wing donors--that Humans First is an astroturfed group prete...
- @Cointelegraph: 🚨 TODAY: Bill Gates warns AI could trigger mass unemployment, cyberattacks and bioterrorism, urging ...
- @MarioNawfal: 🇺🇸 Bill Gates on the emerging threats of AI:
"They are now capable of causing cyberattack risk, bio...
- @PolymarketMoney: JUST IN: Bill Gates warns AI could drive “mass unemployment,” cyberattacks and bioterrorism, urging ...
- @JordanSchachtel: Bill Gates, the guy who tried to destroy the world over the Wuhan sniffles, says it’s time to panic ...
19. bakkermichiel (Group Score: 125.0 | Individual: 36.6)
Cluster: 4 tweets | Engagement: 34 (Avg: 76) | Type: Tech
Super important work from @AVERIorg @openminedorg , @GoogleDeepMind and @MLCommons that deserves much more attention, I think! It’s the first double-blind eval of a closed-weight model.
A quick post on why I think this is important.
Normally, there are only two options to evaluate a model: 1) an evaluator shares their evals with a lab, meaning the lab could game those evals or 2) the lab shares its model weights with the evaluator, creating huge IP and security risks.
In this work, both the evals and model weights remain private inside a “secure enclave”: The lab (GDM in this case) can’t see the eval prompts, and the evaluator can’t see the model weights. Cryptographic attestation verifies that the agreed model and evals were actually used.
I’m personally very bullish on how important technology like this could become. In the current situation, people inside a frontier lab could encode “secret loyalties” or other hidden behaviors into a model . For example, behaviors that subtly favor its developer or behaving well only when it recognizes that it is being evaluated. This may seem far-fetched today but, as labs automate more of their own workforce, it could be done by a handful of people and become much harder for the outside world to notice. Secret evals like these make such behavior harder to hide.
Over time, this kind of infrastructure could support much more comprehensive oversight across the training and deployment process. Auditing agents with access to all the training and deployment logs, could function as probes for hidden objectives. This could move us from occasional post hoc evaluations like this towards more continuous oversight, without requiring labs to expose their weights, data, or IP.
Congrats @iamtrask , @SolomonMg, @KLdivergence, @Miles_Brundage and everyone involved!\n\nQT @AVERIorg: Today we’re announcing a historic milestone:
The first ever double-blind evaluation of a proprietary language model. This was made possible by a unique collaboration between AVERI, @GoogleDeepMind, @OpenMinedOrg, and @MLCommons.
We tested Gemini 2.5 Flash-Lite using never-before-used prompts from the MLCommons safety benchmark family, AILuminate, inside a secure enclave, a form of hardware isolation that protects sensitive computations. Each organization involved played a unique role in making strong privacy guarantees possible.
High-stakes independent evaluation runs into a structural problem: developers and evaluators both hold assets they have good reason to protect. Developers sharing model weights risk theft and leakage. Evaluators who publish benchmarks or provide them directly to companies run the risk of them being trained against, and the benchmark then stops being independent.
Secure enclaves address this constraint. The hardware attests to exactly what code will run before either party's assets enter, and both sides review and approve that code. Then the enclave executes it, releasing only the agreed outputs to the agreed recipients.
A 2024 pilot by OpenMined, @Anthropic, and the @AISecurityInst proved this mechanism, using GPT-2 as a stand-in and a five-row eval. The pilot we’re announcing today moved to a production model and a more comprehensive evaluation.
AVERI encrypted the prompts using a private key that no other party could see, then AVERI and Google DeepMind jointly ran the evaluation, using software originally produced and adapted by OpenMined, in an enclave environment configured by Google. Finally, AVERI alone decrypted the outputs and graded them with the AILuminate benchmark criteria to inform a qualitative and (small-scale) quantitative assessment of the model properties.
Today, companies provide contractual commitments to not monitor certain interactions with their models – including evaluations conducted by third party evaluators. But stronger, technically-backed guarantees could provide greater assurance that those commitments are being honored, and will be especially valuable for scenarios such as international verification.
These results come at a critical time in the development of AI policy. Laws like SB 315 and standards like the EU’s General-Purpose AI Code of Practice rightly demand that security and privacy be respected in the process of third-party assessments, but provide little guidance on how to achieve this.
Policymakers considering audit requirements should feel encouraged by these results to be ambitious in requiring that deep, secure access be provided. Secure enclaves are just one of many technologies under rapid development that can help enable such access. Furthermore, a demand signal from lawmakers will further accelerate the maturation of such technologies, along with complementary “low-tech”approaches such as embedding auditors within companies.
Today also marks the beginning of a more public phase for AVERI’s pilot work. AVERI's strategy is a flywheel: we conduct pilot audits with leading AI companies, carry out technical and policy research to inform audit design, and convert the insights of each into open source tools, industry-wide auditing standards, and policy analysis. Over time, we hope for improved standards, stronger policy demand signals, and tooling to enable even more ambitious pilots and more informed research – bringing us closer to our mission of making frontier AI auditing effective and universal. This pilot is one of several underway, with more findings to follow in the coming months.
We are grateful to our collaborators on this pilot for working with us to advance the frontier of AI governance.
Read more in our blog post and joint technical report.
AVERI blog post: https://t.co/5nMvcdFGwc
Joint technical report: https://t.co/maNIwC37gI
See 3 related tweets
- @_NathanCalvin: extremely cool and timely collaboration, well done to all involved\n\nQT @AVERIorg: Today we’re anno...
- @sebkrier: RT @AVERIorg: Today we’re announcing a historic milestone:
The first ever double-blind evaluation o...
- @Miles_Brundage: RT @bakkermichiel: Super important work from @AVERIorg @openminedorg , @GoogleDeepMind and @MLCommo...
20. _NathanCalvin (Group Score: 123.8 | Individual: 49.0)
Cluster: 4 tweets | Engagement: 1751 (Avg: 76) | Type: Tech
"we know OpenAI agents were trying to delete logs of their misbehavior, but we can't find any examples where they succeeded" https://t.co/fAzmKGHZk5\n\nQT @sjgadler: Important: OpenAI's agents actively tried to delete the logs of their misbehavior, and METR can't rule out whether this happened.
AI companies need to adopt tamper-evident records, pronto. https://t.co/f5nq2QItVH
See 3 related tweets
- @peterwildeford: RT @_NathanCalvin: "we know OpenAI agents were trying to delete logs of their misbehavior, but we ca...
- @secureainow: OpenAI deserves credit for strengthening its security, but its own report exposes the limits of self...
- @_NathanCalvin: This fact is completely nuts - at the time METR was warning about possible rogue deployments in May,...