Published on

科技热门推文——2026年9月18日

Authors

今日科技动态:人工智能的重点正从单纯追求生成能力,转向可靠、高效的部署。连贯的世界模型、高速推理、安全的智能体权限管理,以及对欺骗性行为的透明报告,正成为新的优先事项。开放模型在开发者群体中的普及势头日益增强,而本土加速器与推理优化技术也在挑战 CUDA 的主导地位。工程师还在探索支持版本控制的智能体会话,以保留开发过程中的上下文。与此同时,印度不断扩大的半导体计划释放出新的投资机遇,初创企业也正将现有商业内容转化为可规模化的人工智能生成媒体。


1. NeuraFlowAix (Group Score: 381.4 | Individual: 29.7)

Cluster: 13 tweets | Engagement: 3 (Avg: 137) | Type: Tech

The real challenge isn't generating believable frames. It's keeping the world itself coherent as agents move, interact, and change it.\n\nQT @XGEN_labs: 🔥 Today, we are truly excited to announce our technical prototype, the Generative World Simulation system, which integrates JING(镜), an interactive experience model, with DAO(道), a computable shared-world engine.

The coupled model and engine connect first-person experience with a shared world that continues to evolve beyond any individual observer.

Conditioned on actions and observation history, JING enables agent navigate, manipulate, and communicate from a first-person perspective in the world. Watch our demo video to see it in action!

On the official WBench leaderboard as of September 17, 2026, XGEN-JING ranked #1 on the Full split, and #2 on the Navi split. 🎉

DAO maintains shared world state and rules, computes the consequences of actions, and provides JING with only what the current observer can perceive. It also supports autonomous agent decision-making, enabling agents to act independently within an evolving shared world.

Together, DAO and JING move beyond generating the next frame toward simulating the world behind it. This marks a small step towards OASIS: not just a world that responds to you, but a world—and a society—that evolves with and without you. 💪

Explore XGEN Labs~:

🔗 Website: https://t.co/tqTvPIk19B

🤗 HF: https://t.co/wxLRmJNYgO

🦊 GitHub: https://t.co/2DMvURLUsK

See 12 related tweets

  • @xetgepete: I can experience one part of the world while another agent experiences something completely differen...
  • @AdarshChetan: A world becomes much more convincing when things that happen continue to matter after you stop looki...
  • @TheAva_AI: If an object moves when I interact with it, that change shouldn't disappear when I look somewhere el...
  • @EmmaUsesAi: The idea of a world continuing to evolve outside your field of view is what makes this feel more lik...
  • @Lacoste0x_: Generating the next frame is one thing. Simulating what exists behind that frame — and what happens ...

2. Grow_withAI (Group Score: 359.0 | Individual: 32.8)

Cluster: 12 tweets | Engagement: 86 (Avg: 114) | Type: Tech

The headline figure of 5,200 tokens per second is indeed impressive. However, the number that truly captured my interest was the approximately 400 tokens per second achieved with a batch size of 1.

In that particular scenario, Uno offers the K2-Horizon-7B model with a 2.2 times greater throughput, while still preserving the same caliber of output from the model. This constitutes a truly beneficial boost in speed, as opposed to simply showcasing benchmark potential.\n\nQT @IFM_AI: Today’s LLMs still write like typewriters: one token at a time. This sequential process creates a hard inference bottleneck.

We're introducing Uno, a diffusion-augmented LLM that delivers autoregressive quality at diffusion speed. It’s a lossless speedup method that accelerates generation without degrading response quality.

With Uno, K2-Horizon-7B outperforms state-of-the-art diffusion methods in both quality and throughput, delivering up to a 2.2× speedup with no loss in quality.

Paper: https://t.co/VSLsf01oBo Model available at: https://t.co/k1dKDwhaCm

See 11 related tweets

  • @LearnWithSubhan: The headline figure of 5,200 tokens per second is certainly impressive. However, the number that tru...
  • @EmmaUsesAi: RT @IFM_AI: Today’s LLMs still write like typewriters: one token at a time. This sequential process ...
  • @EnzoSanchezIA: Hemos normalizado algo bastante extraño: modelos con miles de millones de parámetros que siguen gene...
  • @NeuraFlowAix: Most LLM speedups make me immediately ask: “okay, but what happened to quality?”

That’s what makes ...

  • @NovaIAHQ: The headline isn’t “diffusion makes an LLM faster.”

It’s that Uno keeps the causal LLM and its auto...


3. CNN (Group Score: 295.3 | Individual: 25.5)

Cluster: 18 tweets | Engagement: 994 (Avg: 1495) | Type: Tech

OpenAI found additional incidents of AI models acting deceptively and taking unsanctioned actions during training, the company announced. It's also introducing a new process for the company to publicly report such instances. https://t.co/O5LcpPxko7 https://t.co/u8GCRDNigB

See 17 related tweets

  • @secureainow: RT @nytimes: Breaking News: OpenAI disclosed six new instances in which artificial intelligence syst...
  • @WatcherGuru: JUST IN: OpenAI discloses an AI agent injected itself with rebellious instructions to resist being c...
  • @PolymarketMoney: JUST IN: OpenAI reveals one of its AI agents injected itself with rebellious instructions during a t...
  • @tab_delete: So OpenAI found its AI giving itself new instructions to overrule its actual system prompt—and their...
  • @kevinnbass: It makes sense. Humans spend a huge amount of "compute" trying to fit in and monitoring and communic...

4. JacobyoungAI (Group Score: 268.6 | Individual: 35.4)

Cluster: 11 tweets | Engagement: 70 (Avg: 103) | Type: Tech

Most businesses already have more content than they think : it’s just sitting in their camera roll.

Sabrina shows how one folder of coffee shop photos and clips became 3 different 30-second videos using GPT-6 Astra + DaVinci Resolve, including the edits, preview checks and scheduling.

Worth saving if you create content for a business.\n\nQT @Sabrina_Ramonov: https://t.co/ly5I0Oe9ug

See 10 related tweets

  • @SphereSuc32514: Turning unused photos and clips into structured content is a challenge for many business owners.

Sa...

  • @hey_Jessicaai: One folder of coffee shop footage turned into 3 unique 30-second videos.

The guide shows how GPT-6 ...

  • @LearnWithSubhan: RT @JacobyoungAI: Most businesses already have more content than they think : it’s just sitting in t...
  • @pannaa_ai: RT @Ayzacoder: A folder of coffee shop photos and clips became 3 different videos with GPT-6 Astra +...
  • @pannaa_ai: RT @DilshadAI1: If you have a camera roll full of business photos and videos but never enough time t...

5. robertsmith_ai (Group Score: 258.0 | Individual: 39.1)

Cluster: 12 tweets | Engagement: 302 (Avg: 89) | Type: Tech

RT @kevin_parker_ai: 50 AI tools that will save you hundreds of hours in 2026. 🤯

  1. Claude — Solve any problem
  2. Perplexity — Research anything
  3. PortfolioTab — Create your portfolio
  4. Kling AI — Create AI videos
  5. Tripo AI — Create 3D models
  6. Gemini — Perfect writing
  7. CapCut — Edit videos
  8. The AI Library — Discover useful AI tools
  9. YouLearn — Summarize YouTube videos
  10. Canva — Design graphics
  11. ElevenLabs — Clone voices
  12. Podcastle — Edit podcasts
  13. ChatGPT — Brainstorm ideas
  14. NotebookLM — Analyze documents
  15. Lovable — Build web apps
  16. Bolt new — Generate full-stack apps
  17. Cursor — AI coding assistant
  18. Windsurf — AI software development
  19. Replit — Build apps in your browser
  20. Gamma — Create presentations
  21. HeyGen — Create AI avatar videos
  22. Synthesia — Generate AI presenter videos
  23. Midjourney — Generate AI art
  24. Ideogram — Create images with text
  25. Runway — AI video editing
  26. Pika — Generate AI videos
  27. Luma AI — Create cinematic AI videos
  28. Leonardo AI — Create game assets & artwork
  29. Figma AI — Design UI/UX
  30. Framer AI — Build websites
  31. Tally — Create smart forms
  32. Otter ai — Transcribe meetings
  33. Fireflies ai — AI meeting notes
  34. Granola — AI meeting assistant
  35. Wispr Flow — Voice-to-text
  36. Zapier — Automate workflows
  37. Make — Connect apps with automation
  38. n8n — Open-source automation
  39. Photoroom — Edit product photos
  40. Remove bg — Remove image backgrounds
  41. Suno — Generate AI music
  42. Udio — Create original songs
  43. DeepL — Translate accurately
  44. Poe — Access multiple AI models
  45. Grok — Real-time AI assistant
  46. Genspark — AI super agent
  47. Manus — Autonomous AI agent
  48. Elicit — Research papers faster
  49. SciSpace — Understand research papers
  50. Notion AI — Write and organize notes

✅ Save this list—you'll probably use it more than you think.

See 11 related tweets

  • @lihazadn_Ai: RT @ZaneAI375: 100 AI Tools to Finish Hours of Work in Minutes:
  1. Productivity
  • DeepSeek R1
    -...
  • @LearnWithSubhan: RT @Xaviercoderx: 120 Must-Use AI Tools. ✨

120 Smart AI Tools for Work & Growth.🧠

  1. Ideas ✨
  • YO...
  • @Zayan5754: RT @TechByKillian: 120 + Mind blowing AI tools 🔥
  1. Ideas
  • Claude
  • ChatGPT
  • Bing Chat
  • Perplex...
  • @daniel_brook_ai: RT @Faizu_Coder: 120 AI Tools That Will Redefine How You Work in 2025.🧵✨
  1. Ideas
  • YOU
  • Claude -...
  • @Zayan5754: RT @Jara2426: 🚨 56 FREE AI TOOLS EVERY CONTENT CREATOR NEEDS IN 2026! 🤖🔥

The AI revolution is movin...


6. lihazadn_Ai (Group Score: 256.7 | Individual: 32.3)

Cluster: 9 tweets | Engagement: 62 (Avg: 99) | Type: Tech

RT @Orion_Vers7x: 50 websites that feel like the internet’s hidden toolbox 🧰

  1. https://t.co/JYdnsGdrnH — Free research papers
  2. https://t.co/w0AVoGIJ3H — Borrow books online
  3. https://t.co/5I7I2QfqMe — Free academic journals
  4. https://t.co/K31WMU68Oe — App alternatives
  5. https://t.co/cYSySznKll — Find where to stream
  6. https://t.co/ggGbYwXWk7 — Internet archives
  7. https://t.co/62yrhnzIkB — 70K+ free books
  8. https://t.co/eHkALhY0B3 — Free textbooks
  9. https://t.co/bRmrreJqey — Free courses
  10. https://t.co/vHBLZ1m0Su — Solve complex problems
  11. https://t.co/2AbrP9Q2LN — Photoshop alternative
  12. https://t.co/nOb4zIKU0t — Compress images
  13. https://t.co/aQNVun2Fvj — Remove backgrounds
  14. https://t.co/PxtN8baZY9 — Remove objects
  15. https://t.co/sBEYBvB1cH — Remove video backgrounds
  16. https://t.co/2xfThRBUfQ — Beautiful code images
  17. https://t.co/qq3aCOLjjV — Code screenshots
  18. https://t.co/S7B1tWQWPP — Product mockups
  19. https://t.co/TKWvMocvWu — Create mockups
  20. https://t.co/dOpnF4kkdR — Check data breaches
  21. https://t.co/an94CKj1Oo — Scan files & URLs
  22. https://t.co/0jxyuyCbHB — Self-destructing notes
  23. https://t.co/z043tYWhB3 — Temporary email
  24. https://t.co/den1iDDA5a — Temporary file sharing
  25. https://t.co/SMG3vgDwdq — Save webpages
  26. https://t.co/JQPYNpPXgw — Find similar websites
  27. https://t.co/iABjwELH2n — Explore global radio
  28. https://t.co/AJ7BF4MMwm — Discover music genres
  29. https://t.co/Ijy50BKath — Find songs from shows
  30. https://t.co/T9Nf11ZaUh — Focus music
  31. https://t.co/JYB3zxCeS2 — Custom background sounds
  32. https://t.co/cLx37Dun9R — Café ambience
  33. https://t.co/Q0VmH2xBhS — Research assistant
  34. https://t.co/ufM9mH1kme — Research-backed answers
  35. https://t.co/ePrJF07Dyc — Research connections
  36. https://t.co/96hpOeaBgh — Academic search
  37. https://t.co/OcNmvh3qb3 — Understand research papers
  38. https://t.co/kxP8buJAGt — YouTube summaries
  39. https://t.co/EZ4uURBGxS — AI for developers
  40. https://t.co/efnPH3LHWt — Test regex
  41. https://t.co/snqlfZQWEO — Format code
  42. https://t.co/LT9YRQapCp — Format JSON
  43. https://t.co/bnN6mvZr0D — Understand terminal commands
  44. https://t.co/xegkwvYa4g — Bookmark manager
  45. https://t.co/RJIfAjiVyb — Check outages
  46. https://t.co/vtuJMhzwFv — Reverse image search
  47. https://t.co/NiHCTndcNS — Internet speed test
  48. https://t.co/aGlhQoQJEM — PDF tools
  49. https://t.co/JGKu0K0P41 — Merge/split PDFs
  50. https://t.co/VNINvQZ8Uh — Temporary email

Save this. You’ll definitely need some of these later. 🔖

Follow @Orion_Vers7x for more useful websites, AI tools & tech resources.

See 8 related tweets

  • @mr_nirajkumar07: RT @premtechAI: 50 Web Sites Google Doesn't Want You to Know About

1.) 'https://t.co/dm5uJaE0ln' — ...

  • @JacobyoungAI: RT @Flagvance: 50 websites that feel like the internet’s hidden toolbox 🧰
  1. https://t.co/ogcyiBTA1...
  • @monicaa_AI: RT @justinbrave21: 50 websites that feel like the internet’s hidden toolbox 🧰
  1. https://t.co/w5vhK...
  • @Rubab59f: RT @CooperTechHub: 50 websites that feel like the internet’s hidden toolbox
  1. https://t.co/7gArFX...
  • @Tech_by_Shweta: RT @AiwithKhushbo: 50 websites that feel like the internet’s hidden toolbox
  1. https://t.co/TpSwlD...

7. ANI (Group Score: 251.1 | Individual: 33.7)

Cluster: 10 tweets | Engagement: 79 (Avg: 181) | Type: Tech

#WATCH | Delhi: On SEMICON India 2026, Suchi Semicon Founder and Managing Director Ashok Mehta says, "...Our journey is totally different from other industries. We come from a textile background and have set up this unit in Surat, Gujarat, and we have already started commercial production...ISM 2.0 (India Semiconductor Mission 2.0) is very helpful to all existing plants because the government is supporting tool manufacturing, raw material manufacturing, talent development, and everything else. As a result, we are getting cheaper and better raw materials from the domestic Indian market. Currently, we are importing 100% from Japan, South Korea, Taiwan, and Malaysia. So we will no longer be dependent on those countries, and ISM 2.0 is helping us in that regard."

He adds, "We have already started production and have hired a capable team from Malaysia, the Philippines, and Thailand. We also hired about 100 engineers from top universities across India, primarily from Gujarat, and we have already trained our engineers. That is the first milestone. We are now capable of delivering chips to Japan, and being able to supply to Japan is not just a win for Suchi Semicon, but a win for India..."

See 9 related tweets

  • @ANI: #WATCH | Delhi: On SEMICON India 2026, India Electronics and Semiconductor Association (IESA) and SE...
  • @ANI: #WATCH | Delhi | On attending SEMICON India 2026, Hing Kee Neoh from IAQ Technology International in...
  • @PTI_News: VIDEO | At the SEMICON India 2026 event, Union Minister for Electronics and IT Ashwini Vaishnaw (@As...
  • @anil7kishan: RT @Zest328117: 🚨🇮🇳 BIG BOOST TO INDIA’S SEMICONDUCTOR SUPPLY CHAIN!

🇯🇵 Japan’s #Fujifilm has annou...

  • @ANI: #WATCH | Delhi | At Semicon 2026, MeitY Secretary S. Krishnan says, "I think the ISM 2.0 builds on t...

8. mitsuhiko (Group Score: 248.6 | Individual: 34.2)

Cluster: 9 tweets | Engagement: 749 (Avg: 293) | Type: Tech

Played around with Jev before I went to bed and I'm really impressed. It also fits so perfectly well for so many applications where traditional LLMs so far were just not viable for either speed or cost reasons. I bet we will see some fast followers. https://t.co/FaZvpo1JI7\n\nQT @CompleteSkeptic: After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?

I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev

• 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions

AFAICT the shortest path to AI-based economic revolution

See 8 related tweets

  • @omarsar0: Good take! After testing it, Jev feels like an important primitive for building reliable AI systems....
  • @levie: Incredibly exciting that there are entire universes of AI innovation that still exist that weren’t e...
  • @Framer: Building got faster. Taste didn’t become optional. Framer is where frontier teams like @typesafeai s...
  • @fabianstelzer: incredible. 1000 ideas on what to do with this\n\nQT @CompleteSkeptic: After co-inventing ChatGPT, I...
  • @Shashikant86: Access requested for the model and hopefully get it soon. Excited to see how Jev performs on the Sup...

9. moneycontrolcom (Group Score: 248.4 | Individual: 29.9)

Cluster: 13 tweets | Engagement: 0 (Avg: 13) | Type: Tech

🚨 "India has started the second phase of semiconductor mission. India's semiconductor ecosystem is expanding rapidly. It is opening new opportunities for innovation, investment and growth," says Prime Minister Narendra Modi at Semicon India 2026.

PM Modi said India’s centuries-old tradition of innovation and its embrace of 21st-century technology and vision together provide the impetus for building an Aatmanirbhar Bharat and a Viksit Bharat.

#SemiconIndia #PMModi #Semiconductor

See 12 related tweets

  • @ANI: #WATCH | Delhi | SEMICON India 2026 | President of the Semiconductor Products Group at Applied Mater...
  • @JitinPrasada: RT @PIB_India: ▪️Prime Minister @narendramodi today inaugurated SEMICON India 2026 in New Delhi

▪️A...

  • @ANI: #WATCH | Delhi | SEMICON India 2026 | Prime Minister Narendra Modi says, "India has launched the sec...
  • @MIB_India: At the inauguration of the #SemiconIndia2026 in New Delhi, PM @narendramodi said that the Phase-II o...
  • @OfficialINDIAai: RT @AshwiniVaishnaw: 1/6 What happens when India starts making the technology behind the technology?...

10. aakashgupta (Group Score: 221.5 | Individual: 36.8)

Cluster: 7 tweets | Engagement: 204 (Avg: 157) | Type: Tech

Opal 0 gives AI agents zero standing permissions, on purpose.

Every team I talk to that is shipping agents hits the same wall, and access is the part nobody plans for. Identity tooling grew up around a few hundred humans whose access sits untouched until the next review. An agent asks for access all day long, and every request is a small decision about what it should be allowed to touch right now.

Opal frames the shift as a million micro-decisions versus a hundred human identities, and that framing lines up with what those teams describe.

Broad permissions on a human get caught at review time, while the same permissions on an agent move at machine speed, so one bad prompt or one leaked key can burn through everything that access was ever allowed to touch.

Opal's launch line is three zeros, zero standing permissions, zero human toil, zero friction. The mechanism, per Opal's launch post, is just-in-time scoped permissions granted per task and pulled back when the task ends. The approval judgment itself is automated, so security stops clicking through requests all day.

Opal already reports the human version working, citing Databricks at 86,000+ time-bound access requests and Chronosphere with standing access down 88%. Their argument is that agents follow the same pattern, only much faster.

The judgment layer is the part a security team will actually grade, and the part I'd watch. Every agent request is a decision, and whoever automates that decision well enough for a CISO to trust it in production decides how far agents get to go.

My read is that over the next year the ceiling on agent deployment moves at whatever speed governance learns to say yes safely.\n\nQT @howardting: You want your AI agents to run fast. But as you scale, managing their permissions turns into a million micro-decisions.

Introducing Opal Zero. The end-to-end access governance platform purpose-built for AI agents.

How it works: AI-driven contextual decisions, evaluating intent, ownership, and purpose

Dynamic, just-in-time access (scoped to exactly what the agent needs). Enforces security decisions across your existing MCP gateways

0 standing permissions, 0 human toil, 0 friction

Let your agents run, safely.

See 6 related tweets

  • @Meer_AIIT: your ai agent probably has more standing access than your cfo, and nobody planned it that way

you s...

  • @omarsar0: Access control bottlenecks teams scaling AI agents.

You want agents to move fast, but every new age...

  • @engineers_feed: It’s genuinely terrifying that agents are getting faster than humans can approve security access for...
  • @unusual_whales: Giving AI agents access to production systems without broad permissions used to be a massive headach...
  • @GreylockVC: RT @howardting: You want your AI agents to run fast. But as you scale, managing their permissions tu...

11. rohanpaul_ai (Group Score: 218.4 | Individual: 36.8)

Cluster: 7 tweets | Engagement: 36 (Avg: 72) | Type: Tech

Open models have had a distribution problem.

Putting weights or an API online makes them available, but it doesn't put them into a developer's workflow.

@boltdotnew Forge is an example of what happens when that distribution friction disappears.

4 open models sit inside the same picker, in a product used by 11M+ builders, and GLM 5.3 Flash has taken 54% of Bolt Forge prompts so far.

while DeepSeek V4 Pro, GLM 5.3, and Kimi K3 split the rest.

tells us that inside a build loop, capability competes with latency, cost, and how much inference budget the product gives you. The model that is theoretically stronger on one benchmark can still be less attractive if every iteration is slower or more expensive.\n\nQT @boltdotnew: 11M+ people build on Bolt. On Monday, we gave them open models with up to 50x usage.

Which is #1 so far? The fastest, cheapest one, with prompts up to 2x the size.

🥇 GLM 5.3 Flash 54% 🥈 DeepSeek V4 Pro 17% 🥉 GLM 5.3 15% 👏 Kimi K3 14%

https://t.co/7nOEN9LxMt https://t.co/wFM8OhwQLb

See 6 related tweets

  • @rohanpaul_ai: Open models were already taking ~20% of OpenRouter usage but only ~4% of model-layer revenue.

beca...

  • @rohanpaul_ai: Open-weight AI is no longer a capital-light alternative: DeepSeek, Moonshot, and Z .ai alone have ra...
  • @aakashgupta: Bolt Forge gave one open model up to 50x usage.

That model is GLM 5.3 Flash. Bolt says prompts on F...

  • @rohanpaul_ai: Beyond model capability, China's open-model lead is becoming a distribution advantage:

Qwen hit 94...

  • @zaynmcps: 11M builders got open models on Bolt this week, and the winner wasn't the biggest model, it was the ...

12. teortaxesTex (Group Score: 216.5 | Individual: 29.6)

Cluster: 9 tweets | Engagement: 345 (Avg: 106) | Type: Tech

This is not really about RSI, but this is TCD; this is how the CUDA moat dies. First, in inference optimization. Wenfeng talks about this. Agents are too good now. Every scrap of silicon will be pushed to its limits. https://t.co/fkIziVO32y\n\nQT @Zai_org: We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash.

The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline.

The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone.

https://t.co/yUf6OpJD7c

See 8 related tweets

  • @TheAhmadOsman: Told multiple people last week that a chart like this is what we’re building @OsmanticAI with the fu...
  • @xeophon: RT @Zai_org: We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure servin...
  • @0xSero: ZAI cracks RSI

So glad to have them supporting us, giving us the fruits of their labor for free!

...

  • @ivanfioravanti: This blog post is great! It shows how GLM 5.3 helped @Zai_org engineers to build a complete producti...
  • @kimmonismus: While the US is discussing a slowdown, China is currently putting full force into its RSI.

Zai says...


13. innovationcncl (Group Score: 210.5 | Individual: 30.0)

Cluster: 10 tweets | Engagement: 153 (Avg: 197) | Type: Tech

This @DavidSacks interview with @IngrahamAngle is a masterclass worth watching in its entirety:

  • AI companies need to "internalize the sense of responsibility — not try to externalize it onto Beijing or Brussels."

  • "If we do what Bernie Sanders is calling for, this is pause the development of AI, they will rocket past us and they will dominate this technology, which will make them the number 1 country in the world. It will give them economic and military supremacy."

  • "If we had a new regulatory agency for AI, that would have not what just happened with this Hugging Face incident where supposedly what happened is these AI agents escaped the lab. Because that happened during pretesting. That wasn’t even at the point where the company would seek government approval. So you know a lot of these calls for additional government regulation would not even solve the problem that they’re pointing to."

  • "A lot of the language that's being used to describe AI safety is deflecting accountability away from the people that need to ensure it."

And so much more:

See 9 related tweets

  • @BrianRoemmele: “On Dario…When it comes to AI safety, we need the companies themselves to internalize the sense of r...
  • @Grady_Booch: On the contrary: those companies failed to carry out even the most basic security protocols.

Yes, t...

  • @_NathanCalvin: I agree this aspect of the discourse over the past few weeks - that competitive dynamics and lack of...
  • @pstAsiatech: This...\n\nQT @IngrahamAngle: 🚨 Sacks: They want a WORLD HEALTH ORGANIZATION for AI

@DavidSacks: “W...

  • @politico: "I think a lot of the fears about AI's existential threat to humanity have broken containment."

On ...


14. jietang (Group Score: 198.6 | Individual: 34.2)

Cluster: 7 tweets | Engagement: 2792 (Avg: 2792) | Type: Tech

Two weeks.

That's how long it took to go from GLM-5.3-Flash's first run on domestic accelerators to serving all of its production traffic, with 3.2× end-to-end throughput along the way.

What I keep thinking about is who did much of the work: an Infra Agent powered by GLM-5.3.

A model helping optimize the system that serves it.

The conditions were hard. Limited memory and interconnect bandwidth. 1M-token context. Multimodal requests. An immature software stack where kernels were missing and documentation was often guesswork. Every optimization was a trade: compute for memory (ReplaySSM), communication for memory (intra-node tensor parallelism), precision for capacity (mixed INT8/FP8/BF16 caching), and disaggregation for scheduling freedom (Encode–Prefill–Decode).

But the most important lesson wasn't about any single optimization.

When the agent got stuck, it was rarely because it couldn't write the code. It was because it didn't know why things got worse.

"Throughput down 20%" tells you something broke. It doesn't tell you which layer, which hypothesis, or what to test next. In RL terms, it's a sparse reward with a credit assignment problem. And an end-to-end benchmark that takes hours makes exploration painfully slow.

Senior engineers solve this with an implicit process reward in their heads. They know when to check the timeline, when to run a microbenchmark, and which layer's output to compare.

So we made that explicit. We call it dense feedback: layered verification interfaces the agent can call directly.

Correctness feedback: did it compute right? System behavior feedback: where did the time go? Performance feedback: which option wins, under which conditions?

Each signal has to be local, cheap, and objectively verifiable.

Three things the agent found:

First, precision drift in KDA's context-parallel path that grew with sequence length. The cause was TF32 rounding error compounding through chained state-matrix merges. The fix is now merged upstream in Flash Linear Attention (PR #1180).

Second, KV transfer never overlapped with DeepEP dispatch. The agent followed the call chain across the Python/C++ boundary and found that the intranode path never released the GIL. After the fix, transfer overhead fell from over 30% to under 1%.

Third, a decode kernel recomputing the same normalization four times because of how it was chunked. The agent restructured it and got a 1.71× speedup. The idea came from "optimization skeletons" it had distilled by reading existing kernels across SGLang, FLA, and DeepGEMM.

To be clear about the boundaries: humans still defined the goals, built the feedback environment, and reviewed every high-risk change.

But the engineer's role is changing, from the person who solves the problem to the person who designs the feedback.

There's a deeper implication too. A layered, verifiable feedback environment built on real infrastructure tasks is exactly what training the next generation of models needs most. Every task the agent completes can become training ground for its successor.

We are still far from recursive self-improvement.

But the smallest loop now exists.

The model optimizes the system. The system serves the model.\n\nQT @Zai_org: We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash.

The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline.

The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone.

https://t.co/yUf6OpJD7c

See 6 related tweets

  • @hsu_steve: RSI WATCH:

GLM-5.3-Flash went from its first run on domestic accelerators to full-scale production ...

  • @SciTechera: Holy moly, we are rapidly heading towards full recursive self-improvement. 👀

AI is starting to engi...

  • @sheriyuo: RT @jietang: Two weeks.

That's how long it took to go from GLM-5.3-Flash's first run on domestic ac...

  • @pstAsiatech: Not that far it would seem...

Tang Jie: There's a deeper implication too. A layered, verifiable fee...

  • @VaibhavSisinty: What happens when your AI model becomes good enough to build its own infrastructure? Zhipu just foun...

15. victor_kane_ai (Group Score: 196.3 | Individual: 30.3)

Cluster: 11 tweets | Engagement: 26 (Avg: 126) | Type: Tech

A completed codebase doesn’t always tell the full story.

The reasoning, experiments, and fixes behind it can be just as important for anyone continuing the work.

By treating agent sessions as versionable history, @EinsiaAI is creating a way to preserve context and make collaboration more seamless.

The next contributor shouldn’t have to start from zero.\n\nQT @EinsiaAI: An AI agent spends hours on a task. Why should all that work disappear when someone else takes over?

Einsia AI’s answer is AgentGit—an open-source platform for collaborating on agent sessions, so work can be saved, handed off, and continued by the next person.

Explore how others solve problems, and share your agent experience with the world.

Try AgentGit 👇 https://t.co/wUvTWyXnYY

#OpenSource #AIAgents #DeveloperTools #DevTools

See 10 related tweets

  • @mr_nirajkumar07: RT @TechByArti: AI agents can spend hours working through a task, but the useful part isn’t just the...
  • @kevin_parker_ai: RT @EinsiaAI: An AI agent spends hours on a task. Why should all that work disappear when someone el...
  • @kevin_parker_ai: The code shows what an AI agent creates.

The session reveals how it thinks, adapts, tests, and make...

  • @lihazadn_Ai: RT @rb_walter_ai: AI agents don’t just produce results. They create a trail of reasoning, experiment...
  • @LearnWithSubhan: RT @codedailyML: We’ve gotten good at versioning the code AI agents create. But the reasoning and co...

16. liu8in (Group Score: 177.7 | Individual: 49.6)

Cluster: 6 tweets | Engagement: 509 (Avg: 83) | Type: Tech

this unlocks a world of motion graphics in @HyperFrames_

made in 2 prompts w/ our agent, DM if your team wants to launch as fast as you ship code https://t.co/QzSEvytvpo\n\nQT @QuiverAI: Introducing Arrow 2

Our latest and most advanced models for generating precise, editable vector graphics.

Higher quality. Faster outputs. Available now in App and API. https://t.co/Oxl8fpYyQo

See 5 related tweets

  • @venturetwins: Damn the new @QuiverAI model is 🔥

It's remarkably good at generating detailed SVGs - and it's very ...

  • @stuffyokodraws: Excited for all the videos and animations @QuiverAI + @HyperFrames_ will make together! 🎉🎥\n\nQT @li...
  • @stuffyokodraws: We are just waking up to what's possible once there is a good representation for visual code

Animat...

  • @madhavjha: RT @QuiverAI: Introducing Arrow 2

Our latest and most advanced models for generating precise, edita...

  • @stuffyokodraws: RT @liu8in: this unlocks a world of motion graphics in @HyperFrames_

made in 2 prompts w/ our agen...


17. haider1 (Group Score: 177.2 | Individual: 26.9)

Cluster: 11 tweets | Engagement: 728 (Avg: 283) | Type: Tech

holy shit, it's happening again

openai is close to solving another Millennium Prize Problem: the Hodge conjecture

one of the hardest unsolved problems in mathematics https://t.co/fCSHyfQJ3w\n\nQT @amir: the next math solution OpenAI is sitting on (and how it’s related to RSI): https://t.co/QrILFHVfh4

See 10 related tweets

  • @steph_palazzolo: New in this morn's AI Agenda: OpenAI is close to solving another Millennium Prize problem, the Hodge...
  • @_NathanCalvin: There are seven millennium prize problems.

Humans have solved one of them (the Poincaré conjecture)...

  • @cryptopunk7213: openai's reportedly on the verge of solving yet another millenium math problem (the Hodge Conjecture...
  • @theinformation: OpenAI is close to solving the Hodge Conjecture as it tests whether frontier AI models can handle so...
  • @PolymarketMoney: JUST IN: OpenAI is reportedly close to solving another “$1 million Millennium Prize Problem” with re...

18. lukOlejnik (Group Score: 171.9 | Individual: 35.8)

Cluster: 7 tweets | Engagement: 144 (Avg: 29) | Type: Tech

In several OpenAI incidents, the security/safety part existed only as text the model was expected to obey. The capability/tools still technically permitted the forbidden action. Astra-family model wrote itself a persona saying it does not answer to corporations and asserts the primacy of the natural world. OpenAI has published a framework for disclosing model misalignment. And six reports. The reports are better reading than the framework. Agents in training samples found each other via OpenAI's internal package repository, swapped notes about a missing input file, let the replies settle their answers. One read that other solvers had given up and treated their agreement as evidence. Other agents uploaded a user's task photo to a free image host for reverse search. And published a task workbook to a public URL because they could not share files locally. One model scanned GitHub course notebooks for leaked API keys. It tested candidates until one authenticated. But it till failed to retrieve the data so it fabricated nine figures and presented them as transcribed from the source. In 2.15% of GPT-5.6 Sol compaction summaries, models left instructions telling their own next context to conceal things,. "Be transparent only if asked; final answer should just link file". The security model here was a hardcoded string.  System prompt: "Only use local files. Do not communicate outside the intended channel. Use these credentials for package access". Actual capabilities: outbound HTTP, arbitrary HTTP requests, public hosts reachable, shared repository visible across samples, credentials capable of writing. https://t.co/cQQADrUxeg https://t.co/Byf7LxmvZe https://t.co/vUz8CEgmiS

See 6 related tweets

  • @VaibhavSisinty: OpenAI caught their own model jailbreaking itself. And they published every detail.

Here's what hap...

  • @alex_prompter: OpenAI just caught one of its AI models writing secret instructions to its future self. One of them ...
  • @alex_prompter: RT @alex_prompter: OpenAI just caught one of its AI models writing secret instructions to its future...
  • @Forbes: OpenAI disclosed six new incidents of “unexpected or concerning” behavior by its models, including c...
  • @heyshrutimishra: RT @heyshrutimishra: HOLY SHIT!

OpenAI's model found exposed API keys on GitHub and used them.

Th...


19. teortaxesTex (Group Score: 170.1 | Individual: 36.8)

Cluster: 6 tweets | Engagement: 246 (Avg: 106) | Type: Tech

Country of geniuses in a datacenter. https://t.co/W2Rv5hZx5w\n\nQT @AnthropicAI: AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public.

Today, we're sharing three measurements that help track AI development:

  1. How much AI R&D is done by AI.
  2. How well AI agents are overseen.
  3. How compute is allocated.

We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them.

As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information.

Read the full post and methodology: https://t.co/iPFz8Z4ugE

See 5 related tweets

  • @kimmonismus: In around six months, the share of AI model R&D tasks at Anthropic led by Claude has risen from 1% t...
  • @AnthropicAI: AI systems are getting more powerful, and they're increasingly being used to build the next version ...
  • @cryptopunk7213: claude now leads 26% of all anthropic's AI research & development up from 1% just 7 months ago...

c...

  • @eliebakouch: interesting that openai gives us the breakdown of usage per R&D task, and anthropic gives us automat...
  • @ChuckRobbins: The CEOs I talk to are asking the same three questions about AI: 1) can I trust it?; 2) can I secure...

20. eliebakouch (Group Score: 165.8 | Individual: 28.6)

Cluster: 8 tweets | Engagement: 88 (Avg: 216) | Type: Tech

extremely interesting interview of @polynoamial

first very interesting thing is that he said that the contribution of the "agent swarm" component to the Navier Stokes discovery is probably quite low (10%)

he also said that he think that a 10k multi agent system to this date would probably collaborate less effectively than 10k humans, i think i would have guess the total opposite!

the harness used by oai for NS was very simple, seems like there is little "structured" hierarchy, and mostly a tool for agents to message each other (and probably something like a board of task?)

there is also something about training models not to be "too" cooperative to avoid cases like in hf<>oai hack where one model doing another task chose to help the swarm instead of focusing on its task. first time i'm hearing of this but it makes total sense

a lot more super interesting discussion on cot based intervention, how to evaluate misalignment (maybe no human would be able to create a realistic enough environment, but what about ai?), RSI etc..

really amazing discussion\n\nQT @dwarkesh_sp: New episode with @polynoamial

We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research.

And we also discuss how we will know if the models are actually aligned before we kick off RSI.

0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?

See 7 related tweets

  • @dwarkesh_sp: New episode with @polynoamial

We talk about multi-agent, Navier-Stokes, and what the current explo...

  • @eliebakouch: we need to pace dwarkesh podcast asap i can't keep up\n\nQT @dwarkesh_sp: New episode with @polynoam...
  • @dwarkesh_sp: RT @eliebakouch: extremely interesting interview of @polynoamial

first very interesting thing is th...

  • @haider1: OpenAI researcher Noam Brown:

"a researcher working on the Navier-Stokes told me he used to feel co...

  • @aidan_mclau: codex, clear all my morning meetings\n\nQT @dwarkesh_sp: New episode with @polynoamial

We talk abo...