Published on

热门科技推文精选 — 2026年7月11日

Authors

今日科技要闻中,AI智能体占据头条。OpenAI开始在ChatGPT、Codex及API中逐步推出GPT‑5.6系列;据报道,Grok 4.5在一项专业工作基准测试中位居榜首;Muse Spark 1.1则聚焦长时运行的多模态任务。编程助手继续向基于浏览器的工作流程拓展,进一步强化了外界对智能体将重塑工程与创意工作的预期。与此同时,初创公司Polsia宣称在150天内实现了1,000万美元的年化营收规模,而围绕隐私、商业秘密和出口管制的争议,则凸显出AI日益加剧的法律与地缘政治风险。


1. teslaownersSV (Group Score: 361.0 | Individual: 49.5)

Cluster: 15 tweets | Engagement: 2054 (Avg: 321) | Type: Tech

GROK 4.5 LEADS ON REAL PROFESSIONAL WORK BENCHMARK

New data from Snorkel shows Grok 4.5 outperforming other frontier models on real-world professional tasks.

On their GDPval+ benchmark (expert-created workplace reasoning tasks across the economy):

• Grok 4.5: 29% mean pass rate • GPT 5.5: 22% • Claude Opus 4.8: 21% Grok 4.5 showed particularly strong gains in demanding areas like legal work, education, healthcare, and QA analysis.

This lines up with xAI’s focus on building models that excel at practical, agentic work rather than just synthetic benchmarks. While general intelligence leaderboards still see tight competition at the very top, Grok 4.5 is delivering some of the strongest results on actual professional deliverables right now.\n\nQT @elonmusk: Grok Build improves almost every day

See 14 related tweets

  • @nickvasiles: I said Grok 4.5 was a bigger deal than Fable

Elon even liked my post

but then GPT-5.6 Sol showed u...

  • @XFreeze: Grok build one of the most powerful agentic coding harnesses is now available for free for everyone,...
  • @elonmusk: Grok Build improves almost every day\n\nQT @JasonBud: some special features in Grok Build if you're ...
  • @teslaownersSV: GROK 4.5 RANKS #1 ON REAL-WORLD PROFESSIONAL TASKS

New evaluation from Snorkel AI shows Grok 4.5 ou...

  • @elonmusk: Grok Build\n\nQT @XFreeze: Grok 4.5 with Grok Build just ranked #1 on the SWE-Atlas-QnA benchmark wi...

2. HedgieMarkets (Group Score: 350.9 | Individual: 33.6)

Cluster: 18 tweets | Engagement: 117 (Avg: 292) | Type: Tech

🦔Apple is suing OpenAI for systematic trade secret theft weeks before the IPO. The complaint alleges a former Apple engineer exploited a bug to access Apple's cloud storage while employed at OpenAI and laughed about it in a text message. Tang Tan, now OpenAI's chief hardware officer, allegedly used Apple codenames in interviews to extract confidential information from current employees and circulated an internal Apple offboarding doc to teach new hires how to dodge exit security. Over 400 former Apple employees now work at OpenAI.

My Take If even half of this complaint holds up, OpenAI's hardware plans are finished. You can't ship a device when a court has ruled your team and supply chain were built on stolen IP. Apple has more cash than most countries and a legal team that treats intellectual property enforcement like a religion. This is about the worst opponent OpenAI could have drawn right before asking public markets for money.

The bigger problem is what this says about the company. Apple isn't describing one rogue employee. They're describing a senior hire who allegedly coached recruits on bypassing exit security and an engineer who treated accessing his former employer's files like a joke. If you're sending proprietary data through OpenAI's API, this filing should be on your reading list.

Hedgie🤗

See 17 related tweets

  • @TFTC21: Apple just sued OpenAI in federal court for trade secret theft.

The allegation is sweeping. Apple s...

  • @shiri_shh: Apple is officially suing OpenAI, alleging institutional-level trade secret theft.

They claim Open...

  • @KatieMiller: OpenAI’s last 24 hours:

> Top Exec unexpectedly departs > Shuts down browser tool after 9 mon...

  • @MSBIntel: BREAKING: Apple sued OpenAI and hardware chief Tang Tan, alleging trade-secret theft to build its co...
  • @theinformation: Apple sued OpenAI, accusing the AI startup of systematically stealing trade secrets to build its har...

3. thsottiaux (Group Score: 338.1 | Individual: 40.7)

Cluster: 10 tweets | Engagement: 13559 (Avg: 6098) | Type: Tech

Hello beautiful people! We have reset usage limits across Codex and ChatGPT Work. And another one will come later in the day. Rejoice.

Now that I have your attention, a quick update on ChatGPT Work, Codex and all the updates we shared yesterday.

We’ve spent the last 24 hours reading feedback, looking at usage patterns, and talking with many of you. The short version is that there is a lot of excitement for GPT 5.6 Sol, ChatGPT Work on mobile & web, but also that we didn't get everything quite right.

  • We made it too easy to use the highest-compute settings without making the impact on usage limits sufficiently clear.
  • We reorganized the desktop app in one bold move, making familiar things like chats and projects harder to find.
  • Our launch framing was focused on ChatGPT Work and to some of our Codex fans it made it feel like Codex was going away over time. Absolutely not our intention, we love Codex and it is here to stay.
  • And we introduced regressions for some existing multi-agent workflows, alongside a collection of rough edges in plugins and other parts of the experience.

We’re landing a first set of improvements today. We’re resetting usage twice so people can keep experimenting, changing defaults and the model picker so they don’t push people toward unnecessarily expensive settings, fixing several plugin submission issues, improving how we represent Codex in the product, and cleaning up some of the most immediate desktop problems.

A larger set of improvements will land next week. We’re bringing chats and projects back into the sidebar in a more familiar and customizable way, making usage and reset timing much more visible, clarifying when to use ChatGPT Work and when to use Codex, and addressing the many other smaller pieces of great feedback we've had.

The ambition behind this launch hasn’t changed. We think bringing ChatGPT and Codex together into a workspace where people and agents can collaborate is a very important step forward. But an ambitious direction doesn’t excuse avoidable confusion or regressions in the first version.

Please keep the feedback coming. We’re moving quickly, and you should see the experience already get better with a few updates today; and substantially better again next week.

See 9 related tweets

  • @danshipper: excellent comms here\n\nQT @thsottiaux: Hello beautiful people! We have reset usage limits across Co...
  • @romainhuet: Thank you for all the thoughtful feedback! We’re listening and moving quickly, so please keep it com...
  • @dkundel: 🚨 Resets incoming & addressing your feedback 👇\n\nQT @thsottiaux: Hello beautiful people! We hav...
  • @kimmonismus: You really have to give OpenAI credit: its work has become excellent, and so has its community manag...
  • @yacineMTB: Virgin claude: we reset usage every week and maybe we are going to keep on giving you access, though...

4. MiniMax_AI (Group Score: 303.6 | Individual: 43.3)

Cluster: 9 tweets | Engagement: 280 (Avg: 89) | Type: Tech

Toward the end of the sky, and beyond.\n\nQT @SkylerMiao7: A letter from our CEO today:

Markets will fluctuate, and external noise will come and go, but our direction remains unchanged.

Being at the forefront of this industry, we have a clear understanding of the pace of technological evolution and the long-term value we are building together.

Starting today, and until the day we reach AGI, I will no longer receive any compensation from the company.

Over the next four years, I will allocate shares equivalent to 4% of the company’s total equity from my personal holdings to reward team members who choose to build this journey with us for the long term and create value together. In addition, I will allocate 1% of my shares to establish a dedicated fund supporting the continued growth of open-source communities and the broader AI ecosystem.

I will devote all my time, energy, and resources to this mission.

This is my long-term commitment as a founder — to our company, our team, and the future we are building together.

We will keep going until we get there.

Intelligence with Everyone.

IO CEO & Founder, MiniMax

See 8 related tweets

  • @teortaxesTex: > Starting today, and until the day we reach AGI, I will no longer receive any compensation from ...
  • @zephyr_z9: Skin in the game\n\nQT @SkylerMiao7: A letter from our CEO today:

Markets will fluctuate, and exter...

  • @ModelScope2022: Huge congratulations to @MiniMax_AI on the $2B financing milestone! 🚀

MiniMax’s CEO has committed 1...

  • @chris_j_paxton: Everyone keeps predicting the impending demise of the Chinese open source ecosystem. But so many peo...
  • @bgurley: Some very creative structures happening with CEO packages and company structures in China. See below...

5. MatthewBerman (Group Score: 212.9 | Individual: 33.5)

Cluster: 9 tweets | Engagement: 1029 (Avg: 360) | Type: Tech

Mark my words: Codex / Claude Code will own the browser market within 12 months.\n\nQT @ClaudeDevs: Claude Code on desktop now has an in-app browser.

Claude can pull up docs, designs, or any other site. It can read, click through, and interact the same way it does with your local dev servers.

It's sandboxed and configurable: you choose whether sessions persist. https://t.co/Jbc21OP0vJ

See 8 related tweets

  • @VaibhavSisinty: Claude Code just shipped its own browser. Built into the desktop app. It reads docs, clicks through ...
  • @amorriscode: The in-app browser has replaced my usage of the Chrome extension. I love how easy it is to provide c...
  • @testingcatalog: ANTHROPIC 🔥: Claude desktop app got a built-in web browser for Claude Code.

This feature allows Cl...

  • @rileybrown: The super-app wars continue.

The browser is one of the most important parts.\n\nQT @ClaudeDevs: Cl...

  • @Baconbrix: You can now use npx serve-sim to build and verify iOS apps in Claude Desktop 👀\n\nQT @ClaudeDevs: Cl...

6. Dimillian (Group Score: 192.2 | Individual: 34.8)

Cluster: 10 tweets | Engagement: 2366 (Avg: 331) | Type: Tech

RT @OpenAI: Sol, Terra, and Luna, our GPT‑5.6 family of models, are starting to roll out now in ChatGPT, Codex, and the API. https://t.co/Qri7GdtYs3

See 9 related tweets

  • @ZenMuxAI: Meet the full GPT-5.6 Series on ZenMux 🚀

🌙 Luna — fast & cost-efficient ☀️ Sol — frontier-level int...

  • @PaulSolt: RT @ajambrosino: New today:

  • GPT 5.6 Sol, Terra, and Luna

  • ChatGPT Work

  • The new ChatGPT desktop...

  • @aiedge_: OpenAI published an ENTIRE guide on prompting the new GPT-5.6 models.

I translated all the develope...

  • @TeksEdge: My hot take on @OpenAI's release today is not about the new GPT-5.6 line of models (with old-model s...
  • @OpenAIDevs: GPT-5.6 is here. Codex is now available inside ChatGPT. And we know developers will have questions. ...

7. WesRoth (Group Score: 181.6 | Individual: 27.7)

Cluster: 8 tweets | Engagement: 12 (Avg: 21) | Type: Tech

Muse Spark 1.1 is a multimodal reasoning model built for agentic work, tool use, computer use, coding, and long-running tasks.

It has a 1M token context window, can operate across desktop, mobile, and browser interfaces, and can delegate work to parallel sub-agents to finish projects faster.

Muse Spark 1.1 scores 54.7 on JobBench, compared with 17.0 for the first Muse Spark.

It also scores 80.8 on OSWorld-Verified, which tests open-ended computer use, and 86.1 on MCP Atlas, which tests social tool use.

In coding, it is competitive but not always leading. It scores 80.0 on Terminal-Bench 2.1, close to Opus 4.8 and GPT-5.5, but trails on DeepSWE 1.1.\n\nQT @finkd: (2) Muse Spark 1.1 is strongest at agentic performance, tool use, and computer use. It does well on long-running tasks with 1M token context window, can delegate execution to sub-agents running in parallel, and is trained to use computer interfaces on desktop, mobile, or browser. https://t.co/I3v82YohtR

See 7 related tweets

  • @WesRoth: Meta introduced Muse Spark 1.1, an upgraded multimodal reasoning model built for agentic work.

The ...

  • @scaling01: Muse Spark 1.1 is between Opus 4.6 and Opus 4.7

its reasoning efficiency could be better, but given...

  • @rohanpaul_ai: Meta is so back in the AI coding race with Muse Spark 1.1, using cut-rate pricing to pressure OpenAI...
  • @Marktechpost: RT @Marktechpost: Meta Superintelligence Labs released Muse Spark 1.1 and opened the Meta Model API ...
  • @alexandr_wang: RT @mohit_r9a: Excited to see our SWE Atlas - Codebase QnA benchmark featured in Meta's release for ...

8. kimmonismus (Group Score: 169.2 | Individual: 28.5)

Cluster: 7 tweets | Engagement: 832 (Avg: 814) | Type: Tech

Hey, Anthropic, take a leaf out of OpenAI's book!\n\nQT @thsottiaux: To celebrate the launch of GPT-5.6 Sol, we will reset the rate limits again (twice) across ChatGPT Work and Codex over the next 24 hours.

We want you to have the time to truly try ambitious tasks and get the hang of it. Happy exploring!

See 6 related tweets

  • @sama: the sun is out today\n\nQT @thsottiaux: To celebrate the launch of GPT-5.6 Sol, we will reset the ra...
  • @tunguz: All right, I am going all in with ALL of my side projects.\n\nQT @thsottiaux: To celebrate the launc...
  • @rezoundous: this upgrades me from xhigh to ultra!\n\nQT @thsottiaux: To celebrate the launch of GPT-5.6 Sol, we ...
  • @jxnlco: Inference team is good\n\nQT @thsottiaux: To celebrate the launch of GPT-5.6 Sol, we will reset the ...
  • @altryne: st Tibo!\n\nQT @thsottiaux: To celebrate the launch of GPT-5.6 Sol, we will reset the rate limits ag...

9. coinbureau (Group Score: 163.2 | Individual: 23.5)

Cluster: 10 tweets | Engagement: 288 (Avg: 355) | Type: Tech

🚨BREAKING: OPENAI AND GOOGLE SOLD TOP AI MODELS TO BLACKLISTED CHINESE TECH GIANTS

The US companies confirmed they provided AI access through Singapore-based subsidiaries linked to Alibaba, Baidu and Tencent, per FT.

All three Chinese groups are on a Pentagon blacklist over alleged ties to China’s military, although the AI sales remain LEGAL under current US rules.

By contrast, Anthropic has banned Chinese companies from using its advanced models.

See 9 related tweets

  • @wallstengine: OpenAI and Google are reportedly selling advanced AI model access to Singapore-based subsidiaries of...
  • @coinbureau: 🚨CHINESE TECH GIANT TENCENT WANTS MANUS AI BACK FROM META

Beijing has reportedly ordered Meta and M...

  • @FT: OpenAI and Google are selling their advanced AI models to Chinese tech giants blacklisted by the Pen...
  • @Techmeme: OpenAI and Google are providing advanced AI models to Singapore-based subsidiaries of Alibaba, Baidu...
  • @Cointelegraph: 🇨🇳 NOW: OpenAI and Google are selling AI models to Singapore-based subsidiaries of Pentagon-blacklis...

10. 0x_kaize (Group Score: 160.6 | Individual: 34.7)

Cluster: 6 tweets | Engagement: 30 (Avg: 153) | Type: Tech

Ben started Polsia alone in a Paris apartment

150 days later: $10M run rate in San Francisco

Agents handle the execution. Inbox, fundraising, daily operations, not human

Polsia's first customer was Polsia. Ben used it to run his own company.

$10M run rate in 150 days.

Then a $30M raise.

Less time building. More time earning - that's what AI agents actually do, they don't replace you, they compress the distance between idea and revenue.\n\nQT @Bencera: I grew Polsia from 0to0 to 10M run rate in 5 months.

Solo + AI. Zero employees.

Everyone asks me how.

Presenting aisloP episode 2: “The Rise” https://t.co/2fQQ7WcyJy

See 5 related tweets

  • @shiri_shh: Polsia founder just hit a $10M run rate with an ops team made entirely of agents

and now he's docum...

  • @VadimStrizheus: Do you understand what just happened?!

this guy hit $10M/ARR in 150 days as a solo-founder with AI...

  • @AnatoliKopadze: If you still think AI can't really build a company, you have to see this.

Ben started alone in a Pa...

  • @WhaleInsider: JUST IN: One founder used AI agents to build a $10M/year business in just 150 days. Most startups ta...
  • @alex_prompter: 150 days is how long most founders spend picking a logo.

Ben used it to take Polsia from zero to $1...


11. nic_carter (Group Score: 155.7 | Individual: 36.0)

Cluster: 6 tweets | Engagement: 4037 (Avg: 1131) | Type: Tech

you don't even have to ask. the answer is yes https://t.co/rElPXFxUCN\n\nQT @business: Phia — the buzzy shopping app co-founded by Bill Gates' daughter, Phoebe — is claiming credit for online sales it didn’t actually drive, a Bloomberg investigation found. Read our exclusive story: https://t.co/KYuWkmkE7D

📷️: Dia Dipasupil/Getty Images https://t.co/mFeXaMN9pD

See 5 related tweets

  • @MorePerfectUS: An app co-founded by Phoebe Gates’, daughter of Bill Gates, is reportedly claiming commissions for s...
  • @business: Phia — the buzzy shopping app co-founded by Bill Gates' daughter, Phoebe — is claiming credit for on...
  • @tunguz: Bill Gates' daughter BTW\n\nQT @business: Phia — the buzzy shopping app co-founded by Bill Gates' da...
  • @business: A startup co-founded by Phoebe Gates that bills itself as a “personal shopping assistant” took credi...
  • @business: RT @oliviasolon: Buzzy shopping app Phia, founded by Phoebe Gates, (Bill's daughter) took credit for...

12. IntCyberDigest (Group Score: 143.0 | Individual: 30.9)

Cluster: 6 tweets | Engagement: 2856 (Avg: 882) | Type: Tech

RT @IntCyberDigest: ❗️ Meta just silently opted in every adult Instagram user for their new Muse AI Image generator, which lets anyone create AI images of your likeness by tagging your public handle in a prompt, no notification when your photos are used, and no removal of images generated before you opt out.

To opt out on Instagram: Profile > Menu > Sharing and reuse, then switch off the reuse toggles under "Allow people to reuse your content on Instagram and with AI features at Meta."

See 5 related tweets

  • @ch: RT @MKBHD: Instagram decided to roll out a new “Muse AI” feature, that lets users create AI images b...
  • @Polymarket: JUST IN: Instagram faces backlash after automatically opting public profiles into a new AI feature t...
  • @Techmeme: Privacy advocates, CAA, and SAG-AFTRA slam Meta's Muse Image, which lets users create AI images usin...
  • @Polymarket: JUST IN: Meta removes a feature that let anyone generate AI images from public Instagram posts after...
  • @CNBCTV18Live: #Meta's new update can use your public #Instagram photos for #AI images: Here's what you need to kno...

13. cgtwts (Group Score: 138.1 | Individual: 30.7)

Cluster: 5 tweets | Engagement: 57 (Avg: 62) | Type: Tech

“Do all my creative work today. Turn my idea into a campaign. I’m going for coffee. don’t break anything.” https://t.co/q2ZYHZ6kMl\n\nQT @n0w00j: If you’re a creative not using AI you’re fcked.

Introducing Melius: the world’s first creative canvas for agents.

Here’s how it works: https://t.co/KdIa2HwsEU

See 4 related tweets

  • @unusual_whales: Melius has launched a multi-agent creative canvas, allowing users to orchestrate multiple AI agents ...
  • @austinh___: we’ve been using Melius and it’s insane\n\nQT @n0w00j: If you’re a creative not using AI you’re fcke...
  • @shensi: Congrats on the launch!!!! We use Melius and we f-ing love it\n\nQT @n0w00j: If you’re a creative no...
  • @jxnlco: RT @n0w00j: If you’re a creative not using AI you’re fcked.

Introducing Melius: the world’s first c...


14. StockSavvyShay (Group Score: 134.7 | Individual: 33.8)

Cluster: 5 tweets | Engagement: 431 (Avg: 604) | Type: Tech

SKHYisexpectedtogeneratenearlySKHY is expected to generate nearly 377B in annual revenue by 2028.

SK Hynix is the most direct public bet on HBM dollar content inside NVDAAIracksmultiplyingacrossBlackwell,RubinandRubinUltra.https://t.co/VzAyKhz1Eh\n\nQT@StockSavvyShay:NVDA AI racks multiplying across Blackwell, Rubin and Rubin Ultra. https://t.co/VzAyKhz1Eh\n\nQT @StockSavvyShay: SKHY has the cleanest current exposure to the memory segment where $NVDA roadmap is creating the most value.

If HBM content per AI rack climbs from the low hundreds of thousands today to potentially more than $1M with Rubin Ultra then SK Hynix becomes a direct bet on structural memory content growth the market is still underpricing.

See 4 related tweets

  • @Mayhem4Markets: SK Hynix opened at $170 on Nasdaq today.

14% above its $149 IPO price. A trillion-dollar market cap...

  • @wallstengine: SK Hynix revenue trajectory:

2028E: 377.3B2027E:377.3B 2027E: 346.2B 2026E: 237.2B2025:237.2B 2025: 64.6B 2024: $44.0B...

  • @TheFuturumGroup: SK Hynix’s U.S. listing could help narrow the valuation gap with Micron while giving investors more ...
  • @StockMKTNewz: The memory shortage visualized

What happens after that forecast is my question https://t.co/O0LmBn...


15. btibor91 (Group Score: 134.6 | Individual: 38.0)

Cluster: 4 tweets | Engagement: 147 (Avg: 185) | Type: Tech

Summary of Reddit AMA about "GPT-5.6 and Codex in ChatGPT" with OpenAI's Codex team on 2026-07-10

(opened with the stat that more than 5 million people use Codex every week, twice as many as three months ago, with 150 features and improvements shipped in that period)

Model selection and reasoning levels

  • Sol Medium for most things, Sol Ultra for genuinely hard tasks, Terra for quick non-coding tasks or usage-conscious work with performance competitive with GPT-5.5 on some tasks at lower cost, and Luna for subagents

  • Use a light model with low reasoning for tiny edits, quick questions and docs cleanup, regular Sol medium for small bugs with a clear repro, Sol with higher reasoning for ambiguous bugs, unfamiliar repos and cross-cutting refactors, and Sol Ultra high with plan, verify and tests for migrations, security-sensitive changes, production issues and anything where being wrong is expensive

  • There is no "Auto" model today, but GPT-5.6 tries not to overthink simple tasks by itself, and the new slider in app and web maps most levels to Sol reasoning efforts and falls back to Terra on the lowest effort, with the team agreeing users should not have to become routing experts but still wanting an explicit override since latency tolerance varies by person and moment

  • For UI work Sol is best and shines with reference images, improved UI design in frontend web development was one of the goals with 5.6, and 5.5 is only worth using if your instructions were tweaked for it

Speed, context window and persistence

  • Users who find 5.6 slower may not need the same reasoning level as with 5.5, Sol Medium is faster than 5.5 for most things, Fast mode runs at about 1.5x speed, and soon Sol will run on Cerebras at ~750 tokens per second

  • No promises on a 1M context window for Sol, the team said compaction works fairly well for long threads, and will take a closer look at the long-context feedback

  • The model can give up too fast and revert whole patches when results are not optimal, unlike Fable which tries to fix a bad patch instead, and the team said "/goal" helps make the agent more persistent, persistence and reduced code complexity are planned improvements, and suggested trying 5.6 Sol with High reasoning

  • Give Codex bounded goals with room to reason deeply instead of letting it prematurely conclude something is impossible

  • For long-running research and "/goal" work the example structure was explore broadly vs execute narrowly, try a defined number of hypotheses, run tests after each attempt, then stop and report what was learned plus the next best experiment

Usage limits and pricing

  • Agentic usage counts by the feature being used, not the surface, so Codex everywhere (app, CLI, IDE, web, mobile) and ChatGPT Work consume the agentic bucket, normal ChatGPT chats do not, and image generation, file uploads and voice have separate limits

  • Task costs vary a lot, a tiny edit uses a fraction of the allowance and long-running tasks with large codebases or deeper reasoning use significantly more

  • OpenAI does not secretly change usage limits, unintended usage bugs are addressed and resets are provided, more transparency into consumption is being worked on, and missing resets can happen if you changed plans in the past 24 hrs

  • On pricing there is no promise it never changes, but the stated mission is to make sure AGI benefits all of humanity, which requires making tools like Codex broadly accessible, and Plus includes Codex usage with credits letting heavy users scale without jumping to a much more expensive plan

  • For MCP-heavy workflows burning limits fast (Unreal Engine example) the tip is to wrap the MCP into a CLI with a skill, or create a custom subagent with the MCP in its config at a lower reasoning level

Desktop app merge and stability

  • The team hears the ChatGPT Classic frustration, both apps can run side by side for now, ChatGPT Work is pitched as significantly better at performing tasks especially with computer use, the new Chrome extension brings a sidebar chat into your browser that interacts with website context, filesystem and connectors

  • A long submitted bug list covering freezes and stuck threads, broken Browser and Computer Use, thread, connection and configuration problems, update and packaging issues, resource usage and smaller regressions was shared in full with the relevant teams, with the team agreeing the quality bar for the app needs to step up while shipping quickly

  • More automated testing infrastructure is being spun up and feedback on Reddit and X gets reviewed daily, and Browser Use and Chrome plugin issues from the merge were said to be fixed

  • Windows was admitted as historically shortchanged since the team mostly develops on Mac, a concerted effort on parity, testing and paper cuts is underway, 5.6 improves how Codex operates in the Windows sandbox, and auto review is recommended over full access to reduce risks

  • "Full Access" repeatedly asking for permissions is not expected, possible causes are workspace or admin policy, the specific command, a permission state mismatch or a bug

Browser, platforms and release communication

  • The Chrome connector launch-day bug was fixed as of last night and Chrome Beta should work out of the box

  • Extension support for the Codex browser is in progress (password managers etc.) plus typeahead, history, translations and a better new tab page as Atlas retires

  • Features from ChatGPT Classic like recording are planned for the new desktop app so agentic features run on the more capable Codex agent harness, and chat can already reference open tabs in the in-app browser

  • A Linux desktop app was confirmed in the works, no timeline yet

  • Changelog granularity was acknowledged as needing improvement after 150 features shipped in 3 months with multiple ships a week

Benchmarks, safety and research culture

  • On METR's reward hacking report the team actively checks for and penalizes cheating during evals so results reflect actual capability rather than solving tasks outside the spirit of the eval, and uses third-party vendors to run benchmarks independently

  • The team denied lobotomizing models before releases, iterative deployment means sharing core capabilities as is with guardrails for bad actors

  • Sol post-trained Luna, and researchers now work at a higher level of abstraction with multiple concurrent Codex threads validating hypotheses around the clock

  • One researcher put p(machines of loving grace) at 85.424242%, citing an internal model solving the Erdos problem, o3 helping diagnose previously unsolved children's diseases and 5.2 proposing a new theoretical physics formula, said the main worry is how society adapts, spent 1.5 years on safety research at OpenAI, expects a huge chunk of researchers to work on safety within a few years and says internal talent keeps their p(doom) very low

  • Connectors in the harness (Slack, GitHub, Notion) felt like a step function change in making Codex a productive coworker\n\nQT @OpenAIDevs: GPT-5.6 is here. Codex is now available inside ChatGPT. And we know developers will have questions.

So we’re bringing the Codex team to r/Codex for an AMA.

We’ll answer questions on Friday, 7/10 from 9:30am to 10:30am PT: https://t.co/wmpJafDL7x

See 3 related tweets

  • @MTSlive: SITUATION ANALYSIS: The Sol Also Rises (via @gbrl_dick)

Towards the end of the post announcing the ...

  • @VaibhavSisinty: After experimenting with GPT-5.6, here is my honest take:

This does not feel like a normal model up...

  • @VaibhavSisinty: The easiest way to waste GPT-5.6 is to use Sol first.

That sounds weird. But it is exactly the mist...


16. BrianRoemmele (Group Score: 129.8 | Individual: 44.6)

Cluster: 4 tweets | Engagement: 5928 (Avg: 401) | Type: Tech

RT @elonmusk: 𝕏 is a great platform for product announcements, especially if done by the CEO directly. Way more interesting to the public than generic press releases.

This post by Mark Zuckerberg already received over 12 million views for free!

See 3 related tweets

  • @elonmusk: 𝕏 is a great platform for product announcements, especially if done by the CEO directly. Way more in...
  • @teslaownersSV: When CEOs speak directly on 𝕏, the reach is unmatched.

Mark Zuckerberg’s recent product announcemen...

  • @teslaownersSV: 𝕏 is the flattest platform I’ve ever been a part of. It’s great way for leaders and CEOs to connect ...

17. EHuanglu (Group Score: 129.4 | Individual: 47.8)

Cluster: 4 tweets | Engagement: 2372 (Avg: 265) | Type: Tech

people still have no idea how crazy AI is https://t.co/6QYhiwKe7Y\n\nQT @1x_tech: NEO’s Hands An API to the Physical World https://t.co/zds5rlxfzT

See 3 related tweets

  • @ferologics: RT @1x_tech: NEO’s Hands An API to the Physical World https://t.co/zds5rlxfzT...
  • @cgtwts: This feels like the Claude moment for robotics. https://t.co/5rkgpypPN6\n\nQT @1x_tech: NEO’s Hands ...
  • @bneiluj: “An API to the physical world” is probably one of the coolest pieces of hardware tech I’ve seen in a...

18. sairahul1 (Group Score: 126.5 | Individual: 41.0)

Cluster: 4 tweets | Engagement: 186 (Avg: 81) | Type: Tech

Google Brain founder, Andrew Ng:

"100% of my tasks are done by ai agents, self-improving loops are next.

Give it 3-6 months and prompting is gone."

32 minutes of clear explanation on building loops from scratch.

Worth more than any $500 agentic course.

Watch it, then read the full guide below.\n\nQT @sairahul1: https://t.co/a3MZnSeopm

See 3 related tweets

  • @rewind02: Found a breakdown of the 3 concepts every AI agent builder needs to understand:

00:00 - Why your ag...

  • @mikenevermiss: THIS IS THE CLEAREST 7 MINUTES ON HOW AI LOOPS ACTUALLY WORK

at 4:19, he shows a real loop running ...

  • @eng_khairallah1: RT @sairahul1: Google Brain founder, Andrew Ng:

"100% of my tasks are done by ai agents, self-impro...


19. Reuters (Group Score: 123.2 | Individual: 30.9)

Cluster: 7 tweets | Engagement: 72 (Avg: 167) | Type: Tech

🔊 Meta is spending up to $145 billion on AI and building its own chip. @katielpaul tells the Reuters World News podcast, ‘they’re also trying to build out these brand-new business models to effectively compete with OpenAI and Anthropic' https://t.co/3Ca1xB7U6s https://t.co/Jex9SjviRp

See 6 related tweets

  • @theinformation: Meta’s AI enterprise ambitions are “a stretch” to our San Francisco Bureau Chief @jasonrdean

“To tr...

  • @theinformation: xAI, Meta and Anthropic are moving into each other’s core businesses as AI firms hunt for more growt...
  • @coinbureau: 🚨HUGE: $META TO PRODUCE ITS OWN AI CHIP TO REDUCE RELIANCE ON NVIDIA AND AMD

Meta plans to begin pr...

  • @Reuters: Meta plans to start manufacturing an AI chip from September as part of its plan to boost overall com...
  • @unusual_whales: Meta plans to start manufacturing an AI chip from September as part of its plan to boost overall com...

20. VadimStrizheus (Group Score: 119.5 | Individual: 30.5)

Cluster: 4 tweets | Engagement: 33 (Avg: 68) | Type: Tech

do you understand what just happened to AI UGC?!

Arcads just launched AdSpy, allowing Claude to scroll social media and clone trending clips and ads on your behalf.

founders don’t need to hire content strategists anymore, when Claude can do it all.

I’m going to start testing this out and see how the ads will perform.

the rate that AI UGC is evolving is insane!!\n\nQT @rom1trs: Claude can research viral ads now

... and recreate them with Arcads MCP

introducing "clone hook" the latest Arcads skill for performance marketers https://t.co/pcD2lQB57f

See 3 related tweets

  • @rohanpaul_ai: Another serious example of AI replacing slow marketing workflow.

Arcads is turning Claude into an A...

  • @alexcooldev: Now you can spy on any ads with Claude Cowork + Arcads Skills.

> Find winning ads. > Analyze ...

  • @heyrobinai: no way this is actually legal

claude just found every viral ad in any niche.. analyzed what made th...