Published on

科技热推精选 - 2026-06-05

Authors

2026年6月5日科技每日简报

Today's top tech conversations are led by @justjoshinyou13, whose post about 'RT @AnthropicAI: Our internal ...' garnered the highest engagement. Key themes trending across the top stories include https, built, agents, building, model. The community is actively discussing recent developments in AI, engineering practices, and startup strategies.


1. justjoshinyou13 (Group Score: 991.2 | Individual: 56.4)

Cluster: 32 tweets | Engagement: 3284 (Avg: 318) | Type: Tech

RT @AnthropicAI: Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor.

It’s happening faster than we thought, and the implications deserve greater attention. https://t.co/OVVPJO7VQx

See 31 related tweets

  • @scaling01: Claude Mythos speeds up training code of small AI models by 52x

Humans need 4-8 hours to reach a 4x...

  • @scaling01: Anthropic is shipping 3.2x more code per person with Mythos nowadays than with Opus 4.5 around half ...
  • @scaling01: Mythos seems to have improved Claude Code success rate on open-ended task from 40% to around 70% htt...
  • @MTSlive: SITUATION EXPLAINED: Anthropic just published data on how fast AI itself is accelerating AI developm...
  • @scaling01: Anthropic: AI written code is now as good as human written code and will be strictly better within t...

2. lmsysorg (Group Score: 487.7 | Individual: 36.8)

Cluster: 16 tweets | Engagement: 29 (Avg: 20) | Type: Tech

🎉 Meet Nemotron 3 Ultra from @nvidia, a frontier reasoning model, with 550B total params (55B active) built for long-running autonomous agents.

Day-0 support is now live in SGLang, plus Day-0 RL support with Miles: GRPO training on 128 H200s in colocate mode, with DP attention unlocking large-scale EP for Mamba-hybrid MoE.

1️⃣ Hybrid Mamba-Transformer MoE 2️⃣ Same NVFP4 checkpoint runs on Hopper & Blackwell 3️⃣ MTP for faster multi-turn generation 4️⃣ Up to 1M token context 5️⃣ Leads open models on agentic, coding & instruction-following benchmarks

Cookbook: https://t.co/C4qnqjp2xN Run it now with SGLang!\n\nQT @NVIDIAAI: Today we're shipping Nemotron 3 Ultra.

A 550B MoE frontier-intelligence open model built for long-running agents.

It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models. https://t.co/FEXqvfzQFO

See 15 related tweets

  • @cryptopunk7213: this is huge. the US is reclaiming the #1 spot for open source AI models with NVIDIA's new nemotron ...
  • @nvidia: Introducing NVIDIA Nemotron 3 Ultra.

A frontier smart open model built for long-running agents that...

  • @huggingface: RT @NVIDIAAI: Today we're shipping Nemotron 3 Ultra.

A 550B MoE frontier-intelligence open model bu...

  • @yacineMTB: open weights and open data thank you for everything\n\nQT @NVIDIAAI: Today we're shipping Nemotron...
  • @ModelScope2022: Meet Nemotron 3 Ultra @NVIDIAAI 's frontier open reasoning model built for long-running AI agents.🚀 ...

3. Cloudflare (Group Score: 391.6 | Individual: 60.5)

Cluster: 14 tweets | Engagement: 3244 (Avg: 249) | Type: Tech

VoidZero, the team behind Vite, Vitest, Rolldown, Oxc, and Vite+, is joining Cloudflare. Vite stays open source, vendor-agnostic, and built for everyone. https://t.co/DJTpX4Q9Xt

See 13 related tweets

  • @whoiskatrin: RT @voidzerodev: VoidZero is joining Cloudflare.

Our mission stays the same: to make JavaScript dev...

  • @ritakozlov: BIG DAY! @voidzerodev is joining @cloudflare 🚀

before anything else: @vite_js is very much remaini...

  • @Rasmic: We’re entering an era where people won’t care what framework they are using… they just want it to wo...
  • @eastdakota: The best tools to build the agentic future: all native on @Cloudflare. Welcome to the team, VoidZero...
  • @rauchg: Congrats Void team!

We @vercel reaffirm our collaboration on an open platform for the web, with our...


4. StockSavvyShay (Group Score: 263.9 | Individual: 35.9)

Cluster: 12 tweets | Engagement: 1140 (Avg: 662) | Type: Tech

Goldman Sachs projects SpaceX AI revenue could grow 100x to 322Bby2030makingAIthecorecasebehinditsreported322B by 2030 making AI the core case behind its reported 1.8T IPO valuation.

The forecast assumes total revenue reaches $474B by 2030 with AI becoming the largest segment ahead of Starlink and launch. https://t.co/OSjgukCi4h

See 11 related tweets

  • @DeItaone: GOLDMAN SEES SPACEX AI REVENUE EXPLODING TO $322B BY 2030

Goldman Sachs projects SpaceX AI revenue ...

  • @IPONewsroom_: GOLDMAN SACHS PROJECTS SPACEX'S AI REVENUE WILL GROW 100X BY 2030

The bank's projections are the fo...

  • @StockMKTNewz: GOLDMAN SACHS PROJECTS SPACEX'S AI REVENUE WILL GROW 100X BY 2030

*GOLDMAN SACHS IS ONE OF THE BANK...

  • @chiefofautism: sell side working overtime to get retail in on this IPO\n\nQT @DeItaone: GOLDMAN SEES SPACEX AI REVE...
  • @MikeIsaac: alt hed: bankers in charge of securing high valuation for IPO say “this company is good”\n\nQT @DeIt...

5. MTSlive (Group Score: 174.2 | Individual: 49.8)

Cluster: 6 tweets | Engagement: 1898 (Avg: 85) | Type: Tech

SITUATION DETECTED: Sam Altman, Dario Amodei, and Demis Hassabis have signed a joint open letter calling on Congress to mandate screening of synthetic nucleic acid orders, citing AI’s rapidly improving ability to assist with biological research as an urgent biosecurity risk.

See 5 related tweets

  • @deredleritt3r: It's great to see the entire industry come together like this (except for one particular major AI la...
  • @MTSlive: We asked @OliviaHelenS from @IFP and Josh Wentzel from @JoinFAI how they got Sam, Dario, and Demis t...
  • @AlecStapp: RT @fiiiiiist: Today, @IFP and @JoinFAI released an open letter calling for mandatory screening of o...
  • @AIRiskExplorer: According to the Global Synthesis Map, developed by @IBBIS_bio, the US has 107 nucleic acid synthesi...
  • @AIRiskExplorer: Open letter calls US legislators to mandate nucleic acid synthesis screening and recordkeeping, citi...

6. kilocode (Group Score: 166.7 | Individual: 35.1)

Cluster: 8 tweets | Engagement: 106 (Avg: 44) | Type: Tech

<<< NVIDIA Nemotron 3 Ultra is live in Kilo, and it's completely free for a limited time. >>>

550B params (55B active), 1M token context, and the top US open-weights score on the Artificial Analysis Intelligence Index.

Pick it from the model selector and start building. https://t.co/GslpqAe7Ye

See 7 related tweets

  • @TeksEdge: If you're looking for another token provider, @nvidia's recent medium model champion, Nemotron 3 Ult...
  • @TFTC21: NVIDIA announces Nemotron 3 Ultra, an open model built for long-running AI agents that need to plan,...
  • @cline: NVIDIA just released Nemotron 3 Ultra, a 550b-parameter agentic coding model with a 1m context windo...
  • @opencode: Nemotron 3 Ultra is now free on OpenCode

text · 1M context · fully open source

NVIDIA's latest ope...

  • @BrianRoemmele: New open source NVIDIA Nemotron 3 Ultra (BF16 weights) is quite good!

I am testing it now.

It is ...


7. scaling01 (Group Score: 163.3 | Individual: 35.6)

Cluster: 5 tweets | Engagement: 207 (Avg: 219) | Type: Tech

this is brutal https://t.co/BplCoGT5IJ\n\nQT @arena: Introducing Agent Arena: real-world agentic evals at scale.

How do you evaluate agents doing actual work? We measure millions of live sessions where real users accomplish real tasks.

On Arena, models now get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide deck, researching the web, building apps, and analyzing documents.

Every session produces rich signals. Users iterate with the agent turn-by-turn: approving, editing, correcting, praise or expressing frustration. The environment gives feedback too: shell errors, tool failures, recovery attempts, and more.

Our leaderboard measures each model's agentic performance using causal inference across five signals: task success, steerability, error recovery, user praise vs. complaint, and tool hallucination.

This leaderboard snapshot is built from 300K+ tasks, 2M+ tool calls, and 40M lines of code by agents.

Top labs in Agent Arena:

  • #1 @OpenAI: GPT-5.5 (High)
  • #2 @AnthropicAI: Claude-Opus-4.7 (Thinking)
  • #3 @Zai_org: GLM-5.1
  • #4 @GoogleDeepMind: Gemini-3.1-Pro
  • #5 @Kimi_Moonshot: Kimi-K2.6

More analysis in the thread, with the full technical blog below.

See 4 related tweets

  • @AnjneyMidha: Interesting

One of the hardest unsolved problems in frontier systems today is scalable evaluation o...

  • @dejavucoder: arena dot ai introduced agent mode and they are recording live sessions of users who use the agent m...
  • @petergostev: This is our most important eval yet - Agent Arena - it measures real performance of models on real a...
  • @petergostev: RT @ml_angelopoulos: Agent Arena gives every model access to a Claude-Code-like harness and a comput...

8. MTSlive (Group Score: 154.1 | Individual: 48.1)

Cluster: 5 tweets | Engagement: 811 (Avg: 85) | Type: Tech

SITUATION DETECTED: OpenAI has launched a new memory system for ChatGPT called Dreaming, which automatically synthesizes and updates user context in the background across conversations without requiring explicit save requests.

Rolling out to Plus and Pro users in the US today.

See 4 related tweets

  • @AndrewCurran_: Rolling out for both Pro and Plus users this morning. Real memory changes a lot of things.

'Today, ...

  • @mark_k: ChatGPT memory by @OpenAI is getting much more interesting.

The new "dreaming" system is not just a...

  • @MTSlive: ChatGPT be like: https://t.co/tN0lsNl3cI\n\nQT @MTSlive: SITUATION DETECTED: OpenAI has launched a n...
  • @Cointelegraph: ⚡️ NEW: OpenAI has launched a more advanced memory system for ChatGPT that helps the AI retain fresh...

9. Parul_Gautam7 (Group Score: 147.9 | Individual: 32.7)

Cluster: 5 tweets | Engagement: 104 (Avg: 92) | Type: Tech

I've always liked AI avatars for speed, but they often end up feeling repetitive.

What interests me here is the ability to keep the same identity while creating more cinematic and varied scenes.

That feels much closer to how real content gets made.\n\nQT @HeyGen: Cinematic_avatar api is live

keep your likeness, add cinematic range and build your video pipeline via your coding agent

install HeyGen CLI + HyperFrames skill to create launch videos like ours

docs and cli setup in thread ↓ https://t.co/TEpepD94yw

See 4 related tweets

  • @FellMentKE: HeyGen is quietly becoming the video infrastructure layer for the agentic era.

  • An API for your id...

  • @Origin_AI_01: The future of video is programmable.

Keep your identity. Expand your cinematic range. Let AI agents...

  • @TheoBuildsAI: I’ve been experimenting with coding agents recently.

The idea of generating the avatar, creating sc...

  • @FellMentKE: RT @HeyGen: Cinematic_avatar api is live

keep your likeness, add cinematic range and build your vid...


10. levie (Group Score: 139.1 | Individual: 34.9)

Cluster: 5 tweets | Engagement: 843 (Avg: 765) | Type: Tech

The jobs data coming out continues to suggest the opposite of what a lot of people had thought would happen.

Just take engineering, as the prime example of the area with greatest AI impact (and perceived risk). Most companies now have far more software projects than ever before because of AI, and effectively only engineers are going to be the ones doing that work.

You can get by for a while by being non-technical building software, but eventually someone has to understand what the thing is that got built, has to maintain it, has to fix security issues that come up, upgrade the systems beneath it, and so on. That’s all jobs.

Now apply that to a number of other job functions. AI is going to cause companies to hire more in sales because agents can let them process more leads and do more customer research. AI will cause an explosion of new marketing roles because of how much more efficient it is to launch campaigns and target. The list goes on.

AI is going to have the opposite effect that lots of people thought on jobs.\n\nQT @KobeissiLetter: What if AI is actually creating more jobs than it is replacing?

The latest JOLTs data showed that US job openings surged by a massive 731,000 jobs in April.

Markets were expecting no change, resulting in the largest beat in JOLTs history.

As a result, available employment hit 7.6 million for the month, the highest since May 2024.

And, job openings in the professional and business services sector surged by a massive 668,000.

The labor market's bull case from AI is underpriced.

See 4 related tweets

  • @dharmesh: RT @levie: The jobs data coming out continues to suggest the opposite of what a lot of people had th...
  • @owenbjennings: even if it happens to be the case that the number of ppl required to build a given piece of software...
  • @FirstSquawk: AI REMAINS THE LEADING CAUSE OF U.S. JOB CUTS FOR A THIRD STRAIGHT MONTH; 88,000 LAYOFFS HAVE BEEN A...
  • @MTSlive: Is AI actually creating more jobs?

@MattBurtell of @A1Policy:

"AI revenues for OpenAI and Anthropi...


11. esrtweet (Group Score: 137.9 | Individual: 44.1)

Cluster: 7 tweets | Engagement: 1209 (Avg: 459) | Type: Tech

Shorter Sam Altman: the AI bubble is popping.

Make no mistake, it's a hugely useful technology and uptake will continue, even accelerate. But the overinvestment in datacenters that we've been seeing is not sustainable; the business model of the big providers doesn't work, and is floating on VC money.

It's going to get worse. If customers are cutting back on token spend even at the artificially low prices they have now, what do you think they'll do when the big providers dramatically raise their rates in an effort to get to profitability?\n\nQT @BusinessInsider: Sam Altman said AI budgeting has recently become a "huge issue" for some companies, something that "never came up" earlier this year. https://t.co/P2zODBNmDp

See 6 related tweets

  • @Polymarket: JUST IN: Sam Altman says AI budgeting has suddenly become a “huge issue” for companies....
  • @alex: RT @BusinessInsider: Sam Altman said AI budgeting has recently become a "huge issue" for some compan...
  • @godofprompt: We replaced a 60kjuniordeveloperwithaproactiveAIagentthatcasuallyranupa60k junior developer with a proactive AI agent that casually ran up a 150k API bill ...
  • @aiedge_: For the first time EVER, Sam Altman admits AI token spending is halting.

Many companies are trying ...

  • @nalinrajput23: Just Simplified :

AI has moved from the innovation budget to the finance department's problem.

Th...


12. googleaidevs (Group Score: 107.8 | Individual: 25.7)

Cluster: 5 tweets | Engagement: 7 (Avg: 55) | Type: Tech

Play our new open-weights music model, @GoogleMagenta RealTime 2, using a MIDI keyboard, live text prompts, and even hand gestures ✌️ https://t.co/Hgr9gxDsoD\n\nQT @GoogleMagenta: Introducing Magenta RealTime 2 (MRT2): the live music model you can play as an instrument.

MRT2 offers MIDI and prompt controls, and runs natively on a MacBook with <200ms latency.

Open weights. Open source inference engine. Suite of apps and plugins.

Hear what it can do and try it out for yourself below 🧵

See 4 related tweets

  • @TeksEdge: Google is getting into the AI generated music scene with its new open source RealTime 2 music model....
  • @thorwebdev: Magenta RealTime 2: Open & Local Live Music Models 🪩 https://t.co/lT1bILf0nw\n\nQT @osanseviero:...
  • @kazunori_279: RT @GoogleMagenta: Introducing Magenta RealTime 2 (MRT2): the live music model you can play as an in...
  • @KyeGomezB: RT @HuggingPapers: Google just released Magenta RealTime 2 on Hugging Face

The only open-weights mo...


13. wallstengine (Group Score: 100.8 | Individual: 28.3)

Cluster: 5 tweets | Engagement: 143 (Avg: 127) | Type: Tech

$AMZN announced a €10B investment in its European 🇪🇺 fulfillment network and introduced an upgraded AI-powered Proteus warehouse robot.

The new Proteus can respond to conversational prompts, prioritize tasks, plan routes, and move across warehouse floors. It is expected to arrive in Europe in 1H27.

Amazon also plans to roll out its STARK robotic tote-handling system to 15 European sites by 2027 and launch 25+ sub-same-day delivery sites across Europe this year.

Alexa+ is expected to launch in 10 additional countries in 2027.

See 4 related tweets

  • @wallstengine: AMAZON EXPANDS ROBOTICS, FAST DELIVERY AND WORKER TRAINING

$AMZN unveiled a next-gen Proteus robot ...

  • @StockMKTNewz: AMAZON $AMZN JUST ANNOUNCED A MASSIVE EUROPEAN PUSH AT ITS "DELIVERING THE FUTURE" EVENT IN LONDON

...

  • @StockSavvyShay: $AMZN is expanding robotics and faster delivery with new Proteus robots, STARK automation and Amazon...
  • @Techmeme: Amazon unveils its next-gen Proteus warehouse robot, adding AI-powered language capabilities to let ...

14. jerryjliu0 (Group Score: 100.0 | Individual: 34.7)

Cluster: 4 tweets | Engagement: 90 (Avg: 69) | Type: Tech

We're presenting ParseBench at CVPR 2026!

ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k pages of real-world enterprise documents ✅ It has comprehensive evaluation metrics around tables, charts, visual grounding, semantic formatting, and content faithfulness

The core goal is measuring whether models can semantically interpret a document in the right way, without having models overfit to our precise benchmark.

Parsing 100% of PDFs to 100% accuracy is the final boss for document OCR. In general, the latest frontier models have been tuned for coding, math, and scientific reasoning as opposed to precise visual understanding; hope more benchmarks that these will encourage overall progress towards solving this problem!

Poster is below. If you want to learn more come check out our site or 30-page ArXiv paper:

ParseBench: https://t.co/PWczfhp0OX ArXiv: https://t.co/2dEJIaBBkr\n\nQT @llama_index: We're presenting ParseBench at CVPR 2026 today. 🦙

Come learn why document understanding is an AGI-complete problem (an agent can't act on a doc it can't correctly read, and reading a real enterprise table is harder than it looks).

The first doc-parsing benchmark built for AI agents:

2,000+ human-verified pages 167K+ test rules 5 dimensions: tables, charts, faithfulness, formatting, grounding

Fully open source. 📍 Talk TODAY, June 4, 9–10 AM at CVPR. Come say hi 👇 🤗 https://t.co/skla84GVTc 💻 https://t.co/h7SpuTWYVn 📄 https://t.co/VnKcb48oJl

See 3 related tweets

  • @llama_index: We're presenting ParseBench at CVPR 2026 today. 🦙

Come learn why document understanding is an AGI-c...

  • @jerryjliu0: Our team is at CVPR 2026 if you want to come say hi :) https://t.co/exjQNEIALk\n\nQT @jerryjliu0: We...
  • @llama_index: RT @jerryjliu0: We're presenting ParseBench at CVPR 2026!

ParseBench is the most comprehensive doc...


15. MarcoSalzmann80 (Group Score: 98.9 | Individual: 34.5)

Cluster: 3 tweets | Engagement: 0 (Avg: 33) | Type: Tech

🧵 Apple. Google. Nvidia. Intel. Hedera.

The next generation of AI is no longer being built by a single company.

It is emerging as a layered technology stack, where each participant provides a critical piece of the infrastructure.

Reports suggest Apple’s upgraded Siri will rely on Google’s Gemini models running through Google Cloud on Nvidia Blackwell infrastructure.

If accurate, this represents a significant shift.

Even companies with highly integrated ecosystems may increasingly rely on specialized partners to deliver next-generation AI experiences.

Apple’s iPhone is no longer just a smartphone.

It is a computing platform used by hundreds of millions of people worldwide.

If Siri evolves into a true AI assistant, Apple is not simply improving a feature.

It is creating an AI access layer for one of the largest consumer ecosystems on the planet.

The architecture behind that experience is becoming increasingly specialized.

Apple provides the user interface.

Google provides the cloud infrastructure and AI models.

Nvidia provides the computational power.

Intel provides trusted execution environments.

Each layer solves a different problem.

Google’s role extends beyond AI models.

It provides the cloud infrastructure capable of processing AI workloads at global scale.

Google is also a founding member of the Hedera Governing Council, highlighting a long-standing interest in trusted digital infrastructure.

Nvidia’s Blackwell architecture is rapidly becoming foundational infrastructure for advanced AI systems.

But Blackwell is not only about performance.

It also introduces Confidential Computing capabilities designed to protect sensitive workloads during execution.

This is where @EQTYLab enters the picture.

EQTY Lab uses secure hardware environments from Nvidia and Intel to generate cryptographic attestations for AI workloads.

Verifiable Compute roots trust directly in silicon and extends that trust through cryptographic verification.

As AI expands into finance, healthcare, government, defense and enterprise environments, verification becomes increasingly important.

Organizations will need to know not only what an AI system produced, but how it arrived there.

Verifiable Compute addresses that challenge through cryptographic attestations.

Trust has to happen in real time.

Inside the compute process itself.

These attestations can show what happened, where it happened and under which rules it happened.

EQTY Lab’s Verifiable Compute is built with Nvidia and Intel hardware and generates cryptographic attestations for AI workloads.

Those attestations are then anchored on Hedera through the Hedera Consensus Service.

That is the trust layer.

The convergence of AI, cloud infrastructure, secure hardware and cryptographic verification points toward a larger trend.

The future of AI may not be defined solely by the quality of models.

It may also be defined by the ability to verify, govern and trust the systems behind them.

The long-term winners of the AI era may not be limited to model developers.

They may also include the platforms responsible for verification, governance, compliance and auditability.

Everyone is watching the AI.

Few are watching the trust layer underneath.

[ APPLE ] └─> AI Interface & User Experience │ ▼ [ GOOGLE ] └─> Gemini Models & Cloud Infrastructure │ ▼ [ NVIDIA & INTEL ] └─> Secure AI Compute & Trusted Execution │ ▼ [ EQTY LAB ] └─> Verifiable Compute & Cryptographic Attestations │ ▼ [ HEDERA ] └─> Immutable Trust & Verification Layer

$HBAR

👉🏻 https://t.co/SRE4pqFYvp…

👉🏻 https://t.co/fjnHRHxNDI\n\nQT @WatcherGuru: JUST IN: Apple AAPLwilluseGooglesAAPL will use Google's GOOGL Nvidia-powered chips for its overhauled Siri launching in September. https://t.co/qQ5N8WsKuB

See 2 related tweets

  • @shanaka86: JUST IN: The company that taught the world to control everything just admitted it cannot win AI alon...
  • @MarcoSalzmann80: RT @MarcoSalzmann80: 🧵 Apple. Google. Nvidia. Intel. Hedera.

The next generation of AI is no longer...


16. wallstengine (Group Score: 98.4 | Individual: 27.3)

Cluster: 5 tweets | Engagement: 26 (Avg: 127) | Type: Tech

Guide Labs launched Clarity, an invite-only preview for an interpretable AI platform powered by Steerling 8B.

It lets users see concepts driving model outputs, trace responses to training data & steer behavior by amplifying or suppressing concepts without rewriting prompts.\n\nQT @guidelabsai: The first inherently interpretable AI platform is finally here. Welcome to Clarity. https://t.co/IprV5QTMOL

See 4 related tweets

  • @alex_prompter: We spent 3 years learning how to talk around the black box.

Clarity lets you reach inside it.

See ...

  • @cgtwts: POV: you can finally understand what your AI is saying without spending hours double-checking it. ht...
  • @ianmiles: Now you can finally understand what your AI is saying, without spending hours trying to figure it ou...
  • @initialized: RT @guidelabsai: The first inherently interpretable AI platform is finally here. Welcome to Clarity....

17. chandrarsrikant (Group Score: 97.6 | Individual: 23.3)

Cluster: 5 tweets | Engagement: 48 (Avg: 846) | Type: Tech

Google parent is raising $85 billion for AI. Here's how Sundar Pichai plans to spend it

Interestingly, this fundraise, Alphabet's first equity sale since 2006 and the largest equity offering in history, comes just ahead of the highly anticipated mega IPOs of AI leaders SpaceX, Anthropic, and OpenAI.

Alphabet is also an investor in SpaceX and Anthropic.

By @tsuvik

https://t.co/4G0ywHfQn5

See 4 related tweets

  • @chandrarsrikant: Came and scooped up 45Billion,withplansforanother45 Billion, with plans for another 40 Billion. Used the window just ahead of t...
  • @Reuters: Alphabet to raise $84.75 billion in upsized equity offering to fund AI ambitions https://t.co/RbbWrW...
  • @Cointelegraph: 🔥 LATEST: Google's parent Alphabet upsizes its equity offering to $84.75 billion to fund AI infrastr...
  • @moneycontrolcom: #AIWithMC | Google parent is raising $85 billion for AI.

Alphabet’s fundraise, first since 2006, ...


18. sydneyrunkle (Group Score: 96.8 | Individual: 22.5)

Cluster: 6 tweets | Engagement: 43 (Avg: 32) | Type: Tech

transitions like this are why we think it's helpful to have a provider-agnostic harness

we used to talk more about swapping models when the latest and greatest came out -- but the latest and greatest from major providers are expensive!

we're seeing more folks swap to open source models for cost reasons, which is easy with deepagents and/or langchain\n\nQT @Altimor: Pulled the trigger today and switched 100% of Lindy traffic to DeepSeek v4, churning from Anthropic models. Saves us millions of $ and we're actually seeing an increase in performance on many core use cases. Transformative for the business.

See 5 related tweets

  • @pmarca: Interesting.\n\nQT @Altimor: Pulled the trigger today and switched 100% of Lindy traffic to DeepSeek...
  • @dee_bosa: 👀\n\nQT @Altimor: Pulled the trigger today and switched 100% of Lindy traffic to DeepSeek v4, churni...
  • @Thom_Wolf: RT @Altimor: Pulled the trigger today and switched 100% of Lindy traffic to DeepSeek v4, churning fr...
  • @LangChain_OSS: open models 🤝 open harness\n\nQT @sydneyrunkle: transitions like this are why we think it's helpful ...
  • @hwchase17: RT @sydneyrunkle: transitions like this are why we think it's helpful to have a provider-agnostic ha...

19. Scobleizer (Group Score: 94.0 | Individual: 33.6)

Cluster: 3 tweets | Engagement: 17 (Avg: 272) | Type: Tech

Interesting new model shipped today.\n\nQT @hangg70: we made a new model for text-to-image generation and editing. the results are looking good and the leaderboard is looking strong. it turns out that nano banana 2 is not impossible to beat, which felt like the case at the beginning of the year. there are a lot of great models out there that get released often. why should you care about reve 2.0?

to me, there are mainly two reasons. one being that reve is an underdog, reasonably funded but magnitudes less than other big labs, e.g. oai, google, meta, etc. you might be curious about how we managed to make it to the top. two being that reve 2.0 is a decent model, and we as a team are willing to talk openly about some of our learnings and thoughts that could be helpful. in this post, i want to share mine on reve 2.0 and multimodal in general as a person working on it.

first things first, reve 2.0 is a pixel diffusion model with a thing that we call "layout" as the rendering representation. these two things are our research bets that turned out to work amazingly well. pixel diffusion lets us go 4k without sacrificing quality or speed. layout lets us scale better and have better control, which are two sides of the same coin. the field standard has been to use long upsampled prompts for rendering. yet this results in an awkward situation where captioners and users need to describe precise controls with text, which can be inaccurate. this inaccuracy amounts to bad reconstruction and control at test time. it gets worse with scale. and this inherent ambiguity is a curse in current multimodal generators. so what's a layout? a layout is a css of an image, which can be either defined by humans or learned by models. we end up capitalizing a lot on regions, which are good for 2D space. yet this idea naturally generalizes. it turns out to be a standard VLM mid-training task, and that's solvable in good hands. it also brings many good properties in pretraining and post-training, which i am not going to expand on. ideogram independently verified that layout is useful (released on the same day, congrats!). to be clear, these bets are not novel, but to put together a system that makes them work is (and showing it beats nano banana 2).

second, it's nice that these bets, among others, worked out. however, like in many cases, there was a long time when things were underperforming. our competitor models are great, and most likely didn't make many risky bets. it is a big pipelining and engineering problem. why should we risk it? in retrospect, the culture of our team and leadership helped a lot. our priorities didn't swing and have stayed focused during our development. the idea makes sense, the execution is good, if things don't work out it's a bug, let's go find it and try more things. by and large, reve remains a research lab with big computers. this is rare. let me tag some amazing ppl here: @Taesung @m_gharbi @Songwei_Ge @TianweiY James Hong @dima_smirnov_ @theSidlak, ... the list goes on.

third, we spent most of our time improving text-to-image and didn't do much on editing. and our arena ranks show that. to date, we are #2 on text-to-image yet #9 on image editing. it's honestly a bit embarrassing that we didn't do well in editing, as layout promises to do well. but i am confident that this will improve, as we are juggling bandwidth and resources (we are a small team, and hey, come join us!).

fourth, talking about leaderboards and the state of multimodal, i genuinely feel that the gap between labs is shrinking. compared to LLMs, multimodal gen is at least half a year to a year behind. i am talking about architectures and core pipelines. to do good multimodal, you need to do good LLMs. reve has been helped by the OSS community a lot, but we've realized we need to own our language stack. and scaling follows naturally. leaderboards, in turn, are a noisy approximation and average of the real environments that you care about in deployment. they chase scaling and generalizable post-training. reve 2.0 ended up not being driven much by leaderboard evaluation, but relying on our intuition instead.

finally, how can multimodal be more useful? this is a question that keeps me up at night. coding has found its product-market fit and is driving up societal productivity. how can multimodal do that too? to me, we are nailing a single-round rollout that leads to an infinite one. this infinite rollout will drive our digital interaction and creation. for this rollout to be good, it needs to be precise. otherwise rollout efficiency is too low for either humans or agents. we are making bets and concrete progress towards that goal, such as converting images into a css-like layout. if you are interested in this topic, i recommend @stuffyokodraws's post for a high-level digest: https://t.co/mg42EZkd2h. the success of multimodal depends on whether or not it can find a good product-market fit. that's the top question to figure out, then it's the model. it's quite non-linear to be honest, as critical pieces are still missing. but to me it's an area worth pouring my thoughts and efforts into.

give our model a spin, try your tasks, move some boxes. in case you find any bugs, please let me know in a reply or DM. hope it can help you.

See 2 related tweets

  • @mark_k: Reve 2.0 by @reve is interesting because it is not just another prompt-to-pixels image model.

The c...

  • @arena: RT @hangg70: we made a new model for text-to-image generation and editing. the results are looking g...

20. MTSlive (Group Score: 93.3 | Individual: 35.4)

Cluster: 3 tweets | Engagement: 13 (Avg: 85) | Type: Tech

How do we scale biotech while safeguarding against biological risk?

We asked American Wetware CEO @p_maverick_b:

"As we get better at engineering biology, we should also get better at biosecurity."

"Better AI models like AlphaFold mean it's more predictive, 'I am pretty sure this protein binds the receptor I'm trying to drug.' That's a big step forward."

"As the models get better at reasoning and learning the why of what they're trying to do, that improved reasoning also helps biosecurity."

"A very smart AGI version of Claude that has a constitution where it's decided not to develop bioweapons should be able to be more active in screening people who are trying to do bad things."\n\nQT @AndrewCurran_: Sam Altman, Dario Amodei, Demis Hassabis and many others have signed a letter urging Congress to increase security on orders of synthetic nucleic acids - and the equipment needed to make them - as models continue to become increasingly bio-capable. https://t.co/JLw1Iq51Fx

See 2 related tweets

  • @MTSlive: What's the right way to handle AI biosecurity?

We asked @MattBurtell of @A1Policy:

"The letter foc...

  • @MarioNawfal: 🇺🇸 The people building artificial intelligence are warning Congress that their own technology could ...