Published on

热门科技推文 - 2026年7月4日

Authors

2026年7月4日科技每日简报

Today's top tech conversations are led by @petergyang, whose post about 'How I’m getting the most out o...' garnered the highest engagement. Key themes trending across the top stories include https, fable, model, hours, claude. The community is actively discussing recent developments in AI, engineering practices, and startup strategies.


1. petergyang (Group Score: 165.5 | Individual: 37.5)

Cluster: 6 tweets | Engagement: 89 (Avg: 103) | Type: Tech

How I’m getting the most out of Fable before July 7:

  1. Prep context with cheaper models
  2. Plan with Fable, execute with another model
  3. Use lower effort, like Medium, and babysit what Fable is doing

📌 Watch my full tutorial for 5 Fable-worthy use cases to try: https://t.co/XElMEV3FwK\n\nQT @petergyang: Claude Fable 5 is finally back, but you only have until July 7 to use it on your Claude subscription.

I made a new tutorial walking through 5 use cases worth trying Fable on:

→ Find Fable-worthy work → Get life and business advice → Make projects ship-ready → Plan the next big thing → Refactor your project or codebase

As usual, it’s no BS, and I show you Fable’s actual output.

📌 Watch now: https://t.co/XElMEV3FwK

See 5 related tweets

  • @HarryTandy: Amanda Askell, Anthropic technical staff:

"Good prompters should be very experimental."

A second b...

  • @aiedge_: Stop treating Fable 5 like every other AI model.

There are six things you need to know if you want ...

  • @gregisenberg: Quick PSA for anyone using Claude:

Fable 5 is back, but it's ONLY included through July 7.

After ...

  • @alex_prompter: RT @alex_prompter: Fable 5 performs worse when you over-prompt it. Anthropic put that in writing.

t...

  • @_simonsmith: This is more aligned with my experience than some of the claims that Fable has been nerfed. I've exp...

2. cryptopunk7213 (Group Score: 165.2 | Individual: 35.6)

Cluster: 7 tweets | Engagement: 61 (Avg: 51) | Type: Tech

wow, so both meta and spaceX have 10-15T models currently being trained.

people forget ceos like zuck and elon have deep pockets that will scale data center compute quicker than anyone else

if we assume compute = better model then there’s a timeline where meta and spaceX catch up to the frontier

gap is closing with every new release.\n\nQT @alexandr_wang: First, Mark was clearly talking about the industry’s progress on agentic capabilities on the whole.

But, while we’re on the topic: Our next Muse Spark update is coming soon. Big improvements in coding and agentic capabilities to be more competitive with other leading models.

Excited to get these into your hands—will be rolling out to Meta AI and our new API!

See 6 related tweets

  • @kimmonismus: Meta is expected to release an Opus model in the very near future.

Reportedly, Mark Zuckerberg’s di...

  • @wallstengine: RT @alexandr_wang: First, Mark was clearly talking about the industry’s progress on agentic capabili...
  • @negligible_cap: Wang’s comments seem pretty bearish as a whole. If just $META was struggling with agents then that’s...
  • @wallstengine: 👀 https://t.co/wsNfoQ7t2B\n\nQT @alexandr_wang: First, Mark was clearly talking about the industry’s...
  • @GergelyOrosz: The industry’s progress on agentic capabilities is pretty incredible, esp the last few months. Cloud...

3. MTSlive (Group Score: 160.4 | Individual: 27.3)

Cluster: 9 tweets | Engagement: 109 (Avg: 126) | Type: Tech

SITUATION UPDATE: Alibaba has banned employees from using Anthropic’s Claude Code at work after the tool drew scrutiny for features that can identify China-linked users, per Reuters.

The ban follows Anthropic’s accusation that Alibaba was distilling Claude.

See 8 related tweets

  • @Techmeme: Sources: Alibaba has banned employees from using Claude Code and asked them to remove all Claude mod...
  • @FirstSquawk: ALIBABA TO BAN CLAUDE CODE INTERNALLY OVER ALLEGED SECURITY RISKS

Alibaba is reportedly set to ban ...

  • @coinbureau: ⚡UPDATE: CHINA'S ALIBABA BANS STAFF FROM USING CLAUDE CODE

Chinese giant firm Alibaba will ban empl...

  • @Reuters: Alibaba to ban Claude Code in workplace over alleged backdoor risks, source says https://t.co/MPCzdy...
  • @Cointelegraph: 🚨 LATEST: Alibaba is reportedly banning employees from using Anthropic’s Claude Code over alleged ba...

4. BrianRoemmele (Group Score: 158.9 | Individual: 32.4)

Cluster: 6 tweets | Engagement: 186 (Avg: 216) | Type: Tech

HIDDEN AI TEXT OUTPUT WATERMARKS!

When this full breaks open you will see the world in a new way. The hidden AI watermarks on text output will hit some folks really hard.

Out soon only at https://t.co/Ruqey26qLY. https://t.co/4bKsrDKD9x\n\nQT @BrianRoemmele: ⚠️ WARNING⚠️

Two AI companies turned up the text based watermarking technology they use on most paragraph long or longer text output!

There are even serial numbers tracked to you hidden in some simple text outputs.

My exclusive research you will find no place else.

I show you how to find it and remove it.

It’s a big deal.

Exclusive https://t.co/tcKeuiQyql article out soon.

See 5 related tweets

  • @BrianRoemmele: PROJECT ERASER

The History and Neutralization of Invisible Digital Watermarks. From WWII Microdots ...

  • @BrianRoemmele: This is AI watermarking reporting insights you will see no place else.

I have a key and a solution....

  • @BrianRoemmele: DID YOU KNOW YOUR VIBE CODE FROM POPULAR AI MODLES HAS A WATERMARK?

THAT CAN SELF REPORT THE AI MOD...

  • @BrianRoemmele: RT @BrianRoemmele: ⚠️ WARNING⚠️

Two AI companies turned up the text based watermarking technology ...

  • @BrianRoemmele: RT @BrianRoemmele: HIDDEN AI TEXT OUTPUT WATERMARKS!

When this full breaks open you will see the wo...


5. edzitron (Group Score: 153.6 | Individual: 25.4)

Cluster: 7 tweets | Engagement: 916 (Avg: 512) | Type: Tech

Hey everyone we made a model potentially as good as someone else’s. It also takes up way more compute. No real clue how we’ll monetize it, and we haven’t been able to stand up an API service for our current one. Anyway, see ya\n\nQT @CharlesRollet1: SCOOP: Alexandr Wang says Meta's upcoming AI model - codenamed Watermelon - has caught up to OpenAI's GPT-5.5.

"Watermelon uses an order of magnitude more compute than Avocado," he said in an internal meeting, referencing Meta's previous AI model.

https://t.co/XKwQOPkFTy

See 6 related tweets

  • @Techmeme: Sources: Alexandr Wang said Meta's model currently in training, codenamed Watermelon, matches GPT-5....
  • @firstadopter: You all panicked over a fake manufactured sensationalized narrative AGAIN\n\nQT @Techmeme: Sources: ...
  • @Dan_Jeffries1: Open source it and we'll care even more!

We need Meta back in the American open source champion ca...

  • @StockSavvyShay: $META Chief AI Officer Alexandr Wang says Meta’s current model in training codenamed Watermelon is m...
  • @danielnewmanUV: $META has sat out the model race for sometime. But it looks like a comeback is in the making. Still ...

6. chandrarsrikant (Group Score: 152.5 | Individual: 47.5)

Cluster: 4 tweets | Engagement: 1586 (Avg: 291) | Type: Tech

RT @satyanadella: The future of the firm is a learning loop in which human capital and token capital compound.

With our new Frontier Co., our ambition is to help every enterprise build its own AI capability, and to help create a frontier ecosystem where every organization can turn its knowledge, workflows, and judgment into its own AI systems that continuously improve. https://t.co/mvYhkRFyqa

See 3 related tweets

  • @Gorden_Sun: 微软新建部门Microsoft Frontier Company(微软前沿公司),帮助客户做AI前沿转型。虽然微软CEO在文章里说的是超越FDE(Forward Deployed Engineerin...
  • @jgreze: Town is AI for all of the companies and people that can't afford to pay $ millions to Microsoft, Goo...
  • @ashugarg: Every IT Services consultant is now an FDE😜\n\nQT @satyanadella: The future of the firm is a learnin...

7. artman (Group Score: 144.3 | Individual: 33.0)

Cluster: 5 tweets | Engagement: 130 (Avg: 75) | Type: Tech

Humor is hard. In the light of GitHubs recent availability percentage, this is tone-deaf.\n\nQT @github: We heard you. And we agree.

In light of recent developments in physical media, GitHub is proud to announce that you can now obtain your public repo on CD-ROM.

Keep it. Lend it to friends. Pass it on to your children.

Your code is physically yours, forever. Until you lose it, let's be real.

Order yours today. https://t.co/z041pdMH7h

See 4 related tweets

  • @penberg: Can we have decent uptime now that you shipped this important thing?\n\nQT @github: We heard you. An...
  • @mark_k: GitHub offering physical media of your source code repos?

I can't tell if this is serious or a joke...

  • @WesRoth: GitHub has found a way to turn version control into nostalgia. 😎 https://t.co/4VawMrDKUe\n\nQT @gith...
  • @starbuxman: 😂\n\nQT @github: We heard you. And we agree.

In light of recent developments in physical media, Git...


8. guohao_li (Group Score: 141.3 | Individual: 43.1)

Cluster: 4 tweets | Engagement: 112 (Avg: 31) | Type: Tech

such a beautiful graph of agents learning through environment interactions

interesting benchmark: long-running tasks that agents can work on for 12+ hours each. these tasks are substantial even for human experts: recorded expert effort averages 57.2 hours per task, with some taking up to 320 hours

congrats @tikgiau and the team. it would be an absolute nightmare to do async rl on long-running tasks like these. good luck to all the rl infra folks!\n\nQT @tikgiau: Introducing EdgeBench, a benchmark designed to study how agents learn from environments over at least 12~72-hour runs. We find that performance follows a log-sigmoid function of environment interaction time with high precision.

EdgeBench is built with three ingredients:

  • 🌍 Real & Diverse: 134 real-world tasks across 6 task categories, spanning scientific problems, professional knowledge work, software engineering, optimization, formal math, and games.
  • ⏳ Ultra-Long-Horizon: Each task supports 12–72 hours of agent work. Recorded human effort averages 57.2 hours.
  • 🔁 Informative Feedback: Agents receive real-world feedback for continuous improvement.

After 38,000 hours of agent runs on EdgeBench, a scaling law for learning from environments emerges:

  • 📈 As agents interact with task environments over time, their aggregate performance is precisely fit by a log-sigmoid function.
  • 🧠 This phenomenon can be explained by an elegant theory of graph exploration.

We are releasing an initial 51 of the 134 tasks, together with the full evaluation framework, to help advance long-horizon agent research. Check our blog & paper for more findings!

Blog https://t.co/nMOzFsOhbT Paper https://t.co/rZb3eWuvik GitHub https://t.co/oemXd4UrFw Dataset https://t.co/P4SQMrM47o

Details below 👇🧵

See 3 related tweets

  • @rosstaylor90: RT @tikgiau: Introducing EdgeBench, a benchmark designed to study how agents learn from environments...
  • @dejavucoder: blueball bench\n\nQT @tikgiau: Introducing EdgeBench, a benchmark designed to study how agents learn...
  • @rohanpaul_ai: RT @rohanpaul_ai: ByteDance Seed delivered again.

They released EdgeBench, to test whether AI agent...


9. kimmonismus (Group Score: 126.8 | Individual: 33.6)

Cluster: 4 tweets | Engagement: 971 (Avg: 556) | Type: Tech

„I'm told the 5.6 plan limits will be significantly more generous.“

Looks like OpenAI’s efficiency gains pay off a lot\n\nQT @synthwavedd: 🚨 SCOOP:

As previously reported, OpenAI plan to launch GPT-5.6 once back in office next week, with a target window of July 7-9, but want it out as early as possible within that window (so July 7th is the most likely date). This is also perfect timing to catch customers coming from Claude who have just lost their access to Fable 5 in plans, and I'm told the 5.6 plan limits will be significantly more generous. More aggressive safeguards are already being rolled out in preparation for the launch too, although they probably won't be as aggressive as Fable's.

DeepMind have also tentatively set a new launch date for Gemini 3.5 Pro of July 17th. Apparently, this extra time has been spent on a new pretrain (they were planning to keep using the ancient 2.5 Pro base, lol), but it remains to be seen whether it's any good. I'm not hopeful. In other news, work is well underway on a new Nano Banana Pro model based on the new 3.5 Pro base, which I expect to be better received and compete well with GPT-Image 1.

See 3 related tweets

  • @synthwavedd: 🚨 SCOOP:

As previously reported, OpenAI plan to launch GPT-5.6 once back in office next week, with ...

  • @omarsar0: I believe GPT-5.6 could be a huge win for OpenAI.

It's a defining moment for frontier models.

But...

  • @AndrewCurran_: Sounds like Gemini 3.5 Pro arrives around July 17th with a brand-new pretrain. Unless, of course, th...

10. petergostev (Group Score: 107.7 | Individual: 48.3)

Cluster: 3 tweets | Engagement: 2289 (Avg: 245) | Type: Tech

We are now in a position where a tiny proportion of the population uses Fable or soon GPT-5.6, while everyone else's experience of AI is 8-30b-model level - Google's AI Overviews, Meta AI, ChatGPT free tier, maybe MS Copilot at best. People outside of tech must be completely baffled how this is supposed to take their job, and annoyed that hundreds of billions are being poured into it.

See 2 related tweets

  • @kimmonismus: Peter is absolutely right here. Outside our AI bubble, I hardly know anyone who really knows what Fa...
  • @emollick: This is true… but maybe less important than the fact that people don’t try ambitious things with the...

11. ANI (Group Score: 107.1 | Individual: 41.4)

Cluster: 4 tweets | Engagement: 811 (Avg: 106) | Type: Tech

Ministry of Electronics and Information Technology issued notice to Google Android and Apple iOS to remove 7 applications from their App Store for misuse of apps for shutting down batteries in e-rickshaws/vehicles. Apps like BAT-BMS, SMART BMS, LOSSIGY: Sources https://t.co/X5N7opwgli

See 3 related tweets

  • @ANI: Govt has asked to remove BAT-BMS, Epoch-i-ion and Lossigy apps from the app stores on iOS and androi...
  • @chandrarsrikant: 🚨Govt orders removal of BAT-BMS, Lossigy, and Epoch-i-ion apps being misused in remotely disabling b...
  • @moneycontrolcom: 🚨 Govt orders removal of BAT-BMS, Lossigy, and Epoch-i-ion apps being misused in remotely disabling ...

12. matvelloso (Group Score: 104.0 | Individual: 36.9)

Cluster: 4 tweets | Engagement: 48 (Avg: 25) | Type: Tech

Microsoft was doing forward deployment engineering way before forward deployment engineering was invented.

This is their strength. It's how you make AI real.

So, good sign.\n\nQT @StockSavvyShay: MSFTjustannouncedMicrosoftFrontierCompany,anewAIoperatingbusinessbackedbyMSFT just announced Microsoft Frontier Company, a new AI operating business backed by 2.5B and 6,000 employees.

Microsoft says the unit will help customers drive “Frontier Transformation” with industry expertise, change management and enterprise grade AI engineering. https://t.co/FCyLnOdnyy

See 3 related tweets

  • @VaibhavSisinty: Microsoft just committed $2.5 billion and 6,000 engineers to sit inside their customers and build AI...
  • @WesRoth: Microsoft launched Frontier Company, a new operating business backed by a $2.5 billion investment an...
  • @brunoborges: RT @StockSavvyShay: $MSFT just announced Microsoft Frontier Company, a new AI operating business bac...

13. nickvasiles (Group Score: 101.7 | Individual: 37.3)

Cluster: 3 tweets | Engagement: 131 (Avg: 44) | Type: Tech

gary vee was right

analog is back

the anti-trend to AI is nostalgia

in the next few years, i bet we will see a surge in:

wired earbuds ipods macbook pros from 2014 w/ the light up apple logos e-ink displays flip phones single-purpose tech (point and shoot cameras, etc)

you are either building for agents, or you're building for humans who are tired of keeping up with agents\n\nQT @github: We heard you. And we agree.

In light of recent developments in physical media, GitHub is proud to announce that you can now obtain your public repo on CD-ROM.

Keep it. Lend it to friends. Pass it on to your children.

Your code is physically yours, forever. Until you lose it, let's be real.

Order yours today. https://t.co/z041pdMH7h

See 2 related tweets

  • @lukehoban: RT @github: We heard you. And we agree.

In light of recent developments in physical media, GitHub i...

  • @BrianRoemmele: Physical media is vital.

This is brilliant.\n\nQT @github: We heard you. And we agree.

In light o...


14. sgl_project (Group Score: 100.8 | Individual: 38.6)

Cluster: 4 tweets | Engagement: 162 (Avg: 30) | Type: Tech

We've spent months encoding our team's hard-won engineering know-how (benchmarking, profiling, CUDA kernel tuning, production triage) into executable agent skills. Now agents handle the repetitive grind, and developers focus on the hard calls.

And it's working! 3 KDA-Pilot kernel PRs already merged upstream, up to 2.75x kernel speedups on B200, +71.4% serving throughput on Qwen3-Next. Huge effort from the team, and we're just getting started.

Dive in 👇\n\nQT @lmsysorg: 🚀 New blog: Agent-Assisted SGLang Development, the story of how we turn benchmarking, profiling, and kernel optimization know-how into executable agent skills.

Agent-assisted workflows are saving our team massive engineering hours while delivering major gains across the stack: ⚡️ +71.4% throughput & TTFT 456→168ms for Qwen3-Next via allreduce fusion ⚡️ 29–49% TTFT reduction on long-context prompts via router tokenization deduplication ⚡️ Up to 2.32x diffusion denoising speedup via Spectral Progressive Diffusion ⚡️ 10 B200 kernel tasks at 1.13x–2.75x speedups via KDA-Pilot; 3 PRs merged upstream ⚡️ 1.41x faster LTX-2 VAE decode, saving 9.7 GiB peak memory

And rigor is built into every step: benchmarks are fixed before any patching, baseline and candidate share the same ABI, and every change must be backed by profile evidence, eliminating benchmark reward hacking. Each iteration passes a Humanize/RLCR review loop before proceeding.

Read the full blog to see how we're rethinking development workflow 👇

See 3 related tweets

  • @ying11231: On the way to automating open source project maintenance.\n\nQT @lmsysorg: 🚀 New blog: Agent-Assiste...
  • @ying11231: RT @lmsysorg: 🚀 New blog: Agent-Assisted SGLang Development, the story of how we turn benchmarking, ...
  • @richardczl: RT @sgl_project: We've spent months encoding our team's hard-won engineering know-how (benchmarking,...

15. XFreeze (Group Score: 99.6 | Individual: 38.9)

Cluster: 3 tweets | Engagement: 1345 (Avg: 743) | Type: Tech

Another major Grok Build update just landed, packed with new features, extensive bug fixes, and meaningful performance improvements

Release Notes: v0.2.84 — 2026-07-03

Features: • Announcements now update live during active sessions without restart or /new. • Hiding an announcement no longer suppresses later criticals; new ones reappear automatically. • run_terminal_cmd now requires a one-sentence description rationale in every invocation. • ask_user_question timeout policy is now configurable in config.toml and /settings. • Ask-Question timeout can now be toggled from /settings (Agent & Approval). • Thinking/reasoning blocks are now shown by default while the model is working. • Critical announcements now show a red title with a clickable [hide] button and aligned message. • Added remote_fetch option under [features] in config.toml to disable all backend catalog and settings fetches for air-gapped environments.

Bug Fixes: • Images pasted or read from GIF, BMP or TIFF files are now automatically converted so they work with image generation. • Queue panel now shows action buttons on hover and the status bar displays a compact done/total task count. • Hook matchers now correctly see the real MCP tool name instead of the internal dispatcher name. • Copy now succeeds when running inside containers even when the terminal brand cannot be detected. • Tool result previews no longer paint opaque panels in grok --minimal. • grok wrap now correctly handles quoted strings and shell aliases. • Text selection settings now correctly honor explicit keep_text_selection values even when legacy keys remain. • Fixed a freeze that could occur when editing and sending the last message in the queue. • Fixed a startup crash on minimal Linux systems lacking system CA certificates.

Performance: • Grep now stops early on broad searches, returning faster results with far less memory use. • Idle CPU and memory usage after long sessions or resume is now dramatically lower.\n\nQT @XFreeze: Another Grok Build update is here, it brings new usability improvements to make everyday workflows even smoother.

Release Notes: v0.2.83 — 2026-07-02

Features: • Critical announcements now appear in a top banner during active sessions with a hide command. • Pasting the same text again next to a paste chip now expands the chip into editable text instead of duplicating it. • Paste preview now shows a hint explaining how to expand the chip.

See 2 related tweets

  • @mark_k: Grok Build 0.2.84 is out. @xai keeps tightening the terminal experience with a pretty practical rele...
  • @elonmusk: RT @XFreeze: Another major Grok Build update just landed, packed with new features, extensive bug fi...

16. deanwball (Group Score: 95.2 | Individual: 36.8)

Cluster: 3 tweets | Engagement: 525 (Avg: 116) | Type: Tech

Basically I think that, back in 2023 or so, the “consistently wrong about AI” VC and SaaS community was operating under the assumption that AI’s trajectory would mean model capabilities peaking around GPT 5.5/Opus 4.8 capabilities somewhere around 2030, plus robots.

And if that was your assumption, I can totally understand why you think everything commodifies/frontier AI isn’t a legitimate business model, etc.

That is a nice world to believe in! In the real world, however, that community has been wildly wrong for three years, and I would expect them to continue being wrong for more years to come.

They may not be wrong forever! Things eventually commodify. But people have been saying “the models are good enough” since GPT-4, and it’s been untrue. I suspect that will continue to be the case because I think that we remain in the earlier stages of the AI industry, and along the steep part of the trajectory.

More broadly: the notion of “good enough” should gross you out, a little bit. The economy of the future will be about heavy-tailed excellence, not middle-of-the-bell-curve, loser-premise, “good enough”-ness.

See 2 related tweets

  • @WesRoth: Altman is saying the next major economic transition may begin before governments, companies, and wor...
  • @deredleritt3r: In the age of RSI, the claim that models will commoditize looks increasingly dubious. The gap betwe...

17. WesRoth (Group Score: 93.9 | Individual: 24.8)

Cluster: 5 tweets | Engagement: 46 (Avg: 22) | Type: Tech

Once a classifier silently interrupts Fable and hands work to Opus, “Fable 5 performance” becomes impossible to measure cleanly. Developers are no longer benchmarking one model. They are benchmarking a routing system controlled by Anthropic.

Anthropic should publish the fallback rate, the categories most affected, and separate benchmarks for raw Fable 5 versus the guarded product users receive.\n\nQT @bridgemindai: FABLE 5 CAME BACK NERFED.

We re-ran the July 1st version of Claude Fable 5 on BridgeBench.

The results are brutal:

Debugging: 86.2 → 25.9 Refactoring: 73.6 → 38.4 Hallucination: 75.9 → 61.7

The new guardrails are kicking in on way too many tasks and falling back to Opus 4.8.

This is not the model that got banned.

Anthropic owes everyone an explanation.

See 4 related tweets

  • @CosineAI: RT @yangli_: Can we still rely on models that are limited without our input? The latest version of F...
  • @alex_prompter: RT @alex_prompter: Before you rage-quit Fable 5, here’s what’s actually happening and how to avoid i...
  • @WesRoth: Anthropic says only a small fraction of routine coding and debugging work should trigger Fable 5’s u...
  • @WesRoth: RT @WesRoth: Anthropic clarified that not all coding tasks will be routed from Fable 5 to Opus 4.8. ...

18. chamath (Group Score: 91.5 | Individual: 34.3)

Cluster: 5 tweets | Engagement: 9281 (Avg: 1687) | Type: Tech

Tesla is one of the smartest, cracked and most advanced engineering companies in the world.

If they actually did this, then it is likely verifiably true that a dollar above 200/weekiswaste.\n\nQT@Kalshi:JUSTIN:TeslareportedlycapsemployeeAIspendat200/week is waste.\n\nQT @Kalshi: JUST IN: Tesla reportedly caps employee AI spend at 200 per week

See 4 related tweets

  • @rohanpaul_ai: The new AI budget benchmark for software engineers may have just landed at $800/month.

This is Tesl...

  • @Cointelegraph: 🚨 TODAY: Elon Musk's Tesla caps employee AI spending at $200 per week starting July 6, per The Infor...
  • @omarsar0: $200/week is not bad.

It would cover (5-fold) all my engineering & research work. And I do a to...


19. polynoamial (Group Score: 85.5 | Individual: 29.2)

Cluster: 4 tweets | Engagement: 472 (Avg: 472) | Type: Tech

Excellent work from @AISecurityInst investigating the impact of test-time compute budgets for frontier AI model evaluations. They make the case even more convincingly than I could!\n\nQT @AISecurityInst: Most AI agent evaluations boil capability down to one score. But that number hides a key choice: how much compute the agent was allowed to use. New work from our Science of Evaluation team shows why that matters. 🧵 https://t.co/o56NiPnfEj

See 3 related tweets

  • @Miles_Brundage: We could have this at home, American friends (a government agency that is staffed to do, and allowed...
  • @matthewclifford: RT @polynoamial: Excellent work from @AISecurityInst investigating the impact of test-time compute b...
  • @theojaffee: RT @AISecurityInst: An AI agent's performance is best understood as a capability curve over compute,...

20. gokulr (Group Score: 80.4 | Individual: 48.7)

Cluster: 2 tweets | Engagement: 460 (Avg: 90) | Type: Tech

GROWTH FOUNDERS: HIRE A NEAR-PEER

A portfolio founder approaching $100M in ARR asked me about the single biggest and most impactful hire they can make.

I thought for a bit tand said: "Hire a near-peer".

In every generational company I've been part of, the founders hired a near-peer, who was essential to the company's success.

  • Google: Larry and Sergey hired @ericschmidt .
  • Facebook: Mark hired @sherylsandberg .
  • Square: Jack hired @rabois .
  • DoorDash: Tony hired @chrispa
  • Coinbase: Brian hired @emiliemc

Characteristics of a near-peer:

  1. They're so good that the founder will be fine reporting to them if the roles were reversed. (Mark has said publicly that he'd be fine reporting to Sheryl)
  2. Their strengths perfectly complement the founder's strengths; however, they share many cultural attributes with the founder and pass the founder's airport test, since the founder will be spending a ton of time with them (example: Eric being a Computer Scientist, which was culturally very important at Google back in the day)
  3. They're systems builders who have already operated at the scale you're growing into. Pattern recognition on 100Mto100M to 1B is not something you build in real-time.

A near-peer lets the founder focus on what only the founder can do (product, vision, culture). The near-peer handles everything else.

Also, near-peers stay for the long haul. Decade-plus tenure is standard.

The reason this hire matters more than any other: after $50M ARR, the bottleneck shifts from product-market fit to organizational scale. The founder is still the visionary. But the company needs someone who has already scaled a company of that size.

Most founders wait too long. They hire functional VPs first (Sales, Marketing, Engineering, Finance) and hope the collective covers the gap. It rarely does. A stack of VPs reporting into a founder who has never scaled a company creates coordination overhead, and the founder becomes the bottleneck.

The near-peer absorbs that overhead. They turn the founder's vision into daily operating decisions. They give the VPs a leader who has actually run a company this size.

Growth founders: if you're between 50Mand50M and 300M in ARR and every week feels like a coordination tax, the near-peer is the hire that determines whether you build a 10100Bcompanyortopoutat10-100B company or top out at 1B. Start the search now, ideally through warm intros from your venture/angel investors and advisors.

See 1 related tweets

  • @rabois: Excellent advice.\n\nQT @gokulr: GROWTH FOUNDERS: HIRE A NEAR-PEER

A portfolio founder approaching ...