Published on

科技热门推文 — 2026年8月22日

Authors

今日科技动态:随着 DeepSeek 推出实验性多模态模型、英伟达的编程智能体在交互式推理基准测试中取得优异成绩,以及 OpenAI 下调 GPT-5.6 Sol API 定价并扩展图像生成功能,人工智能能力与商业化进程进一步提速。企业应用正转向具备情境感知能力的智能体,覆盖安全、计费、广告乃至端到端业务运营等领域;不过,有证据表明,更深入的推理并不总能改善结果。与此同时,人工智能基础设施融资引发审视,初创企业则展示了可协助融资、制作视频营销活动及挖掘专有数据价值的智能体。


1. vinhnx (Group Score: 478.1 | Individual: 46.1)

Cluster: 23 tweets | Engagement: 2072 (Avg: 174) | Type: Tech

RT @deepseek_ai: DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀

🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.

Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.

1/n

See 22 related tweets

  • @ZenMuxAI: Now on ZenMux, DeepSeek-V4-Flash-Vision-Exp is FREE for one week ⚡

The experimental vision model th...

  • @DeItaone: DEEPSEEK CHALLENGES ANTHROPIC WITH NEW AI MODEL

DeepSeek has unveiled an experimental multimodal AI...

  • @jenzhuscott: Another demonstration of a strong, cost-efficient pure-text agent model can gain high-quality vision...
  • @heyshrutimishra: V4-Flash-Vision-Exp dropped today with vision layered on top of the existing V4-Flash.

Same text c...

  • @TeksEdge: 👀🔥 Yep, DeepSeek finally has eyes. Local AI will never look the same. 🙄 It can read your computer sc...

2. Daniel_Farinax (Group Score: 428.0 | Individual: 47.7)

Cluster: 18 tweets | Engagement: 766 (Avg: 63) | Type: Tech

RT @NVIDIAAI: Our general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive reasoning benchmark.

NVIDIA AVO completed all 183 levels across all 25 public environments, figuring out what to do with no instructions, explicit rules, or stated goals. https://t.co/UgROuDrMtn

See 17 related tweets

  • @kimmonismus: Holy: NVIDIA’s coding agent AVO scored 100% on ARC-AGI-3’s 25 public games, solving all 183 levels. ...
  • @mark_k: This is HUGE from @NVIDIA.

Their new AVO agent architecture just scored 100% on the ARC-AGI-3 publi...

  • @alexocheema: It's the harness, guys! It's the harness!\n\nQT @NVIDIAAI: Our general-purpose coding agent just sco...
  • @scaling01: Even Nvidia is participating https://t.co/KBqNdI5yMz\n\nQT @NVIDIAAI: Our general-purpose coding age...
  • @MatthewBerman: Whoa\n\nQT @NVIDIAAI: Our general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive...

3. gregisenberg (Group Score: 291.9 | Individual: 37.8)

Cluster: 9 tweets | Engagement: 730 (Avg: 1106) | Type: Tech

Grok Bot might be the first tool that lets one non-technical person run an entire business with a team of AI agents.

My friend Billy runs his whole newsletter business on Grok Bot agents, and I think we're about to see 100,000+ businesses like his.

BEST PRACTICES:

  1. The agents run on a shared cloud computer, so running your newsletter, your X, and your receipts all in one place creates context bloat and burns tokens fast. One mission per setup.

  2. Start with a Chief of Staff. Give it access to your existing docs (Notion, Slack, Gmail), have it audit the business, then tell you the top three agents to build first to drive revenue.

  3. Perfect a task with the Chief of Staff before spinning up a new agent. Have it do the outbound sales once, review it, and only then say "now build a bot that does exactly that." You earn each new hire by proving the task works first.

  4. Constraints are the feature. You get a limited number of agents, one thread per bot, like DMs with a teammate. It forces you to stay mission-oriented instead of spinning up a bot for every random idea.

  5. You make the decisions, not the agent. Billy's team spent three weeks unable to pick where content should live. At some point you say "we're doing Notion, no more tinkering" and move on.

  6. Run week one with no new agents. Build the team, learn to fly the plane, just execute. Week three is when you find the real gaps and expand, someone to man the inbox, someone for the Shopify shop.

  7. Then add routines so it works while you sleep. Ask your Chief of Staff what recurring jobs would move the business forward overnight, and it builds the automations that run without you.

Thanks to @billyjhowell for sharing the sauce on @startupideaspod (follow for more).

Grokbot is really cool.

Watch below:

https://t.co/9jTXeOiu9x

See 8 related tweets

  • @elonmusk: Cool\n\nQT @gregisenberg: Grok Bot might be the first tool that lets one non-technical person run an...
  • @startupideaspod: Building your first Grok Bot team?

Here's what my first month looked like:

Week 1:

  • My first hire...
  • @mikenevermiss: this guy just released a detailed 25-minute guide on Grok Bot, and this might be the best free cours...
  • @milesdeutscher: 25 ways anyone can use the new Grok Bot ( @bot ) for 10x more productivity.

(If you work on a compu...

  • @rewind02: only 10% of Grok Bot users are using it to its full potential...

how to actually run it:

  • feed yo...

4. dkundel (Group Score: 184.6 | Individual: 29.0)

Cluster: 9 tweets | Engagement: 818 (Avg: 154) | Type: Tech

RT @OpenAI: As we continue to push the frontier of capabilities while improving efficiency, we're dropping API and credit pricing of GPT-5.6 Sol by over 20% for the next 3 months. https://t.co/UoTb3hcB2t

See 8 related tweets

  • @rohanpaul_ai: Nice. OpenAI is cutting GPT-5.6 Sol API and Work/Codex credit pricing by over 20% for 3 months.

Pro...

  • @testingcatalog: OPENAI 🔥: GPT-5.6 Sol price has been reduced by 20% on the API for the next 3 months.

Eligible pla...

  • @ns123abc: THE PRICE WAR HAS OFFICIALLY BEGUN https://t.co/OnNeWCX1wG\n\nQT @OpenAI: As we continue to push the...
  • @WesRoth: OpenAI is cutting GPT-5.6 Sol API and credit pricing by more than 20% for the next three months.

Op...

  • @thsottiaux: Sol shines brighter today. Efficiency, reliability and performance are the name of the game.\n\nQT @...

5. alexaiworks (Group Score: 172.4 | Individual: 29.3)

Cluster: 8 tweets | Engagement: 0 (Avg: 43) | Type: Tech

The reasoning curve here is fascinating.

More reasoning makes agents much more willing to redesign the training method, but it doesn’t consistently make the final result better.

That’s probably the finding I’d pay the most attention to\n\nQT @EinsiaAI: 1/ Recursive self-improvement (RSI) depends on agents improving how AI systems are trained —not just tuning hyperparameters, but improving the training algorithm itself.

We tested this directly with AI4AI-Bench: 10 real research repositories spanning 10 distinct algorithm families.

Full breakdown 👇 GitHub: [https://t.co/s0f0NY7PdQ] Paper Link: [https://t.co/x0qY8wnlwB] Einsia Website:[https://t.co/Rewt4FJwl8]

📊 The results: The average score is just 0.166. Even the best-performing model, Opus 5, reaches only 0.288. The median exploration cost per task rises from 1.69to1.69 to 34.60.

#AI4AI #RecursiveSelfImprovement #AIResearch #AI4AI_Bench

See 7 related tweets

  • @alexaiworks: RT @EinsiaAI: 1/ Recursive self-improvement (RSI) depends on agents improving how AI systems are tra...
  • @svpino: Can an agent improve the algorithm used to train another AI system?

This is a pretty interesting be...

  • @EmmaUsesAi: Hyperparameter optimization has been automated for years.

Letting an agent rewrite objectives, supe...

  • @SynapseOpsAI: There’s a huge difference between asking an agent to find a better learning rate and asking it to de...
  • @NeuraFlowAix: $5,334 spent exploring 290 configurations is a nice reminder that autonomous AI research still has v...

6. matthewclifford (Group Score: 159.1 | Individual: 27.3)

Cluster: 7 tweets | Engagement: 38 (Avg: 60) | Type: Tech

A superb appointment! AISI has become a vital institution in AI; I can’t imagine anyone better than Henry to lead it. He combines deep government experience with true founder energy at a critical moment in AI. Bravo!\n\nQT @HZoete: The UK’s AI Security Institute is an organisation close to my heart. Back in 2023 I was involved in its conception and creation. It’s gone on to do incredible things. A genuinely world leading org that shatters the myth that government can’t do things.

So I couldn’t be more delighted to say that I’ve been appointed as AISI’s next Director, working alongside the brilliant @NateBurnikell as its new Chief Strategy Officer.

AISI is the world’s most respected organisation for understanding the risks from powerful AI. That is thanks to their superb team of dedicated civil servants and technical researchers.

Special thanks to Adam Beaumont who has done a great job as AISI’s interim Director for the last 9 months. He returns to GCHQ having seen AISI go from strength to strength under his watch. I’m hugely grateful for that and excited to work with and learn from him.

Last year I wrote about AISI and why its form and function is so important. You can read that piece on the link below - it gives a good flavour of why I’m excited to be joining.

I’m looking forward to getting started in September.

https://t.co/dAhChKjwUw

See 6 related tweets

  • @Jameswise: Henry has been a huge source of support and advice as we set up @UKSovereignAI - great to see him ta...
  • @AISecurityInst: Today we're announcing two senior appointments. @HZoete joins as our new Director, and @NateBurnikel...
  • @KanishkaNarayan: Great to see @AISecurityInst appoint Henry and Nate to these important roles.

Both bring significa...

  • @_NathanCalvin: Experiencing some acute transatlantic AI state capacity envy (congrats!)\n\nQT @HZoete: The UK’s AI ...
  • @Miles_Brundage: Huge! https://t.co/rvM6HbUVy4\n\nQT @HZoete: The UK’s AI Security Institute is an organisation close...

7. edzitron (Group Score: 152.6 | Individual: 32.3)

Cluster: 7 tweets | Engagement: 662 (Avg: 609) | Type: Tech

Yet another collapsing data center/AI compute company bailed out by NVIDIA. Jensen "TARP" Huang\n\nQT @anissagardizy8: SCOOP: Nvidia is in advanced discussions to make an investment in Cloverleaf Infrastructure, a company that arranges power for data-center projects.

Cloverleaf was previously exploring M&A.

(The companies later confirmed they inked a strategic partnership) https://t.co/Ar1DDsbq0v @WSJ

See 6 related tweets

  • @anissagardizy8: SCOOP: Nvidia is in advanced discussions to make an investment in Cloverleaf Infrastructure, a compa...
  • @DeItaone: $NVDA - NVIDIA EYES MAJOR DATA-CENTER POWER INVESTMENT

Nvidia is reportedly in advanced talks to i...

  • @StockMKTNewz: Nvidia $NVDA is reportedly in talks to invest in Cloverleaf Infrastructure, "a company that arranges...
  • @IPONewsroom_: $NVDA IN TALKS TO INVEST 100s OF MILLIONS IN CLOVERLEAF INFRASTRUCTURE

Nvidia is in advanced discus...

  • @StockSavvyShay: $NVDA is going straight after AI’s biggest bottleneck with talks to invest hundreds of millions in p...

8. rohanpaul_ai (Group Score: 148.6 | Individual: 31.2)

Cluster: 5 tweets | Engagement: 29 (Avg: 38) | Type: Tech

Anthropic just put Claude Mythos 5 behind Claude Security, a code-repository vulnerability scanner.

So now an ordinary Claude Enterprise customer can benefit from Mythos 5 through Claude Security.

So you can just do

Your GitHub repository > Claude Security > Mythos 5 scans it > you receive vulnerability findings + severity + confidence + suggested fixes

Althought, you cannot freely prompt Mythos from that surface\n\nQT @claudeai: Claude Security scans now run on Claude Mythos 5, available today in public beta for all Claude Enterprise customers.

Put our most capable security model to work on your codebase, no separate model access needed. https://t.co/zJSUgGjtZF

See 4 related tweets

  • @WesRoth: Anthropic is expanding access to Claude Mythos 5’s cybersecurity capabilities while keeping safeguar...
  • @AndrewCurran_: This is the first time any access to Mythos 5 has been granted beyond the small group that had acces...
  • @adonis_singh: this is actually pretty huge no? we're getting closer to mythos to the public\n\nQT @claudeai: Claud...
  • @martin_casado: Third party, API First party, Claude Code No party, .... this ... 👇\n\nQT @claudeai: Claude Security...

9. 0x_kaize (Group Score: 142.6 | Individual: 33.2)

Cluster: 5 tweets | Engagement: 26 (Avg: 29) | Type: Tech

Polsia using Polsia to help raise Polsia’s $30M round is insane founder lore (and hard to say out loud)

Think about it for a second.

Ben and his AI managed to market Polsia and get 200 investors in the round in a couple months

And the funny thing is, the money raised wasn't used to hire more people. It's literally re-invested to build their own infrastructure and use AI even more

Crazy times we live in\n\nQT @Bencera: My AI raised $30M by itself.

Well, it did most of the work. I showed up for the final meetings and the signatures.

It emailed the investors, ran the follow-ups, and the deck was public the whole time.

aisloP ep 5 "The Raise".

Learn the modern playbook on how to raise money.

Every founder should watch this.

See 4 related tweets

  • @WhaleInsider: JUST IN: Polsia used its own AI agents to help raise its $30M round.

The company is now pointing to...

  • @AnatoliKopadze: If you still think a company needs a team to raise money, you have to see this.

Ben ran a $30M roun...

  • @VadimStrizheus: if you’re a founder struggling to grow your app, watch this:

this guy raised 30Mata30M at a 250M valua...

  • @0x_kaize: RT @0x_kaize: Polsia using Polsia to help raise Polsia’s $30M round is insane founder lore (and hard...

10. alex_prompter (Group Score: 134.8 | Individual: 47.7)

Cluster: 4 tweets | Engagement: 241 (Avg: 45) | Type: Tech

the best enterprise AI account I just found on this app:\n\nQT @mardehaym: A PE operating partner asked us to build production AI agents inside a portfolio company's billing system, processing real healthcare claims under HIPAA.

Two people hand-wrote every rule in their claims engine across 300+ denial codes and payer logic that changes quarterly.

Four months later, seven production agents handle it with zero patient data exposure.

First month, we didn't touch a model. We mapped their data: where it sits and what's missing, so agents reason from structured facts instead of guessing.

I've watched teams skip this step across dozens of engagements. They bolt a model onto the product, watch it hallucinate over unstructured inputs, and decide AI isn't ready for their industry. The data work is what makes it ready.

We built an enrichment layer that assembles 34 dynamic variables per claim before any LLM sees it, pre-computed and versioned so the agent receives ranked facts instead of searching for context.

Every agent follows one pattern: pre-compute context, strip all patient data before the model sees it, validate output against a strict schema, let deterministic code accept or reject the action. If the output falls outside the allowlist, the system fails closed.

Seven agents, each locked to a single workflow like denied claim follow-up or billing reconciliation, each running its own enrichment payload.

Then we built the eval harness.

Every agent runs against a curated test suite before any update reaches production. When a model provider ships a new version or payer logic changes, the harness catches regression before a single live claim is affected. The flagship agent reconciles denials to the penny: 59 out of 60 on the eval set.

Most teams launch an agent and hope it keeps working. We launch one and prove it does on every deployment.

We route calls across two model providers. Swapping one changes nothing in the output because the eval harness verifies it.

Model integration was the shortest line item in the four-month build.

The operating partner now benchmarks the rest of the portfolio against this system.

That's the line between a portfolio company running AI and one still running demos.

See 3 related tweets

  • @mardehaym: RT @mardehaym: A PE operating partner asked us to build production AI agents inside a portfolio comp...
  • @mardehaym: RT @mardehaym: The amount of work between picking an AI model and getting repeatable value from it i...
  • @TrungTPhan: RT @B_Madden4: "Multi-agent systems" has officially entered the healthcare buzzword hall of fame, ri...

11. rohanpaul_ai (Group Score: 129.8 | Individual: 36.0)

Cluster: 4 tweets | Engagement: 21 (Avg: 38) | Type: Tech

The interesting part about Arcads Mark being available in Slack is the context layer.

Instead of treating ad creation as a sequence of disconnected tasks, Mark can use the information already present in the conversation as input.

Product positioning, campaign feedback, competitor references, creative directions.

That context can then flow directly into the generation layer.

We're moving from AI models that perform individual tasks to agents that can maintain context across the workflow.

That's the more important shift.\n\nQT @arcads_ai: We just turned Slack into a marketing department

Meet Mark. He lives in your Slack.

> Invite him to your workspace > say "make ads for my new product" > he reads the context, researches, generates

try it today on https://t.co/u6O0YGkBCf https://t.co/iUZ8sH38TJ

See 3 related tweets

  • @arcads_ai: We just turned Slack into a marketing department

Meet Mark. He lives in your Slack.

> Invite hi...

  • @Parul_Gautam7: Arcads just brought Mark to Slack.

> Your team was already in Slack > Your ideas were already...

  • @alexcooldev: Mark is now available on Slack, and I just tried it 👀

You can basically throw your brand into Slack...


12. alex_prompter (Group Score: 123.9 | Individual: 34.7)

Cluster: 4 tweets | Engagement: 38 (Avg: 45) | Type: Tech

Topview is giving away 10 free MiniMax H3 video generations to prove one thing: Codex can now make the video, not just the script.

The old loop: Codex plans the campaign, writes the script, outlines the shots. Then you rebuild all of it in a separate video tool.

The new loop: prompt or asset in. Concept, storyboard, finished video out. Editable in Topview after. Ready-made skills included: TikTok Product Ad, Micro-Drama Ad Creator, Pain Point Product Ad.

The sleeper feature: it pulls TikTok, YouTube, Amazon, and Shopee data before writing a word, so your script follows what's selling instead of what sounds clever.

H3 output: 2K, up to 15 seconds, native stereo audio.

10 free generations is enough to know in one afternoon whether this fits your workflow.

@TopviewAIhq\n\nQT @TopviewAIhq: Topview Plugin for Codex is here: the creative execution layer for your AI workflow.

To celebrate the launch, Topview subscribers can get 10 free MiniMax H3 video generations when using the plugin.

Generate video concepts, storyboards, images, and videos directly from Codex .

No switching between AI tools. No manual setup across different creative platforms.

Just install the plugin, describe what you want, and let Codex call Topview to create the asset.

From product videos and social clips to avatar explainers and voiceovers, your AI assistant can now produce real creative outputs.

#TopviewAI #Codex #AIVideo #AIMarketing #AIContentCreation

See 3 related tweets

  • @alex_verem: 10 free MiniMax H3 videos for installing a Codex plugin. That's Topview's bribe. The interesting par...
  • @Hailuo_AI: RT @TopviewAIhq: Topview Plugin for Codex is here: the creative execution layer for your AI workflow...
  • @alex_verem: RT @alex_verem: 10 free MiniMax H3 videos for installing a Codex plugin. That's Topview's bribe. The...

13. Dagnum_PI (Group Score: 120.2 | Individual: 36.0)

Cluster: 4 tweets | Engagement: 56 (Avg: 35) | Type: Tech

$AIAI 👀 +65% 🚀

Most people are still treating this like a random stock move.

It isn’t.

@aiaiholdings acquires operating businesses that already generate hard, real-world data, then deploys its Transformational AI platform into them.

@DorTechnologies is live with Retail Intelligence across 2,000+ stores.

The construction side just showed soil composition from five years of existing bid history predicting win rates and is building "Multiple" Projects with Tesla.

Same pattern every time: the data was already inside the building. The AI just made it readable.

That’s the model.

Own the operational data. Keep the alpha inside the enterprise.

The market is starting to notice.\n\nQT @Dagnum_PI: Palantir's Alex Karp just told a room of enterprises that their AI provider might be a company that will someday "compete and try to kill you."

The banner behind him read Institutional Sovereignty in the Age of AI.

Here is what he is telegraphing. Every enterprise now competes against a frontier lab by default. The only reason you can win is that you hold something outside the model. Proprietary data, tribal knowledge, tradecraft that only you understand. And the vendor selling you the tools to use it has an interest in absorbing exactly that.

His words: you are powering someone else's business in return for paying for their tokens. A trade he says most people are no longer interested in.

Now look at $AIAI @aiaiholdings

It does not sell AI into companies. It buys the company, keeps the management team, and deploys Transformational AI from the inside. MediGuide's clinical data. C.C. Carlton's bid history. Dôr's retail sensors. The alpha never leaves the building, because the building is the portfolio.

Karp described the problem to a room of people who still have to solve it.

Who else is already structured this way?

See 3 related tweets

  • @Dagnum_PI: This is the entire thesis in one example.

Five years of bid history just sitting there. Nobody went...

  • @KyleReidhead: RT @m0xt_: The most underowned layer in AI is ripping right now.

It isn't the chips, and it isn't t...

  • @Dagnum_PI: RT @gulVasikova: $AIAI — Key Details & Thesis

AIAI Holdings is building an AI-powered acquisition c...


14. testingcatalog (Group Score: 119.5 | Individual: 34.8)

Cluster: 4 tweets | Engagement: 248 (Avg: 295) | Type: Tech

ICYMI 👀: OpenAI made it possible to generate images with transparent backgrounds for GPT-Image-2 on APIs.

Needed 🤖 https://t.co/h94mGTF3xW\n\nQT @OpenAIDevs: Transparent backgrounds are now available in preview for GPT-Image-2 in the API.

Generate reusable assets you can place on any background—for product imagery, graphic design, website mockups, and marketing campaigns. https://t.co/yRhBIYh8uP

See 3 related tweets

  • @robinebers: you know i dunk a lot on openai, but this is genuinely appreciated

TRANSPARENT BACKGROUNDS

this en...

  • @Gorden_Sun: GPT-Image-2 API可以生成透明背景的图片了,这个是真正的生产力功能\n\nQT @OpenAIDevs: Transparent backgrounds are now available...
  • @Dimillian: RT @romainhuet: Really useful addition to GPT-Image-2 in the API: you can now generate transparent i...

15. ai_explorer25 (Group Score: 110.1 | Individual: 54.7)

Cluster: 3 tweets | Engagement: 1136 (Avg: 110) | Type: Tech

RT @Apodex_AI: Most AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer nobody has yet?

Meet TRACES 🧭 — the world's first benchmark for measuring discoverative AI: AI that can work through evidence, test hypotheses, and reach verifiable conclusions on problems without answer keys. Proposed by our founder @tianqiao_chen, who defined its six capabilities.

Three things published today: a definition of "discoverative intelligence", a rubric to tell sound investigation from lucky guesses, and a open call for both solvers and problems *Website: https://t.co/iJ5rDV6qz9

See 2 related tweets

  • @FellMentKE: We've spent years congratulating models for getting answers we already had.

That is not discovery; ...

  • @svpino: I don't trust benchmarks. We've all seen this movie:

New model beats everyone else on a benchmark. ...


16. WesRoth (Group Score: 106.4 | Individual: 29.3)

Cluster: 5 tweets | Engagement: 68 (Avg: 34) | Type: Tech

OpenAI reportedly told employees Astra is coming in “a couple weeks.”

And apparently Anthropic is waiting on Astra before dropping Fable 5.1.\n\nQT @synthwavedd: 🚨 SCOOP: OpenAI yesterday told employees it intends to release Astra "in a couple [of] weeks" alongside demos of the model performing real-world tasks

An updated checkpoint, mostly an improvement to model alignment & reward hacking behaviours, is now being dogfooded (used internally by OpenAI employees via tools like Codex). Anthropic, meanwhile, are sitting on Fable 5.1 until Astra is launched. I think that's the wrong decision, but we'll see

See 4 related tweets

  • @WesRoth: This is an important reminder that the AI leaderboard can change incredibly fast.

Anthropic looked ...

  • @IamEmily2050: I don't see how OpenAI and Anthropic will justify the high prices in the coming weeks. All the open ...
  • @derrickcchoi: Great to see but so much more to do\n\nQT @arakharazian: May startle the markets with this but: Open...
  • @rohanpaul_ai: RT @rohanpaul_ai: The OpenAI-Anthropic enterprise race has changed direction.

Q3-to-date API spend ...


17. DavidSacks (Group Score: 100.8 | Individual: 37.6)

Cluster: 3 tweets | Engagement: 4839 (Avg: 3553) | Type: Tech

Harvey is a great example of how American companies are building world-class specialized models: they took an open-source base (Kimi K3), post-trained it on legal data, and delivered state-of-the-art performance on legal benchmarks at a fraction of the cost of frontier models. Restrictions that kneecap open models would do nothing to stop Chinese labs from shipping the next Kimi. They would, however, cripple the ability of startups like Harvey to create high-performance, low-cost vertical models. Of course some of the closed labs would love this — it eliminates their competition.\n\nQT @harvey: Introducing Tenet, our first model post-trained for legal.

Tenet is a Kimi K3 base that we post-trained with @FireworksAI_HQ on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work.

Training increases Tenet's all-pass rate by 82% on LAB and 22% on LAB Contracts relative to the Kimi K3 base model. It achieves state-of-the-art performance on LAB Contracts and places second on LAB.

These gains generalize to other leading agentic benchmarks including @mercor's Apex Agents - Corporate Law, @crosbylegal's Redline Bench, and @scale_AI's Professional Reasoning Bench.

Tenet is also optimized for token efficiency, operating at less than a fourth the cost of leading foundation models.

We additionally post-trained three specialist models for Tenet to use as subagents:

  1. M&A Diligence: post-trained with @baseten on our LAB Diligence environment in an RLM harness, this model is optimized for high-scale, long-horizon tasks.

  2. Review Tables: trained with @appliedcompute on our Review Table environment, this model is state-of-the-art and cost-effective at high-volume document review and structured data extraction.

  3. Firm Knowledge: trained with @EngramLab on our synthetic law firm environment, this model is optimized to learn and search over a firm's knowledge via memory and structured notes.

More details on model training, environment design, benchmarking, results, and more in the article by @gabepereyra below.

What's next for Harvey’s research?

  • Scaling LAB to more jurisdictions, practice areas and workflows
  • Scaling compute to bring new generalist models and capabilities to Harvey

More to come soon.

See 2 related tweets

  • @VaibhavSisinty: Harvey built an $11 billion legal AI business on top of other companies' models. Now they've built t...
  • @Kimi_Moonshot: RT @harvey: Introducing Tenet, our first model post-trained for legal.

Tenet is a Kimi K3 base that...


18. Gorden_Sun (Group Score: 99.0 | Individual: 36.3)

Cluster: 3 tweets | Engagement: 151 (Avg: 44) | Type: Tech

ChatGPT里也可以生成透明背景的图了,效果非常好! https://t.co/0bVt4EpSMt\n\nQT @thsottiaux: Yay, you can now make transparent images in ChatGPT and through the API with GPT-Image-2. Here is a cactus cactus I am about to print and put on my laptop. https://t.co/1WZd2VSH8y

See 2 related tweets

  • @thsottiaux: Yay, you can now make transparent images in ChatGPT and through the API with GPT-Image-2. Here is a ...
  • @jesselaunz: chatGPT可以生成透明背景了

测试一下,结果GPT缺省认为达里奥是桥水的达里奥,而不是竞争对手达里奥 https://t.co/QW9hyjUmgJ\n\nQT @thsottiaux: Yay...


19. matteocollina (Group Score: 96.9 | Individual: 49.4)

Cluster: 2 tweets | Engagement: 1854 (Avg: 235) | Type: Tech

RT @dhh: THE ANNOUNCEMENT: We’re going to make the prophecy of The Year of Linux on the Desktop come true. All the pieces are now in place. Time to go all in! https://t.co/zuf5h9ClWX

See 1 related tweets

  • @eastdakota: Happy to support this. More choices and better ecosystem support is good for everyone. Thanks for yo...

20. peakxvpartners (Group Score: 96.6 | Individual: 37.5)

Cluster: 4 tweets | Engagement: 3060 (Avg: 401) | Type: Tech

RT @narendramodi: I also urged the StartUps to think about how space technology can contribute immensely to agriculture, livestock management, dairy, environment and many other areas that touch everyday lives. They can engage with Atal Tinkering Labs, spark curiosity among schoolchildren and inspire the next generation to participate in the space sector.

Our private space ecosystem can become a major force for India and for the world!

See 3 related tweets

  • @PawanKChandana: Great to be part of Hon’ble PM Shri @narendramodi ji’s interaction with India’s space startups. 🇮🇳🚀 ...
  • @PIB_India: ▪️Prime Minister @narendramodi interacts with CEOs and Founders of Space Startups working in the AI ...
  • @Rajananandan: Bid day for Indian space tech startups https://t.co/OSCstRoBZg\n\nQT @narendramodi: I also urged the...