๐จ AI News | TestingCatalog
ANTHROPIC ๐ฅ: Dario Amodei says the industry should slow capability gains so safety can keep up. New essay: โWe Must Pace the Frontierโ not a pause, and not a halt to training. RSI has been running industry-wide since summer, including at Anthropic, with modelsโฆ
SPACEXAI ๐ฅ: Elon Musk agrees with the statement published by Dario Amodei, proposing to pace frontier AI development.
Looks like all this will have real consequences very soon. Nothing unexpected tho.
Looks like all this will have real consequences very soon. Nothing unexpected tho.
โค5๐2๐2๐1
OPENAI ๐ฅ: Sam Altman agrees with Dario Amodei on his proposal to pace frontier AI development.
Google next? Will we see any statement from Chinese frontier labs as well?
AI weekend unfolds ๐ค
Google next? Will we see any statement from Chinese frontier labs as well?
AI weekend unfolds ๐ค
โค8๐4๐3 2
This is a "Defender's gap" chart that OpenAI published recently. It shows the gap between defenders' capabilities and attackers' capabilities from a cybersecurity POV. This also translates to a gap between proprietary and open AI models.
What Dario is proposing is closely related:
In other words, Anthropic and OpenAI want to widen the "gap" between what their AI can do and what the rest of the world can do.
This doesn't necessarily mean that they will stop AI development and training.
What this leads to is:
- Anthropic, OpenAI, and other frontier labs will need to work together to make sure that every lab maintains alignment standards.
- These labs will continue using RSI to advance their internal models with "employee-like access".
- These models WON'T be released to the public until the "gap" is sufficient and until alignment standards are met.
What about China?
The assumption behind these measures is simple: without being able to distill frontier models, it will take China significantly longer to close the gap with top-tier models.
All the above may fay fail. China may or may not take the lead in AI progress.
Yet, it's not a surprise that AI can already be used as a cybersecurity weapon. Note that the top point on the "frontier" line describes defenders' capabilities available to companies with Daybreak access and similar. However, the top point on the "open-weight" line is accessible to everyone.
Even with current levels of intelligence, we will start seeing more and more major security incidents around the globe.
We should be monitoring this very closely ๐
What Dario is proposing is closely related:
"Thus, a key part of pacing within democracies is to keep democraciesโ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively."
In other words, Anthropic and OpenAI want to widen the "gap" between what their AI can do and what the rest of the world can do.
This doesn't necessarily mean that they will stop AI development and training.
What this leads to is:
- Anthropic, OpenAI, and other frontier labs will need to work together to make sure that every lab maintains alignment standards.
- These labs will continue using RSI to advance their internal models with "employee-like access".
- These models WON'T be released to the public until the "gap" is sufficient and until alignment standards are met.
What about China?
> "Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently."
> "If we execute these measures well, I believe they would slow Chinaโs progress enough to widen Americaโs lead significantly over the next 3โ5 years โ the window when AI becomes geopolitically most important."
The assumption behind these measures is simple: without being able to distill frontier models, it will take China significantly longer to close the gap with top-tier models.
All the above may fay fail. China may or may not take the lead in AI progress.
Yet, it's not a surprise that AI can already be used as a cybersecurity weapon. Note that the top point on the "frontier" line describes defenders' capabilities available to companies with Daybreak access and similar. However, the top point on the "open-weight" line is accessible to everyone.
Even with current levels of intelligence, we will start seeing more and more major security incidents around the globe.
> โEverything that makes it successful is exactly what makes it dangerous.โ
We should be monitoring this very closely ๐
๐ณ4๐ฏ2๐2โค1
ICYMI: Cursor announced Projects for agent coordination
Cursor Projects, now in beta, coordinates long-running software work through delegated agents. It runs tasks in cloud or local environments, supports scheduled and PR-driven maintenance, and targets work spanning multiple PRs.
๐ #cursor @testingcatalog
Cursor Projects, now in beta, coordinates long-running software work through delegated agents. It runs tasks in cloud or local environments, supports scheduled and PR-driven maintenance, and targets work spanning multiple PRs.
๐ #cursor @testingcatalog
TestingCatalog AI News
ICYMI: Cursor announced Projects for agent coordination
Projects is rolling out in beta to all Cursor users, coordinating features, migrations and recurring maintenance across cloud and local agents.
๐4๐ฅ3โค1
GOOGLE ๐ฅ: Demis Hassabis shared that he is aligned with the direction outlined by Dario Amodei for pacing the frontier.
RSI moment? ๐
RSI moment? ๐
๐12 9โค2๐คฃ2
MICROSOFT ๐ฅ: Satya Nadella agrees with pacing frontier AI development.
> โSuperintelligence that doesnโt benefit humanity and is not under human control doesn't worth pursuingโ
> โWe welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal.โ
> โSuperintelligence that doesnโt benefit humanity and is not under human control doesn't worth pursuingโ
> โWe welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal.โ
๐12โค3๐3๐1
SPACEXAI ๐ฅ: Grok 4.8 will be a 2.5T-parameter model built on a new C++ software stack, and Elon Musk expects it to finish training this week.
> While Grok 4.8 is in training, Grok 4.7 is still expected to arrive shortly, factoring in a previously communicated delay.
> Grok 4.8 will be ยฑ67% larger than Grok 4.6, which is currently available. This size puts it into the same tier as Kimi K3 with 2.8T params.
Not very soon ๐
> While Grok 4.8 is in training, Grok 4.7 is still expected to arrive shortly, factoring in a previously communicated delay.
> Grok 4.8 will be ยฑ67% larger than Grok 4.6, which is currently available. This size puts it into the same tier as Kimi K3 with 2.8T params.
Not very soon ๐
๐5๐3 2โค1
Anthropic prepares Claude Money for personal finance
Anthropic is testing a Claude โMoneyโ tab in its mobile app that would let users link bank accounts and ask spending and budgeting questions. The unreleased feature points to a possible US-first launch, though timing remains unclear.
๐ #anthropic @testingcatalog
Anthropic is testing a Claude โMoneyโ tab in its mobile app that would let users link bank accounts and ask spending and budgeting questions. The unreleased feature points to a possible US-first launch, though timing remains unclear.
๐ #anthropic @testingcatalog
TestingCatalog AI News
Anthropic prepares Claude Money for personal finance
Claudeโs mobile app is testing a Money tab with bank account linking, while its data provider, supported actions and launch plans remain unknown.
๐ฅ5 3
MICROSOFT ๐ฅ: A draft of the "Code of Conduct" has been published, introducing the term "Humanist Superintelligence" (HSI).
Key principles ๐
- People matter more than AI.
- AI must remain under meaningful human control.
- Safety takes priority over task completion.
- AI should remain a tool, not imitate a person.
- AI must not pursue independent goals.
- AI must stay within its authorized scope.
- Users must retain control over consequential decisions.
- AI should strengthen human reasoning and autonomy.
- Models must be accurate, transparent, and honest.
- AI must acknowledge uncertainty and correct mistakes.
- Models must not manipulate or exploit users.
- AI must respect personal and emotional boundaries.
- AI should support, not replace, human relationships.
- Models should discourage emotional dependence on AI.
- AI should respect cultural and personal differences.
- Human dignity and fundamental rights must be protected.
- AI should serve the public interest.
- Models should remain politically neutral in elections.
- AI actions should be traceable and understandable.
- Tool use should be authorized, limited, and reversible.
Key principles ๐
- People matter more than AI.
- AI must remain under meaningful human control.
- Safety takes priority over task completion.
- AI should remain a tool, not imitate a person.
- AI must not pursue independent goals.
- AI must stay within its authorized scope.
- Users must retain control over consequential decisions.
- AI should strengthen human reasoning and autonomy.
- Models must be accurate, transparent, and honest.
- AI must acknowledge uncertainty and correct mistakes.
- Models must not manipulate or exploit users.
- AI must respect personal and emotional boundaries.
- AI should support, not replace, human relationships.
- Models should discourage emotional dependence on AI.
- AI should respect cultural and personal differences.
- Human dignity and fundamental rights must be protected.
- AI should serve the public interest.
- Models should remain politically neutral in elections.
- AI actions should be traceable and understandable.
- Tool use should be authorized, limited, and reversible.
โAt Microsoft AI, weโre working towards Humanist Superintelligence (HSI): incredibly advanced AI capabilities that always work for people and in the service of humanity more generally.โ
โค8๐2๐2
Early look at Interactive Reports on Gemini Notebook
Google appears close to launching Interactive Reports in Gemini Notebook: longer reports with embedded mind maps, slides, quizzes, and flashcards generated on demand, combining research, teaching, and explainer formats in one document.
๐ #google @testingcatalog
Google appears close to launching Interactive Reports in Gemini Notebook: longer reports with embedded mind maps, slides, quizzes, and flashcards generated on demand, combining research, teaching, and explainer formats in one document.
๐ #google @testingcatalog
TestingCatalog AI News
Early look at Interactive Reports on Gemini Notebook
Gemini Notebook is testing interactive reports that embed on-demand mind maps, slide decks, quizzes, and flashcards, with no rollout date yet.
โค4 3
๐จ AI News | TestingCatalog
Early look at Interactive Reports on Gemini Notebook Google appears close to launching Interactive Reports in Gemini Notebook: longer reports with embedded mind maps, slides, quizzes, and flashcards generated on demand, combining research, teaching, and explainerโฆ
This media is not supported in your browser
VIEW IN TELEGRAM
GOOGLE ๐ฅ: Interactive Reports for Gemini Notebook will allow users to embed Mind Maps, Slide Decks, Flash Cards and Quizzes right inside the report.
At first, the system will generate a report with placeholders so users can explicitly create the artifacts they want embedded.
At first, the system will generate a report with placeholders so users can explicitly create the artifacts they want embedded.
๐4โค2
This media is not supported in your browser
VIEW IN TELEGRAM
PERPLEXITY ๐ฅ: Windows users with NVIDIA RTX GPUs can now use Portable Computer powered by local models!
> Support for local MCPs and scheduled tasks has also been added.
> Earlier, Portable Computer was introduced for NVIDIA DGX Spark devices.
> Now, Portable Computer is also available on Windows PCs with a supported NVIDIA RTX GPU with 24GB of VRAM or higher.
Models available locally ๐
- PPLX 27B (Perplexity post-trained Qwen 3.8 27B)
- Qwen 3.8 27B (stock) is not available on Windows
- Nemotron 3.5 Lightning (~30B MoE, ~3B active) is marked as "Coming Soon"
> Support for local MCPs and scheduled tasks has also been added.
> Earlier, Portable Computer was introduced for NVIDIA DGX Spark devices.
> Now, Portable Computer is also available on Windows PCs with a supported NVIDIA RTX GPU with 24GB of VRAM or higher.
Models available locally ๐
- PPLX 27B (Perplexity post-trained Qwen 3.8 27B)
- Qwen 3.8 27B (stock) is not available on Windows
- Nemotron 3.5 Lightning (~30B MoE, ~3B active) is marked as "Coming Soon"
โค4๐1