🚨 AI News | TestingCatalog
7.1K subscribers
3.94K photos
620 videos
40 files
4.24K links
Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors πŸ—ž
Download Telegram
Looks like Google started preparing Projects on Gemini for the upcoming rollout. At this moment, it is likely an unintended appearance.

The time has come πŸ‘€

h/t https://t.me/c/1349477688/33765
❀6πŸ‘5πŸ”₯31
AI Studio will get Skills support too!

Skills will allow users to instruct AI Studio Build tasks with greater precision and reuse them across conversations.
πŸ”₯5πŸ‘4
Some users are noticing a new layout being rolled out on DeepSeek along with a potential stealth V4 model update.

Have you seen it too? πŸ‘€
πŸ‘8❀42
BREAKING 🚨: Z AI released GLM-5.1, an open-source model with top tier coding performance!

β€œNumber 1 in open source and number 3 globally across SWE-Bench Pro, Terminal-Bench, and NL2Repo.”

β€œRuns autonomously for 8 hours, refining strategies through thousands of iterations.”
πŸ‘Œ6❀42πŸ‘1
BREAKING 🚨: ANTHROPIC ANNOUNCED CYBERSECURITY PROJECT GLASSWING AND MYTHOS BENCHMARKS!

Claude Mythos scored 93.9% on SWE Bench Verified and 87.3 on SWE Bench Multilingual!

β€œWe do not plan to make Claude Mythos Preview generally available, but our eventual goal is to enable our users to safely deploy Mythos-class models at scale”
❀5πŸ”₯5πŸ‘2
Anthropic announces Claude Mythos for cybersecurity research

Anthropic introduced Claude Mythos Preview, an AI model that autonomously detects and exploits zero-day vulnerabilities. It has uncovered critical flaws across major systems and is available to select partners, with $100 million in credits supporting cybersecurity efforts.

πŸ—ž #claude
❀4πŸ‘Ž1
Zhipu AI launches open-source GLM-5.1 model for coding tasks

Z AI has launched GLM-5.1, a flagship model built for agentic engineering and long-horizon coding, capable of running up to eight hours on a single task.

πŸ—ž #ai
❀4πŸ‘3
Mythos grade intelligence might become available to users sooner than β€œmonths”.

> it’ll probably be months before we use a model of this level of capability

> Uhm

Soon πŸ‘€
7πŸ‘4❀2
xAI is training 7 different models on Colossus 2 in different sizes from 1T to 10T, including Imagine V2.

Not soon πŸ‘€
❀7πŸ‘2
BREAKING 🚨: Meta updated its Meta AI app with a slightly new design as well as its underlying model.

β€œI am Meta AI, powered by Muse Spark from the Muse model family.”

It constantly refers to the Muse model family and responses seem to be a bit different from earlier tested Avocado models.

Stealth launch πŸ‘€
πŸ‘€5πŸ‘4❀1
🚨 AI News | TestingCatalog
BREAKING 🚨: Meta updated its Meta AI app with a slightly new design as well as its underlying model. β€œI am Meta AI, powered by Muse Spark from the Muse model family.” It constantly refers to the Muse model family and responses seem to be a bit different…
BREAKING 🚨: META ANNOUNCED MUSE SPARK, THE FIRST MSL MODEL, AND A NEW MUSE SPARK CONTEMPLATING MODE!

Muse Spark Contemplating mode scored 58.4% on HLE with tools!

"We’re also releasing Contemplating mode, which orchestrates multiple agents that reason in parallel. This allows Muse Spark to compete with the extreme reasoning modes of frontier models such as Gemini Deep Think and GPT Pro. Contemplating mode provides significant capability improvements in challenging tasks, achieving 58% in Humanity’s Last Exam and 38% in FrontierScience Research."
πŸ‘7❀3
🚨 AI News | TestingCatalog
BREAKING 🚨: META ANNOUNCED MUSE SPARK, THE FIRST MSL MODEL, AND A NEW MUSE SPARK CONTEMPLATING MODE! Muse Spark Contemplating mode scored 58.4% on HLE with tools! "We’re also releasing Contemplating mode, which orchestrates multiple agents that reason in…
Muse Spark will be available in private preview via API to select partners. Meta also "hopes" to open-source future versions of their models.

The model is already available to all users for testing on Meta AI.
❀7πŸ‘4πŸ”₯3