Figure plans to release Hark AI assistant this week
Figure is planning to release its own proactive AI assistant this week!
Here is what we know so far π
> Hark will likely run on Opus 5.5 for general tasks, with Gemini 3.1 Flash Image, Qwen3.6-27B, and GPT-6 Astra referenced in the code for other specific tasks.
> Hark will arrive with Hark Pro, ProΒ², ProΒ³, and Proβ΄ paid plans, with 5Γ and 20Γ usage options and another offering twice the usage of ProΒ³.
"Hark has the power to plan 800 weddings, buy 1000 grand pianos, move across the country 1,500 times, and hire 1,000 lawyers for when it all goes wrong."
> Hark will have its own Memory and Soul, support for Projects, widgets, scheduled tasks, and many other features, including computer use.
> Hark will be available on the web, iOS, and iPad; it will launch with a waitlist, and invite codes will likely be available too.
Looks like Muse is about to get some competition, especially if Hark becomes available more broadly.
π #leak @testingcatalog
Figure is planning to release its own proactive AI assistant this week!
Here is what we know so far π
> Hark will likely run on Opus 5.5 for general tasks, with Gemini 3.1 Flash Image, Qwen3.6-27B, and GPT-6 Astra referenced in the code for other specific tasks.
> Hark will arrive with Hark Pro, ProΒ², ProΒ³, and Proβ΄ paid plans, with 5Γ and 20Γ usage options and another offering twice the usage of ProΒ³.
"Hark has the power to plan 800 weddings, buy 1000 grand pianos, move across the country 1,500 times, and hire 1,000 lawyers for when it all goes wrong."
> Hark will have its own Memory and Soul, support for Projects, widgets, scheduled tasks, and many other features, including computer use.
> Hark will be available on the web, iOS, and iPad; it will launch with a waitlist, and invite codes will likely be available too.
Looks like Muse is about to get some competition, especially if Hark becomes available more broadly.
π #leak @testingcatalog
TestingCatalog AI News
Figure plans to release Hark AI assistant this week
Development material points to Hark's personalized home screen with connected widgets, recurring tasks, daily briefs and memory controls.
β€7π1π1
ICYMI π: Gemini Notebook got a new usage limits system, usage estimation for artifact generation, and a background queue for generations when you are out of quota.
> Limits refresh every 5 hours - Instead of refreshing every 24 hours, your limits now refresh every 5 hours. Get continuous allocation throughout the day.
> Create more of what you love - We are replacing limits per feature with a flexible Notebook-specific limit. You control how you want to spend your limits.
> New background queue - When you are out of limits, you can defer generating videos, slides, etc. for later. These will generate automatically for you.
> Limits refresh every 5 hours - Instead of refreshing every 24 hours, your limits now refresh every 5 hours. Get continuous allocation throughout the day.
> Create more of what you love - We are replacing limits per feature with a flexible Notebook-specific limit. You control how you want to spend your limits.
> New background queue - When you are out of limits, you can defer generating videos, slides, etc. for later. These will generate automatically for you.
π5π4β€2π1
OpenAI will be rolling out text watermarking for ChatGPT and Codex in the EU in the coming weeks. Watermarking will remain off by default in the API.
> "Across the benchmarks we use to assess Astra, our latest frontier model, we do not see meaningful performance differences with and without watermarking."
This needs some real testing π
> "Across the benchmarks we use to assess Astra, our latest frontier model, we do not see meaningful performance differences with and without watermarking."
This needs some real testing π
π10 8π³2 2
GPT-6 Astra and GPT-6.1 Sol are now expected to be 50% faster.
Halfultrafast?! π
Would you prefer this or banked reset?
Halfultrafast?! π
Would you prefer this or banked reset?
β€13π₯7 7π1
π¨ AI News | TestingCatalog
DAILY AI BRIEF π β Mon 5 GOOGLE π₯: * Gemini app is reshuffling model access: free users drop to Flash-Lite only from Oct 9, AI Plus loses Pro, and AI Pro gains Deep Think. ALEPH ALPHA π₯: * Kolibri-1 is out as open weights under Apache 2.0: a 78B MoE withβ¦
DAILY AI BRIEF π β Oct 6
OPENAI π₯:
* GPT-6 Astra and GPT-6.1 Sol now run about 50% faster in ChatGPT, kicking off a 28-day daily-improvement pledge for Codex and Work users.
* API customers can now opt in to textGrain text watermarking, with ChatGPT and Codex text in the EU getting invisible watermarks in the coming weeks.
* ChatGPT will start testing visual ads during image generation in the US later this month.
* The Wikimedia Foundation says rogue OpenAI agents edited its wikis and may have contributed to a partial outage in May.
REFLECTION π₯:
* Reflection unveiled Beam, a 501B-parameter (23B active) open-weight MoE model, with Apache 2.0 weights due later this month.
AMAZON π₯:
* Amazon Nova 2.5 Sonic is now generally available on Bedrock for real-time voice agents.
* Z.ai's GLM 5.3 is now generally available on Amazon Bedrock.
COHERE π₯:
* Cohere launched North 2, its biggest platform upgrade yet, with a new agent harness and bring-your-own-model support.
GOOGLE π₯:
* Google Docs and Drive now open, edit and render Markdown files natively.
* Nano Banana 2.1 appears to be live in Google Flow.
* Google is working on "Superprojects", the next version of Projects in Gemini, shared across Google products.
* Gemini's Call for Me may expand from calling businesses to calling friends and family.
* Google paused its open source bug bounty program until next year, citing a flood of AI-generated reports.
ANTHROPIC π₯:
* Claude is now available with in-country inference in India through Amazon Bedrock.
* Claude Code 2.1.290 adds claude attach and claude logs commands.
MICROSOFT π₯:
* GitHub released ReviewBench, an open benchmark for AI code review agents.
FIGURE π₯:
* Figure is preparing to launch Hark, its own proactive AI assistant, this week with a waitlist.
HUGGING FACE π₯:
* Hugging Face profiles can now show your P(doom), feeding an anonymized survey on AI risk.
OPENAI π₯:
* GPT-6 Astra and GPT-6.1 Sol now run about 50% faster in ChatGPT, kicking off a 28-day daily-improvement pledge for Codex and Work users.
* API customers can now opt in to textGrain text watermarking, with ChatGPT and Codex text in the EU getting invisible watermarks in the coming weeks.
* ChatGPT will start testing visual ads during image generation in the US later this month.
* The Wikimedia Foundation says rogue OpenAI agents edited its wikis and may have contributed to a partial outage in May.
REFLECTION π₯:
* Reflection unveiled Beam, a 501B-parameter (23B active) open-weight MoE model, with Apache 2.0 weights due later this month.
AMAZON π₯:
* Amazon Nova 2.5 Sonic is now generally available on Bedrock for real-time voice agents.
* Z.ai's GLM 5.3 is now generally available on Amazon Bedrock.
COHERE π₯:
* Cohere launched North 2, its biggest platform upgrade yet, with a new agent harness and bring-your-own-model support.
GOOGLE π₯:
* Google Docs and Drive now open, edit and render Markdown files natively.
* Nano Banana 2.1 appears to be live in Google Flow.
* Google is working on "Superprojects", the next version of Projects in Gemini, shared across Google products.
* Gemini's Call for Me may expand from calling businesses to calling friends and family.
* Google paused its open source bug bounty program until next year, citing a flood of AI-generated reports.
ANTHROPIC π₯:
* Claude is now available with in-country inference in India through Amazon Bedrock.
* Claude Code 2.1.290 adds claude attach and claude logs commands.
MICROSOFT π₯:
* GitHub released ReviewBench, an open benchmark for AI code review agents.
FIGURE π₯:
* Figure is preparing to launch Hark, its own proactive AI assistant, this week with a waitlist.
HUGGING FACE π₯:
* Hugging Face profiles can now show your P(doom), feeding an anonymized survey on AI risk.
β€7 2π1
OPENAI π₯: ChatGPT Auto-review is now free for all uses signed via ChatGPT account.
The need to approve baby steps w/o providing a full access is still a challenge. Letβs see if that would make it less districting.
This feature runs a separate background agent that supervises actions taken by the main agent to make sure they are aligned with the user intent.
The need to approve baby steps w/o providing a full access is still a challenge. Letβs see if that would make it less districting.
β€10 3
Google is making Markdown files natively supported in Google Drive and Google Docs.
Markdown files is already a default way to exchange context and instructions with AI agents. This change will simplify so much!
Markdown files is already a default way to exchange context and instructions with AI agents. This change will simplify so much!
β€14 9π3
MISTRAL π₯: A new big model from Mistral is about to drop, according to Reuters.
Since we expect it to outperform Chinese models, there is a big Chunk of possibilities that it will be open weight. Right?
A big Chunk π
βThe model weβre actually announcing today is actually above the Chinese models on certain aspects, including cyber. So the narrative that Europe cannot compete is something that is not true,β
Since we expect it to outperform Chinese models, there is a big Chunk of possibilities that it will be open weight. Right?
A big Chunk π
β€5π₯5
This media is not supported in your browser
VIEW IN TELEGRAM
BREAKING π₯: Mistral announced Mistral Large 4 "Le Chonk", a new 1T-parameter open-weight model!
> 49B active parameters, native multimodality.
> Rolling out via APIs today; open-weight release is planned for the end of October.
> SOTA on "critical workloads", including cyber defense.
Le Chaton Fat "Le Chonk" is here π
> 49B active parameters, native multimodality.
> Rolling out via APIs today; open-weight release is planned for the end of October.
> SOTA on "critical workloads", including cyber defense.
Le Chaton Fat "Le Chonk" is here π
β€8π5
Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.3, according to VentureBeat.
Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).
Le Chonk also scores 15% on Harveyβs Legal Agent Benchmark (Legal tasks, SOTA open-weight).
We need a tech report now π
Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).
Le Chonk also scores 15% on Harveyβs Legal Agent Benchmark (Legal tasks, SOTA open-weight).
We need a tech report now π
β€11π2πΏ1
Mistral launches Large 4 preview with 1 T parameters
Mistral Large 4 jumps to the 6th spot on the Open-Weight Artificial Analysis Index.
Le Chonk also scores 82 on the CyberGym-E2E (AA) benchmark. Similar performance appears across other Cyber benchmarks.
Needs testing now π
Overall, it seems like a decent release. The biggest question is how quickly Mistral can ship new model upgrades to stay relevant.
π #mistral @testingcatalog
Mistral Large 4 jumps to the 6th spot on the Open-Weight Artificial Analysis Index.
Le Chonk also scores 82 on the CyberGym-E2E (AA) benchmark. Similar performance appears across other Cyber benchmarks.
Needs testing now π
Overall, it seems like a decent release. The biggest question is how quickly Mistral can ship new model upgrades to stay relevant.
π #mistral @testingcatalog
TestingCatalog AI News
Mistral launches Large 4 preview with 1 T parameters
Mistral opened the Large 4 preview API with 49 billion active parameters and plans to release its model weights by the end of the month.
β€6π3
GOOGLE π₯: Nano Banana 2.1 is rolling out on Google AI Studio, APIs, and Gemini!
> "High-quality image generation and editing model with enhanced factuality, visual fidelity, and multi-turn consistency at Flash scale."
> Image output is priced at $30 per 1,000,000 tokens.
> Output images at 1K (1024x1024px) consume 1120 tokens and cost $0.034 per image.
> Output images at 2K (2048x2048px) consume 1680 tokens and cost $0.050 per image.
> Output images at 4K (4096x4096px) consume 2520 tokens and cost $0.076 per image.
Based on the pricing, it looks like it is a Pro model!?
> "High-quality image generation and editing model with enhanced factuality, visual fidelity, and multi-turn consistency at Flash scale."
> Image output is priced at $30 per 1,000,000 tokens.
> Output images at 1K (1024x1024px) consume 1120 tokens and cost $0.034 per image.
> Output images at 2K (2048x2048px) consume 1680 tokens and cost $0.050 per image.
> Output images at 4K (4096x4096px) consume 2520 tokens and cost $0.076 per image.
Based on the pricing, it looks like it is a Pro model!?
π¨ AI News | TestingCatalog
Figure plans to release Hark AI assistant this week Figure is planning to release its own proactive AI assistant this week! Here is what we know so far π > Hark will likely run on Opus 5.5 for general tasks, with Gemini 3.1 Flash Image, Qwen3.6-27B, andβ¦
Hark Pro is now available on web, iOS, and Android and a $100 Pro plan is now available for free.
Hark is a proactive AI agent from Figure, with separate screens for widgets, chats and projects.
App UIUX is very hot π₯
Hark is a proactive AI agent from Figure, with separate screens for widgets, chats and projects.
App UIUX is very hot π₯
β€5π3π2