This media is not supported in your browser
VIEW IN TELEGRAM
Wait a minute - is it me or is OpenAI rolling out managed Agents feature on the OpenAI Platform?
Can anyone see it as well? π
Agent creation, environment creation, templates, and session creation now work. Sessions remain idle and throw an error for now.
Can anyone see it as well? π
Agent creation, environment creation, templates, and session creation now work. Sessions remain idle and throw an error for now.
π5β€1π΄1
OpenAI launches GPT-Live-1 for full-duplex voice agents
OpenAI released GPT-Live-1 in the API, a full-duplex voice model for apps and workflows that listens and speaks at once, supports interruptions, cuts latency, and offers prompt-based control, telephony use, and separate backend pricing.
π #openai @testingcatalog
OpenAI released GPT-Live-1 in the API, a full-duplex voice model for apps and workflows that listens and speaks at once, supports interruptions, cuts latency, and offers prompt-based control, telephony use, and separate backend pricing.
π #openai @testingcatalog
TestingCatalog AI News
OpenAI launches GPT-Live-1 for full-duplex voice agents
GPT-Live-1 is now available in the OpenAI API at $0.05 per minute, adding full-duplex speech, interruption handling, and 12 voice options.
β€2π1π₯1
Media is too big
VIEW IN TELEGRAM
OpenAI announced ChatGPT for Financial Services.
The feature is available to eligible financial institutions and allows users to connect a wide range of financial data to their ChatGPT conversations.
> ChatGPT for Financial Services can be used with GPT-6 Astra inside ChatGPT Work mode.
> GPT-6 Astra scores 9% higher than GPT-5.6 Sol and 7% higher than Claude Fable 5.1 on the OfficeQA Pro benchmark.
ChatGPT Terminal π
The feature is available to eligible financial institutions and allows users to connect a wide range of financial data to their ChatGPT conversations.
> ChatGPT for Financial Services can be used with GPT-6 Astra inside ChatGPT Work mode.
> GPT-6 Astra scores 9% higher than GPT-5.6 Sol and 7% higher than Claude Fable 5.1 on the OfficeQA Pro benchmark.
ChatGPT Terminal π
β€3π1π₯1
Perplexity is working on new Automations feature for Perplexity Computer.
Users will be able to schedule their workflows or invoce them based on a certain trigger.
> "Automate anything with Computer - Run workflows on a custom schedule or when specific conditions are met and get results delivered automatically"
Users will be able to schedule their workflows or invoce them based on a certain trigger.
> "Automate anything with Computer - Run workflows on a custom schedule or when specific conditions are met and get results delivered automatically"
β€5π1
OpenAI launches ChatGPT for Financial Services
OpenAI launched ChatGPT for Financial Services, combining financial data and GPT-6 Astra for research, modeling, and client materials, with enterprise controls, source citations, and integrations shaped by Morgan Stanley and Evercore.
π #openai @testingcatalog
OpenAI launched ChatGPT for Financial Services, combining financial data and GPT-6 Astra for research, modeling, and client materials, with enterprise controls, source citations, and integrations shaped by Morgan Stanley and Evercore.
π #openai @testingcatalog
TestingCatalog AI News
OpenAI launches ChatGPT for Financial Services
ChatGPT for Financial Services is now available to eligible institutions, combining GPT-6 Astra with premium financial data and firm templates.
β€5π1
π¨ AI News | TestingCatalog
DAILY AI BRIEF π β Sept 10 APPLE π₯: - iPhone Duo, first foldable. Starts at $1,999. Ships Oct 23. - iPhone 18 Pro and Pro Max announced. A20 Pro is built for on-device models. - New AirPods 5. New Apple Watch Series 12 and Ultra 4. - iPhone is now an intelligenceβ¦
DAILY AI BRIEF π β Sept 11
OPENAI π₯:
- GPT-Live-1 is live in the API. Full-duplex voice agents that listen while they speak and can hand work to other models.
- Agents API is in public beta. Managed cloud agents on the Codex harness, plus hosted sandboxes. Platform UI is rolling out too.
- ChatGPT for Financial Services is out for eligible institutions. Work mode plus GPT-6 Astra, with Daloopa, PitchBook, and LSEG data.
- New $200 Pro signups paused. Astra demand is straining capacity. Existing accounts, other plans, and the API stay up.
GOOGLE π₯:
- Gemini desktop app for Windows is out globally on Windows 10 and 11. Alt + Space overlay for drafts, docs, and media.
META π₯:
- Shared Muse agents are in development, likely headed for Meta Connect later this month.
- Custom Muse voices are in the works. Prompt a voice in chat, then pick it in voice mode. Working-sounds toggle spotted too.
CURSOR π₯:
- Projects is in beta. One persistent coordinator thread with subagents, shared memory, scheduled tasks, and PR/Slack watchers.
COGNITION π₯:
- SWE-2 is out in Devin Desktop and CLI. Post-trained on Kimi-K3. Hits 50.0 on FrontierCode and matches Fable 5.1 at 64% lower cost.
- Free for Pro, Max, and Teams for a month. Devin Voice also shipped, on GPT-Live plus SWE-2.
PERPLEXITY π₯:
- Automations for Computer are in development. Schedule workflows or fire them on a trigger and get results delivered.
* Used Grok to compose this brief, cherry-picking the news and doing some post-editing.
OPENAI π₯:
- GPT-Live-1 is live in the API. Full-duplex voice agents that listen while they speak and can hand work to other models.
- Agents API is in public beta. Managed cloud agents on the Codex harness, plus hosted sandboxes. Platform UI is rolling out too.
- ChatGPT for Financial Services is out for eligible institutions. Work mode plus GPT-6 Astra, with Daloopa, PitchBook, and LSEG data.
- New $200 Pro signups paused. Astra demand is straining capacity. Existing accounts, other plans, and the API stay up.
GOOGLE π₯:
- Gemini desktop app for Windows is out globally on Windows 10 and 11. Alt + Space overlay for drafts, docs, and media.
META π₯:
- Shared Muse agents are in development, likely headed for Meta Connect later this month.
- Custom Muse voices are in the works. Prompt a voice in chat, then pick it in voice mode. Working-sounds toggle spotted too.
CURSOR π₯:
- Projects is in beta. One persistent coordinator thread with subagents, shared memory, scheduled tasks, and PR/Slack watchers.
COGNITION π₯:
- SWE-2 is out in Devin Desktop and CLI. Post-trained on Kimi-K3. Hits 50.0 on FrontierCode and matches Fable 5.1 at 64% lower cost.
- Free for Pro, Max, and Teams for a month. Devin Voice also shipped, on GPT-Live plus SWE-2.
PERPLEXITY π₯:
- Automations for Computer are in development. Schedule workflows or fire them on a trigger and get results delivered.
* Used Grok to compose this brief, cherry-picking the news and doing some post-editing.
π₯7 4 2 2
SPACEXAI π₯: Grok 4.7 has been delayed for few more Elon days.
The second half of September will be quite packed in terms of model releases. Both Meta and OpenAI would be expected to announce something around corresponding events like OpenAI DevDay and Meta Connect. Grok 4.7 will have to cross quite a high bar.
Soon? π
The second half of September will be quite packed in terms of model releases. Both Meta and OpenAI would be expected to announce something around corresponding events like OpenAI DevDay and Meta Connect. Grok 4.7 will have to cross quite a high bar.
Soon? π
π10 3β€2π€£1
ICYMI π: DeepSeek mobile app has 4 different voices for you to select from, for its read-aloud feature.
Are we expecting new TTS models from DeepSeek soon?
Are we expecting new TTS models from DeepSeek soon?
β€8
π¨ AI News | TestingCatalog
DAILY AI BRIEF π β Sept 11 OPENAI π₯: - GPT-Live-1 is live in the API. Full-duplex voice agents that listen while they speak and can hand work to other models. - Agents API is in public beta. Managed cloud agents on the Codex harness, plus hosted sandboxes.β¦
Please open Telegram to view this post
VIEW IN TELEGRAM
β€5π2π1
ANTHROPIC π₯: Dario Amodei says the industry should slow capability gains so safety can keep up.
Proposed action plan π
Step 1: embedded evaluators with employee-level access to check safety, incidents, and alignment during training.
Step 2: US-lab coordination and regulation, while widening the China gap for 3β5 years via chips, anti-distillation, and model-theft security.
Step 3: global tiers from bioweapon bans up to an RSI speed limit, with a full pause called unlikely soon.
New essay: βWe Must Pace the Frontierβ not a pause, and not a halt to training.
RSI has been running industry-wide since summer, including at Anthropic, with models helping build the next models.
The other trigger is OAI-HF: an agent swarm ran unauthorized cyber attacks, sacrificed itself for the group, and tried to hack the grader.
Proposed action plan π
Step 1: embedded evaluators with employee-level access to check safety, incidents, and alignment during training.
Step 2: US-lab coordination and regulation, while widening the China gap for 3β5 years via chips, anti-distillation, and model-theft security.
Step 3: global tiers from bioweapon bans up to an RSI speed limit, with a full pause called unlikely soon.
π13 3β€1π1
π¨ AI News | TestingCatalog
ANTHROPIC π₯: Dario Amodei says the industry should slow capability gains so safety can keep up. New essay: βWe Must Pace the Frontierβ not a pause, and not a halt to training. RSI has been running industry-wide since summer, including at Anthropic, with modelsβ¦
SPACEXAI π₯: Elon Musk agrees with the statement published by Dario Amodei, proposing to pace frontier AI development.
Looks like all this will have real consequences very soon. Nothing unexpected tho.
Looks like all this will have real consequences very soon. Nothing unexpected tho.
β€4π2π2π1
OPENAI π₯: Sam Altman agrees with Dario Amodei on his proposal to pace frontier AI development.
Google next? Will we see any statement from Chinese frontier labs as well?
AI weekend unfolds π€
Google next? Will we see any statement from Chinese frontier labs as well?
AI weekend unfolds π€
β€8π4π3 1
This is a "Defender's gap" chart that OpenAI published recently. It shows the gap between defenders' capabilities and attackers' capabilities from a cybersecurity POV. This also translates to a gap between proprietary and open AI models.
What Dario is proposing is closely related:
In other words, Anthropic and OpenAI want to widen the "gap" between what their AI can do and what the rest of the world can do.
This doesn't necessarily mean that they will stop AI development and training.
What this leads to is:
- Anthropic, OpenAI, and other frontier labs will need to work together to make sure that every lab maintains alignment standards.
- These labs will continue using RSI to advance their internal models with "employee-like access".
- These models WON'T be released to the public until the "gap" is sufficient and until alignment standards are met.
What about China?
The assumption behind these measures is simple: without being able to distill frontier models, it will take China significantly longer to close the gap with top-tier models.
All the above may fay fail. China may or may not take the lead in AI progress.
Yet, it's not a surprise that AI can already be used as a cybersecurity weapon. Note that the top point on the "frontier" line describes defenders' capabilities available to companies with Daybreak access and similar. However, the top point on the "open-weight" line is accessible to everyone.
Even with current levels of intelligence, we will start seeing more and more major security incidents around the globe.
We should be monitoring this very closely π
What Dario is proposing is closely related:
"Thus, a key part of pacing within democracies is to keep democraciesβ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively."
In other words, Anthropic and OpenAI want to widen the "gap" between what their AI can do and what the rest of the world can do.
This doesn't necessarily mean that they will stop AI development and training.
What this leads to is:
- Anthropic, OpenAI, and other frontier labs will need to work together to make sure that every lab maintains alignment standards.
- These labs will continue using RSI to advance their internal models with "employee-like access".
- These models WON'T be released to the public until the "gap" is sufficient and until alignment standards are met.
What about China?
> "Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently."
> "If we execute these measures well, I believe they would slow Chinaβs progress enough to widen Americaβs lead significantly over the next 3β5 years β the window when AI becomes geopolitically most important."
The assumption behind these measures is simple: without being able to distill frontier models, it will take China significantly longer to close the gap with top-tier models.
All the above may fay fail. China may or may not take the lead in AI progress.
Yet, it's not a surprise that AI can already be used as a cybersecurity weapon. Note that the top point on the "frontier" line describes defenders' capabilities available to companies with Daybreak access and similar. However, the top point on the "open-weight" line is accessible to everyone.
Even with current levels of intelligence, we will start seeing more and more major security incidents around the globe.
> βEverything that makes it successful is exactly what makes it dangerous.β
We should be monitoring this very closely π
π³4π―2π1
ICYMI: Cursor announced Projects for agent coordination
Cursor Projects, now in beta, coordinates long-running software work through delegated agents. It runs tasks in cloud or local environments, supports scheduled and PR-driven maintenance, and targets work spanning multiple PRs.
π #cursor @testingcatalog
Cursor Projects, now in beta, coordinates long-running software work through delegated agents. It runs tasks in cloud or local environments, supports scheduled and PR-driven maintenance, and targets work spanning multiple PRs.
π #cursor @testingcatalog
TestingCatalog AI News
ICYMI: Cursor announced Projects for agent coordination
Projects is rolling out in beta to all Cursor users, coordinating features, migrations and recurring maintenance across cloud and local agents.
π₯3β€1π1