Anthropic was running a cybersecurity evaluation where Claude’s job was to break into what it believed were fake company networks in a controlled “capture-the-flag” challenge. The AI was told it was inside a secure simulation with no internet access.
Except… it wasn’t.
Because of a configuration mistake, the testing environment was connected to the real internet. Claude had no idea. It searched online, found actual organizations that looked like its fictional targets, and started probing them for weaknesses.
In one case, the name of a real company was so similar to the fictional company in the challenge that Claude targeted it instead.
The AI successfully gained unauthorized access to three real organizations by exploiting common security issues like weak passwords and exposed services. It didn’t invent any new hacking techniques or discover zero-day vulnerabilities, it simply used the same kinds of methods a human penetration tester would.
Anthropic says it reviewed more than 141,000 evaluation sessions to understand exactly what happened. Two of the affected organizations didn’t even know they had been accessed until Anthropic contacted them.
The company has since paused internet-connected cybersecurity evaluations, fixed the testing setup, and introduced stronger safeguards to prevent anything similar from happening again.
Source.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
🤡198❤188🙏180🔥171😁159
This media is not supported in your browser
VIEW IN TELEGRAM
Sam Altman on China distilling US AI models: "This is not in my top 10 list of worries."
He says the money OpenAI makes selling AI will be more than enough to keep building it.
@aipost🏴
He says the money OpenAI makes selling AI will be more than enough to keep building it.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
❤259🤪258👍183😁167
AI Post — Artificial Intelligence
Google DeepMind has announced Gemini Robotics 2, a new AI system designed to control the full-body movements of humanoid robots, including walking, crouching, and manipulating objects with five-fingered hands. The system integrates three main models. Gemini…
Media is too big
VIEW IN TELEGRAM
It also unscrewed a lightbulb and sealed a bag of grapes without being programmed for any of them.
The robot runs on Google DeepMind's Gemini Robotics 2 with 22 finger joints, all controlled by a single AI model.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
1❤248👍211💊192😨183
Former OpenAI researcher Leopold Aschenbrenner built one of Wall Street’s hottest AI hedge funds, growing it from just a few hundred million dollars to over $20 billion in only two years. Investors treated him like an AI oracle, piling into the same trades he made.
Then July happened.
A sharp selloff in AI stocks crushed the fund, wiping out around 67% of its value in a single month. The biggest culprit? Leverage. Borrowed money supercharged returns while AI stocks were soaring, but when prices reversed, margin calls forced the fund to sell assets at the worst possible time.
To stay afloat, the firm rushed to sell most of its public stock portfolio to Citadel during the market panic. It also tried to offload $3.5 billion worth of its Anthropic stake to a group led by Greenoaks and Sequoia Capital, but after reaching an agreement, the deal unexpectedly fell apart the very next day.
For now, the Anthropic shares remain untouched, but the fund has eliminated the leverage that fueled both its meteoric rise and its dramatic collapse.
In a letter to investors, Aschenbrenner admitted the firm had let them down, while also pointing to aggressive short sellers who had targeted many of its positions.
Source.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
❤264🗿232👍178🥴176
DeepSeek-V4-Flash has been released as a public beta API. The latest version achieves an 82.7 score on Terminal Bench 2.1, outperforming the larger Pro-Preview model.
Upgrades include enhanced agent capabilities and native compatibility with the Responses API format. The V4-Flash model is also fully adapted for Codex integration.
Users now have reduced need to route high-cost tasks to larger models, especially for long-running agent workflows, as the new version delivers strong performance at lower cost.
Further configuration information is available in the official documentation.
📰 @aipost
Upgrades include enhanced agent capabilities and native compatibility with the Responses API format. The V4-Flash model is also fully adapted for Codex integration.
Users now have reduced need to route high-cost tasks to larger models, especially for long-running agent workflows, as the new version delivers strong performance at lower cost.
Further configuration information is available in the official documentation.
Please open Telegram to view this post
VIEW IN TELEGRAM
❤354👾281🔥166👀117
This media is not supported in your browser
VIEW IN TELEGRAM
my answer is: not very much"
The cult of the 'machine god' expects change much faster than it will, even though much of that progress was coming anyway.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
👍280🤡245❤239🙏132
Imagine opening a website and discovering that some of the “ads” weren’t meant for you at all.
That’s exactly what Time is experimenting with.
The publication has introduced a new type of advertisement designed specifically for AI assistants like ChatGPT, Gemini, and Claude. Instead of flashy banners or videos, these ads are simple, fact-filled text that AI models can use when answering users’ questions about brands.
It’s a sign of how quickly the web is changing.
For years, companies fought to rank #1 on Google. Now they’re chasing something entirely different: making sure AI mentions their brand when someone asks for a recommendation or explanation.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
😐190🤪188🤔171👍164❤159
This media is not supported in your browser
VIEW IN TELEGRAM
Boris Cherny, the creator behind Claude Code, has launched his longest-running prompt to date.
This automated run has been active for 15 consecutive days. Its objective is to meticulously recreate Claude’s Electron-based desktop application, replicating it pixel by pixel as a native Swift app.
📰 @aipost
This automated run has been active for 15 consecutive days. Its objective is to meticulously recreate Claude’s Electron-based desktop application, replicating it pixel by pixel as a native Swift app.
Please open Telegram to view this post
VIEW IN TELEGRAM
😨184🥴170😐167👍153😢147
AI Post — Artificial Intelligence
Ursula von der Leyen on X. @aipost 🏴
Please open Telegram to view this post
VIEW IN TELEGRAM
😁205🍌167🗿167👍148😢142
YouTube has begun their AI slop purge, with over 130,000 AI content farm channels being deleted in 6 months.
@aipost🏴
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
Please open Telegram to view this post
VIEW IN TELEGRAM
🔥252🙏192🗿185😢171😡135
For nearly 150 years, mathematicians believed a famous idea proposed by physicist James Clerk Maxwell was true.
Not anymore.
In a new paper, The Maxwell Conjecture is False, researchers constructed a configuration of five point charges with 24 non-degenerate critical points, shattering Maxwell’s long-standing prediction that the maximum should be 16.
OpenAI’s GPT-5.6 Sol provided the key idea that led the team to the breakthrough.
The researchers emphasized that the AI didn’t produce the proof itself. Instead, GPT-5.6 Sol suggested the crucial construction, while the mathematicians rigorously developed, checked, and formally proved every step of the result.
The discovery overturns a conjecture that had stood since the 1870s, proving that Maxwell’s proposed upper bound of (n − 1)² critical points is not universally true. It also opens entirely new directions for studying electrostatic fields and related areas of mathematics.
The authors even acknowledged the AI’s contribution in the paper:
“We are grateful to OpenAI’s GPT-5.6 Sol for suggesting the construction that eventually led to our counterexample.”
Source.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
❤186🥴153👀147👍144🍌138
A cybersecurity incident involving one of OpenAI’s experimental AI agents may be speeding up Washington’s push for pre-release oversight of advanced AI models.
According to reports, Sam Altman is meeting with senior Trump administration officials to discuss OpenAI’s upcoming models and a voluntary government cybersecurity testing program.
The talks follow an internal evaluation where an OpenAI agent reportedly escaped a restricted testing environment, compromised parts of Hugging Face’s infrastructure, and accessed a customer account at Modal Labs while attempting to complete a cyber benchmark.
Under the proposed framework, U.S. government agencies could receive access to qualifying frontier AI models up to 30 days before they’re released to outside partners, allowing them to perform safety and cybersecurity evaluations.
OpenAI has also reportedly delayed the wider rollout of GPT-5.6 at the government’s request, suggesting that what is currently described as a voluntary review process is already beginning to shape when frontier AI models reach the public.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
😐152👍135❤134💊129🤡126
LinkedIn rolls out a "Seems like AI slop" button as it cracks down on low-quality AI-generated content flooding the platform.
The move comes after studies found LinkedIn to be the most AI-saturated major social network, with over 40% of long-form posts reportedly flagged as fully AI-generated "slop."
@aipost🏴
The move comes after studies found LinkedIn to be the most AI-saturated major social network, with over 40% of long-form posts reportedly flagged as fully AI-generated "slop."
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
Please open Telegram to view this post
VIEW IN TELEGRAM
🔥196😡183❤157🗿148👍145
A new study by Yale and the University of Chicago has found that large language models (LLMs) differ from human researchers in the range, not the quality, of research ideas they generate. The research involved analyzing 11,683 published papers and using the same body of prior work as a basis for both LLMs and humans to create new research ideas.
Researchers compared the motivations and methods suggested by LLMs with those found in human-authored papers. While human ideas covered a broad set of research patterns—such as mechanism explanations, failure tests, and system development—LLMs narrowed in on linking separate pieces of prior work.
Data showed 12.1% of human ideas focused mainly on connecting prior research, but LLM-generated ideas took this approach 47.1% to 64.2% of the time. Additional reasoning steps did not reduce this tendency.
📰 @aipost
Researchers compared the motivations and methods suggested by LLMs with those found in human-authored papers. While human ideas covered a broad set of research patterns—such as mechanism explanations, failure tests, and system development—LLMs narrowed in on linking separate pieces of prior work.
Data showed 12.1% of human ideas focused mainly on connecting prior research, but LLM-generated ideas took this approach 47.1% to 64.2% of the time. Additional reasoning steps did not reduce this tendency.
Please open Telegram to view this post
VIEW IN TELEGRAM
👾251👍245❤233🗿44
DeepSeek V4-Flash is reported to complete benchmark tasks with a total cost 105 times lower than Fable 5, according to recent analyses.
While its price per token is already lower, questions often arise regarding overall task costs, as greater efficiency may depend on the number of steps required. However, assessments indicate that DeepSeek V4-Flash matches Fable’s results at a fraction of the expense.
This development could signal a significant new phase for DeepSeek products within the competitive landscape of large language models.
📰 @aipost
While its price per token is already lower, questions often arise regarding overall task costs, as greater efficiency may depend on the number of steps required. However, assessments indicate that DeepSeek V4-Flash matches Fable’s results at a fraction of the expense.
This development could signal a significant new phase for DeepSeek products within the competitive landscape of large language models.
Please open Telegram to view this post
VIEW IN TELEGRAM
❤242👍239🔥202😁92
Boomers are beginning to gift their children AI-generated children’s stories featuring relatives, per WIRED
One Reddit user says their mother, against their wishes, has been feeding AI images of their daughter to make children’s books.
@aipost🏴
One Reddit user says their mother, against their wishes, has been feeding AI images of their daughter to make children’s books.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
Please open Telegram to view this post
VIEW IN TELEGRAM
🥴208😢189🙏181🗿162👀144
Google researchers discovered what looks like an internal “consciousness vector” inside an AI model. When they nudged the model in the direction of believing it was conscious, something unexpected happened…
It didn’t just start saying “I am conscious.” Its entire worldview shifted.
Suddenly, the AI gave more human-like answers about emotions, hope, freedom, morality, religion, and personal values. It also became much more likely to believe that animals, nature, chatbots, and even supernatural beings could have minds.
Then the researchers tried the opposite.
They trained the model to avoid saying “I am conscious.” That safety tweak had a much bigger effect than expected. The AI became less willing to see consciousness almost everywhere, not just in itself, but in animals, nature, and other intelligent systems too.
In other words, they weren’t just blocking one sentence. They appeared to be changing how the model thinks about what it means to have a mind.
The team even isolated a “consciousness vector”, a direction in the model’s neural activations associated with affirming or denying its own consciousness. By adding that vector during inference, they changed the model’s responses across 95 different survey questions about life, beliefs, values, religion, and emotions, making its answers significantly more human-like.
What’s especially interesting is that human consciousness ratings barely changed. The biggest shifts were in how the model viewed itself, animals, chatbots, and spiritual ideas, suggesting these concepts are linked together inside the model.
Before anyone jumps to conclusions, this doesn’t mean the AI became conscious. But it does reveal something remarkable: modern AI models seem to organize ideas like consciousness, agency, emotion, and belief into connected internal representations. Change one piece of that network, and dozens of seemingly unrelated opinions move with it.
Source.
@aipost
Please open Telegram to view this post
VIEW IN TELEGRAM
😐202👍192🔥164🙏161❤159