At IROS 2026, 8 dedicated sessions focus on VLA models, which turn an image and text instruction into robot actions. Force and tactile sensing also appear throughout because cameras alone are not enough for dexterous work.
Diffusion and flow matching now generate robot movements, not images. Humanoid work covers walking and whole-body control, while one competition had robots assembled IKEA furniture themselves.
Reinforcement learning and imitation learning appear in almost every second paper. Robots can already do a lot, but data remains the bottleneck. Unlike language models, they are only starting to build a training pool.
Please open Telegram to view this post
VIEW IN TELEGRAM
Please open Telegram to view this post
VIEW IN TELEGRAM
β€603π189π125π107π73π₯1π’1
The companies open-sourced tools for Huawei Ascend, now the main Chinese AI accelerators used by local startups. The software gap is the hard part: the industry has spent 10+ years writing for Nvidia CUDA.
The release includes DeepGEMM for matrix multiplication and DeepEP for chip-to-chip communication. Their APIs match Nvidia versions, so existing code needs almost no rewriting.
The stack runs on TileLang. Its compiler now handles scheduling and data synchronization instead of requiring manual code. DeepSeek says its training kernels already run on Ascend. This covers kernel writing, not the full CUDA stack.
Please open Telegram to view this post
VIEW IN TELEGRAM
π―289π286β€285π₯283π4π1
The worldβs biggest AI companies signed a safe AI agreement with the White House.
But under Donald Trumpβs signature, βpresident of the United Statesβ became βpresident unites the States.β
Please open Telegram to view this post
VIEW IN TELEGRAM
Please open Telegram to view this post
VIEW IN TELEGRAM
π260π₯235π±234π€198π175π46
This media is not supported in your browser
VIEW IN TELEGRAM
According to Bloomberg, new CEO John Ternus wants Apple to release products more often than at its major autumn and spring presentations. He says faster launches and more experiments are needed to stay competitive in the AI era.
Ternus is also considering fewer management layers between engineers and top executives. Over the past 2 weeks, the hardware development team has shed some engineering program managers. Bloomberg reported earlier cuts in the Siri and Vision Pro teams in August.
The rumored pipeline includes a new HomePod mini and iPad mini. Apple is also reportedly working on a new Apple TV 4K, camera-equipped AirPods, touchscreen MacBooks, plus a special iPhone for its 20th anniversary.
Please open Telegram to view this post
VIEW IN TELEGRAM
π€―216π201π178π170β€137π135π₯114
Anthropic estimates that robots can already perform 74% of physical work tasks in the US under at least some conditions. With LLMs added, the figure reaches 81% of all work.
The most exposed roles are mostly vehicle operators. Waste sorting is another example. Waymo robotaxis and autonomous tractors show how the shift works. Sorting robots target recycling work.
With LLMs alone, transport and moving work had under 15% exposure. Robots raise it to about 90%. But exposure means technology can perform part of a job. Robots are currently cost-competitive with humans in only 0.3% of tasks.
Please open Telegram to view this post
VIEW IN TELEGRAM
π572β€531
Media is too big
VIEW IN TELEGRAM
The model generates clips from text, animates an image, or uses another video as a reference for motion and visual style. Audio is created with the video, including speech, background sound, and effects.
It runs on MiniMax H3, with additional training by HeyGen. Clips run from 5 to 15 seconds in 480p or 768p, with aspect ratios from 21:9 to 9:16.
HeyGenβs blind comparison put the model first against 4 others, based on 4,800 votes. 768p generation costs $0.03 per second, with a 50% discount through the end of October. It is available through API, OpenRouter, Runware, and ComfyUI. The web version will come later.
Please open Telegram to view this post
VIEW IN TELEGRAM
β€308π262π―256π€©188π114π2
This media is not supported in your browser
VIEW IN TELEGRAM
In Fresno, a Big Mac costs $5.69 at one McDonaldβs and $6.89 at another 2 miles away. That is a 21% gap, linked by Reuters to the chainβs AI pricing system.
The system seeks an βoptimal priceβ for each US restaurant and some overseas markets, based on local customersβ willingness to pay. Franchisees see prompts about price sensitivity. McDonaldβs tracks deviations and sends suggestions at least 3 times a year.
Owners formally set the price, but some told Reuters that deviations can trigger calls from headquarters. New franchise standards require them to βconstructively engageβ with approved pricing tools.
Please open Telegram to view this post
VIEW IN TELEGRAM
π€―233β€220π211π192π131π₯105
OpenAI added virtual try-on to ChatGPT. When a clothing product card appears, tap βTry onβ and upload a selfie plus a full-body photo. ChatGPT Images 2.5 generates the result.
The feature can also use clothing screenshots, not just items suggested by the bot. Favorite pieces can be saved to Library. Google added the same feature to Search in 2025.
The catch is privacy. By default, images uploaded from a personal account may be used to train models unless you opt out. That includes full-body photos.
Please open Telegram to view this post
VIEW IN TELEGRAM
π299β€257π³226π₯191π16
Apple is reportedly developing J450, a small cylindrical home security camera with a very low frame rate. AI analyzes what its sensor sees and produces only text descriptions, such as someone entering a room. Face recognition is also expected.
J450 is designed as a companion to Appleβs J490 smart home display. It is part of a broader smart home ecosystem that may include Appleβs own hub, plus security devices from Apple and third parties. Its technology is reportedly similar to what Apple is developing for camera-equipped AirPods.
There will be no video archive to watch later. No video leaves less to leak, while also leaving nothing to review when something happens at home.
Please open Telegram to view this post
VIEW IN TELEGRAM
β€153π€―144π€©128π127π101π97π€48
Karpathy suggests asking AI to explain complex topics in ASD-STE100 simplified English, with short sentences and little ambiguity. For harder subjects, a diagram or a custom HTML page can work better than plain text.
Modern models can already build one-off interfaces for a specific question. Karpathy is especially interested in personal teaching videos, with animation, a 3Blue1Brown-style explanation, and voice from ElevenLabs or local TTS.
As code and AI work become cheaper, temporary interactive artifacts become practical. The next AI interface may be an answer assembled for each task, rather than an endless chat with walls of text.
Please open Telegram to view this post
VIEW IN TELEGRAM
π130β€124π€―111π92