Artificial Intelligence||DL
605 subscribers
22 photos
35 videos
10 files
180 links
Channel for who have a passion for -
* Artificial Intelligence
* Machine Learning
* Deep Learning
* Data Science
* Computer vision
* IT news
Admin: @AIchiman
Download Telegram
😁4
This media is not supported in your browser
VIEW IN TELEGRAM
↗️ TrackVLA++ Visual Trackingβ†˜οΈ

πŸ‘‰TrackVLA++ is a novel Vision-Language-Action model that incorporates spatial reasoning and target identification memory, enabling SOTA performance in both long-horizon and highly crowded tracking scenarios. Model announcedπŸ’™

πŸ‘‰Review https://t.ly/ruYzc
πŸ‘‰Paper https://arxiv.org/pdf/2510.07134
πŸ‘‰Project pku-epic.github.io/TrackVLA-plus-plus-Web/
πŸ‘‰Repo TBA
πŸ”₯1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ’„Pixel-Perfect Depth (SOTA)πŸ’„

πŸ‘‰Pixel-Perfect Depth is a mono-depth estimation model with pixel-space diffusion transformers. New SOTA. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://t.ly/75PGo
πŸ‘‰Paper https://lnkd.in/d8wxFpyY
πŸ‘‰Project https://lnkd.in/dV5HhsqH
πŸ‘‰Repo https://lnkd.in/d9JKFBJq
πŸ‘‰Demo https://lnkd.in/d3wBkKJ9
πŸ”₯1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‘” Universal Image Restoration πŸ‘”

πŸ‘‰LucidFlux by HKUSTGZ is the universal image restoration framework built on a large-scale diffusion transformer that delivers photorealistic restorations of real-world low-quality (LQ) images, outperforming SOTA diffusion-based models across diverse degradations. Repo under custom Non-Commercial LicenseπŸ’™

πŸ‘‰Review https://t.ly/Z5cA3
πŸ‘‰Paper https://arxiv.org/pdf/2509.22414
πŸ‘‰Project https://w2genai-lab.github.io/LucidFlux/
πŸ‘‰Repo https://github.com/W2GenAI-Lab/LucidFlux
πŸ”₯1
This media is not supported in your browser
VIEW IN TELEGRAM
🫧 Detect Anything via MLLM 🫧

πŸ‘‰Rex-Omni is a 3B-multimodal model that unifies visual perception tasks, including object detection, OCR, pointing, key-pointing & visual prompting into a single next point prediction framework. Impressive results. Repo under IDEA License 1.0πŸ’™

πŸ‘‰Review https://t.ly/DCTk_
πŸ‘‰Paper https://lnkd.in/d4VDD-9j
πŸ‘‰Project https://lnkd.in/d6unEyvq
πŸ‘‰Repo https://lnkd.in/dkYJFe-x
πŸ”₯3
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ«™Universal Feature Up-SamplingπŸ«™

πŸ‘‰AnyUp is a novel method for feature up-sampling that can be applied to ANY vision feature at ANY resolution, without encoder-specific training: inference-time feature-agnostic up-sampling architecture to improve up-sampling quality. Repo under CC-4.0πŸ’™

πŸ‘‰Review https://t.ly/HvEw9
πŸ‘‰Paper https://arxiv.org/pdf/2510.12764
πŸ‘‰Project https://wimmerth.github.io/anyup/
πŸ‘‰Repo https://github.com/wimmerth/anyup
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ¦„ City-Tour -> Simulation πŸ¦„

πŸ‘‰UrbanVerse is a novel system to convert real-world urban scenes from city-tour videos into physics-aware, interactive simulation environments, enabling scalable robot learning in urban spaces with real-world generalization. Repo & Data announced πŸ’™

πŸ‘‰Review https://t.ly/UvXNS
πŸ‘‰Paper https://arxiv.org/pdf/2510.15018
πŸ‘‰Project https://urbanverseproject.github.io/
πŸ‘‰Repo TBA
πŸ‘2
🌡All-in-One Dense Keypoints🌡

πŸ‘‰DeepDetect is a novel all-in-one, dense keypoints detector that unifies the strengths of SIFT, ORB, BRISK, FAST, AGAST, Harris, Shi-Tomasi, Canny & Sobel into a neural net. DAMN ROMANTIC. Repo under MITπŸ’™

πŸ‘‰Review https://t.ly/VKGct
πŸ‘‰Paper https://arxiv.org/pdf/2510.17422
πŸ‘‰Repo https://github.com/saktx/DeepDetect
πŸ‘2
This media is not supported in your browser
VIEW IN TELEGRAM
🏜️Omni Driving Navigation Models🏜️

πŸ‘‰OmniNWM is a unified panoramic navigation world model that advances autonomous driving by jointly generating multi-modal states (RGB, semantics, depth, 3D occupancy), enabling precise action control & facilitating closed-loop evaluation through occupancy-based dense rewards. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://t.ly/ktXvz
πŸ‘‰Paper https://lnkd.in/eFKSZnrc
πŸ‘‰Project https://lnkd.in/eSDfccv8
πŸ‘‰Repo https://lnkd.in/efCSvjtp
πŸ”₯1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ¦—Character Mixing GenerationπŸ¦—

πŸ‘‰MBZUAI unveils the first ever video-gen system able to preserve character ID, behavior & original style while generating plausible interactions between characters that have never coexisted - from cartoons (We Bare Bears, Tom & Jerry) to realistic humans (Mr. Bean, Young Sheldon)

πŸ‘‰Review https://t.ly/tN84a
πŸ‘‰Paper https://lnkd.in/dhKMwukv
πŸ‘‰Project https://lnkd.in/dBkJs48h
πŸ‘‰Repo https://lnkd.in/dw_uzgAk
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ¦„Unified Region-Level MLLMπŸ¦„

πŸ‘‰PixeRefers is an unified multimodal LLM framework that supports precise, region-specific understanding in both static images and dynamic videos, overcoming the holistic, scene-level bias of prior MLLMs. SOTA results. Demo, Repo & Dataset availableπŸ’™

πŸ‘‰Review https://t.ly/WH4dQ
πŸ‘‰Paper arxiv.org/pdf/2510.23603
πŸ‘‰Project circleradon.github.io/PixelRefer
πŸ‘‰Repo https://github.com/alibaba-damo-academy/PixelRefer
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‘’Generative View Stitching πŸ‘’

πŸ‘‰GVS is a novel approach that enables collision-free camera-guided video generation for predefined trajectories, it's a non-autoregressive alternative to video length extrapolation. Full repo under MITπŸ’™

πŸ‘‰Review https://t.ly/TiN_5
πŸ‘‰Paper https://arxiv.org/pdf/2510.24718
πŸ‘‰Project https://andrewsonga.github.io/gvs/
πŸ‘‰Repo github.com/andrewsonga/generative_view_stitching
Greetings from the SMART CITY WORLD CONGRESS in Barcellona. If you are around, ping me ;)
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”ͺTracking Object TransformationsπŸ”ͺ

πŸ‘‰"Track Any State": tracking objects through transformations while detecting/describing state changes. Repo & Dataset available under MITπŸ’™

πŸ‘‰Review https://t.ly/NPyW4
πŸ‘‰Paper https://lnkd.in/d4pA3bXJ
πŸ‘‰Project https://lnkd.in/dgbNfCuj
πŸ‘‰Repo https://lnkd.in/dtVWq2z7