Artificial Intelligence || DL
680 subscribers
25 photos
48 videos
10 files
261 links
Channel for who have a passion for -
* Artificial Intelligence
* Machine Learning
* Deep Learning
* Data Science
* Computer vision
* LLMs and NLP
Admin: @idrokdev
Download Telegram
This media is not supported in your browser
VIEW IN TELEGRAM
🌈FlowWAM: flow->action prediction🌈

πŸ‘‰FlowWAM is a novel dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Repo under ApacheπŸ’™

πŸ‘‰Review https://t.ly/FmutT
πŸ‘‰Paper https://arxiv.org/abs/2607.13017
πŸ‘‰Project https://flow-wam.github.io/
πŸ‘‰Repo github.com/YixiangChen515/FlowWAM
This media is not supported in your browser
VIEW IN TELEGRAM
🏯SOTA Music-to-Dance Gen🏯

πŸ‘‰The Tongyi Lab unveils Wan-Dancer, a novel stable minute-scale synthesis at 720p/30fps across five dance genres. Impressive results, new SOTA on long clip by a large margin. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://t.ly/AKY5j
πŸ‘‰Paper https://lnkd.in/d_xA7dwb
πŸ‘‰Project https://lnkd.in/dzfnw2h4
πŸ‘‰Repo https://lnkd.in/d-Zj_cTf
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‘‰Not a render. Not a concept. This is GENE.01 by Generative Bionics, the Italians coolest scaleup strikes back: in just six months, they turned GENE.01 into a fully functional humanoid platform that can walk, sense and interact.

πŸ‘‰Full-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people.

πŸ‘‰More: https://t.ly/F3I3A
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ’’Unified Video Dense PredictionπŸ’’

πŸ‘‰UniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBRπŸ’™

πŸ‘‰Review https://t.ly/oo7et
πŸ‘‰Paper https://arxiv.org/pdf/2607.21592
πŸ‘‰Project https://unid-video.github.io/
πŸ‘‰Repo https://github.com/YihongSun/UniD
This media is not supported in your browser
VIEW IN TELEGRAM
🍿 Dawn of Generative Cinematography 🍿

🟩 #TheOdyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.

πŸ‘‰ Meanwhile, #AI research is heading in the exact opposite direction.

🟩 A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.

πŸ‘‰More https://t.ly/Kd7RV
πŸ‘‰Paper arxiv.org/pdf/2607.24591
πŸ‘‰Project yixuanli98.github.io/cameraanything/
πŸ‘‰Repo github.com/yixuanli98/CameraAnything
πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🐠Dual-branch ID-Tracking🐠

πŸ‘‰TIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT licenseπŸ’™

πŸ‘‰Review https://t.ly/WEDeY
πŸ‘‰Paper https://arxiv.org/pdf/2607.26412
πŸ‘‰Project https://vranlee.github.io/TIDE/
πŸ‘‰Repo https://github.com/vranlee/TIDE
πŸ”₯Unified Points n' LinesπŸ”₯

πŸ‘‰ETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under ApacheπŸ’™

πŸ‘‰Review https://lnkd.in/p/eW8j5JZj
πŸ‘‰Paper https://arxiv.org/pdf/2608.19894
πŸ‘‰Repo https://github.com/francois141/upal
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ† Anyone in 4D is out πŸ†

πŸ‘‰4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/ec4dzGvb
πŸ‘‰Paper https://arxiv.org/pdf/2608.20335
πŸ‘‰Project https://4danyone.github.io
πŸ‘‰Repo github.com/ant-research/4DAnyone
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‹β€πŸŸ©Remesh-Aware Mesh DeformationπŸ‹β€πŸŸ©

πŸ‘‰RADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MITπŸ’™

πŸ‘‰Review https://lnkd.in/p/eK6FZv9c
πŸ‘‰Paper https://arxiv.org/pdf/2608.17182
πŸ‘‰Project https://threedle.github.io/radmesh/
πŸ‘‰Repo https://github.com/threedle/radmesh/
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‘»Emerging Objs from MotionπŸ‘»

πŸ‘‰Motion boundaries provide a strong signal for object-level grouping and can be used to derive pseudo-instance supervision. Suitable for: mono-depth, 3D object detection, 3D occupancy, and end-to-end planning. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/eezZrSJE
πŸ‘‰Paper https://arxiv.org/pdf/2609.04348
πŸ‘‰Project https://tj12342.github.io/object-concepts-from-motion/
πŸ‘‰Repo https://github.com/TJ12342/object-concepts-from-motion/tree/main
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ€McByte++ tracking-by-detectionπŸ€

πŸ‘‰McByte++ is the newer extension of McByte that advances training-free sports MOT toward long-term ID tracking, while simultaneously improving efficiency and runtime performance. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/e4-diVJS
πŸ‘‰Paper https://lnkd.in/e_Vxky-b
πŸ‘‰Repo https://lnkd.in/e8SeCYmk
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯ RelateAnything is gold! πŸ”₯

πŸ‘‰RelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/etAcdFM3
πŸ‘‰Paper https://arxiv.org/pdf/2609.12552
πŸ‘‰Repo https://github.com/Maelic/RelateAnything
πŸ‘‰Project https://maelic.github.io/RelateAnythingProject/