This media is not supported in your browser
VIEW IN TELEGRAM
πFlowWAM: flow->action predictionπ
πFlowWAM is a novel dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Repo under Apacheπ
πReview https://t.ly/FmutT
πPaper https://arxiv.org/abs/2607.13017
πProject https://flow-wam.github.io/
πRepo github.com/YixiangChen515/FlowWAM
πFlowWAM is a novel dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Repo under Apacheπ
πReview https://t.ly/FmutT
πPaper https://arxiv.org/abs/2607.13017
πProject https://flow-wam.github.io/
πRepo github.com/YixiangChen515/FlowWAM
This media is not supported in your browser
VIEW IN TELEGRAM
π―SOTA Music-to-Dance Genπ―
πThe Tongyi Lab unveils Wan-Dancer, a novel stable minute-scale synthesis at 720p/30fps across five dance genres. Impressive results, new SOTA on long clip by a large margin. Repo under Apache 2.0π
πReview https://t.ly/AKY5j
πPaper https://lnkd.in/d_xA7dwb
πProject https://lnkd.in/dzfnw2h4
πRepo https://lnkd.in/d-Zj_cTf
πThe Tongyi Lab unveils Wan-Dancer, a novel stable minute-scale synthesis at 720p/30fps across five dance genres. Impressive results, new SOTA on long clip by a large margin. Repo under Apache 2.0π
πReview https://t.ly/AKY5j
πPaper https://lnkd.in/d_xA7dwb
πProject https://lnkd.in/dzfnw2h4
πRepo https://lnkd.in/d-Zj_cTf
This media is not supported in your browser
VIEW IN TELEGRAM
πNot a render. Not a concept. This is GENE.01 by Generative Bionics, the Italians coolest scaleup strikes back: in just six months, they turned GENE.01 into a fully functional humanoid platform that can walk, sense and interact.
πFull-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people.
πMore: https://t.ly/F3I3A
πFull-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people.
πMore: https://t.ly/F3I3A
This media is not supported in your browser
VIEW IN TELEGRAM
π’Unified Video Dense Predictionπ’
πUniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBRπ
πReview https://t.ly/oo7et
πPaper https://arxiv.org/pdf/2607.21592
πProject https://unid-video.github.io/
πRepo https://github.com/YihongSun/UniD
πUniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBRπ
πReview https://t.ly/oo7et
πPaper https://arxiv.org/pdf/2607.21592
πProject https://unid-video.github.io/
πRepo https://github.com/YihongSun/UniD
This media is not supported in your browser
VIEW IN TELEGRAM
πΏ Dawn of Generative Cinematography πΏ
π© #TheOdyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.
π Meanwhile, #AI research is heading in the exact opposite direction.
π© A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.
πMore https://t.ly/Kd7RV
πPaper arxiv.org/pdf/2607.24591
πProject yixuanli98.github.io/cameraanything/
πRepo github.com/yixuanli98/CameraAnything
π© #TheOdyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.
π Meanwhile, #AI research is heading in the exact opposite direction.
π© A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.
πMore https://t.ly/Kd7RV
πPaper arxiv.org/pdf/2607.24591
πProject yixuanli98.github.io/cameraanything/
πRepo github.com/yixuanli98/CameraAnything
π1
This media is not supported in your browser
VIEW IN TELEGRAM
π Dual-branch ID-Trackingπ
πTIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT licenseπ
πReview https://t.ly/WEDeY
πPaper https://arxiv.org/pdf/2607.26412
πProject https://vranlee.github.io/TIDE/
πRepo https://github.com/vranlee/TIDE
πTIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT licenseπ
πReview https://t.ly/WEDeY
πPaper https://arxiv.org/pdf/2607.26412
πProject https://vranlee.github.io/TIDE/
πRepo https://github.com/vranlee/TIDE
π₯Unified Points n' Linesπ₯
πETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under Apacheπ
πReview https://lnkd.in/p/eW8j5JZj
πPaper https://arxiv.org/pdf/2608.19894
πRepo https://github.com/francois141/upal
πETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under Apacheπ
πReview https://lnkd.in/p/eW8j5JZj
πPaper https://arxiv.org/pdf/2608.19894
πRepo https://github.com/francois141/upal
This media is not supported in your browser
VIEW IN TELEGRAM
π Anyone in 4D is out π
π4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0π
πReview https://lnkd.in/p/ec4dzGvb
πPaper https://arxiv.org/pdf/2608.20335
πProject https://4danyone.github.io
πRepo github.com/ant-research/4DAnyone
π4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0π
πReview https://lnkd.in/p/ec4dzGvb
πPaper https://arxiv.org/pdf/2608.20335
πProject https://4danyone.github.io
πRepo github.com/ant-research/4DAnyone
This media is not supported in your browser
VIEW IN TELEGRAM
πβπ©Remesh-Aware Mesh Deformationπβπ©
πRADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MITπ
πReview https://lnkd.in/p/eK6FZv9c
πPaper https://arxiv.org/pdf/2608.17182
πProject https://threedle.github.io/radmesh/
πRepo https://github.com/threedle/radmesh/
πRADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MITπ
πReview https://lnkd.in/p/eK6FZv9c
πPaper https://arxiv.org/pdf/2608.17182
πProject https://threedle.github.io/radmesh/
πRepo https://github.com/threedle/radmesh/
This media is not supported in your browser
VIEW IN TELEGRAM
π»Emerging Objs from Motionπ»
πMotion boundaries provide a strong signal for object-level grouping and can be used to derive pseudo-instance supervision. Suitable for: mono-depth, 3D object detection, 3D occupancy, and end-to-end planning. Repo under Apache 2.0π
πReview https://lnkd.in/p/eezZrSJE
πPaper https://arxiv.org/pdf/2609.04348
πProject https://tj12342.github.io/object-concepts-from-motion/
πRepo https://github.com/TJ12342/object-concepts-from-motion/tree/main
πMotion boundaries provide a strong signal for object-level grouping and can be used to derive pseudo-instance supervision. Suitable for: mono-depth, 3D object detection, 3D occupancy, and end-to-end planning. Repo under Apache 2.0π
πReview https://lnkd.in/p/eezZrSJE
πPaper https://arxiv.org/pdf/2609.04348
πProject https://tj12342.github.io/object-concepts-from-motion/
πRepo https://github.com/TJ12342/object-concepts-from-motion/tree/main
This media is not supported in your browser
VIEW IN TELEGRAM
πMcByte++ tracking-by-detectionπ
πMcByte++ is the newer extension of McByte that advances training-free sports MOT toward long-term ID tracking, while simultaneously improving efficiency and runtime performance. Repo under Apache 2.0π
πReview https://lnkd.in/p/e4-diVJS
πPaper https://lnkd.in/e_Vxky-b
πRepo https://lnkd.in/e8SeCYmk
πMcByte++ is the newer extension of McByte that advances training-free sports MOT toward long-term ID tracking, while simultaneously improving efficiency and runtime performance. Repo under Apache 2.0π
πReview https://lnkd.in/p/e4-diVJS
πPaper https://lnkd.in/e_Vxky-b
πRepo https://lnkd.in/e8SeCYmk
This media is not supported in your browser
VIEW IN TELEGRAM
π₯ RelateAnything is gold! π₯
πRelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0π
πReview https://lnkd.in/p/etAcdFM3
πPaper https://arxiv.org/pdf/2609.12552
πRepo https://github.com/Maelic/RelateAnything
πProject https://maelic.github.io/RelateAnythingProject/
πRelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0π
πReview https://lnkd.in/p/etAcdFM3
πPaper https://arxiv.org/pdf/2609.12552
πRepo https://github.com/Maelic/RelateAnything
πProject https://maelic.github.io/RelateAnythingProject/