This media is not supported in your browser
VIEW IN TELEGRAM
🦜Streaming 4D Transformer🦜
👉IGGT4D is a novel a streaming instance-grounded geometry transformer for online 4D scene understanding. It processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. Repo/Data announced💙
👉Review https://t.ly/LFrKR
👉Paper https://arxiv.org/pdf/2607.19228
👉Project https://iggt4d.github.io/
👉Repo TBA
👉IGGT4D is a novel a streaming instance-grounded geometry transformer for online 4D scene understanding. It processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. Repo/Data announced💙
👉Review https://t.ly/LFrKR
👉Paper https://arxiv.org/pdf/2607.19228
👉Project https://iggt4d.github.io/
👉Repo TBA
❤3🔥2👏1
🫛Spatially-Aware Class-Agnostic Counting🫛
👉UpCount is reference-free, spatially aware, class-agnostic object counting with an MAE-pretrained ViT, DPT-style feature reassembly, FeatUp-style joint bilateral upsampling & proposal verification. Repo MIT💙
👉Review https://t.ly/dWOc3
👉Paper https://arxiv.org/pdf/2607.16826
👉Repo github.com/r28112072-rgb/upcount
👉UpCount is reference-free, spatially aware, class-agnostic object counting with an MAE-pretrained ViT, DPT-style feature reassembly, FeatUp-style joint bilateral upsampling & proposal verification. Repo MIT💙
👉Review https://t.ly/dWOc3
👉Paper https://arxiv.org/pdf/2607.16826
👉Repo github.com/r28112072-rgb/upcount
🔥5❤2👍1👏1
This media is not supported in your browser
VIEW IN TELEGRAM
💢Unified Video Dense Prediction💢
👉UniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBR💙
👉Review https://t.ly/oo7et
👉Paper https://arxiv.org/pdf/2607.21592
👉Project https://unid-video.github.io/
👉Repo https://github.com/YihongSun/UniD
👉UniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBR💙
👉Review https://t.ly/oo7et
👉Paper https://arxiv.org/pdf/2607.21592
👉Project https://unid-video.github.io/
👉Repo https://github.com/YihongSun/UniD
🔥7❤5👏1
Please open Telegram to view this post
VIEW IN TELEGRAM
This media is not supported in your browser
VIEW IN TELEGRAM
💄MagicMakeup Transfer💄
👉Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Authors: Zhejiang University & vivo BlueImage Lab. Repo for non commercial💙
👉Review https://t.ly/JYpCr
👉Paper harxiv.org/pdf/2607.20924
👉Project vivocameraresearch.github.io/magicmakeup
👉Repo github.com/vivoCameraResearch/Magic-Makeup
👉Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Authors: Zhejiang University & vivo BlueImage Lab. Repo for non commercial💙
👉Review https://t.ly/JYpCr
👉Paper harxiv.org/pdf/2607.20924
👉Project vivocameraresearch.github.io/magicmakeup
👉Repo github.com/vivoCameraResearch/Magic-Makeup
❤3👍1🔥1
This media is not supported in your browser
VIEW IN TELEGRAM
🔎MicroZoom at Extreme Scale🔎
👉MicroZoom by UWA synthesizes gigapixel-resolution images grounded in consumer-grade microscope close-ups at magnification levels up to 350×. Impressive. Repo under MIT💙
👉Review https://t.ly/hgJD7
👉Paper https://arxiv.org/pdf/2607.24729
👉Project https://microzoom-sr.github.io/
👉Repo github.com/MicroZoom-SR/MicroZoom-SR.github.io
👉MicroZoom by UWA synthesizes gigapixel-resolution images grounded in consumer-grade microscope close-ups at magnification levels up to 350×. Impressive. Repo under MIT💙
👉Review https://t.ly/hgJD7
👉Paper https://arxiv.org/pdf/2607.24729
👉Project https://microzoom-sr.github.io/
👉Repo github.com/MicroZoom-SR/MicroZoom-SR.github.io
🤯5❤3🔥1👏1
This media is not supported in your browser
VIEW IN TELEGRAM
🍿 Dawn of Generative Cinematography 🍿
🟩 The Odyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.
👉 Meanwhile, AI research is heading in the exact opposite direction.
🟩 A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.
👉More: https://t.ly/g-qUh
👉Paper arxiv.org/pdf/2607.24591
👉Project yixuanli98.github.io/cameraanything/
👉Repo github.com/yixuanli98/CameraAnything
🟩 The Odyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.
👉 Meanwhile, AI research is heading in the exact opposite direction.
🟩 A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.
👉More: https://t.ly/g-qUh
👉Paper arxiv.org/pdf/2607.24591
👉Project yixuanli98.github.io/cameraanything/
👉Repo github.com/yixuanli98/CameraAnything
🔥7❤5👍1👏1
This media is not supported in your browser
VIEW IN TELEGRAM
🔥Decoder-only Any-to-Any Model🔥
👉MODUS unifies any-to-any multimodal generation with one decoder, two experts, and zero task heads. Impressive work. Repo under Apache💙
👉Review https://t.ly/-2QKT
👉Paper https://lnkd.in/dhfBQGhB
👉Project https://lnkd.in/dPD_ECXk
👉Repo https://lnkd.in/dbDHw24u
👉MODUS unifies any-to-any multimodal generation with one decoder, two experts, and zero task heads. Impressive work. Repo under Apache💙
👉Review https://t.ly/-2QKT
👉Paper https://lnkd.in/dhfBQGhB
👉Project https://lnkd.in/dPD_ECXk
👉Repo https://lnkd.in/dbDHw24u
🔥5❤3👍1🍾1
This media is not supported in your browser
VIEW IN TELEGRAM
🐠Dual-branch ID-Tracking🐠
👉TIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT license💙
👉Review https://t.ly/WEDeY
👉Paper https://arxiv.org/pdf/2607.26412
👉Project https://vranlee.github.io/TIDE/
👉Repo https://github.com/vranlee/TIDE
👉TIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT license💙
👉Review https://t.ly/WEDeY
👉Paper https://arxiv.org/pdf/2607.26412
👉Project https://vranlee.github.io/TIDE/
👉Repo https://github.com/vranlee/TIDE
🔥9❤6👍1👏1
🔥Unified Points n' Lines🔥
👉ETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under Apache💙
👉Review https://lnkd.in/p/eW8j5JZj
👉Paper https://arxiv.org/pdf/2608.19894
👉Repo https://github.com/francois141/upal
👉ETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under Apache💙
👉Review https://lnkd.in/p/eW8j5JZj
👉Paper https://arxiv.org/pdf/2608.19894
👉Repo https://github.com/francois141/upal
❤10🔥6👍3👏1🍾1
This media is not supported in your browser
VIEW IN TELEGRAM
🐆 Anyone in 4D is out 🐆
👉4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0💙
👉Review https://lnkd.in/p/ec4dzGvb
👉Paper https://arxiv.org/pdf/2608.20335
👉Project https://4danyone.github.io
👉Repo github.com/ant-research/4DAnyone
👉4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0💙
👉Review https://lnkd.in/p/ec4dzGvb
👉Paper https://arxiv.org/pdf/2608.20335
👉Project https://4danyone.github.io
👉Repo github.com/ant-research/4DAnyone
🔥7❤2👍1👏1
This media is not supported in your browser
VIEW IN TELEGRAM
🍋🟩Remesh-Aware Mesh Deformation🍋🟩
👉RADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MIT💙
👉Review https://lnkd.in/p/eK6FZv9c
👉Paper https://arxiv.org/pdf/2608.17182
👉Project https://threedle.github.io/radmesh/
👉Repo https://github.com/threedle/radmesh/
👉RADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MIT💙
👉Review https://lnkd.in/p/eK6FZv9c
👉Paper https://arxiv.org/pdf/2608.17182
👉Project https://threedle.github.io/radmesh/
👉Repo https://github.com/threedle/radmesh/
🔥3❤2👍1👏1